Building Resilient AI Agents in .NET: Memory, Interruptions, and Long-Running Workflows
An AI agent that answers a question is relatively easy to build.
An AI agent that can complete useful work reliably is a different problem.
Imagine an agent that receives a request to review inventory, identify low-stock products, prepare a purchase recommendation, ask a manager for approval, update the system, and generate a summary.
The workflow may take seconds, minutes, or much longer depending on the tools involved.
What happens if the user closes the browser?
What happens if the application restarts?
What if the agent needs information from a human before continuing?
What if a tool fails halfway through the process?
And what happens when the agent needs information from an earlier interaction rather than only the current conversation?
These are not primarily language-model problems. They are application architecture problems.
This is where concepts such as agent memory, workflow state, checkpoints, interruption handling, human-in-the-loop interactions, and durable execution become important.
Microsoft Agent Framework provides capabilities for building these types of workflows in .NET and Python. Its workflow model can save execution state at checkpoints, pause for external requests, and resume execution later. Microsoft also provides durable workflow integration for scenarios where execution needs to survive process failures and restarts.
This article focuses on how .NET developers can think about resilient AI agents and how these capabilities fit together.
Why a Working AI Agent Can Still Fail in Production
A simple agent might look like this:
User
↓
AI Agent
↓
LLM
↓
Response
For a question-answering application, that may be enough.
A production agent usually looks more like:
User
↓
AI Agent
↓
Plan
↓
Tool
↓
Database
↓
Another Tool
↓
Human Approval
↓
External API
↓
Final Result
Every additional step introduces another possible failure point.
For example, consider an inventory assistant:
1. Read current stock
2. Find low-stock products
3. Check recent sales
4. Prepare purchase recommendation
5. Ask manager for approval
6. Create purchase order
7. Notify supplier
8. Generate report
If step 6 fails after the manager has already approved the purchase, the application needs to know what has already happened.
Simply sending the original prompt to the AI again may produce a completely different execution.
A resilient agent needs something more reliable than conversation history.
It needs state.
Memory and State Are Not the Same Thing
The words "memory" and "state" are often used interchangeably in AI applications, but they solve different problems.
Agent Memory
Memory helps an agent retain useful information across interactions.
For example:
User prefers monthly reports.
User manages the Dhaka warehouse.
User normally works with BDT.
Memory can be conversational, persistent, semantic, or application-specific depending on the architecture.
Workflow State
State describes what is currently happening inside a process.
For example:
Inventory analysis ✓
Sales analysis ✓
Purchase recommendation ✓
Manager approval ✓
Purchase order pending
Supplier notification pending
The distinction is important.
Memory answers:
What should the agent remember?
Workflow state answers:
Where is the operation right now?
A production system may need both.
Why Conversation History Is Not Enough
Suppose an agent has a conversation like this:
User:
Create a purchase order for the low-stock products.
Agent:
I found 12 products below their reorder levels.
User:
Proceed.
Agent:
The purchase order has been prepared.
Now imagine the application crashes before the purchase order is actually created.
When the application restarts, simply replaying the conversation does not guarantee that the agent knows which operations were successfully completed.
The conversation says:
The purchase order has been prepared.
But it may not tell the application whether the database transaction actually succeeded.
This is why resilient agent architectures separate:
Conversation history
Agent memory
Business state
Workflow state
External side effects
What Makes an AI Agent Resilient?
A resilient agent should be able to handle situations such as:
Application restart
Temporary tool failure
Network interruption
User interruption
Human approval
Long-running execution
Partial completion
External API failure
Multiple application instances
Delayed responses
A useful architecture is:
┌──────────────────┐
│ User │
└────────┬─────────┘
↓
┌──────────────────┐
│ AI Agent │
└────────┬─────────┘
↓
┌──────────────────┐
│ Workflow │
└────────┬─────────┘
↓
┌─────────────────────────┐
│ State + Checkpoint │
└───────────┬─────────────┘
↓
┌────────────────────────────────┐
│ Tools / DB / APIs / Services │
└────────────────────────────────┘
The AI handles reasoning.
The application controls execution.
That separation is extremely important.
Checkpoints: Saving Progress During Agent Execution
A checkpoint is a saved representation of workflow state.
Instead of treating the entire operation as one large execution, a workflow can save progress after meaningful stages.
For example:
Start
↓
Analyze inventory
↓
Checkpoint
↓
Generate recommendation
↓
Checkpoint
↓
Request approval
↓
Checkpoint
↓
Create purchase order
↓
Checkpoint
↓
Complete
If the process stops after the approval step, the application can resume from a known state instead of starting everything again.
Microsoft Agent Framework workflows create checkpoints at workflow superstep boundaries, and the saved state can include executor state, pending messages, requests, responses, and shared state.
This makes checkpoints useful for long-running workflows where losing all progress would be expensive.
A Simple .NET Workflow Concept
The .NET workflow API uses WorkflowBuilder to connect typed executors into an explicit execution graph.
Conceptually:
var workflow = new WorkflowBuilder(startExecutor)
.AddEdge(startExecutor, analysisExecutor)
.AddEdge(analysisExecutor, approvalExecutor)
.AddEdge(approvalExecutor, purchaseExecutor)
.Build();
The exact workflow depends on the application's requirements, but the architectural idea is important.
Instead of allowing the model to decide everything dynamically, developers can define critical execution boundaries in application code.
The agent can still perform reasoning inside individual steps.
The workflow controls the overall process.
Agent Reasoning vs Workflow Control
This distinction becomes especially useful in business applications.
Consider:
AI Agent:
"What products should be reordered?"
This is a reasoning problem.
But:
Can the purchase order be created?
Has the manager approved it?
Was the order already created?
Should the supplier be notified?
These are application-control problems.
You generally don't want the LLM to be the only source of truth for these decisions.
A safer architecture is:
LLM
↓
Recommendation
↓
Application validation
↓
Business rules
↓
Human approval
↓
Tool execution
This keeps AI reasoning inside controlled application boundaries.
Human-in-the-Loop Changes the Workflow
Some operations should not execute automatically.
For example:
AI:
I recommend purchasing 500 units.
System:
Manager approval required.
Manager:
Approve.
System:
Continue workflow.
This creates an interruption in the workflow.
The workflow cannot simply continue immediately.
It must wait for an external response.
Microsoft Agent Framework supports this through request and response handling. Executors can send requests outside the workflow and wait for responses, including human approvals. Pending requests can also be preserved in checkpoints so they can be restored when the workflow resumes.
This is particularly useful for:
Purchase approvals
Financial transactions
Customer communication
Production deployments
Account changes
Data deletion
High-value orders
Interruptions Are Not Always Failures
An important design principle is that an interruption does not necessarily mean something went wrong.
Sometimes the workflow is intentionally waiting.
For example:
Agent
↓
Prepare deployment
↓
Request approval
↓
WAIT
↓
Human approves
↓
Continue deployment
The waiting period may be minutes or hours.
That is not a failed workflow.
It is a workflow in a pending state.
Treating waiting as a first-class state makes agent applications much easier to reason about.
Resume Instead of Restart
Consider this workflow:
Step 1 → Complete
Step 2 → Complete
Step 3 → Waiting for approval
The user approves the operation two hours later.
A resilient implementation should resume from step 3.
It should not:
Restart Step 1
Restart Step 2
Restart Step 3
because repeating earlier side effects could cause duplicate operations.
For example, you don't want:
Create Purchase Order #1001
Create Purchase Order #1002
simply because an agent restarted.
The application should know which operations have already completed.
Idempotency Becomes Important
Resilient workflows also require careful handling of side effects.
Suppose the agent calls:
await orderService.CreatePurchaseOrderAsync(request);
If the network connection fails immediately after the server creates the order, the agent may not know whether the operation succeeded.
Retrying blindly could create a duplicate order.
A better approach is to use an idempotency key:
Workflow ID:
WF-2026-00982
Purchase Operation:
PO-CREATE-WF-2026-00982
The application can then recognize that the operation has already been processed.
This is not an AI-specific technique.
It is a standard distributed-systems principle that becomes especially important when AI agents control business operations.
Long-Running Agents Need Durable State
An in-memory workflow is useful for development and short-lived operations.
But imagine an agent performing a process that lasts several hours.
The application process could restart.
The server could be redeployed.
A container could be replaced.
A machine could fail.
If the state only exists in memory, the workflow may disappear.
For production scenarios, durable state becomes important.
Microsoft's Agent Framework documentation describes durable workflow execution using Durable Task integration, where workflow state can survive process restarts and failures, support long-running orchestration, and run across distributed infrastructure.
The architecture becomes:
AI Agent
↓
Workflow
↓
┌─────────────────┐
│ Durable State │
└────────┬────────┘
↓
┌─────────────────┐
│ Business Tools │
└─────────────────┘
In-Memory vs Durable Execution
For development:
In-memory state
↓
Simple
↓
Fast testing
For production:
Persistent checkpoint
↓
Restartable workflow
↓
Distributed execution
Microsoft provides different checkpoint storage options depending on the deployment scenario. The current documentation includes in-memory and file-based options as well as Cosmos DB-backed checkpoint storage.
The important point is not to use the most complicated option automatically.
Use the smallest architecture that satisfies the application's reliability requirements.
A Real-World .NET Example: Inventory Agent
Let's consider an inventory management system.
A user asks:
"Check our inventory and prepare a purchase order for products that need restocking."
The agent could execute:
1. Load inventory
2. Analyze reorder levels
3. Check recent sales
4. Calculate recommended quantities
5. Prepare purchase order
6. Request manager approval
7. Create purchase order
8. Notify supplier
9. Generate summary
A resilient implementation might look like:
User Request
↓
Inventory Agent
↓
Analyze Inventory
↓
Checkpoint
↓
Prepare Recommendation
↓
Checkpoint
↓
Manager Approval
↓
┌──────────┴──────────┐
│ │
Rejected Approved
│ │
Finish ↓
Create Order
↓
Checkpoint
↓
Notify Supplier
↓
Finish
Now suppose the application restarts while waiting for approval.
The workflow does not need to start from the beginning.
It can restore the saved state and continue after approval.
Memory Can Make the Agent More Useful
The same inventory agent may also use memory.
For example:
User preference:
Show purchase recommendations in BDT.
Business context:
Dhaka warehouse uses supplier group A.
Previous preference:
Purchase orders require manager approval.
This information can improve future interactions.
However, memory should not replace authoritative business data.
For example, the agent should not rely on memory to determine:
Current stock = 120 units
if the actual inventory database says:
Current stock = 87 units
The database is the source of truth.
A good architecture is:
Memory
↓
Preferences / context
Database
↓
Current business data
Workflow State
↓
Current execution status
Each serves a different purpose.
Memory Should Be Selective
It is tempting to save everything an agent sees.
That can create problems.
You generally don't want to persist every temporary thought, tool response, or irrelevant conversation fragment.
Instead, consider storing information that provides continuing value.
Examples:
User preferences
Business configuration
Useful long-term context
Approved workflow settings
Important historical facts
Avoid treating the entire conversation as permanent memory by default.
This also reduces storage requirements and makes privacy management easier.
What Happens When a Tool Fails?
Suppose the inventory agent reaches:
Create Purchase Order
↓
ERP API
↓
Timeout
The agent should not automatically assume:
Purchase order failed.
The API may have created the order before the response was lost.
A resilient system can use:
Idempotency keys
Operation status checks
Retry policies
Checkpoints
Transaction boundaries
Explicit failure states
The AI can explain the failure, but the application should determine whether the operation actually completed.
Retry Is Not the Same as Recovery
These two concepts are often confused.
Retry
Try the same operation again.
API call
↓
Timeout
↓
Retry
Recovery
Determine where the workflow stopped and continue from the correct state.
Step 1 ✓
Step 2 ✓
Step 3 ✓
Step 4 failed
↓
Recover
↓
Retry or compensate Step 4
For simple operations, retry may be enough.
For complex workflows, checkpoint-based recovery is much more appropriate.
Designing an ASP.NET Core Application
A practical ASP.NET Core architecture could look like:
ASP.NET Core API
↓
Agent Service
↓
Workflow
↓
Executors
┌─────┼───────────┐
↓ ↓ ↓
DB External API Approval
↓
Checkpoint Storage
The API should not contain the entire agent execution logic.
For example:
[ApiController]
[Route("api/agent")]
public class AgentController : ControllerBase
{
private readonly IInventoryAgent _agent;
public AgentController(IInventoryAgent agent)
{
_agent = agent;
}
[HttpPost("purchase-analysis")]
public async Task<IActionResult> Analyze(
PurchaseRequest request)
{
var result = await _agent.AnalyzeAsync(request);
return Ok(result);
}
}
The controller handles HTTP.
The agent service handles agent-related operations.
The workflow handles execution.
The business services handle actual business rules.
This separation keeps the architecture maintainable.
Don't Put Business Rules Inside the Prompt
A common mistake is to write something like:
You are an inventory manager.
Never create a purchase order above $10,000.
Always require approval for expensive orders.
and assume that this is sufficient enforcement.
It isn't.
A prompt can guide model behavior, but critical business rules should exist in application code.
For example:
if (purchaseOrder.Total > approvalLimit)
{
return ApprovalRequired();
}
The model can recommend.
The application decides.
This is especially important when the agent can execute tools that modify real data.
Human Approval Should Be a Real Application Boundary
Instead of:
Prompt:
Ask the user before creating an order.
design an explicit approval state:
PurchasePrepared
↓
ApprovalRequired
↓
Waiting
↓
Approved / Rejected
Now the application knows exactly where the workflow is.
It can show the pending request in a dashboard.
It can send an email or notification.
It can wait for the manager.
It can restore the request after a restart.
That is much more reliable than relying only on conversation instructions.
Testing Interrupted Agents
Testing an AI agent should not only mean checking whether the final answer is correct.
Test what happens when the workflow is interrupted.
For example:
Scenario 1
Agent starts
→ Application stops
→ Application restarts
→ Workflow resumes
Scenario 2
Agent requests approval
→ User closes browser
→ User returns later
→ Approval submitted
→ Workflow continues
Scenario 3
External API timeout
→ Retry
→ API succeeds
→ Workflow continues
Scenario 4
Database operation succeeds
→ Network response lost
→ Workflow retries
→ Duplicate operation prevented
Scenario 5
Manager rejects purchase
→ Workflow terminates safely
→ No supplier notification
These tests often reveal more architectural problems than testing only successful executions.
Observability Matters
When an agent performs multiple steps, logging only the final response is not enough.
Useful telemetry can include:
Workflow ID
Agent ID
User ID
Step
Executor
Tool
Start time
Duration
Status
Retry count
Checkpoint ID
Approval status
Error
For example:
Workflow: WF-2026-00982
Step: CreatePurchaseOrder
Status: Waiting
Reason: ManagerApproval
This makes production troubleshooting much easier.
Microsoft Agent Framework also exposes workflow and executor events that can be used for monitoring and application-level observability.
When Should You Use a Workflow?
Not every AI feature needs a workflow.
For example:
User → AI → Answer
doesn't require a complex orchestration layer.
A workflow becomes more useful when you have:
Multiple steps
Multiple tools
Business rules
Human approval
Long-running execution
External services
State transitions
Retry/recovery requirements
Audit requirements
A simple rule is:
If losing the agent's progress would matter, start thinking about workflow state and recovery.
When Should You Use Durable Execution?
Durable execution becomes especially relevant when:
Execution can last a long time
OR
Process restarts are possible
OR
Multiple machines may execute the workflow
OR
Human interaction can pause execution
OR
Losing progress is expensive
For a short chatbot request, this may be unnecessary.
For an AI-driven business process that can run for hours or days, it can become an important architectural requirement.
A Practical Architecture for Production
A production-oriented .NET AI application might eventually look like:
┌──────────────┐
│ User │
└──────┬───────┘
↓
┌──────────────┐
│ ASP.NET Core │
│ API │
└──────┬───────┘
↓
┌──────────────┐
│ AI Agent │
└──────┬───────┘
↓
┌──────────────┐
│ Workflow │
└──────┬───────┘
↓
┌─────────────────────────────────┐
│ │
↓ ↓
Agent Memory Workflow State
│ │
↓ ↓
Memory Provider Checkpoint Storage
│ │
└──────────────┬──────────────────┘
↓
┌─────────────────┐
│ Business Tools │
└───────┬─────────┘
↓
┌──────────────────────────┐
│ Database / APIs / Services│
└──────────────────────────┘
The exact technologies can change, but the responsibilities should remain clear.
Common Mistakes to Avoid
1. Treating conversation history as workflow state
Conversation history doesn't reliably represent completed business operations.
2. Letting the LLM control critical business rules
Use application code for authorization, validation, limits, and state transitions.
3. Retrying side effects blindly
A retry can create duplicate records or transactions.
4. Saving everything as memory
Not every piece of context deserves long-term persistence.
5. Using in-memory state for long-running production workflows
A process restart can destroy in-memory progress.
6. Ignoring human approval as a workflow state
Approval is an explicit application event, not merely a sentence in a prompt.
7. Testing only successful execution
Production failures often happen between steps rather than inside the final response.
8. Building a complex multi-agent system too early
Start with the simplest workflow that satisfies the business requirement.
Microsoft's current workflow guidance similarly emphasizes using the simplest orchestration pattern that meets the application's requirements rather than adding complexity unnecessarily.
A Practical Checklist
Before putting a long-running AI agent into production, ask:
Does the agent need persistent memory?
What information should actually be remembered?
What is the authoritative source for business data?
What happens if the application restarts?
Where are checkpoints created?
Can a workflow pause for human approval?
Can a paused workflow resume later?
Are external operations idempotent?
What happens when an API times out?
Can duplicate operations occur?
Can the workflow be audited?
Are critical business rules enforced outside the prompt?
Can developers see where a workflow failed?
Is the state storage appropriate for multiple application instances?
Have interrupted workflows been tested?
If these questions have clear answers, the agent is much closer to being production-ready.
Final Thoughts
Building an AI agent is not only about connecting an LLM to a set of tools.
The difficult part begins when the agent needs to perform work that takes time, depends on external systems, requires human approval, or must survive interruptions.
That is where memory, workflow state, checkpoints, idempotency, human-in-the-loop interactions, and durable execution become important.
For .NET developers, Microsoft Agent Framework provides building blocks for these scenarios. Workflows can define explicit execution paths, checkpoints can preserve progress, request/response mechanisms can support human interaction, and durable execution can help long-running processes survive failures and restarts.
The most important architectural principle is simple:
Let the AI reason, but let the application control execution.
The model can recommend what should happen next. Your .NET application should remain responsible for authorization, validation, business rules, state, persistence, and side effects.
That separation makes AI agents easier to test, easier to monitor, and much safer to operate in real-world applications.
For a simple chatbot, you may not need any of this.
For an AI agent that can perform real business work and continue where it left off after an interruption, resilience is no longer an optional feature. It becomes part of the architecture.
Comments 0