DOTNET EXPERT BLOG

Building Resilient AI Agents in .NET: Memory, Interruptions, and Long-Running Workflows

9/25/2026 8:02:36 AM Noor All Safaet Loading... 0


An AI agent that answers a question is relatively easy to build.

An AI agent that can complete useful work reliably is a different problem.

Imagine an agent that receives a request to review inventory, identify low-stock products, prepare a purchase recommendation, ask a manager for approval, update the system, and generate a summary.

The workflow may take seconds, minutes, or much longer depending on the tools involved.

What happens if the user closes the browser?

What happens if the application restarts?

What if the agent needs information from a human before continuing?

What if a tool fails halfway through the process?

And what happens when the agent needs information from an earlier interaction rather than only the current conversation?

These are not primarily language-model problems. They are application architecture problems.

This is where concepts such as agent memory, workflow state, checkpoints, interruption handling, human-in-the-loop interactions, and durable execution become important.

Microsoft Agent Framework provides capabilities for building these types of workflows in .NET and Python. Its workflow model can save execution state at checkpoints, pause for external requests, and resume execution later. Microsoft also provides durable workflow integration for scenarios where execution needs to survive process failures and restarts.

This article focuses on how .NET developers can think about resilient AI agents and how these capabilities fit together.

Why a Working AI Agent Can Still Fail in Production

A simple agent might look like this:

User
  ↓
AI Agent
  ↓
LLM
  ↓
Response

For a question-answering application, that may be enough.

A production agent usually looks more like:

User
  ↓
AI Agent
  ↓
Plan
  ↓
Tool
  ↓
Database
  ↓
Another Tool
  ↓
Human Approval
  ↓
External API
  ↓
Final Result

Every additional step introduces another possible failure point.

For example, consider an inventory assistant:

1. Read current stock
2. Find low-stock products
3. Check recent sales
4. Prepare purchase recommendation
5. Ask manager for approval
6. Create purchase order
7. Notify supplier
8. Generate report

If step 6 fails after the manager has already approved the purchase, the application needs to know what has already happened.

Simply sending the original prompt to the AI again may produce a completely different execution.

A resilient agent needs something more reliable than conversation history.

It needs state.

Memory and State Are Not the Same Thing

The words "memory" and "state" are often used interchangeably in AI applications, but they solve different problems.

Agent Memory

Memory helps an agent retain useful information across interactions.

For example:

User prefers monthly reports.
User manages the Dhaka warehouse.
User normally works with BDT.

Memory can be conversational, persistent, semantic, or application-specific depending on the architecture.

Workflow State

State describes what is currently happening inside a process.

For example:

Inventory analysis       ✓
Sales analysis           ✓
Purchase recommendation ✓
Manager approval         ✓
Purchase order           pending
Supplier notification    pending

The distinction is important.

Memory answers:

What should the agent remember?

Workflow state answers:

Where is the operation right now?

A production system may need both.

Why Conversation History Is Not Enough

Suppose an agent has a conversation like this:

User:
Create a purchase order for the low-stock products.

Agent:
I found 12 products below their reorder levels.

User:
Proceed.

Agent:
The purchase order has been prepared.

Now imagine the application crashes before the purchase order is actually created.

When the application restarts, simply replaying the conversation does not guarantee that the agent knows which operations were successfully completed.

The conversation says:

The purchase order has been prepared.

But it may not tell the application whether the database transaction actually succeeded.

This is why resilient agent architectures separate:

  • Conversation history

  • Agent memory

  • Business state

  • Workflow state

  • External side effects

What Makes an AI Agent Resilient?

A resilient agent should be able to handle situations such as:

  • Application restart

  • Temporary tool failure

  • Network interruption

  • User interruption

  • Human approval

  • Long-running execution

  • Partial completion

  • External API failure

  • Multiple application instances

  • Delayed responses

A useful architecture is:

                 ┌──────────────────┐
                 │       User       │
                 └────────┬─────────┘
                          ↓
                 ┌──────────────────┐
                 │    AI Agent      │
                 └────────┬─────────┘
                          ↓
                 ┌──────────────────┐
                 │     Workflow     │
                 └────────┬─────────┘
                          ↓
             ┌─────────────────────────┐
             │ State + Checkpoint      │
             └───────────┬─────────────┘
                         ↓
        ┌────────────────────────────────┐
        │ Tools / DB / APIs / Services   │
        └────────────────────────────────┘

The AI handles reasoning.

The application controls execution.

That separation is extremely important.

Checkpoints: Saving Progress During Agent Execution

A checkpoint is a saved representation of workflow state.

Instead of treating the entire operation as one large execution, a workflow can save progress after meaningful stages.

For example:

Start
  ↓
Analyze inventory
  ↓
Checkpoint
  ↓
Generate recommendation
  ↓
Checkpoint
  ↓
Request approval
  ↓
Checkpoint
  ↓
Create purchase order
  ↓
Checkpoint
  ↓
Complete

If the process stops after the approval step, the application can resume from a known state instead of starting everything again.

Microsoft Agent Framework workflows create checkpoints at workflow superstep boundaries, and the saved state can include executor state, pending messages, requests, responses, and shared state.

This makes checkpoints useful for long-running workflows where losing all progress would be expensive.

A Simple .NET Workflow Concept

The .NET workflow API uses WorkflowBuilder to connect typed executors into an explicit execution graph.

Conceptually:

var workflow = new WorkflowBuilder(startExecutor)
    .AddEdge(startExecutor, analysisExecutor)
    .AddEdge(analysisExecutor, approvalExecutor)
    .AddEdge(approvalExecutor, purchaseExecutor)
    .Build();

The exact workflow depends on the application's requirements, but the architectural idea is important.

Instead of allowing the model to decide everything dynamically, developers can define critical execution boundaries in application code.

The agent can still perform reasoning inside individual steps.

The workflow controls the overall process.

Agent Reasoning vs Workflow Control

This distinction becomes especially useful in business applications.

Consider:

AI Agent:
"What products should be reordered?"

This is a reasoning problem.

But:

Can the purchase order be created?
Has the manager approved it?
Was the order already created?
Should the supplier be notified?

These are application-control problems.

You generally don't want the LLM to be the only source of truth for these decisions.

A safer architecture is:

LLM
 ↓
Recommendation
 ↓
Application validation
 ↓
Business rules
 ↓
Human approval
 ↓
Tool execution

This keeps AI reasoning inside controlled application boundaries.

Human-in-the-Loop Changes the Workflow

Some operations should not execute automatically.

For example:

AI:
I recommend purchasing 500 units.

System:
Manager approval required.

Manager:
Approve.

System:
Continue workflow.

This creates an interruption in the workflow.

The workflow cannot simply continue immediately.

It must wait for an external response.

Microsoft Agent Framework supports this through request and response handling. Executors can send requests outside the workflow and wait for responses, including human approvals. Pending requests can also be preserved in checkpoints so they can be restored when the workflow resumes.

This is particularly useful for:

  • Purchase approvals

  • Financial transactions

  • Customer communication

  • Production deployments

  • Account changes

  • Data deletion

  • High-value orders

Interruptions Are Not Always Failures

An important design principle is that an interruption does not necessarily mean something went wrong.

Sometimes the workflow is intentionally waiting.

For example:

Agent
  ↓
Prepare deployment
  ↓
Request approval
  ↓
WAIT
  ↓
Human approves
  ↓
Continue deployment

The waiting period may be minutes or hours.

That is not a failed workflow.

It is a workflow in a pending state.

Treating waiting as a first-class state makes agent applications much easier to reason about.

Resume Instead of Restart

Consider this workflow:

Step 1 → Complete
Step 2 → Complete
Step 3 → Waiting for approval

The user approves the operation two hours later.

A resilient implementation should resume from step 3.

It should not:

Restart Step 1
Restart Step 2
Restart Step 3

because repeating earlier side effects could cause duplicate operations.

For example, you don't want:

Create Purchase Order #1001
Create Purchase Order #1002

simply because an agent restarted.

The application should know which operations have already completed.

Idempotency Becomes Important

Resilient workflows also require careful handling of side effects.

Suppose the agent calls:

await orderService.CreatePurchaseOrderAsync(request);

If the network connection fails immediately after the server creates the order, the agent may not know whether the operation succeeded.

Retrying blindly could create a duplicate order.

A better approach is to use an idempotency key:

Workflow ID:
WF-2026-00982

Purchase Operation:
PO-CREATE-WF-2026-00982

The application can then recognize that the operation has already been processed.

This is not an AI-specific technique.

It is a standard distributed-systems principle that becomes especially important when AI agents control business operations.

Long-Running Agents Need Durable State

An in-memory workflow is useful for development and short-lived operations.

But imagine an agent performing a process that lasts several hours.

The application process could restart.

The server could be redeployed.

A container could be replaced.

A machine could fail.

If the state only exists in memory, the workflow may disappear.

For production scenarios, durable state becomes important.

Microsoft's Agent Framework documentation describes durable workflow execution using Durable Task integration, where workflow state can survive process restarts and failures, support long-running orchestration, and run across distributed infrastructure.

The architecture becomes:

             AI Agent
                 ↓
             Workflow
                 ↓
       ┌─────────────────┐
       │ Durable State   │
       └────────┬────────┘
                ↓
       ┌─────────────────┐
       │ Business Tools  │
       └─────────────────┘

In-Memory vs Durable Execution

For development:

In-memory state
      ↓
Simple
      ↓
Fast testing

For production:

Persistent checkpoint
      ↓
Restartable workflow
      ↓
Distributed execution

Microsoft provides different checkpoint storage options depending on the deployment scenario. The current documentation includes in-memory and file-based options as well as Cosmos DB-backed checkpoint storage.

The important point is not to use the most complicated option automatically.

Use the smallest architecture that satisfies the application's reliability requirements.

A Real-World .NET Example: Inventory Agent

Let's consider an inventory management system.

A user asks:

"Check our inventory and prepare a purchase order for products that need restocking."

The agent could execute:

1. Load inventory
2. Analyze reorder levels
3. Check recent sales
4. Calculate recommended quantities
5. Prepare purchase order
6. Request manager approval
7. Create purchase order
8. Notify supplier
9. Generate summary

A resilient implementation might look like:

                User Request
                     ↓
              Inventory Agent
                     ↓
             Analyze Inventory
                     ↓
                Checkpoint
                     ↓
          Prepare Recommendation
                     ↓
                Checkpoint
                     ↓
             Manager Approval
                     ↓
          ┌──────────┴──────────┐
          │                     │
       Rejected              Approved
          │                     │
        Finish                  ↓
                           Create Order
                                ↓
                           Checkpoint
                                ↓
                       Notify Supplier
                                ↓
                             Finish

Now suppose the application restarts while waiting for approval.

The workflow does not need to start from the beginning.

It can restore the saved state and continue after approval.

Memory Can Make the Agent More Useful

The same inventory agent may also use memory.

For example:

User preference:
Show purchase recommendations in BDT.

Business context:
Dhaka warehouse uses supplier group A.

Previous preference:
Purchase orders require manager approval.

This information can improve future interactions.

However, memory should not replace authoritative business data.

For example, the agent should not rely on memory to determine:

Current stock = 120 units

if the actual inventory database says:

Current stock = 87 units

The database is the source of truth.

A good architecture is:

Memory
   ↓
Preferences / context

Database
   ↓
Current business data

Workflow State
   ↓
Current execution status

Each serves a different purpose.

Memory Should Be Selective

It is tempting to save everything an agent sees.

That can create problems.

You generally don't want to persist every temporary thought, tool response, or irrelevant conversation fragment.

Instead, consider storing information that provides continuing value.

Examples:

User preferences
Business configuration
Useful long-term context
Approved workflow settings
Important historical facts

Avoid treating the entire conversation as permanent memory by default.

This also reduces storage requirements and makes privacy management easier.

What Happens When a Tool Fails?

Suppose the inventory agent reaches:

Create Purchase Order
       ↓
ERP API
       ↓
Timeout

The agent should not automatically assume:

Purchase order failed.

The API may have created the order before the response was lost.

A resilient system can use:

  1. Idempotency keys

  2. Operation status checks

  3. Retry policies

  4. Checkpoints

  5. Transaction boundaries

  6. Explicit failure states

The AI can explain the failure, but the application should determine whether the operation actually completed.

Retry Is Not the Same as Recovery

These two concepts are often confused.

Retry

Try the same operation again.

API call
 ↓
Timeout
 ↓
Retry

Recovery

Determine where the workflow stopped and continue from the correct state.

Step 1 ✓
Step 2 ✓
Step 3 ✓
Step 4 failed
       ↓
Recover
       ↓
Retry or compensate Step 4

For simple operations, retry may be enough.

For complex workflows, checkpoint-based recovery is much more appropriate.

Designing an ASP.NET Core Application

A practical ASP.NET Core architecture could look like:

ASP.NET Core API
       ↓
Agent Service
       ↓
Workflow
       ↓
Executors
 ┌─────┼───────────┐
 ↓     ↓           ↓
DB   External API  Approval
       ↓
Checkpoint Storage

The API should not contain the entire agent execution logic.

For example:

[ApiController]
[Route("api/agent")]
public class AgentController : ControllerBase
{
    private readonly IInventoryAgent _agent;

    public AgentController(IInventoryAgent agent)
    {
        _agent = agent;
    }

    [HttpPost("purchase-analysis")]
    public async Task<IActionResult> Analyze(
        PurchaseRequest request)
    {
        var result = await _agent.AnalyzeAsync(request);

        return Ok(result);
    }
}

The controller handles HTTP.

The agent service handles agent-related operations.

The workflow handles execution.

The business services handle actual business rules.

This separation keeps the architecture maintainable.

Don't Put Business Rules Inside the Prompt

A common mistake is to write something like:

You are an inventory manager.

Never create a purchase order above $10,000.

Always require approval for expensive orders.

and assume that this is sufficient enforcement.

It isn't.

A prompt can guide model behavior, but critical business rules should exist in application code.

For example:

if (purchaseOrder.Total > approvalLimit)
{
    return ApprovalRequired();
}

The model can recommend.

The application decides.

This is especially important when the agent can execute tools that modify real data.

Human Approval Should Be a Real Application Boundary

Instead of:

Prompt:
Ask the user before creating an order.

design an explicit approval state:

PurchasePrepared
      ↓
ApprovalRequired
      ↓
Waiting
      ↓
Approved / Rejected

Now the application knows exactly where the workflow is.

It can show the pending request in a dashboard.

It can send an email or notification.

It can wait for the manager.

It can restore the request after a restart.

That is much more reliable than relying only on conversation instructions.

Testing Interrupted Agents

Testing an AI agent should not only mean checking whether the final answer is correct.

Test what happens when the workflow is interrupted.

For example:

Scenario 1

Agent starts
→ Application stops
→ Application restarts
→ Workflow resumes

Scenario 2

Agent requests approval
→ User closes browser
→ User returns later
→ Approval submitted
→ Workflow continues

Scenario 3

External API timeout
→ Retry
→ API succeeds
→ Workflow continues

Scenario 4

Database operation succeeds
→ Network response lost
→ Workflow retries
→ Duplicate operation prevented

Scenario 5

Manager rejects purchase
→ Workflow terminates safely
→ No supplier notification

These tests often reveal more architectural problems than testing only successful executions.

Observability Matters

When an agent performs multiple steps, logging only the final response is not enough.

Useful telemetry can include:

Workflow ID
Agent ID
User ID
Step
Executor
Tool
Start time
Duration
Status
Retry count
Checkpoint ID
Approval status
Error

For example:

Workflow: WF-2026-00982
Step: CreatePurchaseOrder
Status: Waiting
Reason: ManagerApproval

This makes production troubleshooting much easier.

Microsoft Agent Framework also exposes workflow and executor events that can be used for monitoring and application-level observability.

When Should You Use a Workflow?

Not every AI feature needs a workflow.

For example:

User → AI → Answer

doesn't require a complex orchestration layer.

A workflow becomes more useful when you have:

  • Multiple steps

  • Multiple tools

  • Business rules

  • Human approval

  • Long-running execution

  • External services

  • State transitions

  • Retry/recovery requirements

  • Audit requirements

A simple rule is:

If losing the agent's progress would matter, start thinking about workflow state and recovery.

When Should You Use Durable Execution?

Durable execution becomes especially relevant when:

Execution can last a long time
            OR
Process restarts are possible
            OR
Multiple machines may execute the workflow
            OR
Human interaction can pause execution
            OR
Losing progress is expensive

For a short chatbot request, this may be unnecessary.

For an AI-driven business process that can run for hours or days, it can become an important architectural requirement.

A Practical Architecture for Production

A production-oriented .NET AI application might eventually look like:

                    ┌──────────────┐
                    │     User     │
                    └──────┬───────┘
                           ↓
                    ┌──────────────┐
                    │ ASP.NET Core │
                    │      API     │
                    └──────┬───────┘
                           ↓
                    ┌──────────────┐
                    │ AI Agent     │
                    └──────┬───────┘
                           ↓
                    ┌──────────────┐
                    │   Workflow   │
                    └──────┬───────┘
                           ↓
        ┌─────────────────────────────────┐
        │                                 │
        ↓                                 ↓
  Agent Memory                     Workflow State
        │                                 │
        ↓                                 ↓
 Memory Provider                  Checkpoint Storage
        │                                 │
        └──────────────┬──────────────────┘
                       ↓
              ┌─────────────────┐
              │ Business Tools  │
              └───────┬─────────┘
                      ↓
        ┌──────────────────────────┐
        │ Database / APIs / Services│
        └──────────────────────────┘

The exact technologies can change, but the responsibilities should remain clear.

Common Mistakes to Avoid

1. Treating conversation history as workflow state

Conversation history doesn't reliably represent completed business operations.

2. Letting the LLM control critical business rules

Use application code for authorization, validation, limits, and state transitions.

3. Retrying side effects blindly

A retry can create duplicate records or transactions.

4. Saving everything as memory

Not every piece of context deserves long-term persistence.

5. Using in-memory state for long-running production workflows

A process restart can destroy in-memory progress.

6. Ignoring human approval as a workflow state

Approval is an explicit application event, not merely a sentence in a prompt.

7. Testing only successful execution

Production failures often happen between steps rather than inside the final response.

8. Building a complex multi-agent system too early

Start with the simplest workflow that satisfies the business requirement.

Microsoft's current workflow guidance similarly emphasizes using the simplest orchestration pattern that meets the application's requirements rather than adding complexity unnecessarily.

A Practical Checklist

Before putting a long-running AI agent into production, ask:

  • Does the agent need persistent memory?

  • What information should actually be remembered?

  • What is the authoritative source for business data?

  • What happens if the application restarts?

  • Where are checkpoints created?

  • Can a workflow pause for human approval?

  • Can a paused workflow resume later?

  • Are external operations idempotent?

  • What happens when an API times out?

  • Can duplicate operations occur?

  • Can the workflow be audited?

  • Are critical business rules enforced outside the prompt?

  • Can developers see where a workflow failed?

  • Is the state storage appropriate for multiple application instances?

  • Have interrupted workflows been tested?

If these questions have clear answers, the agent is much closer to being production-ready.

Final Thoughts

Building an AI agent is not only about connecting an LLM to a set of tools.

The difficult part begins when the agent needs to perform work that takes time, depends on external systems, requires human approval, or must survive interruptions.

That is where memory, workflow state, checkpoints, idempotency, human-in-the-loop interactions, and durable execution become important.

For .NET developers, Microsoft Agent Framework provides building blocks for these scenarios. Workflows can define explicit execution paths, checkpoints can preserve progress, request/response mechanisms can support human interaction, and durable execution can help long-running processes survive failures and restarts.

The most important architectural principle is simple:

Let the AI reason, but let the application control execution.

The model can recommend what should happen next. Your .NET application should remain responsible for authorization, validation, business rules, state, persistence, and side effects.

That separation makes AI agents easier to test, easier to monitor, and much safer to operate in real-world applications.

For a simple chatbot, you may not need any of this.

For an AI agent that can perform real business work and continue where it left off after an interruption, resilience is no longer an optional feature. It becomes part of the architecture.

Comments 0