DOTNET EXPERT BLOG

Add Conversation Memory to a C# AI Application with ASP.NET Core | Day 4

9/29/2026 4:42:25 AM Noor All Safaet Loading... 0

Our AI application can answer questions, but it still has one major limitation:

It does not remember the conversation.

Suppose a user sends:

My name is Alex.

The AI responds:

Nice to meet you, Alex.

Then the user asks:

What is my name?

If we send only the latest message to the model, it has no access to the earlier conversation.

Today, we will fix that.

In Day 4 of our C# AI tutorial series, we will add conversation memory to the ASP.NET Core application we built on Day 3.

We will learn how to:

  • create conversation IDs

  • store chat history

  • keep conversations isolated

  • send previous messages to the model

  • store AI responses back in history

  • handle concurrent requests safely

  • limit conversation history

  • understand the limitations of in-memory storage

  • prepare the architecture for persistent storage later

By the end of this tutorial, our API will support real multi-turn conversations.


Previously in This Series

So far, we have built the application step by step.

Day 1: Build Your First AI Application with C# and .NET

We connected a C# application to an AI model.

Day 2: Understanding IChatClient in .NET

We learned why Microsoft.Extensions.AI provides a useful abstraction between our application and AI providers.

Day 3: Build a Reusable AI Chat Service in ASP.NET Core

We moved the AI integration into an ASP.NET Core Web API using dependency injection, validation, secure configuration, and an application-level AI service.

Our previous request flow was:

HTTP Request
      ↓
AiController
      ↓
IAiChatService
      ↓
IChatClient
      ↓
AI Model
      ↓
Response

Today we will add:

Conversation History

to that flow.


What Does Conversation Memory Mean?

The phrase AI memory can describe several different techniques.

For today's application, conversation memory means:

The application stores previous messages and sends relevant conversation history to the model with each new request.

The model is not automatically remembering every HTTP request our application previously sent.

Our application manages that context.

Consider:

User:
My name is Alex.

Assistant:
Nice to meet you, Alex.

User:
What is my name?

For the final question, our application can send:

System:
You are a helpful AI assistant.

User:
My name is Alex.

Assistant:
Nice to meet you, Alex.

User:
What is my name?

The model now has the context needed to answer:

Your name is Alex.

That is the basic conversation-memory pattern we will implement.


Conversation Memory vs Long-Term Memory

Before writing code, we should separate two concepts.

Conversation Memory

Conversation memory contains context from the current conversation.

For example:

User:
My company sells computer accessories.

Assistant:
How can I help?

User:
What type of products did I say we sell?

The previous message provides the context needed to understand the latest question.

Long-Term Memory

Long-term memory may contain information that survives beyond one conversation.

Examples include:

  • user preferences

  • business information

  • saved instructions

  • profile information

  • previous project context

Long-term memory normally requires persistent storage or another dedicated memory strategy.

We are not implementing long-term memory today.

Day 4 focuses specifically on multi-turn conversation history.


Why Day 3 Had No Memory

On Day 3, we created a new message collection for every request:

var messages = new List<ChatMessage>
{
    new(
        ChatRole.System,
        SystemPrompt),

    new(
        ChatRole.User,
        message)
};

Then:

var response =
    await _chatClient.GetResponseAsync(
        messages,
        cancellationToken: cancellationToken);

When the request finished, that local collection disappeared.

The next HTTP request created another collection.

Conceptually:

Request 1
   ↓
Temporary Messages
   ↓
AI Response
   ↓
Request Ends

Then:

Request 2
   ↓
New Messages

We need a way to connect those requests.


Our New Architecture

The application will now work like this:

Client
  ↓
Conversation ID
  ↓
ASP.NET Core API
  ↓
Conversation Store
  ↓
Previous Messages
  ↓
AI Service
  ↓
IChatClient
  ↓
AI Model
  ↓
Assistant Response
  ↓
Updated Conversation Store

For Day 4, our conversation store will use application memory.

This keeps the implementation easy to understand while preserving a storage abstraction that we can replace later.


Step 1: Update the Request Model

On Day 3, our request looked like:

{
  "message": "Explain dependency injection."
}

Now we also need to identify the conversation.

Update ChatRequest.cs:

using System.ComponentModel.DataAnnotations;

public sealed class ChatRequest
{
    public Guid? ConversationId { get; set; }

    [Required]
    [StringLength(
        4000,
        MinimumLength = 1,
        ErrorMessage = "Message must contain between 1 and 4000 characters.")]
    public string Message { get; set; } = string.Empty;
}

For the first message, the client can omit ConversationId:

{
  "message": "My name is Alex."
}

The server creates a new conversation.

For the next message, the client sends the returned ID:

{
  "conversationId": "5f7d95d7-bf12-4d32-9c42-7366d50c8fc3",
  "message": "What is my name?"
}

Now the server knows which conversation should be continued.


Step 2: Update the Response Model

The API must return the conversation ID so the client can reuse it.

Update ChatResponse.cs:

public sealed class ChatResponse
{
    public Guid ConversationId { get; set; }

    public string Answer { get; set; } = string.Empty;
}

The first response can now look like:

{
  "conversationId": "5f7d95d7-bf12-4d32-9c42-7366d50c8fc3",
  "answer": "Nice to meet you, Alex."
}

The client stores that ID and sends it with the next message.


Step 3: Create the Conversation Model

Create:

Models/AiConversation.cs

Add:

using Microsoft.Extensions.AI;

public sealed class AiConversation
{
    public Guid Id { get; init; } = Guid.NewGuid();

    public List<ChatMessage> Messages { get; } = [];

    public SemaphoreSlim Lock { get; } = new(1, 1);

    public DateTimeOffset CreatedAt { get; init; }
        = DateTimeOffset.UtcNow;

    public DateTimeOffset LastActivityAt { get; set; }
        = DateTimeOffset.UtcNow;
}

Each conversation now contains:

  • a unique ID

  • message history

  • a synchronization lock

  • creation time

  • last activity time

The lock will help us prevent two simultaneous requests from modifying the same conversation history at the same time.


Step 4: Create a Conversation Store

We do not want AiChatService to know whether conversations eventually live in:

Application Memory
Redis
SQL Server
PostgreSQL
Cosmos DB
Another Store

Create:

Services/IConversationStore.cs

Add:

public interface IConversationStore
{
    AiConversation Create();

    AiConversation? Get(Guid conversationId);

    void Remove(Guid conversationId);
}

The interface is intentionally small.

The AI service works with conversations without needing to understand their storage mechanism.


Step 5: Implement the In-Memory Store

Create:

Services/InMemoryConversationStore.cs

Add:

using System.Collections.Concurrent;

public sealed class InMemoryConversationStore
    : IConversationStore
{
    private readonly ConcurrentDictionary<Guid, AiConversation>
        _conversations = new();

    public AiConversation Create()
    {
        var conversation = new AiConversation();

        _conversations[conversation.Id] = conversation;

        return conversation;
    }

    public AiConversation? Get(Guid conversationId)
    {
        _conversations.TryGetValue(
            conversationId,
            out var conversation);

        return conversation;
    }

    public void Remove(Guid conversationId)
    {
        _conversations.TryRemove(
            conversationId,
            out _);
    }
}

We use:

ConcurrentDictionary

because an ASP.NET Core application can process many requests concurrently.

However, there is an important distinction.

ConcurrentDictionary protects dictionary operations.

It does not automatically make this mutable collection thread-safe:

conversation.Messages

That is why each conversation has its own synchronization lock.


Step 6: Register the Conversation Store

Open Program.cs.

Register:

builder.Services.AddSingleton<
    IConversationStore,
    InMemoryConversationStore>();

We use a singleton because the store must remain available across multiple HTTP requests.

If a new store were created for every request, the previous conversation history would disappear.

Conceptually:

Application Process
      ↓
Conversation Store
      ↓
├── Conversation A
├── Conversation B
└── Conversation C

Each conversation contains its own message history.


Step 7: Create the AI Result Model

Our previous AI service returned only:

string

Now it must return both:

Conversation ID
AI Answer

Create:

Models/AiChatResult.cs

Add:

public sealed record AiChatResult(
    Guid ConversationId,
    string Answer);

Now update IAiChatService:

public interface IAiChatService
{
    Task<AiChatResult> AskAsync(
        Guid? conversationId,
        string message,
        CancellationToken cancellationToken = default);
}

The service can now either create a new conversation or continue an existing one.


Step 8: Update AiChatService

Now we can implement the actual conversation-memory behavior.

using Microsoft.Extensions.AI;

public sealed class AiChatService : IAiChatService
{
    private const string SystemPrompt =
        """
        You are a helpful AI assistant.
        Give clear, concise, and accurate answers.
        Use previous conversation messages when relevant.
        """;

    private const int MaxConversationTurns = 10;

    private readonly IChatClient _chatClient;
    private readonly IConversationStore _conversationStore;
    private readonly ILogger<AiChatService> _logger;

    public AiChatService(
        IChatClient chatClient,
        IConversationStore conversationStore,
        ILogger<AiChatService> logger)
    {
        _chatClient = chatClient;
        _conversationStore = conversationStore;
        _logger = logger;
    }

    public async Task<AiChatResult> AskAsync(
        Guid? conversationId,
        string message,
        CancellationToken cancellationToken = default)
    {
        AiConversation conversation =
            GetOrCreateConversation(conversationId);

        await conversation.Lock.WaitAsync(
            cancellationToken);

        try
        {
            conversation.Messages.Add(
                new ChatMessage(
                    ChatRole.User,
                    message));

            var requestMessages =
                BuildRequestMessages(
                    conversation.Messages);

            _logger.LogInformation(
                "Sending AI request for conversation {ConversationId}.",
                conversation.Id);

            ChatResponse response =
                await _chatClient.GetResponseAsync(
                    requestMessages,
                    cancellationToken: cancellationToken);

            conversation.Messages.AddMessages(
                response);

            TrimConversation(
                conversation.Messages);

            conversation.LastActivityAt =
                DateTimeOffset.UtcNow;

            return new AiChatResult(
                conversation.Id,
                response.Text);
        }
        finally
        {
            conversation.Lock.Release();
        }
    }

    private AiConversation GetOrCreateConversation(
        Guid? conversationId)
    {
        if (!conversationId.HasValue)
        {
            return _conversationStore.Create();
        }

        return _conversationStore.Get(
                   conversationId.Value)
               ?? throw new KeyNotFoundException(
                   "Conversation was not found.");
    }

    private static List<ChatMessage> BuildRequestMessages(
        IReadOnlyList<ChatMessage> history)
    {
        var messages = new List<ChatMessage>
        {
            new(
                ChatRole.System,
                SystemPrompt)
        };

        messages.AddRange(history);

        return messages;
    }

    private static void TrimConversation(
        List<ChatMessage> messages)
    {
        int maxMessages =
            MaxConversationTurns * 2;

        while (messages.Count > maxMessages)
        {
            RemoveOldestTurn(messages);
        }
    }

    private static void RemoveOldestTurn(
        List<ChatMessage> messages)
    {
        if (messages.Count == 0)
        {
            return;
        }

        messages.RemoveAt(0);

        if (messages.Count > 0 &&
            messages[0].Role == ChatRole.Assistant)
        {
            messages.RemoveAt(0);
        }
    }
}

There are several important ideas here.

Let's look at them individually.


Creating or Loading a Conversation

The helper:

GetOrCreateConversation(...)

handles two scenarios.

If no ID is supplied:

return _conversationStore.Create();

A new conversation begins.

If an ID is supplied:

return _conversationStore.Get(
           conversationId.Value)
       ?? throw new KeyNotFoundException(
           "Conversation was not found.");

the existing history is loaded.

This keeps the main AskAsync method easier to follow.


Adding the User Message

Before calling the model, we store:

conversation.Messages.Add(
    new ChatMessage(
        ChatRole.User,
        message));

Suppose the first request is:

My name is Alex.

Our stored history becomes:

User:
My name is Alex.

The system instruction is deliberately not stored inside every conversation.

Instead, we prepend it when constructing the model request.

This avoids duplicating the same system message in every conversation history.


Building the Request Context

Our helper:

BuildRequestMessages(...)

creates:

System Instruction
       +
Conversation History

For example:

System:
You are a helpful AI assistant.

User:
My name is Alex.

Assistant:
Nice to meet you, Alex.

User:
What is my name?

That complete collection is sent to:

_chatClient.GetResponseAsync(...)

The model can now use previous messages as context.


Store the Assistant Response

After receiving:

ChatResponse response =
    await _chatClient.GetResponseAsync(...);

we call:

conversation.Messages.AddMessages(
    response);

A simple implementation could instead create an assistant message manually from:

response.Text

However, adding the response messages preserves the response structure more naturally.

That becomes useful as we later explore richer AI interactions.

After the first turn, our history may contain:

User:
My name is Alex.

Assistant:
Nice to meet you, Alex.

Why We Changed the History Limit to Turns

The earlier implementation used:

MaxConversationMessages = 20;

and removed individual messages from the beginning.

That is simple, but it can produce an awkward history.

Imagine:

User A
Assistant A
User B
Assistant B

If we remove only one message, we could end up with:

Assistant A
User B
Assistant B

The assistant response has lost the user message it was answering.

A better mental model is a conversation turn.

For our basic text-chat scenario:

User
  ↓
Assistant

represents one turn.

We therefore use:

private const int MaxConversationTurns = 10;

and remove older turns rather than blindly removing one message at a time.

This produces cleaner history for the model.


A Note About More Advanced Responses

Our pair-aware trimming works well for the simple text conversation we are building today.

Later, however, a single interaction may contain more than:

User
Assistant

For example, function calling may introduce:

User
Assistant Tool Call
Tool Result
Assistant

At that point, trimming should understand complete interaction boundaries rather than assuming every turn contains exactly two messages.

That is one reason memory management becomes more sophisticated as AI applications gain tools and agents.

For Day 4, our simple approach is appropriate.


Step 9: Update the Controller

Update AiController:

using Microsoft.AspNetCore.Mvc;

[ApiController]
[Route("api/[controller]")]
public sealed class AiController : ControllerBase
{
    private readonly IAiChatService _aiChatService;

    public AiController(
        IAiChatService aiChatService)
    {
        _aiChatService = aiChatService;
    }

    [HttpPost("chat")]
    public async Task<ActionResult<ChatResponse>> ChatAsync(
        [FromBody] ChatRequest request,
        CancellationToken cancellationToken)
    {
        try
        {
            AiChatResult result =
                await _aiChatService.AskAsync(
                    request.ConversationId,
                    request.Message,
                    cancellationToken);

            return Ok(new ChatResponse
            {
                ConversationId =
                    result.ConversationId,

                Answer =
                    result.Answer
            });
        }
        catch (KeyNotFoundException)
        {
            return NotFound(
                "Conversation was not found.");
        }
    }
}

The controller remains intentionally small.

Conversation-history management belongs to the application service and conversation store rather than the HTTP layer.


Step 10: Test the First Message

Run:

dotnet run

Send:

POST /api/ai/chat

with:

{
  "message": "My name is Alex."
}

The API may return:

{
  "conversationId": "5f7d95d7-bf12-4d32-9c42-7366d50c8fc3",
  "answer": "Nice to meet you, Alex."
}

Save the returned:

conversationId

because the next request needs it.


Step 11: Continue the Conversation

Now send:

{
  "conversationId": "5f7d95d7-bf12-4d32-9c42-7366d50c8fc3",
  "message": "What is my name?"
}

The stored history contains the previous exchange.

The model can therefore answer something similar to:

{
  "conversationId": "5f7d95d7-bf12-4d32-9c42-7366d50c8fc3",
  "answer": "Your name is Alex."
}

We now have a multi-turn conversation.


Step 12: Test Conversation Isolation

Create another conversation without sending the previous ID:

{
  "message": "My favorite programming language is C#."
}

The server returns another conversation ID.

Conceptually, our store now contains:

Conversation A
├── User: My name is Alex.
├── Assistant: Nice to meet you, Alex.
├── User: What is my name?
└── Assistant: Your name is Alex.

Conversation B
├── User: My favorite programming language is C#.
└── Assistant: ...

Conversation A and Conversation B have separate histories.

This is essential for any multi-user AI application.


Why We Need Per-Conversation Locking

Imagine two requests for the same conversation arrive at almost the same time.

Without synchronization, both could try to modify:

conversation.Messages

concurrently.

This can cause:

  • incorrect ordering

  • inconsistent context

  • concurrent collection modification

  • responses based on stale history

We therefore use:

SemaphoreSlim

for each conversation.

Before processing:

await conversation.Lock.WaitAsync(
    cancellationToken);

After processing:

conversation.Lock.Release();

The release occurs inside finally, so an exception does not permanently lock the conversation.

Importantly, the lock belongs to one conversation.

A request for Conversation A does not block Conversation B.


How Much Conversation History Should You Send?

A common question is:

How many previous messages should I send to the model?

There is no universal number.

Our tutorial keeps:

10 conversation turns

because it is simple enough to demonstrate the technique.

A real application should consider several factors.

Model Context Window

Every model has limits on how much context it can process.

Your history must leave enough room for:

System instructions
Current user message
Retrieved data
Tool information
Generated response

Conversation history is only one part of that budget.

Message Size

Ten short messages may be tiny.

Ten messages containing long reports may be huge.

That is why message count alone is not an ideal production strategy.

Cost

More input can mean more token usage and higher cost.

Latency

Large prompts can increase request time.

Relevance

A message from 100 turns ago may no longer help answer the current question.

For production systems, a useful strategy is often:

Recent Messages
      +
Important Summary
      +
Relevant Retrieved Context

rather than sending the entire conversation forever.


Better Memory Strategies

Our current implementation uses a simple sliding window.

As the application grows, several alternatives become useful.

1. Recent-Turn Window

Keep the latest:

5
10
20

conversation turns.

Simple and predictable.

2. Token-Based Window

Keep messages until a defined token budget is reached.

This is more accurate than counting messages because message lengths vary.

3. Conversation Summarization

Older conversation:

Old Messages
      ↓
Summary

Then send:

System Prompt
+
Conversation Summary
+
Recent Messages
+
Current Message

This preserves useful older context without sending everything.

4. Retrieval-Based Memory

Store historical information and retrieve only the pieces relevant to the current question.

This becomes useful for long-running assistants.

There is no single best memory strategy.

The right choice depends on your application.


Conversation ID Is Not Authorization

Our demo identifies conversations using a GUID.

That does not mean a GUID should be treated as proof of ownership.

Suppose one authenticated user somehow obtains another user's conversation ID.

A production application should still verify ownership.

A persisted conversation might eventually contain:

Conversation
├── Id
├── UserId
├── TenantId
├── CreatedAt
└── LastActivityAt

Retrieval should then validate:

ConversationId
+
Current User

and, in multi-tenant applications:

ConversationId
+
Current User
+
Tenant

A difficult-to-guess identifier is not a substitute for authorization.


In-Memory Storage Is Temporary

Our current conversation store exists inside the ASP.NET Core process.

That means:

Application Running
      ↓
Conversation Available

But after:

Application Restart

the process memory is cleared.

The conversations disappear.

This implementation is therefore suitable for:

  • tutorials

  • prototypes

  • local development

  • proof-of-concept applications

For durable production conversations, persistent storage is usually required.


What Happens with Multiple Servers?

Consider:

             Load Balancer
            /      |      \
           /       |       \
     Server A   Server B   Server C

Suppose Conversation A exists only in Server A's memory.

The next HTTP request might reach Server B.

Server B does not automatically know what Server A stored.

For multi-instance deployments, conversation state may need shared storage such as:

Redis
SQL Server
PostgreSQL
Cosmos DB
Distributed Storage

The correct choice depends on whether conversation history is temporary or durable business data.


Cache and Persistence Are Different

This distinction matters.

A distributed cache can be useful for fast temporary conversation state.

But suppose users expect conversations to survive:

  • deployments

  • server restarts

  • cache eviction

  • infrastructure failures

  • several days or months of inactivity

Then the conversation may need durable persistence.

For example:

Conversation
      ↓
Database
      ↓
Conversation Messages

with an optional cache layer:

Application
     ↓
Redis
     ↓
Database

Do not automatically treat important conversation history as disposable cache data.


Conversation Expiration

Our dictionary currently keeps conversations until they are manually removed or the application stops.

That is not a good unlimited production strategy.

We already track:

LastActivityAt

which could support policies such as:

Remove temporary conversations
after 30 minutes of inactivity

or:

Archive conversations
after several days

A production system should define:

  • expiration policy

  • retention period

  • cleanup mechanism

  • maximum storage

  • deletion behavior

Without cleanup, an in-memory store can continue growing as new conversations are created.


What Happens If the AI Request Fails?

Our implementation stores the user's message before sending the request to the provider:

conversation.Messages.Add(
    new ChatMessage(
        ChatRole.User,
        message));

Suppose the provider then fails.

The history contains the user message but no assistant response.

For this tutorial, that behavior is acceptable.

In a larger system, you may want to track message state:

Pending
Completed
Failed

A persistent message model might eventually contain:

Message
├── Id
├── ConversationId
├── Role
├── Content
├── Status
├── CreatedAt
└── Error

This becomes useful for retries, auditing, and recovering from partial failures.


Conversation Memory and Data Privacy

Conversation history may contain sensitive information.

Users can paste:

  • customer data

  • financial information

  • API keys

  • passwords

  • access tokens

  • internal business information

  • private documents

Once conversation history is stored, memory becomes a data-management concern as well as an AI feature.

Depending on the application, you may need:

  • access controls

  • redaction

  • encryption

  • retention policies

  • deletion workflows

  • audit logs

  • data classification

Do not store sensitive information simply because the model can process it.


Conversation Memory Is Not RAG

Conversation memory answers:

What has already been said in this conversation?

RAG answers:

What external information should be retrieved to answer this question?

For example:

Conversation Memory

User:
We were discussing Product A.

User:
Which product were we discussing?

The answer can come from conversation history.

But:

User:
What is Product A's current stock?

may require retrieving current information from the inventory database.

An advanced application can combine both:

User Question
      ↓
Conversation Context
      +
Retrieved Business Data
      ↓
AI Model

We will explore retrieval later in this series.


Practical Inventory Example

Imagine our assistant belongs to an inventory application.

The conversation begins:

User:
I want to review Wireless Mouse stock.

Then:

User:
Current stock is 4 and reorder level is 10.

Later:

User:
What product were we discussing?

The phrase:

the product

only makes sense because the previous conversation mentioned:

Wireless Mouse

Conversation memory makes follow-up questions such as these possible:

What about that product?

Explain it again.

Why is it low?

What did I tell you earlier?

Without previous context, the model cannot reliably resolve those references.


Suggested Project Structure

After Day 4:

CSharpAiApi
│
├── Controllers
│   └── AiController.cs
│
├── Models
│   ├── AiChatResult.cs
│   ├── AiConversation.cs
│   ├── ChatRequest.cs
│   └── ChatResponse.cs
│
├── Services
│   ├── IAiChatService.cs
│   ├── AiChatService.cs
│   ├── IConversationStore.cs
│   └── InMemoryConversationStore.cs
│
├── Program.cs
│
└── appsettings.json

The main responsibilities are now separated:

IChatClient
    ↓
AI Communication

IAiChatService
    ↓
Application AI Behavior

IConversationStore
    ↓
Conversation State

This makes it easier to replace the storage strategy later without rewriting the controller or AI provider integration.


Common Error: The AI Does Not Remember

Make sure the second request sends the same:

conversationId

If you send:

{
  "message": "What is my name?"
}

without the previous ID, the server starts another conversation.

The new conversation has no knowledge of the first one.


Common Error: Conversation Not Found

If you receive:

Conversation was not found.

possible causes include:

  • incorrect conversation ID

  • application restart

  • conversation removed

  • conversation expired

  • request reached another application instance

The current implementation does not persist conversations across application restarts.


Common Error: Context Keeps Growing

Do not append conversation history forever.

Our tutorial uses:

MaxConversationTurns = 10;

to demonstrate a simple sliding window.

Production systems should use a limit appropriate for their model, cost, latency, and application requirements.


Common Error: Assuming Message Count Equals Token Count

It does not.

This:

20 short messages

can consume far less context than:

5 very large messages

Message-count limits are easy to understand, but token-aware limits are more accurate for production systems.


Common Error: Treating a GUID as Security

A conversation ID identifies a conversation.

It does not prove that the current user owns it.

Authenticated applications must enforce authorization separately.


What We Built Today

We started with a stateless AI API.

Now the application supports multi-turn conversations.

We added:

  • conversation IDs

  • isolated conversation history

  • user and assistant messages

  • an in-memory conversation store

  • concurrent dictionary access

  • per-conversation synchronization

  • conversation timestamps

  • pair-aware history trimming

  • conversation-not-found handling

  • clear boundaries for future persistent storage

Our application now follows:

Client
  ↓
Conversation ID
  ↓
Conversation Store
  ↓
Chat History
  ↓
AI Service
  ↓
IChatClient
  ↓
AI Model
  ↓
Response
  ↓
Updated History

The important point is that the application owns the conversation state.


Day 4 Checklist

Before moving to Day 5, make sure you understand:

  • why separate AI requests do not automatically share context

  • what conversation memory means

  • the difference between conversation memory and long-term memory

  • why each conversation needs an identifier

  • how previous ChatMessage objects are sent to IChatClient

  • why assistant responses are stored back in history

  • what AddMessages(response) does

  • why different conversations must remain isolated

  • why concurrent requests need synchronization

  • why history should be trimmed by meaningful interaction boundaries

  • why message count is not the same as token count

  • why in-memory storage is temporary

  • why multi-server deployments need shared storage

  • why conversation IDs do not replace authorization

If these concepts are clear, our assistant is ready for the next improvement.


What's Next: Day 5

Our assistant can now remember previous messages.

But there is still a user-experience problem.

With:

GetResponseAsync(...)

the application normally waits for the response operation to complete before returning the final result.

For longer answers, users may spend several seconds staring at an empty screen.

In Day 5: Streaming AI Responses in C# with IChatClient and ASP.NET Core, we will improve that experience.

Instead of:

Request
   ↓
Wait
   ↓
Wait
   ↓
Complete Response

we will build:

Request
   ↓
First Chunk
   ↓
Next Chunk
   ↓
Next Chunk
   ↓
Complete Response

We will explore:

  • GetStreamingResponseAsync

  • IAsyncEnumerable

  • await foreach

  • ASP.NET Core response streaming

  • cancellation during streaming

  • collecting the complete streamed response

  • updating conversation history

  • handling partially completed streams

That will make our AI application feel much more responsive.


Final Thoughts

Conversation memory changes an AI application from isolated question-and-answer requests into a real multi-turn experience.

But implementing memory correctly involves more than keeping messages in a list.

A web application must consider:

Conversation Identity
History
Concurrency
Context Limits
Storage
Authorization
Failure Handling
Privacy

Today we implemented a deliberately simple architecture while still addressing those important boundaries.

The AI model generates the response.

But the application controls what the model remembers by controlling the context it sends.

That distinction is fundamental.

Day 1 connected C# to an AI model.

Day 2 introduced IChatClient.

Day 3 built a reusable ASP.NET Core AI service.

Day 4 added conversation memory with isolated, controlled chat history.

Next, we will make those conversations feel faster and more natural by streaming AI responses as they are generated.

Comments 0