Add Conversation Memory to a C# AI Application with ASP.NET Core | Day 4
Our AI application can answer questions, but it still has one major limitation:
It does not remember the conversation.
Suppose a user sends:
My name is Alex.The AI responds:
Nice to meet you, Alex.Then the user asks:
What is my name?If we send only the latest message to the model, it has no access to the earlier conversation.
Today, we will fix that.
In Day 4 of our C# AI tutorial series, we will add conversation memory to the ASP.NET Core application we built on Day 3.
We will learn how to:
create conversation IDs
store chat history
keep conversations isolated
send previous messages to the model
store AI responses back in history
handle concurrent requests safely
limit conversation history
understand the limitations of in-memory storage
prepare the architecture for persistent storage later
By the end of this tutorial, our API will support real multi-turn conversations.
Previously in This Series
So far, we have built the application step by step.
Day 1: Build Your First AI Application with C# and .NET
We connected a C# application to an AI model.
Day 2: Understanding IChatClient in .NET
We learned why Microsoft.Extensions.AI provides a useful abstraction between our application and AI providers.
Day 3: Build a Reusable AI Chat Service in ASP.NET Core
We moved the AI integration into an ASP.NET Core Web API using dependency injection, validation, secure configuration, and an application-level AI service.
Our previous request flow was:
HTTP Request
↓
AiController
↓
IAiChatService
↓
IChatClient
↓
AI Model
↓
ResponseToday we will add:
Conversation Historyto that flow.
What Does Conversation Memory Mean?
The phrase AI memory can describe several different techniques.
For today's application, conversation memory means:
The application stores previous messages and sends relevant conversation history to the model with each new request.
The model is not automatically remembering every HTTP request our application previously sent.
Our application manages that context.
Consider:
User:
My name is Alex.
Assistant:
Nice to meet you, Alex.
User:
What is my name?For the final question, our application can send:
System:
You are a helpful AI assistant.
User:
My name is Alex.
Assistant:
Nice to meet you, Alex.
User:
What is my name?The model now has the context needed to answer:
Your name is Alex.That is the basic conversation-memory pattern we will implement.
Conversation Memory vs Long-Term Memory
Before writing code, we should separate two concepts.
Conversation Memory
Conversation memory contains context from the current conversation.
For example:
User:
My company sells computer accessories.
Assistant:
How can I help?
User:
What type of products did I say we sell?The previous message provides the context needed to understand the latest question.
Long-Term Memory
Long-term memory may contain information that survives beyond one conversation.
Examples include:
user preferences
business information
saved instructions
profile information
previous project context
Long-term memory normally requires persistent storage or another dedicated memory strategy.
We are not implementing long-term memory today.
Day 4 focuses specifically on multi-turn conversation history.
Why Day 3 Had No Memory
On Day 3, we created a new message collection for every request:
var messages = new List<ChatMessage>
{
new(
ChatRole.System,
SystemPrompt),
new(
ChatRole.User,
message)
};Then:
var response =
await _chatClient.GetResponseAsync(
messages,
cancellationToken: cancellationToken);When the request finished, that local collection disappeared.
The next HTTP request created another collection.
Conceptually:
Request 1
↓
Temporary Messages
↓
AI Response
↓
Request EndsThen:
Request 2
↓
New MessagesWe need a way to connect those requests.
Our New Architecture
The application will now work like this:
Client
↓
Conversation ID
↓
ASP.NET Core API
↓
Conversation Store
↓
Previous Messages
↓
AI Service
↓
IChatClient
↓
AI Model
↓
Assistant Response
↓
Updated Conversation StoreFor Day 4, our conversation store will use application memory.
This keeps the implementation easy to understand while preserving a storage abstraction that we can replace later.
Step 1: Update the Request Model
On Day 3, our request looked like:
{
"message": "Explain dependency injection."
}Now we also need to identify the conversation.
Update ChatRequest.cs:
using System.ComponentModel.DataAnnotations;
public sealed class ChatRequest
{
public Guid? ConversationId { get; set; }
[Required]
[StringLength(
4000,
MinimumLength = 1,
ErrorMessage = "Message must contain between 1 and 4000 characters.")]
public string Message { get; set; } = string.Empty;
}For the first message, the client can omit ConversationId:
{
"message": "My name is Alex."
}The server creates a new conversation.
For the next message, the client sends the returned ID:
{
"conversationId": "5f7d95d7-bf12-4d32-9c42-7366d50c8fc3",
"message": "What is my name?"
}Now the server knows which conversation should be continued.
Step 2: Update the Response Model
The API must return the conversation ID so the client can reuse it.
Update ChatResponse.cs:
public sealed class ChatResponse
{
public Guid ConversationId { get; set; }
public string Answer { get; set; } = string.Empty;
}The first response can now look like:
{
"conversationId": "5f7d95d7-bf12-4d32-9c42-7366d50c8fc3",
"answer": "Nice to meet you, Alex."
}The client stores that ID and sends it with the next message.
Step 3: Create the Conversation Model
Create:
Models/AiConversation.csAdd:
using Microsoft.Extensions.AI;
public sealed class AiConversation
{
public Guid Id { get; init; } = Guid.NewGuid();
public List<ChatMessage> Messages { get; } = [];
public SemaphoreSlim Lock { get; } = new(1, 1);
public DateTimeOffset CreatedAt { get; init; }
= DateTimeOffset.UtcNow;
public DateTimeOffset LastActivityAt { get; set; }
= DateTimeOffset.UtcNow;
}Each conversation now contains:
a unique ID
message history
a synchronization lock
creation time
last activity time
The lock will help us prevent two simultaneous requests from modifying the same conversation history at the same time.
Step 4: Create a Conversation Store
We do not want AiChatService to know whether conversations eventually live in:
Application Memory
Redis
SQL Server
PostgreSQL
Cosmos DB
Another StoreCreate:
Services/IConversationStore.csAdd:
public interface IConversationStore
{
AiConversation Create();
AiConversation? Get(Guid conversationId);
void Remove(Guid conversationId);
}The interface is intentionally small.
The AI service works with conversations without needing to understand their storage mechanism.
Step 5: Implement the In-Memory Store
Create:
Services/InMemoryConversationStore.csAdd:
using System.Collections.Concurrent;
public sealed class InMemoryConversationStore
: IConversationStore
{
private readonly ConcurrentDictionary<Guid, AiConversation>
_conversations = new();
public AiConversation Create()
{
var conversation = new AiConversation();
_conversations[conversation.Id] = conversation;
return conversation;
}
public AiConversation? Get(Guid conversationId)
{
_conversations.TryGetValue(
conversationId,
out var conversation);
return conversation;
}
public void Remove(Guid conversationId)
{
_conversations.TryRemove(
conversationId,
out _);
}
}We use:
ConcurrentDictionarybecause an ASP.NET Core application can process many requests concurrently.
However, there is an important distinction.
ConcurrentDictionary protects dictionary operations.
It does not automatically make this mutable collection thread-safe:
conversation.MessagesThat is why each conversation has its own synchronization lock.
Step 6: Register the Conversation Store
Open Program.cs.
Register:
builder.Services.AddSingleton<
IConversationStore,
InMemoryConversationStore>();We use a singleton because the store must remain available across multiple HTTP requests.
If a new store were created for every request, the previous conversation history would disappear.
Conceptually:
Application Process
↓
Conversation Store
↓
├── Conversation A
├── Conversation B
└── Conversation CEach conversation contains its own message history.
Step 7: Create the AI Result Model
Our previous AI service returned only:
stringNow it must return both:
Conversation ID
AI AnswerCreate:
Models/AiChatResult.csAdd:
public sealed record AiChatResult(
Guid ConversationId,
string Answer);Now update IAiChatService:
public interface IAiChatService
{
Task<AiChatResult> AskAsync(
Guid? conversationId,
string message,
CancellationToken cancellationToken = default);
}The service can now either create a new conversation or continue an existing one.
Step 8: Update AiChatService
Now we can implement the actual conversation-memory behavior.
using Microsoft.Extensions.AI;
public sealed class AiChatService : IAiChatService
{
private const string SystemPrompt =
"""
You are a helpful AI assistant.
Give clear, concise, and accurate answers.
Use previous conversation messages when relevant.
""";
private const int MaxConversationTurns = 10;
private readonly IChatClient _chatClient;
private readonly IConversationStore _conversationStore;
private readonly ILogger<AiChatService> _logger;
public AiChatService(
IChatClient chatClient,
IConversationStore conversationStore,
ILogger<AiChatService> logger)
{
_chatClient = chatClient;
_conversationStore = conversationStore;
_logger = logger;
}
public async Task<AiChatResult> AskAsync(
Guid? conversationId,
string message,
CancellationToken cancellationToken = default)
{
AiConversation conversation =
GetOrCreateConversation(conversationId);
await conversation.Lock.WaitAsync(
cancellationToken);
try
{
conversation.Messages.Add(
new ChatMessage(
ChatRole.User,
message));
var requestMessages =
BuildRequestMessages(
conversation.Messages);
_logger.LogInformation(
"Sending AI request for conversation {ConversationId}.",
conversation.Id);
ChatResponse response =
await _chatClient.GetResponseAsync(
requestMessages,
cancellationToken: cancellationToken);
conversation.Messages.AddMessages(
response);
TrimConversation(
conversation.Messages);
conversation.LastActivityAt =
DateTimeOffset.UtcNow;
return new AiChatResult(
conversation.Id,
response.Text);
}
finally
{
conversation.Lock.Release();
}
}
private AiConversation GetOrCreateConversation(
Guid? conversationId)
{
if (!conversationId.HasValue)
{
return _conversationStore.Create();
}
return _conversationStore.Get(
conversationId.Value)
?? throw new KeyNotFoundException(
"Conversation was not found.");
}
private static List<ChatMessage> BuildRequestMessages(
IReadOnlyList<ChatMessage> history)
{
var messages = new List<ChatMessage>
{
new(
ChatRole.System,
SystemPrompt)
};
messages.AddRange(history);
return messages;
}
private static void TrimConversation(
List<ChatMessage> messages)
{
int maxMessages =
MaxConversationTurns * 2;
while (messages.Count > maxMessages)
{
RemoveOldestTurn(messages);
}
}
private static void RemoveOldestTurn(
List<ChatMessage> messages)
{
if (messages.Count == 0)
{
return;
}
messages.RemoveAt(0);
if (messages.Count > 0 &&
messages[0].Role == ChatRole.Assistant)
{
messages.RemoveAt(0);
}
}
}There are several important ideas here.
Let's look at them individually.
Creating or Loading a Conversation
The helper:
GetOrCreateConversation(...)handles two scenarios.
If no ID is supplied:
return _conversationStore.Create();A new conversation begins.
If an ID is supplied:
return _conversationStore.Get(
conversationId.Value)
?? throw new KeyNotFoundException(
"Conversation was not found.");the existing history is loaded.
This keeps the main AskAsync method easier to follow.
Adding the User Message
Before calling the model, we store:
conversation.Messages.Add(
new ChatMessage(
ChatRole.User,
message));Suppose the first request is:
My name is Alex.Our stored history becomes:
User:
My name is Alex.The system instruction is deliberately not stored inside every conversation.
Instead, we prepend it when constructing the model request.
This avoids duplicating the same system message in every conversation history.
Building the Request Context
Our helper:
BuildRequestMessages(...)creates:
System Instruction
+
Conversation HistoryFor example:
System:
You are a helpful AI assistant.
User:
My name is Alex.
Assistant:
Nice to meet you, Alex.
User:
What is my name?That complete collection is sent to:
_chatClient.GetResponseAsync(...)The model can now use previous messages as context.
Store the Assistant Response
After receiving:
ChatResponse response =
await _chatClient.GetResponseAsync(...);we call:
conversation.Messages.AddMessages(
response);A simple implementation could instead create an assistant message manually from:
response.TextHowever, adding the response messages preserves the response structure more naturally.
That becomes useful as we later explore richer AI interactions.
After the first turn, our history may contain:
User:
My name is Alex.
Assistant:
Nice to meet you, Alex.Why We Changed the History Limit to Turns
The earlier implementation used:
MaxConversationMessages = 20;and removed individual messages from the beginning.
That is simple, but it can produce an awkward history.
Imagine:
User A
Assistant A
User B
Assistant BIf we remove only one message, we could end up with:
Assistant A
User B
Assistant BThe assistant response has lost the user message it was answering.
A better mental model is a conversation turn.
For our basic text-chat scenario:
User
↓
Assistantrepresents one turn.
We therefore use:
private const int MaxConversationTurns = 10;and remove older turns rather than blindly removing one message at a time.
This produces cleaner history for the model.
A Note About More Advanced Responses
Our pair-aware trimming works well for the simple text conversation we are building today.
Later, however, a single interaction may contain more than:
User
AssistantFor example, function calling may introduce:
User
Assistant Tool Call
Tool Result
AssistantAt that point, trimming should understand complete interaction boundaries rather than assuming every turn contains exactly two messages.
That is one reason memory management becomes more sophisticated as AI applications gain tools and agents.
For Day 4, our simple approach is appropriate.
Step 9: Update the Controller
Update AiController:
using Microsoft.AspNetCore.Mvc;
[ApiController]
[Route("api/[controller]")]
public sealed class AiController : ControllerBase
{
private readonly IAiChatService _aiChatService;
public AiController(
IAiChatService aiChatService)
{
_aiChatService = aiChatService;
}
[HttpPost("chat")]
public async Task<ActionResult<ChatResponse>> ChatAsync(
[FromBody] ChatRequest request,
CancellationToken cancellationToken)
{
try
{
AiChatResult result =
await _aiChatService.AskAsync(
request.ConversationId,
request.Message,
cancellationToken);
return Ok(new ChatResponse
{
ConversationId =
result.ConversationId,
Answer =
result.Answer
});
}
catch (KeyNotFoundException)
{
return NotFound(
"Conversation was not found.");
}
}
}The controller remains intentionally small.
Conversation-history management belongs to the application service and conversation store rather than the HTTP layer.
Step 10: Test the First Message
Run:
dotnet runSend:
POST /api/ai/chatwith:
{
"message": "My name is Alex."
}The API may return:
{
"conversationId": "5f7d95d7-bf12-4d32-9c42-7366d50c8fc3",
"answer": "Nice to meet you, Alex."
}Save the returned:
conversationIdbecause the next request needs it.
Step 11: Continue the Conversation
Now send:
{
"conversationId": "5f7d95d7-bf12-4d32-9c42-7366d50c8fc3",
"message": "What is my name?"
}The stored history contains the previous exchange.
The model can therefore answer something similar to:
{
"conversationId": "5f7d95d7-bf12-4d32-9c42-7366d50c8fc3",
"answer": "Your name is Alex."
}We now have a multi-turn conversation.
Step 12: Test Conversation Isolation
Create another conversation without sending the previous ID:
{
"message": "My favorite programming language is C#."
}The server returns another conversation ID.
Conceptually, our store now contains:
Conversation A
├── User: My name is Alex.
├── Assistant: Nice to meet you, Alex.
├── User: What is my name?
└── Assistant: Your name is Alex.
Conversation B
├── User: My favorite programming language is C#.
└── Assistant: ...Conversation A and Conversation B have separate histories.
This is essential for any multi-user AI application.
Why We Need Per-Conversation Locking
Imagine two requests for the same conversation arrive at almost the same time.
Without synchronization, both could try to modify:
conversation.Messagesconcurrently.
This can cause:
incorrect ordering
inconsistent context
concurrent collection modification
responses based on stale history
We therefore use:
SemaphoreSlimfor each conversation.
Before processing:
await conversation.Lock.WaitAsync(
cancellationToken);After processing:
conversation.Lock.Release();The release occurs inside finally, so an exception does not permanently lock the conversation.
Importantly, the lock belongs to one conversation.
A request for Conversation A does not block Conversation B.
How Much Conversation History Should You Send?
A common question is:
How many previous messages should I send to the model?
There is no universal number.
Our tutorial keeps:
10 conversation turnsbecause it is simple enough to demonstrate the technique.
A real application should consider several factors.
Model Context Window
Every model has limits on how much context it can process.
Your history must leave enough room for:
System instructions
Current user message
Retrieved data
Tool information
Generated responseConversation history is only one part of that budget.
Message Size
Ten short messages may be tiny.
Ten messages containing long reports may be huge.
That is why message count alone is not an ideal production strategy.
Cost
More input can mean more token usage and higher cost.
Latency
Large prompts can increase request time.
Relevance
A message from 100 turns ago may no longer help answer the current question.
For production systems, a useful strategy is often:
Recent Messages
+
Important Summary
+
Relevant Retrieved Contextrather than sending the entire conversation forever.
Better Memory Strategies
Our current implementation uses a simple sliding window.
As the application grows, several alternatives become useful.
1. Recent-Turn Window
Keep the latest:
5
10
20conversation turns.
Simple and predictable.
2. Token-Based Window
Keep messages until a defined token budget is reached.
This is more accurate than counting messages because message lengths vary.
3. Conversation Summarization
Older conversation:
Old Messages
↓
SummaryThen send:
System Prompt
+
Conversation Summary
+
Recent Messages
+
Current MessageThis preserves useful older context without sending everything.
4. Retrieval-Based Memory
Store historical information and retrieve only the pieces relevant to the current question.
This becomes useful for long-running assistants.
There is no single best memory strategy.
The right choice depends on your application.
Conversation ID Is Not Authorization
Our demo identifies conversations using a GUID.
That does not mean a GUID should be treated as proof of ownership.
Suppose one authenticated user somehow obtains another user's conversation ID.
A production application should still verify ownership.
A persisted conversation might eventually contain:
Conversation
├── Id
├── UserId
├── TenantId
├── CreatedAt
└── LastActivityAtRetrieval should then validate:
ConversationId
+
Current Userand, in multi-tenant applications:
ConversationId
+
Current User
+
TenantA difficult-to-guess identifier is not a substitute for authorization.
In-Memory Storage Is Temporary
Our current conversation store exists inside the ASP.NET Core process.
That means:
Application Running
↓
Conversation AvailableBut after:
Application Restartthe process memory is cleared.
The conversations disappear.
This implementation is therefore suitable for:
tutorials
prototypes
local development
proof-of-concept applications
For durable production conversations, persistent storage is usually required.
What Happens with Multiple Servers?
Consider:
Load Balancer
/ | \
/ | \
Server A Server B Server CSuppose Conversation A exists only in Server A's memory.
The next HTTP request might reach Server B.
Server B does not automatically know what Server A stored.
For multi-instance deployments, conversation state may need shared storage such as:
Redis
SQL Server
PostgreSQL
Cosmos DB
Distributed StorageThe correct choice depends on whether conversation history is temporary or durable business data.
Cache and Persistence Are Different
This distinction matters.
A distributed cache can be useful for fast temporary conversation state.
But suppose users expect conversations to survive:
deployments
server restarts
cache eviction
infrastructure failures
several days or months of inactivity
Then the conversation may need durable persistence.
For example:
Conversation
↓
Database
↓
Conversation Messageswith an optional cache layer:
Application
↓
Redis
↓
DatabaseDo not automatically treat important conversation history as disposable cache data.
Conversation Expiration
Our dictionary currently keeps conversations until they are manually removed or the application stops.
That is not a good unlimited production strategy.
We already track:
LastActivityAtwhich could support policies such as:
Remove temporary conversations
after 30 minutes of inactivityor:
Archive conversations
after several daysA production system should define:
expiration policy
retention period
cleanup mechanism
maximum storage
deletion behavior
Without cleanup, an in-memory store can continue growing as new conversations are created.
What Happens If the AI Request Fails?
Our implementation stores the user's message before sending the request to the provider:
conversation.Messages.Add(
new ChatMessage(
ChatRole.User,
message));Suppose the provider then fails.
The history contains the user message but no assistant response.
For this tutorial, that behavior is acceptable.
In a larger system, you may want to track message state:
Pending
Completed
FailedA persistent message model might eventually contain:
Message
├── Id
├── ConversationId
├── Role
├── Content
├── Status
├── CreatedAt
└── ErrorThis becomes useful for retries, auditing, and recovering from partial failures.
Conversation Memory and Data Privacy
Conversation history may contain sensitive information.
Users can paste:
customer data
financial information
API keys
passwords
access tokens
internal business information
private documents
Once conversation history is stored, memory becomes a data-management concern as well as an AI feature.
Depending on the application, you may need:
access controls
redaction
encryption
retention policies
deletion workflows
audit logs
data classification
Do not store sensitive information simply because the model can process it.
Conversation Memory Is Not RAG
Conversation memory answers:
What has already been said in this conversation?
RAG answers:
What external information should be retrieved to answer this question?
For example:
Conversation Memory
User:
We were discussing Product A.
User:
Which product were we discussing?The answer can come from conversation history.
But:
User:
What is Product A's current stock?may require retrieving current information from the inventory database.
An advanced application can combine both:
User Question
↓
Conversation Context
+
Retrieved Business Data
↓
AI ModelWe will explore retrieval later in this series.
Practical Inventory Example
Imagine our assistant belongs to an inventory application.
The conversation begins:
User:
I want to review Wireless Mouse stock.Then:
User:
Current stock is 4 and reorder level is 10.Later:
User:
What product were we discussing?The phrase:
the productonly makes sense because the previous conversation mentioned:
Wireless MouseConversation memory makes follow-up questions such as these possible:
What about that product?
Explain it again.
Why is it low?
What did I tell you earlier?Without previous context, the model cannot reliably resolve those references.
Suggested Project Structure
After Day 4:
CSharpAiApi
│
├── Controllers
│ └── AiController.cs
│
├── Models
│ ├── AiChatResult.cs
│ ├── AiConversation.cs
│ ├── ChatRequest.cs
│ └── ChatResponse.cs
│
├── Services
│ ├── IAiChatService.cs
│ ├── AiChatService.cs
│ ├── IConversationStore.cs
│ └── InMemoryConversationStore.cs
│
├── Program.cs
│
└── appsettings.jsonThe main responsibilities are now separated:
IChatClient
↓
AI Communication
IAiChatService
↓
Application AI Behavior
IConversationStore
↓
Conversation StateThis makes it easier to replace the storage strategy later without rewriting the controller or AI provider integration.
Common Error: The AI Does Not Remember
Make sure the second request sends the same:
conversationIdIf you send:
{
"message": "What is my name?"
}without the previous ID, the server starts another conversation.
The new conversation has no knowledge of the first one.
Common Error: Conversation Not Found
If you receive:
Conversation was not found.possible causes include:
incorrect conversation ID
application restart
conversation removed
conversation expired
request reached another application instance
The current implementation does not persist conversations across application restarts.
Common Error: Context Keeps Growing
Do not append conversation history forever.
Our tutorial uses:
MaxConversationTurns = 10;to demonstrate a simple sliding window.
Production systems should use a limit appropriate for their model, cost, latency, and application requirements.
Common Error: Assuming Message Count Equals Token Count
It does not.
This:
20 short messagescan consume far less context than:
5 very large messagesMessage-count limits are easy to understand, but token-aware limits are more accurate for production systems.
Common Error: Treating a GUID as Security
A conversation ID identifies a conversation.
It does not prove that the current user owns it.
Authenticated applications must enforce authorization separately.
What We Built Today
We started with a stateless AI API.
Now the application supports multi-turn conversations.
We added:
conversation IDs
isolated conversation history
user and assistant messages
an in-memory conversation store
concurrent dictionary access
per-conversation synchronization
conversation timestamps
pair-aware history trimming
conversation-not-found handling
clear boundaries for future persistent storage
Our application now follows:
Client
↓
Conversation ID
↓
Conversation Store
↓
Chat History
↓
AI Service
↓
IChatClient
↓
AI Model
↓
Response
↓
Updated HistoryThe important point is that the application owns the conversation state.
Day 4 Checklist
Before moving to Day 5, make sure you understand:
why separate AI requests do not automatically share context
what conversation memory means
the difference between conversation memory and long-term memory
why each conversation needs an identifier
how previous
ChatMessageobjects are sent toIChatClientwhy assistant responses are stored back in history
what
AddMessages(response)doeswhy different conversations must remain isolated
why concurrent requests need synchronization
why history should be trimmed by meaningful interaction boundaries
why message count is not the same as token count
why in-memory storage is temporary
why multi-server deployments need shared storage
why conversation IDs do not replace authorization
If these concepts are clear, our assistant is ready for the next improvement.
What's Next: Day 5
Our assistant can now remember previous messages.
But there is still a user-experience problem.
With:
GetResponseAsync(...)the application normally waits for the response operation to complete before returning the final result.
For longer answers, users may spend several seconds staring at an empty screen.
In Day 5: Streaming AI Responses in C# with IChatClient and ASP.NET Core, we will improve that experience.
Instead of:
Request
↓
Wait
↓
Wait
↓
Complete Responsewe will build:
Request
↓
First Chunk
↓
Next Chunk
↓
Next Chunk
↓
Complete ResponseWe will explore:
GetStreamingResponseAsyncIAsyncEnumerableawait foreachASP.NET Core response streaming
cancellation during streaming
collecting the complete streamed response
updating conversation history
handling partially completed streams
That will make our AI application feel much more responsive.
Final Thoughts
Conversation memory changes an AI application from isolated question-and-answer requests into a real multi-turn experience.
But implementing memory correctly involves more than keeping messages in a list.
A web application must consider:
Conversation Identity
History
Concurrency
Context Limits
Storage
Authorization
Failure Handling
PrivacyToday we implemented a deliberately simple architecture while still addressing those important boundaries.
The AI model generates the response.
But the application controls what the model remembers by controlling the context it sends.
That distinction is fundamental.
Day 1 connected C# to an AI model.
Day 2 introduced IChatClient.
Day 3 built a reusable ASP.NET Core AI service.
Day 4 added conversation memory with isolated, controlled chat history.
Next, we will make those conversations feel faster and more natural by streaming AI responses as they are generated.
Comments 0