DOTNET EXPERT BLOG

Stream AI Responses in C# with IChatClient and ASP.NET Core | Day 5

9/30/2026 12:12:01 PM Noor All Safaet Loading... 0

Our AI assistant can now remember conversations.

But there is still a user-experience problem.

Suppose the model generates a long response.

With the approach we used previously:

ChatResponse response =
    await _chatClient.GetResponseAsync(
        messages,
        cancellationToken: cancellationToken);

the application waits for the response operation to complete before returning the final answer.

From the user's perspective:

Send Message
     ↓
Wait
     ↓
Wait
     ↓
Complete Answer

Modern AI chat applications usually provide a more responsive experience.

They display the answer gradually as it is generated.

That is AI response streaming.

In Day 5, we will add streaming to the ASP.NET Core AI application we have been building.

We will learn how to:

  • use GetStreamingResponseAsync

  • understand IAsyncEnumerable<T>

  • process ChatResponseUpdate

  • use await foreach

  • stream AI output through ASP.NET Core

  • read the stream from JavaScript

  • preserve conversation memory

  • propagate cancellation

  • handle interrupted streams

  • understand production buffering issues

By the end, our browser will begin displaying the AI answer before the complete response has been generated.


Previously in This Series

Our project has evolved gradually.

Day 1: Build Your First AI Application with C# and .NET

We connected a .NET application to an AI model.

Day 2: Understanding IChatClient

We introduced the provider-independent AI abstraction from Microsoft.Extensions.AI.

Day 3: Build a Reusable AI Chat Service in ASP.NET Core

We created an ASP.NET Core API with dependency injection and a reusable application service.

Day 4: Add Conversation Memory to a C# AI Application

We introduced:

  • conversation IDs

  • chat history

  • isolated conversations

  • in-memory storage

  • history trimming

  • concurrency protection

Our application currently follows:

Client
  ↓
Conversation
  ↓
Chat History
  ↓
IChatClient
  ↓
AI Model
  ↓
Complete Response

Today we will change the final part:

AI Model
   ↓
Update
   ↓
Update
   ↓
Update
   ↓
Client

What Is AI Response Streaming?

A normal request waits for the response operation to finish:

ChatResponse response =
    await _chatClient.GetResponseAsync(
        messages);

Streaming works differently.

Instead of receiving one final ChatResponse, we process updates as they become available:

await foreach (
    ChatResponseUpdate update
    in _chatClient.GetStreamingResponseAsync(
        messages))
{
    Console.Write(update.Text);
}

Conceptually:

AI Model
   ↓
"Dependency "
   ↓
"injection "
   ↓
"is "
   ↓
"a useful..."

The client can start displaying content immediately.

Streaming mainly improves perceived responsiveness.

It does not necessarily make the model generate the entire answer faster.


GetResponseAsync vs GetStreamingResponseAsync

For a normal response:

ChatResponse response =
    await _chatClient.GetResponseAsync(
        messages);

the flow is:

Request
   ↓
Generation
   ↓
Complete ChatResponse

This is useful when:

  • responses are short

  • you need the complete result first

  • you are doing background processing

  • you need to validate the complete result

  • the endpoint returns standard JSON

For streaming:

await foreach (
    ChatResponseUpdate update
    in _chatClient.GetStreamingResponseAsync(
        messages))
{
    // Process update
}

the flow becomes:

Request
   ↓
Update
   ↓
Update
   ↓
Update
   ↓
Complete

This is particularly useful for interactive AI chat interfaces.


Understanding IAsyncEnumerable<T>

Most .NET developers are familiar with:

Task<T>

Streaming introduces another useful asynchronous abstraction:

IAsyncEnumerable<T>

It represents values that can become available asynchronously over time.

We consume it using:

await foreach (var item in GetItemsAsync())
{
    Console.WriteLine(item);
}

That makes it a natural fit for AI streaming.


What Is ChatResponseUpdate?

GetStreamingResponseAsync produces:

ChatResponseUpdate

objects.

For example:

await foreach (
    ChatResponseUpdate update
    in _chatClient.GetStreamingResponseAsync(
        "Explain dependency injection."))
{
    Console.Write(update.Text);
}

One important detail:

Do not assume one update equals one token or one word.

An update may contain:

  • part of a word

  • one word

  • several words

  • other response information

The exact chunk boundaries can depend on the provider and implementation.

Your application should simply process updates in the order they arrive.


Our Day 5 Architecture

We need streaming without breaking the conversation memory from Day 4.

Our flow becomes:

Browser
   ↓
ASP.NET Core
   ↓
Conversation Store
   ↓
Previous Messages
   ↓
IChatClient
   ↓
GetStreamingResponseAsync
   ↓
ChatResponseUpdate
   ↓
HTTP Response
   ↓
Browser Displays Text

At the same time, we collect the updates.

After streaming completes:

Collected Updates
      ↓
ChatResponse
      ↓
Conversation History

This gives us both streaming and memory.


Step 1: Create AiStreamContext

The client needs the conversation ID before it can continue the conversation later.

Create:

Models/AiStreamContext.cs

Add:

using Microsoft.Extensions.AI;

public sealed record AiStreamContext(
    Guid ConversationId,
    IAsyncEnumerable<ChatResponseUpdate> Updates);

This gives us two things:

ConversationId
+
Streaming Updates

Step 2: Update IAiChatService

Update the interface:

public interface IAiChatService
{
    Task<AiChatResult> AskAsync(
        Guid? conversationId,
        string message,
        CancellationToken cancellationToken = default);

    AiStreamContext StartStream(
        Guid? conversationId,
        string message,
        CancellationToken cancellationToken = default);
}

Notice that:

StartStream(...)

does not return:

Task<AiStreamContext>

There is no asynchronous work required merely to create or retrieve the conversation and construct the async stream.

The actual asynchronous operation begins when the returned stream is enumerated.

This keeps the API cleaner than wrapping the result in:

Task.FromResult(...)

without a real asynchronous operation.


Step 3: Implement StartStream

Inside AiChatService:

public AiStreamContext StartStream(
    Guid? conversationId,
    string message,
    CancellationToken cancellationToken = default)
{
    AiConversation conversation =
        GetOrCreateConversation(
            conversationId);

    IAsyncEnumerable<ChatResponseUpdate> updates =
        StreamConversationAsync(
            conversation,
            message,
            cancellationToken);

    return new AiStreamContext(
        conversation.Id,
        updates);
}

This method performs two jobs:

Find/Create Conversation
        ↓
Create Streaming Sequence

The controller immediately receives the conversation ID.

The actual AI request happens when:

Updates

is enumerated.


Step 4: Implement the Streaming Iterator

Add:

using System.Runtime.CompilerServices;
using Microsoft.Extensions.AI;

Then implement:

private async IAsyncEnumerable<ChatResponseUpdate>
    StreamConversationAsync(
        AiConversation conversation,
        string message,
        [EnumeratorCancellation]
        CancellationToken cancellationToken = default)
{
    await conversation.Lock.WaitAsync(
        cancellationToken);

    var updates =
        new List<ChatResponseUpdate>();

    try
    {
        conversation.Messages.Add(
            new ChatMessage(
                ChatRole.User,
                message));

        var requestMessages =
            BuildRequestMessages(
                conversation.Messages);

        _logger.LogInformation(
            "Starting streamed AI request for conversation {ConversationId}.",
            conversation.Id);

        await foreach (
            ChatResponseUpdate update
            in _chatClient.GetStreamingResponseAsync(
                requestMessages,
                cancellationToken: cancellationToken))
        {
            updates.Add(update);

            yield return update;
        }

        ChatResponse response =
            updates.ToChatResponse();

        conversation.Messages.AddMessages(
            response);

        TrimConversation(
            conversation.Messages);

        conversation.LastActivityAt =
            DateTimeOffset.UtcNow;

        _logger.LogInformation(
            "Completed streamed AI request for conversation {ConversationId}.",
            conversation.Id);
    }
    finally
    {
        conversation.Lock.Release();
    }
}

This is the core of our Day 5 implementation.


Why [EnumeratorCancellation] Matters

The method returns:

IAsyncEnumerable<ChatResponseUpdate>

and accepts a cancellation token.

Using:

[EnumeratorCancellation]

helps cancellation participate correctly in asynchronous enumeration.

Conceptually:

Browser disconnects
       ↓
HTTP request cancelled
       ↓
CancellationToken
       ↓
Async iterator
       ↓
AI streaming request

This allows unnecessary work to stop when cancellation is supported through the stack.


Why We Collect Updates

Inside our loop:

updates.Add(update);

yield return update;

each update has two purposes.

yield return sends it toward the caller:

AI
 ↓
Update
 ↓
Controller
 ↓
Browser

while:

updates.Add(update);

keeps a copy for conversation memory.

After streaming completes:

ChatResponse response =
    updates.ToChatResponse();

reconstructs the final response.

Then:

conversation.Messages.AddMessages(
    response);

adds it to the conversation history.

Without this step, the browser would see the AI answer, but our application-managed memory would not contain the completed assistant response.


Step 5: Add the Streaming Endpoint

Add this action to AiController:

[HttpPost("stream")]
public async Task StreamAsync(
    [FromBody] ChatRequest request,
    CancellationToken cancellationToken)
{
    AiStreamContext stream;

    try
    {
        stream =
            _aiChatService.StartStream(
                request.ConversationId,
                request.Message,
                cancellationToken);
    }
    catch (KeyNotFoundException)
    {
        Response.StatusCode =
            StatusCodes.Status404NotFound;

        await Response.WriteAsync(
            "Conversation was not found.",
            cancellationToken);

        return;
    }

    Response.ContentType =
        "text/plain; charset=utf-8";

    Response.Headers["X-Conversation-Id"] =
        stream.ConversationId.ToString();

    await foreach (
        ChatResponseUpdate update
        in stream.Updates
            .WithCancellation(cancellationToken))
    {
        if (string.IsNullOrEmpty(update.Text))
        {
            continue;
        }

        await Response.WriteAsync(
            update.Text,
            cancellationToken);

        await Response.Body.FlushAsync(
            cancellationToken);
    }
}

Our application now exposes:

POST /api/ai/stream

The normal:

POST /api/ai/chat

endpoint can remain available for non-streaming requests.


Why We Set Headers Before Streaming

Before writing the first response chunk, we set:

Response.ContentType =
    "text/plain; charset=utf-8";

and:

Response.Headers["X-Conversation-Id"] =
    stream.ConversationId.ToString();

This matters because once the response starts, HTTP headers may already be committed.

Important metadata should therefore be determined before streaming begins.


Why Call FlushAsync?

Inside the loop:

await Response.WriteAsync(
    update.Text,
    cancellationToken);

await Response.Body.FlushAsync(
    cancellationToken);

WriteAsync writes the text to the response pipeline.

FlushAsync asks the pipeline to flush buffered data toward the client.

Conceptually:

AI Update
    ↓
WriteAsync
    ↓
Response Pipeline
    ↓
FlushAsync
    ↓
Client

However, other layers may still affect buffering.

These can include:

  • reverse proxies

  • response compression

  • gateways

  • hosting infrastructure

  • browsers

  • HTTP clients

So production streaming should always be tested through the actual deployment path.


Step 6: Test with cURL

Run the application:

dotnet run

Then:

curl -N -X POST "https://localhost:7001/api/ai/stream" \
  -H "Content-Type: application/json" \
  -d "{\"message\":\"Explain dependency injection in ASP.NET Core.\"}"

Replace the port with your actual development port.

The:

-N

option disables cURL's output buffering.

You should see the response appear progressively rather than waiting for the complete answer.


Step 7: Build a Simple Browser Streaming Client

Server-side streaming is only half of the feature.

Now let's read the stream in a browser.

Create a simple HTML page:

<!DOCTYPE html>
<html>
<head>
    <meta charset="utf-8" />
    <title>C# AI Streaming Demo</title>
</head>
<body>

    <h1>AI Chat</h1>

    <textarea
        id="message"
        rows="5"
        cols="60"
        placeholder="Ask something..."></textarea>

    <br />

    <button id="sendButton">
        Send
    </button>

    <h2>Response</h2>

    <pre id="output"></pre>

    <script src="chat.js"></script>

</body>
</html>

Now create:

chat.js

Step 8: Read the Stream with JavaScript fetch()

Add:

let conversationId = null;

const sendButton =
    document.getElementById("sendButton");

const messageInput =
    document.getElementById("message");

const output =
    document.getElementById("output");

sendButton.addEventListener(
    "click",
    sendMessage);

async function sendMessage() {

    const message =
        messageInput.value.trim();

    if (!message) {
        return;
    }

    output.textContent = "";

    const response =
        await fetch("/api/ai/stream", {
            method: "POST",

            headers: {
                "Content-Type": "application/json"
            },

            body: JSON.stringify({
                conversationId,
                message
            })
        });

    if (!response.ok) {
        const error =
            await response.text();

        output.textContent =
            `Request failed: ${error}`;

        return;
    }

    const returnedConversationId =
        response.headers.get(
            "X-Conversation-Id");

    if (returnedConversationId) {
        conversationId =
            returnedConversationId;
    }

    if (!response.body) {
        output.textContent =
            "Streaming is not supported.";

        return;
    }

    const reader =
        response.body.getReader();

    const decoder =
        new TextDecoder();

    while (true) {

        const { value, done } =
            await reader.read();

        if (done) {
            break;
        }

        const text =
            decoder.decode(
                value,
                { stream: true });

        output.textContent += text;
    }

    output.textContent +=
        decoder.decode();
}

Now we have a complete browser-to-server streaming flow.


How the Browser Streaming Code Works

The important line is:

const reader =
    response.body.getReader();

Instead of doing:

await response.text();

which waits for the response body to complete, we access the response stream.

Then:

await reader.read();

returns data as it becomes available.

The loop continues until:

done

is true.

Each binary chunk is decoded:

decoder.decode(
    value,
    { stream: true });

and immediately appended:

output.textContent += text;

Now the user sees the response grow while the model is still generating it.


The Complete Streaming Flow

We can now see the full architecture:

Browser
  │
  │ POST /api/ai/stream
  ↓
AiController
  ↓
IAiChatService
  ↓
Conversation Store
  ↓
IChatClient
  ↓
AI Provider
  ↓
AI Model
  │
  │ ChatResponseUpdate
  ↓
AiChatService
  │
  │ yield return
  ↓
Controller
  │
  │ WriteAsync + FlushAsync
  ↓
HTTP Response
  ↓
fetch()
  ↓
ReadableStream
  ↓
TextDecoder
  ↓
Browser UI

After generation finishes:

Collected Updates
       ↓
ToChatResponse()
       ↓
Conversation History

This gives us both:

responsive streaming + conversation memory.


Test Conversation Memory from the Browser

First send:

My favorite programming language is C#.

The browser stores the returned:

X-Conversation-Id

in:

conversationId

Now send:

What programming language did I say I liked?

The request automatically includes:

{
    conversationId,
    message
}

The server retrieves the existing history.

The AI should therefore understand the follow-up question.

We now have a working browser-based multi-turn streaming chat.


Add Client-Side Cancellation

Streaming also gives us a natural place to support a Stop Generating button.

Add:

<button id="stopButton">
    Stop Generating
</button>

Then update JavaScript:

let conversationId = null;
let abortController = null;

const sendButton =
    document.getElementById("sendButton");

const stopButton =
    document.getElementById("stopButton");

const messageInput =
    document.getElementById("message");

const output =
    document.getElementById("output");

sendButton.addEventListener(
    "click",
    sendMessage);

stopButton.addEventListener(
    "click",
    () => {
        abortController?.abort();
    });

async function sendMessage() {

    const message =
        messageInput.value.trim();

    if (!message) {
        return;
    }

    output.textContent = "";

    abortController =
        new AbortController();

    try {

        const response =
            await fetch("/api/ai/stream", {
                method: "POST",

                headers: {
                    "Content-Type":
                        "application/json"
                },

                body: JSON.stringify({
                    conversationId,
                    message
                }),

                signal:
                    abortController.signal
            });

        if (!response.ok) {
            const error =
                await response.text();

            output.textContent =
                `Request failed: ${error}`;

            return;
        }

        const returnedConversationId =
            response.headers.get(
                "X-Conversation-Id");

        if (returnedConversationId) {
            conversationId =
                returnedConversationId;
        }

        if (!response.body) {
            output.textContent =
                "Streaming is not supported.";

            return;
        }

        const reader =
            response.body.getReader();

        const decoder =
            new TextDecoder();

        while (true) {

            const { value, done } =
                await reader.read();

            if (done) {
                break;
            }

            output.textContent +=
                decoder.decode(
                    value,
                    { stream: true });
        }

        output.textContent +=
            decoder.decode();
    }
    catch (error) {

        if (error.name === "AbortError") {
            output.textContent +=
                "\n\n[Generation stopped]";
        }
        else {
            console.error(error);

            output.textContent +=
                "\n\n[Streaming failed]";
        }
    }
    finally {
        abortController = null;
    }
}

Now clicking:

Stop Generating

calls:

abortController.abort();

which cancels the browser request.

That cancellation can propagate through ASP.NET Core to our AI streaming operation.


What Should We Store After Cancellation?

This introduces an important design decision.

Suppose the user sees:

Dependency injection is a design pattern that...

and clicks Stop.

Should that partial assistant response be saved?

There are several approaches.

Option 1: Save Only Completed Responses

Our current implementation does this.

The assistant response is added to conversation history only after:

updates.ToChatResponse()

runs after successful stream completion.

This keeps stored history clean.

However, the UI may contain partial text that is not stored.

Option 2: Save the Partial Response

This makes persisted history more closely match what the user saw.

But the stored response might end halfway through a sentence.

Option 3: Store Message Status

A more advanced application can persist:

Assistant Message
├── Content
├── Status
│   ├── Completed
│   ├── Cancelled
│   └── Failed
├── CreatedAt
└── CompletedAt

That is often a stronger design for durable conversation systems.

For Day 5, we will keep the simpler behavior:

Only completed streamed responses are added to conversation history.


What Happens If Streaming Fails Halfway Through?

Streaming changes normal HTTP error handling.

Before response content is sent:

Error
 ↓
Set HTTP status
 ↓
Return error body

But after several chunks have already reached the client:

Chunk
 ↓
Chunk
 ↓
Chunk
 ↓
Provider Failure

the response has already started.

At that point, you generally cannot replace the entire response with a normal JSON error document.

The client needs to understand that the stream ended unexpectedly.

This is one reason more advanced applications often use structured streaming protocols.


Raw Text vs Server-Sent Events

Our tutorial uses:

text/plain

because it clearly demonstrates the mechanics.

But raw text cannot easily represent different event types.

For example:

Content
Error
Completed
Usage
Metadata
Tool Call
Conversation ID

A structured streaming protocol such as Server-Sent Events (SSE) can represent these separately.

Conceptually:

event: content
data: Dependency injection

event: content
data: is useful because...

event: completed
data: {}

SSE is a natural next step for browser-based AI streaming when the client needs more than plain generated text.


Do We Need SignalR or WebSockets?

Not necessarily.

If the flow is:

Browser sends request
       ↓
Server streams response

HTTP streaming may be enough.

SignalR or WebSockets become more useful when the application needs ongoing bidirectional communication such as:

  • server notifications

  • collaborative sessions

  • multiple asynchronous server events

  • long-lived connections

  • client events while generation is running

Choose the transport based on the communication pattern rather than simply because the application contains AI.


Streaming and Conversation Memory Solve Different Problems

Streaming controls:

How does the user receive the response?

Conversation memory controls:

What previous context does the model receive?

Our request still contains:

System Prompt
+
Conversation History
+
Current User Message

Streaming does not remove context limits.

The history strategy from Day 4 remains important.


Streaming Does Not Automatically Reduce AI Cost

Streaming mainly changes response delivery.

The model may still process the same input and generate the same output.

Cost can still depend on:

  • selected model

  • input tokens

  • output tokens

  • conversation history

  • provider pricing

  • generated response length

Use streaming primarily to improve responsiveness, not as a cost optimization.


Avoid Expensive Work for Every Update

Do not do this without a specific reason:

await foreach (var update in updates)
{
    await database.SaveChangesAsync();
}

A single AI response may produce many updates.

That could cause many database operations.

A better default is:

Receive Update
      ↓
Send to Client
      ↓
Collect
      ↓
Stream Completes
      ↓
Persist Logical Response

Transport chunks are not normally the same thing as application messages.


Avoid Logging Every Update

Similarly, avoid:

_logger.LogInformation(
    "Chunk: {Chunk}",
    update.Text);

for every response update.

That can cause:

  • excessive logs

  • unnecessary I/O

  • higher logging costs

  • sensitive content exposure

Useful operational logging can instead record:

Stream Started
Stream Completed
Stream Cancelled
Stream Failed
Duration
Conversation ID

without logging every fragment of generated text.


Production Buffering Considerations

A stream that works on localhost can behave differently after deployment.

The real response path might be:

ASP.NET Core
      ↓
IIS / Reverse Proxy
      ↓
Compression
      ↓
Gateway / CDN
      ↓
Browser

Any of those layers may influence buffering.

If the full answer suddenly appears at once in production, investigate the entire HTTP path.

Do not immediately assume that:

GetStreamingResponseAsync

is not working.


Practical Inventory Example

Suppose our AI assistant is part of an inventory system.

The user asks:

Analyze Wireless Mouse stock.

Current stock: 4
Reorder level: 10

Explain the situation and recommend the next step.

Without streaming:

Request
   ↓
Wait
   ↓
Complete Analysis

With streaming, the user might first see:

Current stock is below...

then:

Current stock is below the configured reorder level...

and finally the complete recommendation.

The business logic has not changed.

The experience has.

That difference can be significant in AI-powered interfaces.


Suggested Project Structure After Day 5

Our project can now look like:

CSharpAiApi
│
├── Controllers
│   └── AiController.cs
│
├── Models
│   ├── AiChatResult.cs
│   ├── AiConversation.cs
│   ├── AiStreamContext.cs
│   ├── ChatRequest.cs
│   └── ChatResponse.cs
│
├── Services
│   ├── IAiChatService.cs
│   ├── AiChatService.cs
│   ├── IConversationStore.cs
│   └── InMemoryConversationStore.cs
│
├── wwwroot
│   ├── index.html
│   └── chat.js
│
├── Program.cs
└── appsettings.json

If you serve the HTML directly from wwwroot, make sure static files are enabled in your ASP.NET Core application according to the project setup you are using.


Common Error: Browser Receives Everything at Once

Check:

  • GetStreamingResponseAsync is being used

  • the controller writes each update

  • the response is flushed

  • the browser reads response.body

  • the client is not calling only response.text()

  • a proxy is not buffering

  • compression is not delaying small chunks

  • the provider supports streaming

Testing with:

curl -N

can help determine whether the issue is server-side or browser-side.


Common Error: Conversation Memory Stops Working

Make sure the streamed updates are collected:

updates.Add(update);

then reconstructed:

ChatResponse response =
    updates.ToChatResponse();

and finally stored:

conversation.Messages.AddMessages(
    response);

Streaming content to the browser alone does not update conversation memory.


Common Error: ToChatResponse() Is Missing

Verify:

using Microsoft.Extensions.AI;

and check the version of Microsoft.Extensions.AI installed in the project.

AI-related .NET APIs are evolving, so code from newer documentation may not exactly match an older package version.


Common Error: Stop Button Does Nothing

Check that JavaScript sends:

signal: abortController.signal

and that the ASP.NET Core cancellation token continues through the service into:

GetStreamingResponseAsync(...)

Cancellation must flow through the whole operation.


Common Error: Changing Headers After Streaming Starts

Set:

Content-Type
Conversation ID
Status
Other headers

before the first call to:

Response.WriteAsync(...)

Once response data has started going to the client, changing headers may no longer be possible.


Production Checklist

Before exposing an AI streaming endpoint publicly, review:

  • authentication

  • authorization

  • conversation ownership

  • rate limiting

  • input limits

  • API key protection

  • cancellation

  • timeout handling

  • partial-response policy

  • provider failure handling

  • data privacy

  • logging

  • conversation persistence

  • context limits

  • proxy buffering

  • response compression

  • AI usage and cost

Streaming improves user experience, but it also introduces additional operational concerns.


What We Built Today

Our application has progressed from:

Request
   ↓
Wait
   ↓
Complete Response

to:

Request
   ↓
AI Update
   ↓
HTTP Chunk
   ↓
Browser
   ↓
AI Update
   ↓
HTTP Chunk
   ↓
Browser

We implemented:

  • GetStreamingResponseAsync

  • IAsyncEnumerable<ChatResponseUpdate>

  • await foreach

  • ASP.NET Core HTTP streaming

  • WriteAsync

  • FlushAsync

  • browser ReadableStream

  • JavaScript TextDecoder

  • client-side cancellation

  • conversation ID handling

  • conversation-memory preservation

  • partial-stream considerations

We now have a working end-to-end streaming architecture rather than only a server-side example.


Day 5 Checklist

Before moving to Day 6, make sure you understand:

  • GetResponseAsync vs GetStreamingResponseAsync

  • what IAsyncEnumerable<T> does

  • what ChatResponseUpdate represents

  • how await foreach processes updates

  • why one update should not be treated as one token

  • how ASP.NET Core streams response content

  • why FlushAsync matters

  • how JavaScript reads response.body

  • how TextDecoder converts chunks into text

  • how the conversation ID is preserved

  • why updates must be reconstructed for memory

  • how browser cancellation works

  • what happens when streaming fails halfway through

  • why proxy buffering can affect production behavior

  • when SSE or SignalR may be more appropriate

If those concepts are clear, our application is ready for the next step.


What's Next: Day 6

So far, our AI returns natural-language text.

That works well when a human will read the answer.

Applications often need something more predictable.

Suppose our inventory application asks the model to analyze:

Product: Wireless Mouse
Current Stock: 4
Reorder Level: 10

Instead of receiving:

The stock appears low and you should probably reorder soon...

our C# application may need:

{
  "stockStatus": "Low",
  "riskLevel": "High",
  "recommendedReorderQuantity": 20,
  "explanation": "Current stock is below the reorder level."
}

That is structured AI output.

In Day 6: Structured AI Output in C# with IChatClient, we will explore:

  • strongly typed C# results

  • JSON-based AI output

  • structured response schemas

  • deserialization

  • validation

  • why parsing free-form AI text is fragile

  • using AI results safely in application logic

This will move our project beyond chat and toward AI-powered business features.


Final Thoughts

Streaming looks like a presentation feature, but implementing it correctly affects several layers.

We changed:

AI Provider
ASP.NET Core
HTTP Response
Browser
Conversation Memory
Cancellation

The most important idea is that streaming changes how the response travels, not who owns the application state.

Our ASP.NET Core application still owns:

  • conversation identity

  • history

  • validation

  • authorization

  • persistence decisions

  • failure handling

The model generates the content.

The application controls how that content is delivered and stored.

Day 1 connected C# to an AI model.

Day 2 introduced IChatClient.

Day 3 created a reusable ASP.NET Core AI service.

Day 4 added conversation memory.

Day 5 added end-to-end AI response streaming from the model all the way to the browser.

Next, we will make AI responses easier and safer for C# code to consume using structured output.

Comments 0