Steg 9 (R4): unified migrate-corpus.mjs --write over engineering/governance/ infrastructure/security. 327 filer mutert, verified=null, prosa byte-identisk (fra første ## seksjon), advisor urørt (0 endringer). To applier-fixes oppdaget under kjøring (TDD, RED→GREEN): - insertHeaderFields: anker faller nå tilbake når en meta-linje selv passerer 500B (2 filer pakket et avsnitt i **Status:** → Type/Source landet utenfor scan-vinduet, applierens post-write-assertion fanget + restaurerte). - normalizeStaleVerified: fjerner nå ALLE stale non-date **Verified:** i 500B-vinduet, inkl. stray body-dup rett under --- (9 mlops-genaiops-filer var ellers falskt "verified"/fresh, droppet fra worklist). Operatør-godkjent utvidelse av carve-out; kun stray metadata-linjer, aldri prosa. test-transform-criterion: precondition oppdatert til post-migrasjons-sannhet (fila bærer nå Source). Suite 728/728 grønn.
17 KiB
Response Chunking Strategies
Last updated: 2026-06-24 Status: GA Category: Performance & Scalability Type: reference Source: https://learn.microsoft.com/azure/application-gateway/use-server-sent-events
Innhold
- Introduksjon
- Kjernekomponenter
- Streaming med Server-Sent Events
- Semantic Chunking Approaches
- Client-Side Reassembly
- Error Handling in Chunks
- Norsk offentlig sektor
- Beslutningsrammeverk
- Referanser
- For Cosmo
Introduksjon
Response chunking handler om hvordan store AI-modellresponser fra Azure OpenAI brytes opp og leveres til klienter. Det finnes to hovedtilnærminger: streaming via Server-Sent Events (SSE) der modellens output leveres token-for-token i sanntid, og chunking av store responser der output deles opp i semantisk meningsfulle blokker for videre prosessering.
Streaming er den mest brukte chunking-strategien for Azure OpenAI. Når stream: true settes i API-kallet, returnerer tjenesten delta-oppdateringer som Server-Sent Events ettersom tokens genereres. Dette gir brukeren umiddelbar feedback (time-to-first-token typisk 200-500ms) i stedet for å vente på hele responsen (som kan ta 5-30 sekunder for lange output). For programmatisk prosessering der hele responsen trengs, er chunking av det endelige resultatet i semantisk koherente blokker viktig for downstream-systemer.
For norsk offentlig sektor der AI brukes til å generere lange dokumenter (saksframlegg, utredninger, rapporter), er response chunking avgjørende for å levere god brukeropplevelse og for å kunne prosessere store responser effektivt i saksbehandlingssystemer.
Kjernekomponenter
| Komponent | Formål | Teknologi |
|---|---|---|
| Server-Sent Events (SSE) | Real-time streaming av tokens | HTTP SSE |
| stream_options | Konfigurer streaming-oppførsel | Azure OpenAI API |
| Application Gateway | SSE proxy og load balancing | Azure App Gateway |
| API Management | SSE-støtte med policy-basert routing | Azure APIM |
| SignalR | Real-time push til web-klienter | Azure SignalR |
Streaming med Server-Sent Events
Python streaming-implementasjon
from openai import AzureOpenAI
import sys
client = AzureOpenAI(
azure_endpoint="https://my-aoai.openai.azure.com",
api_key="...",
api_version="2024-10-21"
)
def stream_chat_completion(messages: list[dict], model: str = "gpt-4o"):
"""Stream response with real-time token delivery."""
collected_content = []
stream = client.chat.completions.create(
model=model,
messages=messages,
stream=True,
stream_options={"include_usage": True}, # Få token-bruk til slutt
max_tokens=2000
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
token = chunk.choices[0].delta.content
collected_content.append(token)
sys.stdout.write(token)
sys.stdout.flush()
# Siste chunk inneholder usage
if hasattr(chunk, 'usage') and chunk.usage:
return {
"content": "".join(collected_content),
"prompt_tokens": chunk.usage.prompt_tokens,
"completion_tokens": chunk.usage.completion_tokens,
"total_tokens": chunk.usage.total_tokens
}
return {"content": "".join(collected_content)}
# Asynkron streaming
async def async_stream_completion(
client: AsyncAzureOpenAI,
messages: list[dict],
model: str = "gpt-4o",
on_token: callable = None
):
"""Async stream with callback per token."""
chunks = []
async with client.chat.completions.create(
model=model,
messages=messages,
stream=True,
stream_options={"include_usage": True}
) as stream:
async for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
token = chunk.choices[0].delta.content
chunks.append(token)
if on_token:
await on_token(token)
return "".join(chunks)
.NET streaming med IAsyncEnumerable
using Azure.AI.OpenAI;
using OpenAI.Chat;
public class StreamingService
{
private readonly AzureOpenAIClient _client;
public async IAsyncEnumerable<string> StreamCompletionAsync(
string deploymentName,
IList<ChatMessage> messages,
int maxTokens = 2000)
{
var chatClient = _client.GetChatClient(deploymentName);
var options = new ChatCompletionOptions
{
MaxOutputTokenCount = maxTokens
};
// Stream deltas
await foreach (var update in
chatClient.CompleteChatStreamingAsync(messages, options))
{
foreach (var part in update.ContentUpdate)
{
if (!string.IsNullOrEmpty(part.Text))
{
yield return part.Text;
}
}
}
}
// Bruk i ASP.NET controller
public async Task StreamToClient(
HttpContext context,
string deploymentName,
IList<ChatMessage> messages)
{
context.Response.ContentType = "text/event-stream";
context.Response.Headers.Append("Cache-Control", "no-cache");
context.Response.Headers.Append("Connection", "keep-alive");
var writer = new StreamWriter(context.Response.Body);
await foreach (var token in StreamCompletionAsync(
deploymentName, messages))
{
await writer.WriteAsync($"data: {token}\n\n");
await writer.FlushAsync();
}
await writer.WriteAsync("data: [DONE]\n\n");
await writer.FlushAsync();
}
}
Semantic Chunking Approaches
Chunk store responser i meningsfulle blokker
import re
from dataclasses import dataclass
@dataclass
class SemanticChunk:
index: int
content: str
chunk_type: str # "heading", "paragraph", "code", "list", "table"
token_count: int
def semantic_chunk_response(
response_text: str,
max_chunk_tokens: int = 500,
model: str = "gpt-4o"
) -> list[SemanticChunk]:
"""Split AI response into semantically coherent chunks."""
import tiktoken
enc = tiktoken.encoding_for_model(model)
chunks = []
current_chunk = []
current_tokens = 0
chunk_type = "paragraph"
# Del på naturlige grenser
lines = response_text.split('\n')
for line in lines:
line_tokens = len(enc.encode(line))
# Identifiser chunk-type
if line.startswith('#'):
chunk_type = "heading"
elif line.startswith('```'):
chunk_type = "code"
elif line.startswith('- ') or line.startswith('* '):
chunk_type = "list"
elif line.startswith('|'):
chunk_type = "table"
else:
chunk_type = "paragraph"
# Ny chunk ved heading eller ved token-grense
if (line.startswith('#') and current_chunk) or \
(current_tokens + line_tokens > max_chunk_tokens and current_chunk):
chunks.append(SemanticChunk(
index=len(chunks),
content='\n'.join(current_chunk),
chunk_type=chunk_type,
token_count=current_tokens
))
current_chunk = []
current_tokens = 0
current_chunk.append(line)
current_tokens += line_tokens
# Siste chunk
if current_chunk:
chunks.append(SemanticChunk(
index=len(chunks),
content='\n'.join(current_chunk),
chunk_type=chunk_type,
token_count=current_tokens
))
return chunks
Streaming accumulator med chunk-deteksjon
class StreamingChunkAccumulator:
"""Accumulate streaming tokens into semantic chunks."""
def __init__(
self,
on_chunk_complete: callable = None,
chunk_boundary_pattern: str = r'\n#{1,3}\s'
):
self.buffer = []
self.chunks = []
self.on_chunk_complete = on_chunk_complete
self.boundary_pattern = re.compile(chunk_boundary_pattern)
async def feed_token(self, token: str):
"""Feed a streaming token to the accumulator."""
self.buffer.append(token)
# Sjekk om vi har nådd en chunk-grense
current_text = ''.join(self.buffer)
if self.boundary_pattern.search(current_text):
# Del på grensen
parts = self.boundary_pattern.split(current_text, maxsplit=1)
if len(parts) > 1:
completed = parts[0]
remaining = current_text[len(completed):]
if completed.strip():
chunk = SemanticChunk(
index=len(self.chunks),
content=completed.strip(),
chunk_type=self._detect_type(completed),
token_count=len(completed.split()) # Estimat
)
self.chunks.append(chunk)
if self.on_chunk_complete:
await self.on_chunk_complete(chunk)
self.buffer = [remaining]
def finalize(self) -> list[SemanticChunk]:
"""Finalize and return all chunks."""
remaining = ''.join(self.buffer).strip()
if remaining:
self.chunks.append(SemanticChunk(
index=len(self.chunks),
content=remaining,
chunk_type=self._detect_type(remaining),
token_count=len(remaining.split())
))
return self.chunks
def _detect_type(self, text: str) -> str:
if text.startswith('```'):
return "code"
if text.startswith('#'):
return "heading"
if text.startswith('- ') or text.startswith('* '):
return "list"
return "paragraph"
Client-Side Reassembly
Web-klient med progressiv rendering
// TypeScript: Client-side SSE consumption with chunk assembly
interface StreamChunk {
content: string;
isComplete: boolean;
tokenCount: number;
}
class AIResponseAssembler {
private chunks: string[] = [];
private onUpdate: (text: string) => void;
private onComplete: (text: string, stats: object) => void;
constructor(
onUpdate: (text: string) => void,
onComplete: (text: string, stats: object) => void
) {
this.onUpdate = onUpdate;
this.onComplete = onComplete;
}
async streamFromEndpoint(url: string, body: object): Promise<void> {
const response = await fetch(url, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ ...body, stream: true }),
});
if (!response.body) throw new Error('No response body');
const reader = response.body
.pipeThrough(new TextDecoderStream())
.getReader();
let fullText = '';
let buffer = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += value;
const lines = buffer.split('\n');
buffer = lines.pop() || '';
for (const line of lines) {
if (line.startsWith('data: ')) {
const data = line.slice(6);
if (data === '[DONE]') {
this.onComplete(fullText, {
totalChunks: this.chunks.length,
totalLength: fullText.length
});
return;
}
try {
const parsed = JSON.parse(data);
const token = parsed.choices?.[0]?.delta?.content || '';
if (token) {
fullText += token;
this.chunks.push(token);
this.onUpdate(fullText);
}
} catch { /* skip malformed */ }
}
}
}
}
}
Error Handling in Chunks
Robust feilhåndtering for streaming
class ResilientStreamProcessor:
"""Handle errors during streaming response."""
def __init__(self, client: AsyncAzureOpenAI, max_retries: int = 3):
self.client = client
self.max_retries = max_retries
async def stream_with_recovery(
self,
messages: list[dict],
model: str = "gpt-4o",
max_tokens: int = 2000
) -> dict:
"""Stream with automatic recovery on failure."""
accumulated = []
total_tokens_generated = 0
for attempt in range(self.max_retries):
try:
stream = await self.client.chat.completions.create(
model=model,
messages=messages,
stream=True,
max_tokens=max_tokens - total_tokens_generated
)
async for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
token = chunk.choices[0].delta.content
accumulated.append(token)
total_tokens_generated += 1
# Sjekk for finish_reason
if chunk.choices and chunk.choices[0].finish_reason:
return {
"content": "".join(accumulated),
"finish_reason": chunk.choices[0].finish_reason,
"attempts": attempt + 1,
"recovered": attempt > 0
}
# Stream fullført uten finish_reason
return {
"content": "".join(accumulated),
"finish_reason": "stop",
"attempts": attempt + 1,
"recovered": attempt > 0
}
except Exception as e:
if attempt < self.max_retries - 1:
# Fortsett fra der vi stoppet
partial = "".join(accumulated)
if partial:
# Legg til partial output som assistant-melding
messages = messages + [
{"role": "assistant", "content": partial},
{"role": "user", "content": "Fortsett fra der du stoppet."}
]
await asyncio.sleep(2 ** attempt)
else:
return {
"content": "".join(accumulated),
"finish_reason": "error",
"error": str(e),
"attempts": attempt + 1
}
Norsk offentlig sektor
- Universell utforming: Streaming gir bedre brukeropplevelse for skjermlesere og sakte nettverk — bruker ser innhold progressivt i stedet for å vente.
- Saksbehandlingssystemer: Chunk store AI-responser i semantiske blokker (overskrifter, avsnitt, tabeller) for enkel integrasjon i saksbehandlingsdokumenter.
- Logging og audit: Ved streaming, logg den komplette responsen etter fullføring for arkiverings- og revisjonskrav.
- Application Gateway: Konfigurer response buffer disabled for SSE-støtte gjennom Azure Application Gateway eller API Management.
Beslutningsrammeverk
| Scenario | Anbefaling | Begrunnelse |
|---|---|---|
| Interaktiv chat UI | SSE streaming | Umiddelbar bruker-feedback |
| Batch dokumentprosessering | Ikke-streaming + semantic chunking | Enklere feilhåndtering |
| API-til-API integrasjon | Ikke-streaming | Enklere å parse komplett respons |
| Lang respons (>2000 tokens) | Streaming + chunk accumulator | Reduser opplevd ventetid |
| Kritisk pålitelighet | Streaming med recovery | Gjenoppta ved feil |
Referanser
- Azure OpenAI streaming — Streaming API
- Server-Sent Events with Application Gateway — SSE proxy
- API Management SSE configuration — APIM SSE
- Server-Sent Events with App Gateway for Containers — Container SSE
For Cosmo
- Bruk denne referansen når kunden implementerer streaming i AI-applikasjoner, trenger å chunke store responser, eller har feilhåndteringsproblemer med SSE.
- Streaming er alltid anbefalt for brukervendte applikasjoner — time-to-first-token reduseres fra sekunder til millisekunder.
- Konfigurer
stream_options: { include_usage: true }for å få token-bruk i siste chunk — uten dette mangler kostnadssporing. - Ved bruk av Application Gateway eller API Management som proxy: deaktiver response buffering for SSE-kompatibilitet.
- Implementer alltid recovery-logikk for streaming — nettverksavbrudd er uunngåelig i produksjon, og delvis generert output bør gjenbrukes.