Skip to content

Streaming & SSE

  • SSE is the default transport for LLM responses. One-way server→client, plain text over a single long-lived HTTP response, with the browser’s EventSource handling reconnection for free.
  • Generation is inherently a stream. A model emits one token per forward pass; making the user wait for the whole completion throws away the latency win. Stream tokens as they’re produced and time-to-first-token becomes the metric that matters.
  • SSE vs WebSockets is a one-way vs two-way decision. If the server just pushes and the client just listens (tokens, progress, notifications), SSE is simpler, cheaper, and firewall-friendly. If you need true bidirectional, low-latency, binary frames (collab, games), use WebSockets.
  • The wire format is trivial on purpose: data: lines, blank-line-delimited events, optional event:/id:/retry:. That simplicity is why it survives proxies and reconnects cleanly.
  • Backpressure and disconnects are the hard parts. A client that stops reading, a proxy that buffers, a dropped mobile connection — the platform must handle cancellation (stop the GPU work) and resumption (Last-Event-ID).
  • Build a /stream endpoint that emits one SSE event per word of a sentence; consume it in the browser with EventSource and watch them arrive incrementally.
  • Wire it to a real model: stream vLLM/OpenAI tokens through your server to the client. Measure TTFT vs total time.
  • Kill the connection mid-stream and confirm EventSource auto-reconnects; implement Last-Event-ID resume.
  • Re-implement the same feature over WebSockets and over gRPC server-streaming; list what each gained and lost.

A request/response cycle assumes the answer exists all at once. But a generated answer comes into being token by token, and a long job makes progress step by step — the result is a sequence over time, not a value. Streaming transports keep one connection open and push pieces as they’re ready, so the user sees the first token in tens of milliseconds instead of waiting seconds for the last one. SSE is the minimal viable version of this: a normal HTTP response that simply never ends, framed as a list of text events. Most “AI streaming” is exactly this, and its constraints (one-way, text, HTTP) are features, not limitations.

Key ideas:

  • Perceived latency: autoregressive decoding produces ~1 token per forward pass. Buffering the whole completion means the user waits for all of it; streaming means they read along with generation. Time-to-first-token (TTFT) becomes the experience.
  • Progress and cancellation: a stream lets the user stop a bad generation early — which lets the platform free the GPU immediately instead of finishing wasted work.
  • Memory: streaming a large result (or a long-running job’s logs) avoids buffering it all server-side.
  • This is the last mile of the inference path: model → serving engine → gateway → stream to client.

A never-ending HTTP response of text events

The protocol (Content-Type: text/event-stream):

data: Hello
data: world
id: 42
event: token
data: {"token": "!", "done": false}

Key ideas:

  • Fields: data: (the payload, can repeat for multi-line), event: (named event type the client can listen for), id: (event id, echoed back as Last-Event-ID on reconnect), retry: (reconnect delay in ms). A blank line dispatches the event.
  • Client: const es = new EventSource("/stream"); es.onmessage = e => append(e.data); — the browser handles the connection, parsing, and automatic reconnection.
  • One-way: server → client only. The client opens it with a normal GET; to send data it uses a separate request.
  • Plain HTTP: works through proxies, CDNs, and firewalls that mishandle WebSocket upgrades. Over HTTP/2 it isn’t limited by the old 6-connections-per-domain cap.
  • Text only: SSE carries UTF-8 text. Binary must be base64-encoded (a reason to prefer WebSockets for binary).

3. SSE vs WebSockets vs Long-Polling vs gRPC Streaming

Section titled “3. SSE vs WebSockets vs Long-Polling vs gRPC Streaming”
TransportDirectionFormatReconnectBest for
SSEServer → clientTextBuilt-in (EventSource)LLM tokens, notifications, progress, live feeds
WebSocketsBidirectionalText or binaryManualChat, collaboration, games, trading
Long-pollingServer → client (faked)AnyPer-requestLegacy fallback when SSE/WS unavailable
gRPC streamingUni/bidiBinary (Protobuf)ManualService-to-service streams (not browsers)

Key distinctions:

  • SSE vs WebSockets: SSE is simpler (it’s just HTTP), one-way, auto-reconnecting, text. WebSockets are bidirectional, binary-capable, lower-overhead per message, but you manage reconnection and they need a protocol upgrade some infra mishandles. For streaming model output, SSE is almost always the right default.
  • SSE vs gRPC streaming: gRPC server-streaming is the same idea for backend-to-backend (binary, typed, HTTP/2). Browsers can’t speak it directly. The common pattern: gRPC stream internally → SSE to the browser at the gateway.
  • Long-polling: the fallback — client requests, server holds until data, responds, client re-requests. Higher latency and overhead; use only when SSE/WebSockets aren’t available.

The canonical AI-platform streaming path

Key ideas:

  • The flow: client opens an SSE request → gateway calls the inference engine (vLLM/TGI) which yields tokens via its own streaming API (often gRPC server-stream or an async generator) → gateway re-emits each token as an SSE data: event → final data: [DONE] sentinel closes it.
  • Framing the payload: most APIs send JSON deltas ({"choices":[{"delta":{"content":"Hel"}}]}) per event, ending with a [DONE] marker — the de-facto convention (OpenAI-style).
  • Cancellation is critical: when the client disconnects (closes the tab, hits stop), propagate the cancellation down to the serving engine so it evicts the request and frees KV-cache/GPU immediately (ties to continuous batching in LLM Systems).
  • Tool calls / structured events: use named SSE event: types (token, tool_call, error, done) so the client can branch without sniffing payloads.

Key ideas:

  • Buffering proxies: nginx/load balancers may buffer the response and defeat streaming — disable buffering (X-Accel-Buffering: no), set the right timeouts, and use HTTP/1.1 chunked or HTTP/2.
  • Heartbeats / keep-alive: send a comment line (: ping) periodically so idle connections and intermediaries don’t time out.
  • Resumption: on reconnect the browser sends Last-Event-ID; the server can replay from that id if it kept a buffer (often unnecessary for fresh generations, essential for durable event feeds).
  • Connection limits & backpressure: each stream is a held-open connection and (for generation) a pinned GPU slot. Cap concurrency, apply timeouts, and shed load — a slow client shouldn’t pin a GPU forever.
  • Scaling: long-lived connections complicate load balancing and graceful shutdown (drain, don’t cut). Sticky routing or stateless re-attach; HTTP/2 multiplexing helps fan-out.

SituationUse
Stream LLM tokens to a browserSSE
Server-push notifications / progress / live feedSSE
Bidirectional, low-latency, binary (chat/collab/games)WebSockets
Backend-to-backend streaminggRPC server-streaming
Infra blocks SSE and WSLong-polling (fallback)
Need to cancel a generationPropagate disconnect → evict from serving engine
  • A result that arrives over time is a stream, not a value — push the pieces as they’re ready instead of buffering the whole answer.
  • One-way vs two-way is the decision — SSE when the server pushes and the client listens; WebSockets only when you truly need bidirectional/binary.
  • Cancellation must propagate — a client disconnect should free the upstream resource (the GPU slot), or you leak work.
  • Backpressure is mandatory — a held-open connection is a held resource; bound concurrency, time out, and shed load.
  • Internal stream → edge stream — gRPC server-streaming between services, SSE at the last mile; pick the transport per hop.
ConceptConnected TrackApplication
Token-by-token decode, TTFT, continuous batchingLLM Systems & InferenceWhy output is a stream; cancellation frees GPU
gRPC streaming, HTTP/2, protocol choiceRPC & ProtocolsInternal stream → SSE at the edge
Load balancing, timeouts, connection handlingSystem DesignServing long-lived connections
Ingress, draining, graceful shutdownCloud NativeSSE through K8s ingress
Tail latency, TTFT measurementObservabilityInstrumenting the stream
Event streams, replay, delivery semanticsData EngineeringDurable feeds vs ephemeral generation
CompanyThe pattern they lean onInstance
OpenAI / AnthropicServer-push token streamingSSE is the public API surface
Any GenAI product teamStream + cancel for TTFT-driven UXSSE with stop-generation
Vercel / CloudflareEdge streaming / AI gatewaySSE at the CDN
Slack / FigmaBidirectional real-timeWebSockets (the contrast case)