FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

Streaming LLM output to the UI: SSE first, WebSocket only when you need it

Most LLM chat features only need the server to push tokens to the browser. Reaching for WebSocket out of habit can leave you with infrastructure trade-offs the client never needed.

In brief

  • OpenAI's streaming documentation centres on HTTP streaming with stream=true over SSE, which makes SSE the natural starting point for LLM chat.
  • Use WebSocket only when the interface genuinely needs to send signals back up mid-stream, and accept that it has no backpressure and cannot be cached through a CDN or proxy.
  • Polling wastes resources because most requests come back empty; keep it as a fallback.
ShareLinkedInFacebookX
GraphicSSE flow for LLM chat: reconnects never re-call the model
  1. 1POST creates the chat turnThe question goes in a separate request because SSE can't send data from the client
  2. 2Call model with stream=trueBackend calls the LLM exactly once and writes each text chunk to the turn's buffer
  3. 3GET returns text/event-streamStream endpoint only reads the buffer, writing each data: line and then flushing
  4. 4EventSource renders text liveThe browser receives each event and appends it to the answer pane
  5. 5Auto-reconnect on network lossBrowser waits per the retry field, then rereads the buffer; the model isn't called again

Only the POST request calls the model. The SSE stream just reads the buffer, so a browser auto-reconnect costs no extra calls.

Graphic: FDE Times

OpenAI’s streaming guide states plainly that it focuses on HTTP streaming with stream=true over server-sent events. The model provider has already chosen how tokens leave its side. What remains is getting those tokens from your backend to the client’s screen, and this is where many teams choose badly.

Picture the first week at a banking client. They want an internal assistant that answers questions about procedures, and staff do not want to stare at a blank screen while the model thinks. An engineer used to building realtime apps will set up WebSocket straight away.

But if the interface only needs to receive text as it flows down, that choice brings infrastructure trade-offs the problem does not require.

For an FDE, choosing the streaming channel is the first architectural decision the client will actually see, because it determines whether the demo feels smooth. This guide covers three options in the order you should try them: SSE first, WebSocket when there is a reason, polling when nothing else works.

What you will build, and what you need

You will build one endpoint that accepts a question and calls the LLM with stream=true, plus a second endpoint that pushes each chunk of text to the browser over SSE. The browser uses EventSource to render the text as it arrives. You will then rewrite the client side with WebSocket for comparison.

You need any backend you are comfortable with, an API key from an LLM provider that supports streaming, and a browser with DevTools. The code below is a sketch with error handling and authentication stripped out. Treat it as a skeleton to port to your framework, not something to paste into production.

Step 1: Separate calling the model from reading the stream

According to MDN, the server-side script sending events must respond with the MIME type text/event-stream. Without this header EventSource will not work, even though the data still crosses the network.

MDN also describes SSE as a one-way connection: the client cannot send events back to the server. So the user’s question travels in a separate POST request that creates the conversation turn. The GET endpoint only reads back what the model has already generated and never calls the model itself.

# Simplified sketch, not tied to any framework
POST /chat/turns                 # create the turn, call the model exactly once
  turn_id = new_id()
  run_in_background:
      for chunk in llm.call(messages, stream=true):
          buffer[turn_id].append(chunk)
  return turn_id

GET /chat/stream?id=TURN_ID      # only reads the buffer, never calls the model
  set header Content-Type: text/event-stream
  for chunk in buffer[TURN_ID], wait for new chunks until done:
      write "data: " + chunk + "\n\n"
      flush

How to check: open the Network tab in DevTools, call the GET endpoint and look at the Content-Type header. If the response appears all at once at the end instead of trickling in, the backend or a proxy layer may be buffering data before sending it. Check the flush call first.

Step 2: The browser side takes only a few lines

Once the POST returns a turnId, the browser opens the stream to receive the answer.

// Simplified sketch
const es = new EventSource('/chat/stream?id=' + turnId);
es.onmessage = (e) => { output.textContent += e.data; };

The most important check is to switch off the network for a few seconds mid-stream and then switch it back on. According to MDN, the browser reconnects automatically by default when the connection closes, and the wait time is controlled by the retry field. Open the backend logs and confirm that the model was not called again.

Note that the sketch above rereads the buffer from the start on reconnection, so text on screen may be printed twice. In a real build you need to clear what was already rendered before rereading, or track the position already read.

Phil Sturgeon, who writes the APIs You Won’t Hate blog, observes that SSE works very well inside an HTTP/REST API for sending updates. Because SSE is still HTTP, the stream endpoint can sit alongside your other APIs.

But as Step 1 warned, the client’s proxy can still buffer the stream, so test on their actual infrastructure before the demo.

Step 3: When is it worth moving to WebSocket?

WebSocket opens a two-way session between browser and server. roadmap.sh describes it as a persistent full-duplex channel over a single TCP connection. Ably draws the line in a similar way: SSE suits one-way updates pushed by the server, while WebSocket suits two-way communication such as games or chat.

// Simplified sketch
const ws = new WebSocket('wss://example.com/chat');
ws.onmessage = (e) => { output.textContent += e.data; };
ws.send(JSON.stringify({ type: 'question', text: q }));

The deciding question is this: while the model is answering, does the interface need to send anything up? An agent that pauses to ask the user for confirmation before taking an action, or a screen where several people watch the same answer stream, are legitimate reasons. A “Stop” button, by contrast, usually needs only a separate request to cancel that turn.

Switching to WebSocket costs you a few things. MDN notes that the standard WebSocket API does not support backpressure, so if the client processes data more slowly than tokens arrive, data piles up. roadmap.sh points out that CDNs and proxies cannot cache WebSocket connections.

For these reasons, roadmap.sh recommends SSE over HTTP rather than WebSocket when you only need the server to push data down.

Polling: a fallback, not a default

With polling, the client periodically sends a request asking the server whether more text is available. Hookdeck observes that this wastes resources, and that you constantly have to weigh whether the next call will return anything. With an LLM, most requests will come back empty or with only a few extra tokens.

Polling still has a place when the client’s environment does not let long-lived connections survive. In that case, use it as a conditional fallback mode and tell the client clearly that the experience will be choppier.

SSE WebSocket Polling
Direction Server → client Two-way Client asks periodically
Weakness to remember 6-connection limit without HTTP/2 No backpressure, cannot be cached through a CDN/proxy Wastes resources, many empty requests
Use when Chat, single-turn answers Agent needs to ask back, multiple users sharing a view Long-lived connections are blocked

Three failures that wreck a demo

The first comes from automatic reconnection itself. If the stream endpoint both accepts the request and calls the model, every browser reconnection can generate a fresh answer from scratch, wasting tokens and printing duplicate text. The safer approach is to separate creating the conversation turn from reading the stream, as in Step 1.

The second is the connection limit. MDN warns that when not running over HTTP/2, SSE is limited to a very low 6 connections per browser. If a user opens several tabs of the same chat page, the seventh tab may hang waiting for a connection. Ask the client’s infrastructure team whether the server runs HTTP/2 before the demo.

The third is more about process than technology. OpenAI warns that streaming model output in production makes content moderation harder, because a partial answer is difficult to evaluate.

With financial or healthcare clients, ask early whether content must pass through a filter before it is displayed, because the answer may force you to change the whole design.

Three questions before the first line of code

On a client site, the hard part is not writing EventSource. Before writing the first line of code, ask all three questions: does the interface need to send anything up mid-stream, does the infrastructure have HTTP/2 and which proxy layers sit in between, and does the content need moderation?

When reading job descriptions for FDE or AI engineer roles, watch for phrases such as “streaming”, “realtime UI” or “production LLM app”. On your CV, instead of writing “used WebSocket”, write one line that gives the reasoning: chose SSE for one-way chat, separated the conversation turn from the read stream so reconnections never call the model again.

A line like that shows you understand the trade-offs, not just the tools.

Choose the simplest option that still meets the need, and have your reasoning ready when someone asks why you did not use WebSocket.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
7 sources
Read next on the roadmap · Stage 3: Applied AIClassical ML or deep learning: choose by the shape of the client's data, not by fashionOn a ten-thousand-row table, tree-based models are still hard to beat. On images and text the balance flips, and a good FDE should be able to explain why in the first meeting.