Streaming
Receive completions incrementally via server-sent events to cut perceived latency and render tokens as they arrive.
Set stream: true and AnyRouter returns a stream of server-sent events (SSE) instead of one buffered reply. Every Chat Completions and Messages request supports it. This is the fastest way to make long generations feel responsive.
Why stream?
- Lower perceived latency — render the first token the moment the model emits it, instead of waiting for the full response.
- Long outputs feel responsive — users see progress on multi-second generations.
- Early cancellation — abort mid-flight if the user changes their mind, and pay only for tokens already emitted.
Before you start
- An AnyRouter API key (prefix
sk-ar-). Create one in the dashboard.
Make a streaming request
Set "stream": true in the request body. Pick your client below.
curl -N https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer sk-ar-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-5.4-mini",
"stream": true,
"messages": [{"role": "user", "content": "Count to 10"}]
}'
The -N flag disables curl's output buffering so chunks appear as they arrive.
const stream = await client.chat.completions.create({
model: "openai/gpt-5.4-mini",
messages: [{ role: "user", content: "Count to 10" }],
stream: true,
})
for await (const chunk of stream) {
const token = chunk.choices[0]?.delta?.content ?? ""
process.stdout.write(token)
}
The official OpenAI SDK parses the SSE stream for you.
stream = client.chat.completions.create(
model="openai/gpt-5.4-mini",
messages=[{"role": "user", "content": "Count to 10"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
Read the event stream
Each line is an SSE data: event carrying a JSON payload:
data: {"id":"chatcmpl-abc","choices":[{"delta":{"content":"Hel"}}]}
data: {"id":"chatcmpl-abc","choices":[{"delta":{"content":"lo"}}]}
data: {"id":"chatcmpl-abc","choices":[{"delta":{"content":"!"},"finish_reason":"stop"}]}
data: [DONE]
delta.content— the token(s) added in this event.delta.role— set only on the first event.finish_reason— set on the final event.[DONE]— sentinel. No more events follow.
Cancellation
Close the HTTP connection to cancel a stream. AnyRouter propagates the abort to the upstream provider, so you only pay for tokens emitted up to the cancel point.
Wire cancellation into your UI: when the user navigates away or hits a stop button, call controller.abort() on your AbortController. AnyRouter bills only for tokens streamed before the abort.
Verify
You are streaming correctly when tokens print incrementally rather than appearing all at once, and the final event carries finish_reason followed by data: [DONE].