Pre-launchAPI access is by invitation. Prices shown are estimates.Request access →
Contact sales
ModelsPricingDedicatedDocsEnterpriseContact sales
Inference

Docs / Inference

Streaming

Receive tokens as they are generated with server-sent events.

View .md

Set stream: true to receive tokens as server-sent events (SSE) instead of waiting for the full response.

stream = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4.1-Flash",
    messages=[{"role": "user", "content": "Write a haiku about rivers."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print("\n", chunk.usage)

What to expect

  • Each event is a data: line containing a chat completion chunk; the stream ends with data: [DONE].
  • With stream_options.include_usage, the final chunk carries the usage object. Kurrens always reports usage for streamed requests.
  • Reasoning models may send SSE comment lines (: keep-alive) while they think. Clients should ignore them — they keep the connection from timing out.
  • If an error occurs mid-stream, a final event contains an error object before the stream closes.

Cancelling

Closing the connection cancels generation. You’re billed only for tokens generated before cancellation.

    Type to search titles, headings, and page text.

    ↑↓ to move · ↵ to open