Streaming

Stream modes, assembling chunks into messages, cancelling, and recovering from a dropped connection.

A run is one turn of the agent: it reads the thread, calls the model and any tools, and appends the reply. runs.stream runs it and sends back Server-Sent Events as it goes.

for chunk in client.runs.stream(
    thread_id,
    "agent_platform",
    input={"messages": [{"role": "user", "content": text}]},
    stream_mode=["messages-tuple"],
):
    if chunk.event != "messages":
        continue
    message, metadata = chunk.data
    ...

Stream modes

Pass one or more modes as a list:

ModeEmitsUse it for
messages-tuple(message chunk, metadata) pairs: model tokens and tool calls as they're producedchat UIs. The default choice
valuesthe full state after each stepdebugging, or a UI that redraws everything
updatesonly what each step changedserver-side consumers that keep their own state
eventsraw LangChain callback eventsdeep debugging. High volume, unstable shape
customnothing in this deploymentforward compatibility only

Chunks are deltas

Each messages chunk carries the next piece of a message, not the text so far. Several messages can be in flight in one run (the model's reply, a tool's result), so keep one buffer per message id and append:

from collections import defaultdict

buffers: dict[str, str] = defaultdict(str)

for chunk in client.runs.stream(...):
    if chunk.event != "messages":
        continue
    message, _ = chunk.data
    if isinstance(message.get("content"), str):
        buffers[message["id"]] += message["content"]

Replace instead of appending and the message visibly collapses to its last token.

Two more things to render correctly:

  • An assistant message with empty content and a populated tool_calls is a real step: the agent calling a tool. Show it as a tool call, not an empty bubble.
  • content isn't always a string. Some model providers send a list of content blocks. Handle both, or render only string parts.

Waiting instead of streaming

When nothing is watching token by token, such as a background job, runs.wait runs the turn and returns the final state in one call:

state = client.runs.wait(
    thread_id,
    "agent_platform",
    input={"messages": [{"role": "user", "content": text}]},
)

Cancelling

A chat that can't stop a long answer feels broken. Start the run with runs.create to get its ID, then cancel it:

run = client.runs.create(thread_id, "agent_platform", input={...})
client.runs.cancel(thread_id, run["run_id"])

A dropped connection isn't a lost run

The run keeps going on the server whether or not anyone is listening, and the platform keeps a replay buffer for reconnecting with Last-Event-ID. If a stream throws halfway:

  1. Don't send the message again. That starts a second turn.
  2. Read the thread's state and show what's there:
state = client.threads.get_state(thread_id)
messages = state["values"]["messages"]

A run can succeed and still say no

A finished stream with status: "success" doesn't always mean the agent answered. Rate limits, budget ceilings and content rules end the run with an ordinary assistant message explaining why. See runs that succeed but say no.

On this page