Streaming
Stream modes, assembling chunks into messages, cancelling, and recovering from a dropped connection.
A run is one turn of the agent: it reads the thread, calls the model and any tools, and appends the reply. runs.stream runs it and sends back Server-Sent Events as it goes.
for chunk in client.runs.stream(
thread_id,
"agent_platform",
input={"messages": [{"role": "user", "content": text}]},
stream_mode=["messages-tuple"],
):
if chunk.event != "messages":
continue
message, metadata = chunk.data
...Stream modes
Pass one or more modes as a list:
| Mode | Emits | Use it for |
|---|---|---|
messages-tuple | (message chunk, metadata) pairs: model tokens and tool calls as they're produced | chat UIs. The default choice |
values | the full state after each step | debugging, or a UI that redraws everything |
updates | only what each step changed | server-side consumers that keep their own state |
events | raw LangChain callback events | deep debugging. High volume, unstable shape |
custom | nothing in this deployment | forward compatibility only |
Chunks are deltas
Each messages chunk carries the next piece of a message, not the text so far. Several messages can be in flight in one run (the model's reply, a tool's result), so keep one buffer per message id and append:
from collections import defaultdict
buffers: dict[str, str] = defaultdict(str)
for chunk in client.runs.stream(...):
if chunk.event != "messages":
continue
message, _ = chunk.data
if isinstance(message.get("content"), str):
buffers[message["id"]] += message["content"]Replace instead of appending and the message visibly collapses to its last token.
Two more things to render correctly:
- An assistant message with empty
contentand a populatedtool_callsis a real step: the agent calling a tool. Show it as a tool call, not an empty bubble. contentisn't always a string. Some model providers send a list of content blocks. Handle both, or render only string parts.
Waiting instead of streaming
When nothing is watching token by token, such as a background job, runs.wait runs the turn and returns the final state in one call:
state = client.runs.wait(
thread_id,
"agent_platform",
input={"messages": [{"role": "user", "content": text}]},
)Cancelling
A chat that can't stop a long answer feels broken. Start the run with runs.create to get its ID, then cancel it:
run = client.runs.create(thread_id, "agent_platform", input={...})
client.runs.cancel(thread_id, run["run_id"])A dropped connection isn't a lost run
The run keeps going on the server whether or not anyone is listening, and the platform keeps a replay buffer for reconnecting with Last-Event-ID. If a stream throws halfway:
- Don't send the message again. That starts a second turn.
- Read the thread's state and show what's there:
state = client.threads.get_state(thread_id)
messages = state["values"]["messages"]A run can succeed and still say no
A finished stream with status: "success" doesn't always mean the agent answered. Rate limits, budget ceilings and content rules end the run with an ordinary assistant message explaining why. See runs that succeed but say no.