Errors and limits
What fails before a run starts, what fails inside one, the refusals that look like success, and every limit.
Before a run starts
Only two kinds of failure happen before the run exists:
| Status | Cause | What to do |
|---|---|---|
401 | missing, invalid or expired key, or a missing / non-UUID X-User-Id | fix the configuration. Only token expired changes over time. See Authentication |
422 | a configurable field outside the allowed set | fix the code that builds the config. Not a user-facing error |
503 | the platform couldn't fetch the console's public keys | retry later. The check fails closed |
Error bodies share one shape:
{ "error": "validation_error", "message": "…", "details": null }error follows the status: bad_request, unauthorized, forbidden, not_found, conflict, validation_error, internal_error, service_unavailable. A 500 never includes internal detail.
During and after a run
| Symptom | Cause | What to do |
|---|---|---|
the run finishes with status: "error" | something failed inside it: a tool, the model call, a middleware | show a failed turn. The request that started it succeeded |
| the stream throws halfway | the connection dropped. Not necessarily a failed run | re-read threads.get_state(). Never resend the message |
404 from get_state or title generation | the thread was deleted, or belongs to another user | go to the empty state; don't render a broken conversation |
threads.search() is missing a thread you know exists | it belongs to another user or organization | that's isolation working. Check which user the client acts for |
409 from title generation | the thread has no turns yet | call it after a turn |
503 from title generation or starters | suggestions are switched off for this deployment | render without titles or chips |
Streaming and polling code should always check the run's status. "The request succeeded" isn't the same as "the turn produced an answer".
Error types in the client
The two halves of the client raise different errors:
agents,skills,usageandsuggestionsraise the AICE hierarchy (AuthenticationError,NotFoundErrorand so on) and get AICE's retries.threads,runs,assistantsandcronsraise whatever the LangGraph SDK raises, with its own retry policy.
If you want one error type across both, catch the LangGraph SDK's errors at the call site and translate them.
Runs that succeed but say no
This is the most commonly missed behaviour, because nothing signals it. Five safeguards can end a run by writing an ordinary assistant message and stopping. The status is success: no error, no special event.
| Condition | What the user sees |
|---|---|
| organization over its request rate | This workspace is sending requests faster than its limit allows. Please wait a moment and try again. |
| the run hit its token or cost ceiling | This conversation has reached its size limit and cannot continue. Please start a new conversation; a summary of what we covered will help you pick up where we left off. |
| organization used up its period budget | This workspace has reached its usage limit for the current period. Please contact your administrator to continue. |
| the user's message matched a content rule | That request cannot be processed because it matches a content rule for this workspace. |
| the reply matched a content rule | The response was withheld because it matched a content rule for this workspace. |
Don't match on these strings to drive UI. They're wording, not an API, and can change. They're listed so you recognise them in a transcript and can explain them in support.
Two related behaviours:
- Rate limiting never returns
429. It fails open if its backing store is unreachable, and otherwise refuses with the message above. - Malformed tool calls are repaired, not refused. Invalid or duplicate-ID calls are dropped and a note like
[dropped 2 invalid tool call(s): …]is appended to the reply.
Limits
| Limit | Default | Set by |
|---|---|---|
| tokens per run | 250,000 | budget.max_tokens |
| cost per run | $2.00 | budget.max_usd |
| tool calls per run | 60 | budget.max_tool_calls |
| model calls per run | 40 | budget.max_model_calls |
| time per tool call | 20 s | budget.tool_timeout_seconds |
| budget alert | 80% of the token or cost ceiling | deployment only |
| clock leeway on keys | 30 s | deployment only |
| suggestions per set | 3 | deployment only |
| suggestion title | 1–80 characters | fixed |
| suggestion prompt | 1–500 characters | fixed |
| starters refresh after | 6 hours | deployment only |
You can lower the budget limits for a single run; see Run configuration.