Errors and limits

What fails before a run starts, what fails inside one, the refusals that look like success, and every limit.

Before a run starts

Only two kinds of failure happen before the run exists:

StatusCauseWhat to do
401missing, invalid or expired key, or a missing / non-UUID X-User-Idfix the configuration. Only token expired changes over time. See Authentication
422a configurable field outside the allowed setfix the code that builds the config. Not a user-facing error
503the platform couldn't fetch the console's public keysretry later. The check fails closed

Error bodies share one shape:

{ "error": "validation_error", "message": "…", "details": null }

error follows the status: bad_request, unauthorized, forbidden, not_found, conflict, validation_error, internal_error, service_unavailable. A 500 never includes internal detail.

During and after a run

SymptomCauseWhat to do
the run finishes with status: "error"something failed inside it: a tool, the model call, a middlewareshow a failed turn. The request that started it succeeded
the stream throws halfwaythe connection dropped. Not necessarily a failed runre-read threads.get_state(). Never resend the message
404 from get_state or title generationthe thread was deleted, or belongs to another usergo to the empty state; don't render a broken conversation
threads.search() is missing a thread you know existsit belongs to another user or organizationthat's isolation working. Check which user the client acts for
409 from title generationthe thread has no turns yetcall it after a turn
503 from title generation or starterssuggestions are switched off for this deploymentrender without titles or chips

Streaming and polling code should always check the run's status. "The request succeeded" isn't the same as "the turn produced an answer".

Error types in the client

The two halves of the client raise different errors:

  • agents, skills, usage and suggestions raise the AICE hierarchy (AuthenticationError, NotFoundError and so on) and get AICE's retries.
  • threads, runs, assistants and crons raise whatever the LangGraph SDK raises, with its own retry policy.

If you want one error type across both, catch the LangGraph SDK's errors at the call site and translate them.

Runs that succeed but say no

This is the most commonly missed behaviour, because nothing signals it. Five safeguards can end a run by writing an ordinary assistant message and stopping. The status is success: no error, no special event.

ConditionWhat the user sees
organization over its request rateThis workspace is sending requests faster than its limit allows. Please wait a moment and try again.
the run hit its token or cost ceilingThis conversation has reached its size limit and cannot continue. Please start a new conversation; a summary of what we covered will help you pick up where we left off.
organization used up its period budgetThis workspace has reached its usage limit for the current period. Please contact your administrator to continue.
the user's message matched a content ruleThat request cannot be processed because it matches a content rule for this workspace.
the reply matched a content ruleThe response was withheld because it matched a content rule for this workspace.

Don't match on these strings to drive UI. They're wording, not an API, and can change. They're listed so you recognise them in a transcript and can explain them in support.

Two related behaviours:

  • Rate limiting never returns 429. It fails open if its backing store is unreachable, and otherwise refuses with the message above.
  • Malformed tool calls are repaired, not refused. Invalid or duplicate-ID calls are dropped and a note like [dropped 2 invalid tool call(s): …] is appended to the reply.

Limits

LimitDefaultSet by
tokens per run250,000budget.max_tokens
cost per run$2.00budget.max_usd
tool calls per run60budget.max_tool_calls
model calls per run40budget.max_model_calls
time per tool call20 sbudget.tool_timeout_seconds
budget alert80% of the token or cost ceilingdeployment only
clock leeway on keys30 sdeployment only
suggestions per set3deployment only
suggestion title1–80 charactersfixed
suggestion prompt1–500 charactersfixed
starters refresh after6 hoursdeployment only

You can lower the budget limits for a single run; see Run configuration.

On this page