Errors and retries
One error hierarchy across every client, and what the transport retries for you.
Every service client raises the same typed errors, from aice-core / @aiceafrica/core. Catch the base class to handle all of them, or a subclass to handle one case.
The hierarchy
AiceError base
├── APIConnectionError no response at all: DNS, refused, timed out
├── AuthenticationError 401 key missing, invalid or expired
├── AiceForbiddenError 403 authenticated, but not allowed
├── NotFoundError 404
├── ValidationError 400 / 422, with field-level details
├── RateLimitError 429, with the server's Retry-After
├── ServerError 5xx
└── InsufficientScopeError raised in the client, before any requestAiceForbiddenError is named that way so it doesn't shadow Python's built-in PermissionError.
Every error carries the HTTP status, the service's error code, the request ID the server stamped on the response, and any details. Quote the request ID in a support request: it pins down the exact call.
| Field | Python | TypeScript |
|---|---|---|
| HTTP status | err.status_code | err.status |
| Service error code | err.code | err.code |
| Request ID | err.request_id | err.requestId |
| Details | err.details | err.details |
| Retry-After (429 only) | err.retry_after | err.retryAfter |
from aice_core import AiceError, NotFoundError, ValidationError
try:
client.recommend("user-1", n=10)
except NotFoundError:
... # the user doesn't exist yet: create it, or show an empty state
except ValidationError as err:
log.warning("bad request", extra={"details": err.details, "request_id": err.request_id})
except AiceError as err:
log.error("AICE call failed", extra={"status": err.status_code, "request_id": err.request_id})
raiseWhat gets retried
The transport retries for you. You don't need your own retry loop around an SDK call.
- Retried: network errors, every
5xx, and429. - Never retried: any other
4xx. Retrying a bad request only repeats it. - Backoff: exponential with jitter, capped at 10 seconds. A
Retry-Afterheader wins when present. - Budget: 3 retries after the first attempt. Change it with
max_retries/maxRetrieson the client.
Retries and writes
A retried POST could create something twice. So POST and PATCH requests get an automatically generated Idempotency-Key header, and the transport only retries them when that key is present. GET, PUT and DELETE are safe to repeat and always retry.
Some calls opt out on purpose, making exactly one attempt, because their service doesn't yet deduplicate on the key and a duplicate would be costly:
| Call | Why it isn't retried |
|---|---|
Crawler jobs.create | a duplicate starts a second crawl |
Data Brain connections.process, connections.cleanup | processing twice, or a destructive cleanup twice |
KYC verify.start, face.create_capture_session, admin.accept, admin.recheck | compliance actions and paid vendor sessions |
Sandbox session.create | each session holds quota |
For these, a transient failure reaches you immediately. Decide at the call site whether to try again.
Agent Platform is different
The Agent Platform client has two halves. Its management calls (agents, skills, usage, suggestions) raise the errors above. Its conversation calls (threads, runs, assistants, crons) come straight from the LangGraph SDK, which has its own error types and retry policy. And some successful runs are refusals, with no error at all. See Agent Platform errors.