Observability

Give your own Python services traces, metrics and logs over OpenTelemetry in a few lines.

aice-observability isn't an API client. It's the OpenTelemetry setup every AICE service uses, packaged so your own Python services report the same way: traces, metrics and logs, sent to any OTLP endpoint.

pip install aice-observability

Set it up

from fastapi import FastAPI
from observability import Observability

app = FastAPI()

obs = (
    Observability()
    .service("orders-api", version="2.0.0")
    .start()                      # before any instrument_*() call
    .instrument_fastapi(app)
    .instrument_sqlalchemy(engine)
)

Call .start() first. OpenTelemetry's instrumentation attaches to whatever provider is registered at the moment it runs. Instrument before .start() and it silently records nothing.

Outgoing httpx calls are traced automatically once .start() runs.

MethodInstruments
instrument_fastapi(app)FastAPI requests
instrument_flask(app)Flask requests
instrument_django()Django requests
instrument_sqlalchemy(engine)SQLAlchemy queries
instrument_psycopg2()psycopg2 queries
instrument_redis()Redis commands
shutdown()flushes and stops exporting. Call on exit

Configure it

Settings come from the environment:

VariableDefaultMeaning
OTEL_EXPORTER_OTLP_ENDPOINTwhere to send telemetry, over OTLP HTTP. Required unless everything is off
OTEL_EXPORTER_OTLP_HEADERSauth headers as key=value,key=value, for hosted backends
OTEL_SERVICE_NAMEthe service's name in dashboards. Required
OTEL_ENVIRONMENTdevelopmentdevelopment, staging, production…
OTEL_TRACES_ENABLEDtruesend traces
OTEL_METRICS_ENABLEDtruesend metrics
OTEL_LOGS_ENABLEDtruesend logs
OTEL_LOG_LEVELINFOthe lowest log level captured
.env
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318
OTEL_SERVICE_NAME=orders-api
OTEL_ENVIRONMENT=production

What you get

The standard dashboards are built around questions rather than raw metrics:

DashboardAnswers
Service overviewIs it healthy? How much traffic, how fast, how many errors?
InfrastructureCPU, memory, disk, network, threads, garbage collection, restarts
HTTP / APIper-endpoint rate, latency percentiles, errors, and the slowest and largest requests
Databasequery rate and duration, slow and failed queries, pool saturation. Only when a database is instrumented
External servicescalls your service makes to others
Logs and tracescorrelated by trace ID

On this page