Data Brain

Connect databases and files, inspect them, and query them without handing credentials to a model.

Data Brain sits between your data and anything that asks questions of it. It connects to a source, preprocesses and inspects it, and answers queries, so an agent can use the data without ever holding the connection details.

Use the REST API for now

The aice-data-brain / @aiceafrica/data-brain clients call /connections/... and /descriptors, which the current service doesn't serve, and they send a different body shape. Until they're updated, call the endpoints below directly.

Base URL

http://localhost:6300 in the local stack. Every connection route is under /v1/plugin.

Connectors

connection_type is one of:

DatabasesFilesAPIs
postgresql, mysql, mariadb, oracledb, couchdb, duckdbcsv_v2, excel, parquet, jsondatadog (only when the deployment has its API collection)

GET /v1/datasource-descriptors returns every connector with the fields its configuration needs. Build your connection form from it.

The lifecycle

Five actions, each a POST with the connection's configuration under data:

EndpointDoes
/v1/plugin/{connection_type}/preprocessvalidates the configuration and prepares the source
/v1/plugin/{connection_type}/processprocesses it, for example loading a file
/v1/plugin/{connection_type}/inspectdescribes what's there: tables, columns, types
/v1/plugin/{connection_type}/queryruns a query. The only one with a second field, query
/v1/plugin/{connection_type}/cleanupreleases what process created
curl -X POST "$AICE_DATA_BRAIN_URL/v1/plugin/postgresql/inspect" \
  -H "Content-Type: application/json" \
  -H "X-Workspace-Id: $WORKSPACE_ID" -H "X-Connection-Id: $CONNECTION_ID" \
  -d '{"data": {"host": "db.internal", "port": 5432, "database": "sales", "user": "reader", "password": "…"}}'
curl -X POST "$AICE_DATA_BRAIN_URL/v1/plugin/postgresql/query" \
  -H "Content-Type: application/json" \
  -d '{
    "data":  { "id": "conn_1", "workspace_id": "ws_1", "host": "db.internal", "database": "sales" },
    "query": { "query": "SELECT region, SUM(total) FROM orders GROUP BY region" }
  }'

The fields inside data and query depend on the connector; the descriptors list them. The configuration fields above are illustrative.

Headers

HeaderPurpose
X-Workspace-Id, X-Connection-Idoptional. Let the service pool connections and keep file-based connectors inside a workspace

Caching

Query results are cached when data includes both workspace_id and id. Without them, every query runs fresh. The cache deliberately doesn't use the headers above.

Errors

A malformed body returns 400 with {statusCode, error, message}, not 422. An unknown connection_type returns 404.

Access

The REST routes don't check a key today. Run Data Brain on a private network, reachable only from your backend and the Agent Platform. Agents reach it through its MCP endpoint at /v1/mcp, which uses signed requests between services.

GET /health reports the service and each connector's status.

On this page