Knowledge

Upload documents, organise them into scopes, and search them by meaning.

The Knowledge service turns documents (PDF, DOCX, and scans through OCR) into searchable text and vectors. Scopes let you tag documents so each user only searches what they're allowed to see.

pip install aice-knowledge

Create a client

import os
from aice_knowledge import KnowledgeClient

client = KnowledgeClient(
    base_url=os.environ["AICE_KNOWLEDGE_URL"],  # http://localhost:6100/knowledge
    api_key=os.environ["AICE_API_KEY"],
    user_id=current_user.id,
    user_scopes=["finance"],                    # optional; see below
)

Upload

Uploading queues the document for parsing and embedding. It becomes searchable once processing finishes.

with open("contract.pdf", "rb") as f:
    doc = client.documents.upload(file=f.read(), filename="contract.pdf")

Check progress with documents.get(id), and re-run a failed document with documents.retry(id). Of the optional metadata, the service reads only source_uri.

Known issue: batch upload

documents.upload_batch / uploadBatch names its files file0, file1 and so on, but the service expects them all under files, so the call fails validation. Until that's fixed, upload in a loop, or post the files yourself:

curl -X POST "$AICE_KNOWLEDGE_URL/documents/batch" \
  -H "Authorization: Bearer $AICE_API_KEY" -H "X-User-Id: $USER_ID" \
  -F "files=@a.pdf" -F "files=@b.pdf"
hits = client.search("payment terms", top_k=5)

The request accepts query, top_k, scope_ids (search only these scopes) and explain. Pass the others the same way as top_k.

The client also sends a limit field, which the service ignores. Use top_k to set how many results come back.

Scopes

Scopes are how you control visibility. There are three pieces:

  1. A scope type is a kind of grouping, like "department".
  2. Scopes are its values, like "finance" or "legal". They can be nested.
  3. Tags link documents to scopes.
dept = client.scope_types.create({"name": "department"})
client.scopes.create({
    "scope_type_id": dept["id"],
    "scopes": [{"name": "finance", "parent_scope_id": None}],
})
client.document_scopes.link({"tags": [{"document_id": doc["id"], "scope_id": finance_id}]})
CallBody
scope_types.create{name}
scopes.create{scope_type_id, scopes: [{name, parent_scope_id}]}
scopes.delete{scope_ids: [...]}
document_scopes.link / .delete{tags: [{document_id, scope_id}]}
document_scopes.listquery document_ids (required)

List calls return an object wrapping the list, such as {"scope_types": [...]}, not a bare list.

What a user can see

The acting user's X-User-Scopes decides what they find:

user_scopes / userScopesSearches
not setevery document in the project
["finance", "legal"]documents tagged with those scopes
[]nothing that's tagged

Set it from your own permissions for the user. The service trusts it because only your backend holds the key.

Sharing

documents.share_create(id, {"tenant_id": other_project_id}) shares a document with another project (the field is still called tenant_id), and share_list(id) lists who has it. To stop sharing, share_delete(id, tenant_id) takes the other project's ID as its second argument (the SDK calls it share_id).

Reference

PythonTypeScriptEndpoint
documents.uploaddocuments.uploadPOST /documents
documents.upload_batchdocuments.uploadBatchPOST /documents/batch (see known issue)
documents.listdocuments.listGET /documents (status, batch_id, limit, offset via params)
documents.getdocuments.getGET /documents/{id}
documents.retrydocuments.retryPOST /documents/{id}/retry
documents.deletedocuments.deleteDELETE /documents/{id}
documents.share_create / share_list / share_deleteshareCreate / shareList / shareDelete/documents/{id}/share
searchsearchPOST /search
scope_types.*scopeTypes.*/scope-types
scopes.*scopes.*/scopes
document_scopes.*documentScopes.*/document-scopes

On this page