Knowledge
Upload documents, organise them into scopes, and search them by meaning.
The Knowledge service turns documents (PDF, DOCX, and scans through OCR) into searchable text and vectors. Scopes let you tag documents so each user only searches what they're allowed to see.
pip install aice-knowledgeCreate a client
import os
from aice_knowledge import KnowledgeClient
client = KnowledgeClient(
base_url=os.environ["AICE_KNOWLEDGE_URL"], # http://localhost:6100/knowledge
api_key=os.environ["AICE_API_KEY"],
user_id=current_user.id,
user_scopes=["finance"], # optional; see below
)Upload
Uploading queues the document for parsing and embedding. It becomes searchable once processing finishes.
with open("contract.pdf", "rb") as f:
doc = client.documents.upload(file=f.read(), filename="contract.pdf")Check progress with documents.get(id), and re-run a failed document with documents.retry(id). Of the optional metadata, the service reads only source_uri.
Known issue: batch upload
documents.upload_batch / uploadBatch names its files file0, file1 and so on, but the service expects them all under files, so the call fails validation. Until that's fixed, upload in a loop, or post the files yourself:
curl -X POST "$AICE_KNOWLEDGE_URL/documents/batch" \
-H "Authorization: Bearer $AICE_API_KEY" -H "X-User-Id: $USER_ID" \
-F "files=@a.pdf" -F "files=@b.pdf"Search
hits = client.search("payment terms", top_k=5)The request accepts query, top_k, scope_ids (search only these scopes) and explain. Pass the others the same way as top_k.
The client also sends a limit field, which the service ignores. Use top_k to set how many results come back.
Scopes
Scopes are how you control visibility. There are three pieces:
- A scope type is a kind of grouping, like "department".
- Scopes are its values, like "finance" or "legal". They can be nested.
- Tags link documents to scopes.
dept = client.scope_types.create({"name": "department"})
client.scopes.create({
"scope_type_id": dept["id"],
"scopes": [{"name": "finance", "parent_scope_id": None}],
})
client.document_scopes.link({"tags": [{"document_id": doc["id"], "scope_id": finance_id}]})| Call | Body |
|---|---|
scope_types.create | {name} |
scopes.create | {scope_type_id, scopes: [{name, parent_scope_id}]} |
scopes.delete | {scope_ids: [...]} |
document_scopes.link / .delete | {tags: [{document_id, scope_id}]} |
document_scopes.list | query document_ids (required) |
List calls return an object wrapping the list, such as {"scope_types": [...]}, not a bare list.
What a user can see
The acting user's X-User-Scopes decides what they find:
user_scopes / userScopes | Searches |
|---|---|
| not set | every document in the project |
["finance", "legal"] | documents tagged with those scopes |
[] | nothing that's tagged |
Set it from your own permissions for the user. The service trusts it because only your backend holds the key.
Sharing
documents.share_create(id, {"tenant_id": other_project_id}) shares a document with another project (the field is still called tenant_id), and share_list(id) lists who has it. To stop sharing, share_delete(id, tenant_id) takes the other project's ID as its second argument (the SDK calls it share_id).
Reference
| Python | TypeScript | Endpoint |
|---|---|---|
documents.upload | documents.upload | POST /documents |
documents.upload_batch | documents.uploadBatch | POST /documents/batch (see known issue) |
documents.list | documents.list | GET /documents (status, batch_id, limit, offset via params) |
documents.get | documents.get | GET /documents/{id} |
documents.retry | documents.retry | POST /documents/{id}/retry |
documents.delete | documents.delete | DELETE /documents/{id} |
documents.share_create / share_list / share_delete | shareCreate / shareList / shareDelete | /documents/{id}/share |
search | search | POST /search |
scope_types.* | scopeTypes.* | /scope-types |
scopes.* | scopes.* | /scopes |
document_scopes.* | documentScopes.* | /document-scopes |