Crawler
Crawl a site within a page budget, review what came back, and publish the approved pages to Knowledge.
The Crawler collects pages from a website (or from a web search), cleans them, and holds them for review. Pages you approve can then be published into Knowledge for search.
pip install aice-crawlerCreate a client
import os
from aice_crawler import CrawlerClient
client = CrawlerClient(
base_url=os.environ["AICE_CRAWLER_URL"], # http://localhost:6200/crawl
api_key=os.environ["AICE_API_KEY"],
user_id=current_user.id,
)A crawl, start to finish
Create the job
job = client.jobs.create(
start_urls=["https://docs.example.com"],
config={
"seed_urls": ["https://docs.example.com"],
"max_pages": 200,
"max_depth": 2,
"scope": "host",
},
)Put the job's fields in config
The client names two fields start_urls and page_budget, but the service expects seed_urls and max_pages, and ignores the others. Everything in config is sent as-is, so set the real fields there, as above. start_urls is still required by the Python signature.
| Field | Default | Meaning |
|---|---|---|
seed_urls | pages to start from. This or search_query is required | |
search_query | start from web search results instead | |
max_search_results | 5 | results to take from the search (1–25) |
max_pages | deployment default | the page budget |
max_depth | deployment default | how many links deep to follow (0–6) |
scope | deployment default | host, domain or any. any is limited to depth 1 |
mode | raw | raw, or llm to clean each page with an instruction |
llm_instruction | required when mode is llm | |
review_mode | manual | manual holds pages for review; auto approves them |
name | a label for the job |
URLs that can't be crawled are rejected when you create the job, not when it runs.
Start it and read the pages
client.jobs.start(job["id"])
pages = client.jobs.list_pages(job["id"])List calls return an object wrapping the list, such as {"pages": [...]}.
Approve pages
Approval is per page. The request needs the IDs, from 1 to 500 at a time:
curl -X POST "$AICE_CRAWLER_URL/jobs/$JOB_ID/approve" \
-H "Authorization: Bearer $AICE_API_KEY" -H "X-User-Id: $USER_ID" \
-H "Content-Type: application/json" \
-d '{"page_ids": ["pg_1", "pg_2"], "approved": true}'Known issue: jobs.approve
jobs.approve(job_id) sends no body, so the service rejects it. Send the request yourself, as above, until the client takes page IDs. Send "approved": false to un-approve.
Publish
client.jobs.publish(job["id"])Approved pages go to Knowledge and become searchable once processed.
Re-crawling on a schedule
client.jobs.schedule(job["id"], {"interval_seconds": 86400})
client.jobs.schedule(job["id"], {"interval_seconds": None}) # cancelThere's a minimum interval, so a schedule can't hammer a site. first_run_at sets when it starts. jobs.list_runs shows each run, and jobs.reclean re-processes pages without fetching them again.
Organization policy
policy.get() returns the project's crawl limits, and policy.update(...) changes them. Updating needs the admin, owner or service_role role.
| Field | Range |
|---|---|
max_concurrent_jobs | 1–50 |
max_jobs_per_day | 1–1,000 |
max_pages_per_day | 1–1,000,000 |
agent_allowed_domains | up to 200 domains agents may crawl |
Reference
| Python | TypeScript | Endpoint |
|---|---|---|
jobs.create | jobs.create | POST /jobs (not retried; see Errors) |
jobs.list / get | jobs.list / get | GET /jobs, GET /jobs/{id} |
jobs.start | jobs.start | POST /jobs/{id}/start |
jobs.list_pages | jobs.listPages | GET /jobs/{id}/pages |
pages.get | pages.get | GET /pages/{id} |
jobs.approve | jobs.approve | POST /jobs/{id}/approve (see known issue) |
jobs.publish | jobs.publish | POST /jobs/{id}/publish |
jobs.schedule | jobs.schedule | PUT /jobs/{id}/schedule |
jobs.list_runs | jobs.listRuns | GET /jobs/{id}/runs |
jobs.reclean | jobs.reclean | POST /jobs/{id}/reclean |
policy.get / update | policy.get / update | GET, PUT /policy |