Crawler

Crawl a site within a page budget, review what came back, and publish the approved pages to Knowledge.

The Crawler collects pages from a website (or from a web search), cleans them, and holds them for review. Pages you approve can then be published into Knowledge for search.

pip install aice-crawler

Create a client

import os
from aice_crawler import CrawlerClient

client = CrawlerClient(
    base_url=os.environ["AICE_CRAWLER_URL"],  # http://localhost:6200/crawl
    api_key=os.environ["AICE_API_KEY"],
    user_id=current_user.id,
)

A crawl, start to finish

Create the job

job = client.jobs.create(
    start_urls=["https://docs.example.com"],
    config={
        "seed_urls": ["https://docs.example.com"],
        "max_pages": 200,
        "max_depth": 2,
        "scope": "host",
    },
)

Put the job's fields in config

The client names two fields start_urls and page_budget, but the service expects seed_urls and max_pages, and ignores the others. Everything in config is sent as-is, so set the real fields there, as above. start_urls is still required by the Python signature.

FieldDefaultMeaning
seed_urlspages to start from. This or search_query is required
search_querystart from web search results instead
max_search_results5results to take from the search (1–25)
max_pagesdeployment defaultthe page budget
max_depthdeployment defaulthow many links deep to follow (0–6)
scopedeployment defaulthost, domain or any. any is limited to depth 1
moderawraw, or llm to clean each page with an instruction
llm_instructionrequired when mode is llm
review_modemanualmanual holds pages for review; auto approves them
namea label for the job

URLs that can't be crawled are rejected when you create the job, not when it runs.

Start it and read the pages

client.jobs.start(job["id"])
pages = client.jobs.list_pages(job["id"])

List calls return an object wrapping the list, such as {"pages": [...]}.

Approve pages

Approval is per page. The request needs the IDs, from 1 to 500 at a time:

curl -X POST "$AICE_CRAWLER_URL/jobs/$JOB_ID/approve" \
  -H "Authorization: Bearer $AICE_API_KEY" -H "X-User-Id: $USER_ID" \
  -H "Content-Type: application/json" \
  -d '{"page_ids": ["pg_1", "pg_2"], "approved": true}'

Known issue: jobs.approve

jobs.approve(job_id) sends no body, so the service rejects it. Send the request yourself, as above, until the client takes page IDs. Send "approved": false to un-approve.

Publish

client.jobs.publish(job["id"])

Approved pages go to Knowledge and become searchable once processed.

Re-crawling on a schedule

client.jobs.schedule(job["id"], {"interval_seconds": 86400})
client.jobs.schedule(job["id"], {"interval_seconds": None})  # cancel

There's a minimum interval, so a schedule can't hammer a site. first_run_at sets when it starts. jobs.list_runs shows each run, and jobs.reclean re-processes pages without fetching them again.

Organization policy

policy.get() returns the project's crawl limits, and policy.update(...) changes them. Updating needs the admin, owner or service_role role.

FieldRange
max_concurrent_jobs1–50
max_jobs_per_day1–1,000
max_pages_per_day1–1,000,000
agent_allowed_domainsup to 200 domains agents may crawl

Reference

PythonTypeScriptEndpoint
jobs.createjobs.createPOST /jobs (not retried; see Errors)
jobs.list / getjobs.list / getGET /jobs, GET /jobs/{id}
jobs.startjobs.startPOST /jobs/{id}/start
jobs.list_pagesjobs.listPagesGET /jobs/{id}/pages
pages.getpages.getGET /pages/{id}
jobs.approvejobs.approvePOST /jobs/{id}/approve (see known issue)
jobs.publishjobs.publishPOST /jobs/{id}/publish
jobs.schedulejobs.schedulePUT /jobs/{id}/schedule
jobs.list_runsjobs.listRunsGET /jobs/{id}/runs
jobs.recleanjobs.recleanPOST /jobs/{id}/reclean
policy.get / updatepolicy.get / updateGET, PUT /policy

On this page