Poll the status of a bulk extraction job
Return the current state of a bulk job: `status`, how many URLs have `completed` out of `total`, and the `results` collected so far. Results stream in as they finish, so this is safe to poll while `status` is `processing`. Each result carries its own `status`, so check per item rather than assuming a finished job means every URL succeeded. Jobs are scoped to the key that created them; polling someone else's job returns 403. Unmetered.
Authorization
BearerAuth Your API key, sent as Authorization: Bearer wm_.... Keys are wm_ followed by 40 hex characters. Create one with POST /register (free, no auth) or POST /keys.
In: header
Path Parameters
uuidResponse Body
application/json
application/json
application/json
application/json
application/json
application/json
curl -X GET "https://example.com/bulk/497f6eca-6276-4993-bfeb-53cbbbba6f08"{ "completed": 0, "created_at": "2019-08-24T14:15:22Z", "finished_at": "2019-08-24T14:15:22Z", "job_id": "453bd7d7-5355-4d6d-a38e-d9e7eb218c3f", "kind": "bulk", "results": [ { "blocks": [ { "level": 0, "text": "string", "type": "heading" } ], "chunks": [ { "end_token": 0, "start_token": 0, "text": "string" } ], "error": "string", "html": "string", "links": [ "string" ], "markdown": "string", "metadata": { "author": "string", "date": "string", "retrieved_at": "2019-08-24T14:15:22Z", "title": "string", "url": "string" }, "metrics": { "content_bytes": 0, "input_tokens": 0, "output_tokens": 0, "reduction_pct": 0, "tokens_saved": 0 }, "url": "string" } ], "status": "string", "total": 0}Submit multiple URLs for concurrent extraction
Queue a list of URLs for concurrent extraction and return a `job_id` immediately — this endpoint is asynchronous and does not wait for the work to finish. Poll `GET /bulk/{job_id}` for progress and results, or supply `webhook_url` to be notified on completion instead of polling. The response includes `webhook_signing_secret` the first time you use a webhook; store it, as it is shown once. Individual URLs can fail without failing the job: each entry in `results` carries its own `status`, so a partial success is normal and is reported as a 200. The per-request URL cap depends on your plan. Send `Idempotency-Key` to make retries safe — the same key with the same body returns the original job rather than queuing a second one. Metered as one request per URL submitted.
Crawl a site starting from a root URL
Follow links from a starting URL and extract every page reached, up to `max_depth` and `max_pages`. Returns a `job_id` immediately; crawling is asynchronous. Poll `GET /crawl/{job_id}` or supply `webhook_url`. The crawler stays on the starting host, honours `robots.txt`, and applies your key's domain policy to every link it follows — not just the root. A crawl stops early rather than failing when it hits a ceiling: `truncated` is set with a `truncated_reason` of `page_cap_reached` or `quota_exhausted`, and the pages gathered so far are still returned. `max_depth` is capped by plan. Metered per page extracted, not per page discovered.