# Web Scraping (/docs/web)
Web Scraping [#web-scraping]
Scrape, search, map, extract, crawl, and monitor any public web page. Content-bearing responses return the unified [WebPage schema](/docs/schema/web-page.md) — clean markdown/HTML, metadata, structured extraction, and fetch provenance under `data.page` — while search and map return a `WebPageList` (`data.items[]`).
This is the only platform on the API that sells **stateful** work: long-running jobs you poll, monitors that keep checking on a schedule, and browser sessions you drive step by step. Everything else here is one request, one answer.
Base URL: `/v1/web/...`
Getting Started [#getting-started]
1\. Scrape a page to markdown [#1-scrape-a-page-to-markdown]
```bash
curl "https://www.socialcrawl.dev/v1/web/scrape?url=https://example.com" \
-H "x-api-key: sc_your_api_key_here"
```
The page content arrives at `data.page.content.markdown` (add `formats=markdown,screenshot` for a screenshot URL under `data.page.media`).
2\. Search the web [#2-search-the-web]
```bash
curl "https://www.socialcrawl.dev/v1/web/search?query=social%20media%20data%20api&limit=10" \
-H "x-api-key: sc_your_api_key_here"
```
Web, news, and image sources merge into one `data.items[]` list, each row tagged with `source_type`.
3\. Pull typed fields instead of prose [#3-pull-typed-fields-instead-of-prose]
```bash
curl "https://www.socialcrawl.dev/v1/web/extract?url=https://example.com&prompt=the%20page%20title%20and%20any%20pricing%20mentioned" \
-H "x-api-key: sc_your_api_key_here"
```
`GET /v1/web/extract` needs **either** a `schema` (JSON) **or** a `prompt` — sending neither is a `400` before billing. Typed fields land under `data.page.extraction`.
The four synchronous endpoints [#the-four-synchronous-endpoints]
These four answer immediately and cover most integrations.
* **`GET /v1/web/scrape`** — one known URL to clean markdown or HTML, with the resolved URL, status code, and fetch metadata. **1cr normally.** It steps to **5cr** when you ask for work the cheap path cannot do: `proxy=auto`, `proxy=enhanced`, or `pdf_parse=true`. The `data.page.fetch.proxy_tier` field tells you which path actually ran.
* **`GET /v1/web/search`** — ranked web, news, and image results as one normalised list when you do not yet know which URLs you need. **2cr per 10 results**, so `limit=50` is 10cr. Adding `include_content=true` scrapes each result inline and adds 1cr per result on top, which is how a `limit=50` content search reaches 60cr. Filter with `include_domains` / `exclude_domains`, `country`, `time_range`, and `categories` before you widen `limit`.
* **`GET /v1/web/map`** — every URL a site exposes, without fetching any page bodies. **1cr.** Run it first to decide what is worth crawling; `search` (a substring filter over the discovered URLs) and `include_subdomains` narrow it.
* **`GET /v1/web/extract`** — structured fields from one page, shaped by your `schema` or `prompt`. **5cr flat.**
The usual chain is `map` to find the pages, `scrape` or `extract` on the ones that matter. Reach for `crawl` only when you want the site walked for you.
```bash
# 1. What pages does this site have?
curl "https://www.socialcrawl.dev/v1/web/map?url=https://example.com&search=pricing" \
-H "x-api-key: sc_your_api_key_here"
# 2. Pull the one that matters
curl "https://www.socialcrawl.dev/v1/web/scrape?url=https://example.com/pricing&formats=markdown" \
-H "x-api-key: sc_your_api_key_here"
```
Async jobs: submit, poll, settle [#async-jobs-submit-poll-settle]
`crawl`, `batch-scrape`, and `agent` are too slow to answer inline, so they return a job id and you poll one unified surface. All three ids are prefixed `job_`, and all three read through the same four routes — there is no per-kind status path.
```bash
# 1. Submit — returns data.job_id, e.g. "job_9f3k2n8d1"
curl -X POST "https://www.socialcrawl.dev/v1/web/crawl" \
-H "x-api-key: sc_your_api_key_here" \
-H "content-type: application/json" \
-d '{"url":"https://example.com","limit":25,"max_depth":2}'
# 2. Poll until data.status is terminal
curl "https://www.socialcrawl.dev/v1/web/jobs/job_9f3k2n8d1" \
-H "x-api-key: sc_your_api_key_here"
# 3. Lost the id? List everything this key has running
curl "https://www.socialcrawl.dev/v1/web/jobs?limit=20" \
-H "x-api-key: sc_your_api_key_here"
# 4. Stop it early and get the unspent hold back
curl -X DELETE "https://www.socialcrawl.dev/v1/web/jobs/job_9f3k2n8d1" \
-H "x-api-key: sc_your_api_key_here"
```
* **`POST /v1/web/crawl`** walks a site from a root URL and scrapes what it finds. `limit` caps pages, `max_depth` caps how far from the root it goes, and `include_paths` / `exclude_paths` keep it in the section you care about.
* **`POST /v1/web/batch-scrape`** takes a `urls` array when you already know every page you want. `ignore_invalid_urls` keeps one bad entry from rejecting the whole submission.
* **`POST /v1/web/agent`** takes a `url` and a plain-language `prompt` and drives a browser toward the answer. Use it when the data is behind clicking and navigating rather than at a URL you can name.
* **`GET /v1/web/jobs`** lists this key's jobs with cursor pagination; **`GET /v1/web/jobs/{job_id}`** returns status, progress, credits held and charged, and the result once terminal; **`DELETE /v1/web/jobs/{job_id}`** cancels and refunds the unspent hold. Reading and cancelling are free.
* `POST /v1/web/crawl` and `POST /v1/web/batch-scrape` accept a `webhook_url` if you would rather be told than poll. See [Webhooks](/docs/webhooks.md) for the payload.
**Credits are held, then settled.** A crawl holds 1 credit per page of `limit` (so `limit=25` holds 25) and charges only for the pages actually fetched. A batch scrape holds one credit per URL you submit. An agent job holds a flat 25 and settles against the work it did, with a 5-credit floor. Whatever is held and not spent comes back automatically — including on cancel.
The crawl hold is taken against
`limit`
, not against how big the site turns out to be. Submitting
`limit=1000`
reserves 1000 credits the moment the job is accepted, even if the site only has 30 pages. Run
`GET /v1/web/map`
first and set
`limit`
from what you find.
**Per-page failures.** Append `/errors` to a job read — `GET /v1/web/jobs/{job_id}`, then the same path with `/errors` on the end — to list the individual pages that failed inside an otherwise successful job. That sub-route is a convenience mount rather than a registry entry, so it does not appear in the roster below, in `openapi.json`, or in the `llms-*.txt` files. It works, but treat it as unversioned: do not generate an SDK method for it.
Monitors: watch a page on a schedule [#monitors-watch-a-page-on-a-schedule]
A monitor re-runs a scrape or a search on a cadence and records every result, so you can ask what changed rather than diffing it yourself.
```bash
# Create one — returns data.monitor_id, e.g. "wm_5d1p8s3k7"
curl -X POST "https://www.socialcrawl.dev/v1/web/monitors" \
-H "x-api-key: sc_your_api_key_here" \
-H "content-type: application/json" \
-d '{"url":"https://example.com/pricing","name":"Pricing page","mode":"scrape","cadence_minutes":15}'
# Read what it has found
curl "https://www.socialcrawl.dev/v1/web/monitors/wm_5d1p8s3k7/checks?limit=20" \
-H "x-api-key: sc_your_api_key_here"
# Pause it without losing history
curl -X PATCH "https://www.socialcrawl.dev/v1/web/monitors/wm_5d1p8s3k7" \
-H "x-api-key: sc_your_api_key_here" \
-H "content-type: application/json" \
-d '{"status":"paused"}'
```
* **`POST /v1/web/monitors`** creates one. `mode` is `scrape` (watch a page) or `search` (watch a query, via `query`). Schedule it with `cadence_minutes`, `schedule_text` ("every 15 minutes"), or `schedule_cron` — the last two are mutually exclusive. `retention_days` controls how long checks are kept, `webhook_url` pushes each result, and `judge_enabled` turns on the change assessment against `goal`.
* **`cadence_minutes` is 5-60, or a whole number of hours up to 1440** (120, 180, and so on). Hour cadences are scheduled as cron under the hood. Values between those two bands are rejected.
* **`GET /v1/web/monitors`** lists them (cursor-paginated), **`GET /v1/web/monitors/{monitor_id}`** returns one monitor's settings plus `last_check_at` and `next_check_at`, and **`GET /v1/web/monitors/{monitor_id}/checks`** is the history — the results, not the configuration. That last one is the path a webhook's `run_id` refers back to.
* **`PATCH /v1/web/monitors/{monitor_id}`** changes `status` (pause/resume) or the schedule, keeping the check history. **`DELETE /v1/web/monitors/{monitor_id}`** stops future checks; past checks stay readable.
* **Creating, listing, reading, updating, and deleting a monitor are all free.** You pay when a check runs: the underlying scrape or search cost plus 1 credit of orchestration. A paused monitor bills nothing.
Browser sessions: drive a live page [#browser-sessions-drive-a-live-page]
When content only appears after clicking, typing, or logging in, open a session and send it instructions.
```bash
# 1. Open — returns data.session_id, e.g. "ws_3g7x1v5m2"
curl -X POST "https://www.socialcrawl.dev/v1/web/sessions" \
-H "x-api-key: sc_your_api_key_here" \
-H "content-type: application/json" \
-d '{"url":"https://example.com","ttl_seconds":120}'
# 2. Drive it
curl -X POST "https://www.socialcrawl.dev/v1/web/sessions/ws_3g7x1v5m2/execute" \
-H "x-api-key: sc_your_api_key_here" \
-H "content-type: application/json" \
-d '{"code":"await page.click(\"#load-more\"); console.log(await page.title());","language":"node"}'
# 3. Close it
curl -X DELETE "https://www.socialcrawl.dev/v1/web/sessions/ws_3g7x1v5m2" \
-H "x-api-key: sc_your_api_key_here"
```
* **`POST /v1/web/sessions`** opens a short-lived browser at `url` and returns a `ws_`-prefixed id. `ttl_seconds` sets the lifetime and `activity_ttl_seconds` the idle timeout; `stream_web_view=true` gives you a viewable stream.
* **`POST /v1/web/sessions/{session_id}/execute`** runs code inside it. `language` is `node`, `python`, or `bash`. With `node` you get the Playwright `page`, `browser`, and `chromium` globals and the final expression lands in `result`; `bash` and `python` report through `stdout`. The response carries `success`, `stdout`, `stderr`, `exitCode`, and a `killed` flag. Executing is free — you already paid for the session.
* **`GET /v1/web/sessions`** and **`GET /v1/web/sessions/{session_id}`** show what is still open and when it expires. **`DELETE /v1/web/sessions/{session_id}`** closes it and settles the hold.
* **A session holds 5 credits minimum, and 20 credits per hour of `ttl_seconds`** — a 120-second session holds the 5-credit floor, a two-hour session holds 40. Close it as soon as you are done rather than waiting for expiry; the unused hold refunds on close.
Document parse [#document-parse]
**`POST /v1/web/parse`** takes a `multipart/form-data` upload in the `file` field and returns the same `data.page` shape as a scrape, with `page_count` and `total_page_count` filled in. PDFs and office documents both work. **1cr.** Use it when you hold the file itself; for a document sitting at a URL, `scrape` with `pdf_parse=true` gets there in one call.
```bash
curl -X POST "https://www.socialcrawl.dev/v1/web/parse" \
-H "x-api-key: sc_your_api_key_here" \
-F "file=@report.pdf" \
-F "filename=report.pdf"
```
Endpoints [#endpoints]
Notes [#notes]
* Method-aware surface: the sync data endpoints (`scrape`, `search`, `map`, `extract`) and all reads are `GET`; async jobs, monitors, sessions, and parse use `POST`/`PATCH`/`DELETE` as listed above
* Authentication via `x-api-key` header
* Content responses follow the unified [WebPage schema](/docs/schema/web-page.md); job/monitor/session resources return envelope-wrapped status objects with cursor pagination on lists
* Job ids are `job_…`, monitor ids `wm_…`, session ids `ws_…`. All three are scoped to the API key that created them, so listing only ever shows your own
* `POST /v1/web/crawl` and `POST /v1/web/batch-scrape` honour an `Idempotency-Key` header — resubmitting the same key returns the original job rather than starting a second one
* Variable-cost work (crawl pages, PDF parsing, enhanced proxy, agent runs, session time) is metered honestly: held, settled to actual usage, and refunded when unused