Web Scraping
Scrape, search, map, extract, crawl, and monitor any public web page. Clean markdown and structured data in the unified WebPage schema
Scrape, search, map, extract, crawl, and monitor any public web page. Content-bearing responses return the unified WebPage schema: clean markdown or HTML, metadata, structured extraction, and fetch provenance under data.page, while search and map return a WebPageList under data.items[].
This is the only platform on the API that sells stateful work: long-running jobs you poll, monitors that keep checking on a schedule, and browser sessions you drive step by step. Everything else here is one request, one answer.
Base URL: /v1/web/...
The crawl hold is taken against limit, not against how big the site turns
out to be. Submitting limit=1000 reserves 1000 credits the moment the job is
accepted, even if the site only has 30 pages. Run GET /v1/web/map first and
set limit from what you find.
Quickstart
1. Scrape a page to markdown
curl "https://www.socialcrawl.dev/v1/web/scrape?url=https://example.com" \
-H "x-api-key: YOUR_API_KEY"The page content arrives at data.page.content.markdown. Add formats=markdown,screenshot for a screenshot URL under data.page.media.
2. Pull typed fields instead of prose
curl "https://www.socialcrawl.dev/v1/web/extract?url=https://example.com&prompt=the%20page%20title%20and%20any%20pricing%20mentioned" \
-H "x-api-key: YOUR_API_KEY"extract needs either a schema (JSON) or a prompt. Sending neither is a 400 before billing. Typed fields land under data.page.extraction.
Synchronous endpoints
These four answer immediately and cover most integrations.
| Endpoint | Credits | What it returns | Key parameters |
|---|---|---|---|
GET /v1/web/scrape | 1, or 5 for heavy work | One known URL as clean markdown or HTML, with the resolved URL, status code, and fetch metadata | url, formats, only_main_content, wait_for, mobile, timeout, max_age, location_country, include_tags, exclude_tags, proxy, pdf_parse, block_ads |
GET /v1/web/search | 2 per 10 results | Ranked web, news, and image results as one normalised list, each row tagged with source_type | query, sources, categories, limit, country, location, time_range, sort_by_date, include_domains, exclude_domains, include_content |
GET /v1/web/map | 1 | Every URL a site exposes, without fetching any page bodies | url, search, limit, sitemap, include_subdomains, ignore_query_parameters, fresh |
GET /v1/web/extract | 5 flat | Structured fields from one page, shaped by your schema or prompt | url, plus schema or prompt, only_main_content, timeout, max_age, proxy |
scrape steps from 1 to 5 credits when you ask for work the cheap path cannot do: proxy=auto, proxy=enhanced, or pdf_parse=true. The data.page.fetch.proxy_tier field tells you which path actually ran.
search is 2 credits per 10 results, so limit=50 is 10 credits. Adding include_content=true scrapes each result inline and adds 1 credit per result on top, which is how a limit=50 content search reaches 60 credits. Filter with include_domains / exclude_domains, country, time_range, and categories before you widen limit.
The usual chain is map to find the pages, then scrape or extract on the ones that matter. Reach for crawl only when you want the site walked for you.
# 1. What pages does this site have?
curl "https://www.socialcrawl.dev/v1/web/map?url=https://example.com&search=pricing" \
-H "x-api-key: YOUR_API_KEY"
# 2. Pull the one that matters
curl "https://www.socialcrawl.dev/v1/web/scrape?url=https://example.com/pricing&formats=markdown" \
-H "x-api-key: YOUR_API_KEY"Async jobs: submit, poll, settle
crawl, batch-scrape, and agent are too slow to answer inline, so they return a job id and you poll one unified surface. All three ids are prefixed job_, and all three read through the same four routes. There is no per-kind status path.
| Endpoint | Credits | What it does | Key parameters |
|---|---|---|---|
POST /v1/web/crawl | 1 per page crawled | Walks a site from a root URL and scrapes what it finds | url, limit, max_depth, allow_backward_links, allow_external_links, include_paths, exclude_paths, formats, webhook_url |
POST /v1/web/batch-scrape | 1 per URL submitted | Scrapes a urls array you already know | urls, ignore_invalid_urls, formats, only_main_content, proxy, webhook_url |
POST /v1/web/agent | Holds 25, floor 5 | Drives a browser toward an answer from a plain-language prompt | url, prompt, model |
GET /v1/web/jobs | 0 | Lists this key's jobs, cursor-paginated | limit, cursor |
GET /v1/web/jobs/{job_id} | 0 | Status, progress, credits held and charged, and the result once terminal | job_id |
DELETE /v1/web/jobs/{job_id} | 0 | Cancels the job and refunds the unspent hold | job_id |
limit caps pages on a crawl and max_depth caps how far from the root it goes, with include_paths / exclude_paths keeping it in the section you care about. ignore_invalid_urls keeps one bad entry from rejecting a whole batch submission. Use agent when the data is behind clicking and navigating rather than at a URL you can name.
# 1. Submit, returns data.job_id, e.g. "job_9f3k2n8d1"
curl -X POST "https://www.socialcrawl.dev/v1/web/crawl" \
-H "x-api-key: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{"url":"https://example.com","limit":25,"max_depth":2}'
# 2. Poll until data.status is terminal
curl "https://www.socialcrawl.dev/v1/web/jobs/job_9f3k2n8d1" \
-H "x-api-key: YOUR_API_KEY"
# 3. Lost the id? List everything this key has running
curl "https://www.socialcrawl.dev/v1/web/jobs?limit=20" \
-H "x-api-key: YOUR_API_KEY"
# 4. Stop it early and get the unspent hold back
curl -X DELETE "https://www.socialcrawl.dev/v1/web/jobs/job_9f3k2n8d1" \
-H "x-api-key: YOUR_API_KEY"How async jobs are billed
Credits are held, then settled. A crawl holds 1 credit per page of limit (so limit=25 holds 25) and charges only for the pages actually fetched. The default limit is 10, and the maximum is 10,000. A batch scrape holds one credit per URL you submit. An agent job holds a flat 25 and settles against the work it did, with a 5-credit floor. Whatever is held and not spent comes back automatically, including on cancel.
Reading and cancelling a job are free. POST /v1/web/crawl and POST /v1/web/batch-scrape also accept a webhook_url if you would rather be told than poll. See Webhooks for the payload.
Per-page failures. Append /errors to a job read (GET /v1/web/jobs/{job_id}, then the same path with /errors on the end) to list the individual pages that failed inside an otherwise successful job. That sub-route is a convenience mount rather than a registry entry, so it does not appear in the roster below or in the machine-readable contracts. It works, but treat it as unversioned: do not generate an SDK method for it.
Monitors: watch a page on a schedule
A monitor re-runs a scrape or a search on a cadence and records every result, so you can ask what changed rather than diffing it yourself.
| Endpoint | Credits | What it does | Key parameters |
|---|---|---|---|
POST /v1/web/monitors | 0 to create | Creates a monitor | url, name, mode, cadence_minutes, schedule_text, schedule_cron, timezone, query, goal, retention_days, judge_enabled, webhook_url |
GET /v1/web/monitors | 0 | Lists them, cursor-paginated | cursor, limit |
GET /v1/web/monitors/{monitor_id} | 0 | One monitor's settings plus last_check_at and next_check_at | monitor_id |
GET /v1/web/monitors/{monitor_id}/checks | 0 | The check history: the results, not the configuration | monitor_id, limit |
PATCH /v1/web/monitors/{monitor_id} | 0 | Changes status (pause/resume) or the schedule, keeping the check history | monitor_id, status, cadence_minutes, schedule_text, schedule_cron, timezone |
DELETE /v1/web/monitors/{monitor_id} | 0 | Stops future checks; past checks stay readable | monitor_id |
# Create one, returns data.monitor_id, e.g. "wm_5d1p8s3k7"
curl -X POST "https://www.socialcrawl.dev/v1/web/monitors" \
-H "x-api-key: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{"url":"https://example.com/pricing","name":"Pricing page","mode":"scrape","cadence_minutes":15}'
# Read what it has found
curl "https://www.socialcrawl.dev/v1/web/monitors/wm_5d1p8s3k7/checks?limit=20" \
-H "x-api-key: YOUR_API_KEY"
# Pause it without losing history
curl -X PATCH "https://www.socialcrawl.dev/v1/web/monitors/wm_5d1p8s3k7" \
-H "x-api-key: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{"status":"paused"}'modeisscrape(watch a page) orsearch(watch a query, viaquery). Schedule it withcadence_minutes,schedule_text("every 15 minutes"), orschedule_cron. The last two are mutually exclusive.retention_dayscontrols how long checks are kept,webhook_urlpushes each result, andjudge_enabledturns on the change assessment againstgoal.cadence_minutesis 5-60, or a whole number of hours up to 1440 (120, 180, and so on). Hour cadences are scheduled as cron under the hood. Values between those two bands are rejected.- Creating, listing, reading, updating, and deleting a monitor are all free. You pay when a check runs: the underlying scrape or search cost plus 1 credit of orchestration. A paused monitor bills nothing.
GET /v1/web/monitors/{monitor_id}/checksis the path a webhook'srun_idrefers back to.
Browser sessions: drive a live page
When content only appears after clicking, typing, or logging in, open a session and send it instructions.
| Endpoint | Credits | What it does | Key parameters |
|---|---|---|---|
POST /v1/web/sessions | 20 per browser-hour, min 5 | Opens a short-lived browser at url and returns a ws_-prefixed id | url, ttl_seconds, activity_ttl_seconds, stream_web_view |
POST /v1/web/sessions/{session_id}/execute | 0 | Runs code inside the session | session_id, code, language, timeout |
GET /v1/web/sessions | 0 | What is still open, cursor-paginated | limit, cursor |
GET /v1/web/sessions/{session_id} | 0 | One session and when it expires | session_id |
DELETE /v1/web/sessions/{session_id} | 0 | Closes it and settles the hold | session_id |
# 1. Open, returns data.session_id, e.g. "ws_3g7x1v5m2"
curl -X POST "https://www.socialcrawl.dev/v1/web/sessions" \
-H "x-api-key: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{"url":"https://example.com","ttl_seconds":120}'
# 2. Drive it
curl -X POST "https://www.socialcrawl.dev/v1/web/sessions/ws_3g7x1v5m2/execute" \
-H "x-api-key: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{"code":"await page.click(\"#load-more\"); console.log(await page.title());","language":"node"}'
# 3. Close it
curl -X DELETE "https://www.socialcrawl.dev/v1/web/sessions/ws_3g7x1v5m2" \
-H "x-api-key: YOUR_API_KEY"- A session holds 20 credits per browser-hour of
ttl_seconds, with a minimum of 5.ttl_secondsdefaults to 60 and clamps to 30-3600, so a two-minute session holds the 5-credit floor and the one-hour maximum holds 20. Close it as soon as you are done rather than waiting for expiry. The unused hold refunds on close. ttl_secondssets the lifetime andactivity_ttl_secondsthe idle timeout.stream_web_view=truegives you a viewable stream.languageonexecuteisnode,python, orbash. Withnodeyou get the Playwrightpage,browser, andchromiumglobals and the final expression lands inresult.bashandpythonreport throughstdout. The response carriessuccess,stdout,stderr,exitCode, and akilledflag. Executing is free: you already paid for the session.
Document parse
| Endpoint | Credits | What it returns | Key parameters |
|---|---|---|---|
POST /v1/web/parse | 1 | The same data.page shape as a scrape, with page_count and total_page_count set | file (multipart), filename, mime_type, url |
PDFs and office documents both work. Use it when you hold the file itself. For a document sitting at a URL, scrape with pdf_parse=true gets there in one call.
curl -X POST "https://www.socialcrawl.dev/v1/web/parse" \
-H "x-api-key: YOUR_API_KEY" \
-F "file=@report.pdf" \
-F "filename=report.pdf"All endpoints
22 endpoints available.
| Endpoint | Path | Credit Tier |
|---|---|---|
| Start an async batch scrape | /v1/web/batch-scrape | standard (1cr) |
| Start an async web crawl | /v1/web/crawl | standard (1-10000cr)metered |
| List async web jobs | /v1/web/jobs | standard (0cr) |
| Get an async web job | /v1/web/jobs/{job_id} | standard (0cr) |
| Cancel an async web job | /v1/web/jobs/{job_id} | standard (0cr) |
| Map URLs on a site | /v1/web/map | standard (1cr) |
| Create a web monitor | /v1/web/monitors | standard (0cr) |
| List web monitors | /v1/web/monitors | standard (0cr) |
| Get a web monitor | /v1/web/monitors/{monitor_id} | standard (0cr) |
| Update a web monitor | /v1/web/monitors/{monitor_id} | standard (0cr) |
| Delete a web monitor | /v1/web/monitors/{monitor_id} | standard (0cr) |
| List web monitor checks | /v1/web/monitors/{monitor_id}/checks | standard (0cr) |
| Parse an uploaded document | /v1/web/parse | standard (1cr) |
| Scrape a web page | /v1/web/scrape | standard (1-5cr)metered |
| Search the web | /v1/web/search | standard (2-120cr)metered |
| List interactive web sessions | /v1/web/sessions | standard (0cr) |
| Get an interactive web session | /v1/web/sessions/{session_id} | standard (0cr) |
| Close an interactive web session | /v1/web/sessions/{session_id} | standard (0cr) |
| Execute an interaction in a web session | /v1/web/sessions/{session_id}/execute | standard (0cr) |
| Extract structured data from a web page | /v1/web/extract | advanced (5cr)metered |
| Create an interactive web session | /v1/web/sessions | advanced (5cr) |
| Start an async web agent job | /v1/web/agent | premium (25cr) |
Platform notes
- Method-aware surface. The sync data endpoints (
scrape,search,map,extract) and all reads areGET. Async jobs, monitors, sessions, and parse usePOST/PATCH/DELETEas listed above. - Variable-cost work is metered honestly. Crawl pages, PDF parsing, enhanced proxy, agent runs, and session time are held, settled to actual usage, and refunded when unused.
- Some sites cannot be scraped. When
/v1/web/scrapecannot fetch a site at all, it returns404 RESOURCE_NOT_FOUNDwitherror.details.reason: "site_not_supported"andretry_will_succeed: false, and charges nothing. Retrying returns the same answer. When the URL is on a platform SocialCrawl reads directly,error.details.suggestionnames the endpoint to use instead, for example/v1/instagram/profileor/v1/reddit/post. - A markdown scrape can be served by a second extractor. When the primary scraper cannot fetch a page, a markdown-only request is retried on a second extractor at the same 1-credit price. That extractor returns the whole page, so
only_main_contentdoes not apply to it. - Job ids are
job_…, monitor idswm_…, session idsws_…. All three are scoped to the API key that created them, so listing only ever shows your own. POST /v1/web/crawlandPOST /v1/web/batch-scrapehonour anIdempotency-Keyheader. Resubmitting the same key returns the original job rather than starting a second one.- Content responses follow the unified WebPage schema. Job, monitor, and session lists return the top-level
paginationblock. Send its opaquenext_cursorback verbatim ascursor; the public token preserves the database keyset, and a malformed or stale token safely restarts at the first page.
