SocialCrawl

Web Scraping

Scrape, search, map, extract, crawl, and monitor any public web page. Clean markdown and structured data in the unified WebPage schema

Scrape, search, map, extract, crawl, and monitor any public web page. Content-bearing responses return the unified WebPage schema: clean markdown or HTML, metadata, structured extraction, and fetch provenance under data.page, while search and map return a WebPageList under data.items[].

This is the only platform on the API that sells stateful work: long-running jobs you poll, monitors that keep checking on a schedule, and browser sessions you drive step by step. Everything else here is one request, one answer.

Base URL: /v1/web/...

The crawl hold is taken against limit, not against how big the site turns out to be. Submitting limit=1000 reserves 1000 credits the moment the job is accepted, even if the site only has 30 pages. Run GET /v1/web/map first and set limit from what you find.

Quickstart

1. Scrape a page to markdown

cURL
curl "https://www.socialcrawl.dev/v1/web/scrape?url=https://example.com" \
  -H "x-api-key: YOUR_API_KEY"

The page content arrives at data.page.content.markdown. Add formats=markdown,screenshot for a screenshot URL under data.page.media.

2. Pull typed fields instead of prose

cURL
curl "https://www.socialcrawl.dev/v1/web/extract?url=https://example.com&prompt=the%20page%20title%20and%20any%20pricing%20mentioned" \
  -H "x-api-key: YOUR_API_KEY"

extract needs either a schema (JSON) or a prompt. Sending neither is a 400 before billing. Typed fields land under data.page.extraction.

Synchronous endpoints

These four answer immediately and cover most integrations.

EndpointCreditsWhat it returnsKey parameters
GET /v1/web/scrape1, or 5 for heavy workOne known URL as clean markdown or HTML, with the resolved URL, status code, and fetch metadataurl, formats, only_main_content, wait_for, mobile, timeout, max_age, location_country, include_tags, exclude_tags, proxy, pdf_parse, block_ads
GET /v1/web/search2 per 10 resultsRanked web, news, and image results as one normalised list, each row tagged with source_typequery, sources, categories, limit, country, location, time_range, sort_by_date, include_domains, exclude_domains, include_content
GET /v1/web/map1Every URL a site exposes, without fetching any page bodiesurl, search, limit, sitemap, include_subdomains, ignore_query_parameters, fresh
GET /v1/web/extract5 flatStructured fields from one page, shaped by your schema or prompturl, plus schema or prompt, only_main_content, timeout, max_age, proxy

scrape steps from 1 to 5 credits when you ask for work the cheap path cannot do: proxy=auto, proxy=enhanced, or pdf_parse=true. The data.page.fetch.proxy_tier field tells you which path actually ran.

search is 2 credits per 10 results, so limit=50 is 10 credits. Adding include_content=true scrapes each result inline and adds 1 credit per result on top, which is how a limit=50 content search reaches 60 credits. Filter with include_domains / exclude_domains, country, time_range, and categories before you widen limit.

The usual chain is map to find the pages, then scrape or extract on the ones that matter. Reach for crawl only when you want the site walked for you.

cURL
# 1. What pages does this site have?
curl "https://www.socialcrawl.dev/v1/web/map?url=https://example.com&search=pricing" \
  -H "x-api-key: YOUR_API_KEY"

# 2. Pull the one that matters
curl "https://www.socialcrawl.dev/v1/web/scrape?url=https://example.com/pricing&formats=markdown" \
  -H "x-api-key: YOUR_API_KEY"

Async jobs: submit, poll, settle

crawl, batch-scrape, and agent are too slow to answer inline, so they return a job id and you poll one unified surface. All three ids are prefixed job_, and all three read through the same four routes. There is no per-kind status path.

EndpointCreditsWhat it doesKey parameters
POST /v1/web/crawl1 per page crawledWalks a site from a root URL and scrapes what it findsurl, limit, max_depth, allow_backward_links, allow_external_links, include_paths, exclude_paths, formats, webhook_url
POST /v1/web/batch-scrape1 per URL submittedScrapes a urls array you already knowurls, ignore_invalid_urls, formats, only_main_content, proxy, webhook_url
POST /v1/web/agentHolds 25, floor 5Drives a browser toward an answer from a plain-language prompturl, prompt, model
GET /v1/web/jobs0Lists this key's jobs, cursor-paginatedlimit, cursor
GET /v1/web/jobs/{job_id}0Status, progress, credits held and charged, and the result once terminaljob_id
DELETE /v1/web/jobs/{job_id}0Cancels the job and refunds the unspent holdjob_id

limit caps pages on a crawl and max_depth caps how far from the root it goes, with include_paths / exclude_paths keeping it in the section you care about. ignore_invalid_urls keeps one bad entry from rejecting a whole batch submission. Use agent when the data is behind clicking and navigating rather than at a URL you can name.

cURL
# 1. Submit, returns data.job_id, e.g. "job_9f3k2n8d1"
curl -X POST "https://www.socialcrawl.dev/v1/web/crawl" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{"url":"https://example.com","limit":25,"max_depth":2}'

# 2. Poll until data.status is terminal
curl "https://www.socialcrawl.dev/v1/web/jobs/job_9f3k2n8d1" \
  -H "x-api-key: YOUR_API_KEY"

# 3. Lost the id? List everything this key has running
curl "https://www.socialcrawl.dev/v1/web/jobs?limit=20" \
  -H "x-api-key: YOUR_API_KEY"

# 4. Stop it early and get the unspent hold back
curl -X DELETE "https://www.socialcrawl.dev/v1/web/jobs/job_9f3k2n8d1" \
  -H "x-api-key: YOUR_API_KEY"

How async jobs are billed

Credits are held, then settled. A crawl holds 1 credit per page of limit (so limit=25 holds 25) and charges only for the pages actually fetched. The default limit is 10, and the maximum is 10,000. A batch scrape holds one credit per URL you submit. An agent job holds a flat 25 and settles against the work it did, with a 5-credit floor. Whatever is held and not spent comes back automatically, including on cancel.

Reading and cancelling a job are free. POST /v1/web/crawl and POST /v1/web/batch-scrape also accept a webhook_url if you would rather be told than poll. See Webhooks for the payload.

Per-page failures. Append /errors to a job read (GET /v1/web/jobs/{job_id}, then the same path with /errors on the end) to list the individual pages that failed inside an otherwise successful job. That sub-route is a convenience mount rather than a registry entry, so it does not appear in the roster below or in the machine-readable contracts. It works, but treat it as unversioned: do not generate an SDK method for it.

Monitors: watch a page on a schedule

A monitor re-runs a scrape or a search on a cadence and records every result, so you can ask what changed rather than diffing it yourself.

EndpointCreditsWhat it doesKey parameters
POST /v1/web/monitors0 to createCreates a monitorurl, name, mode, cadence_minutes, schedule_text, schedule_cron, timezone, query, goal, retention_days, judge_enabled, webhook_url
GET /v1/web/monitors0Lists them, cursor-paginatedcursor, limit
GET /v1/web/monitors/{monitor_id}0One monitor's settings plus last_check_at and next_check_atmonitor_id
GET /v1/web/monitors/{monitor_id}/checks0The check history: the results, not the configurationmonitor_id, limit
PATCH /v1/web/monitors/{monitor_id}0Changes status (pause/resume) or the schedule, keeping the check historymonitor_id, status, cadence_minutes, schedule_text, schedule_cron, timezone
DELETE /v1/web/monitors/{monitor_id}0Stops future checks; past checks stay readablemonitor_id
cURL
# Create one, returns data.monitor_id, e.g. "wm_5d1p8s3k7"
curl -X POST "https://www.socialcrawl.dev/v1/web/monitors" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{"url":"https://example.com/pricing","name":"Pricing page","mode":"scrape","cadence_minutes":15}'

# Read what it has found
curl "https://www.socialcrawl.dev/v1/web/monitors/wm_5d1p8s3k7/checks?limit=20" \
  -H "x-api-key: YOUR_API_KEY"

# Pause it without losing history
curl -X PATCH "https://www.socialcrawl.dev/v1/web/monitors/wm_5d1p8s3k7" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{"status":"paused"}'
  • mode is scrape (watch a page) or search (watch a query, via query). Schedule it with cadence_minutes, schedule_text ("every 15 minutes"), or schedule_cron. The last two are mutually exclusive. retention_days controls how long checks are kept, webhook_url pushes each result, and judge_enabled turns on the change assessment against goal.
  • cadence_minutes is 5-60, or a whole number of hours up to 1440 (120, 180, and so on). Hour cadences are scheduled as cron under the hood. Values between those two bands are rejected.
  • Creating, listing, reading, updating, and deleting a monitor are all free. You pay when a check runs: the underlying scrape or search cost plus 1 credit of orchestration. A paused monitor bills nothing.
  • GET /v1/web/monitors/{monitor_id}/checks is the path a webhook's run_id refers back to.

Browser sessions: drive a live page

When content only appears after clicking, typing, or logging in, open a session and send it instructions.

EndpointCreditsWhat it doesKey parameters
POST /v1/web/sessions20 per browser-hour, min 5Opens a short-lived browser at url and returns a ws_-prefixed idurl, ttl_seconds, activity_ttl_seconds, stream_web_view
POST /v1/web/sessions/{session_id}/execute0Runs code inside the sessionsession_id, code, language, timeout
GET /v1/web/sessions0What is still open, cursor-paginatedlimit, cursor
GET /v1/web/sessions/{session_id}0One session and when it expiressession_id
DELETE /v1/web/sessions/{session_id}0Closes it and settles the holdsession_id
cURL
# 1. Open, returns data.session_id, e.g. "ws_3g7x1v5m2"
curl -X POST "https://www.socialcrawl.dev/v1/web/sessions" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{"url":"https://example.com","ttl_seconds":120}'

# 2. Drive it
curl -X POST "https://www.socialcrawl.dev/v1/web/sessions/ws_3g7x1v5m2/execute" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{"code":"await page.click(\"#load-more\"); console.log(await page.title());","language":"node"}'

# 3. Close it
curl -X DELETE "https://www.socialcrawl.dev/v1/web/sessions/ws_3g7x1v5m2" \
  -H "x-api-key: YOUR_API_KEY"
  • A session holds 20 credits per browser-hour of ttl_seconds, with a minimum of 5. ttl_seconds defaults to 60 and clamps to 30-3600, so a two-minute session holds the 5-credit floor and the one-hour maximum holds 20. Close it as soon as you are done rather than waiting for expiry. The unused hold refunds on close.
  • ttl_seconds sets the lifetime and activity_ttl_seconds the idle timeout. stream_web_view=true gives you a viewable stream.
  • language on execute is node, python, or bash. With node you get the Playwright page, browser, and chromium globals and the final expression lands in result. bash and python report through stdout. The response carries success, stdout, stderr, exitCode, and a killed flag. Executing is free: you already paid for the session.

Document parse

EndpointCreditsWhat it returnsKey parameters
POST /v1/web/parse1The same data.page shape as a scrape, with page_count and total_page_count setfile (multipart), filename, mime_type, url

PDFs and office documents both work. Use it when you hold the file itself. For a document sitting at a URL, scrape with pdf_parse=true gets there in one call.

cURL
curl -X POST "https://www.socialcrawl.dev/v1/web/parse" \
  -H "x-api-key: YOUR_API_KEY" \
  -F "file=@report.pdf" \
  -F "filename=report.pdf"

All endpoints

22 endpoints available.

EndpointPathCredit Tier
Start an async batch scrape/v1/web/batch-scrapestandard (1cr)
Start an async web crawl/v1/web/crawlstandard (1-10000cr)metered
List async web jobs/v1/web/jobsstandard (0cr)
Get an async web job/v1/web/jobs/{job_id}standard (0cr)
Cancel an async web job/v1/web/jobs/{job_id}standard (0cr)
Map URLs on a site/v1/web/mapstandard (1cr)
Create a web monitor/v1/web/monitorsstandard (0cr)
List web monitors/v1/web/monitorsstandard (0cr)
Get a web monitor/v1/web/monitors/{monitor_id}standard (0cr)
Update a web monitor/v1/web/monitors/{monitor_id}standard (0cr)
Delete a web monitor/v1/web/monitors/{monitor_id}standard (0cr)
List web monitor checks/v1/web/monitors/{monitor_id}/checksstandard (0cr)
Parse an uploaded document/v1/web/parsestandard (1cr)
Scrape a web page/v1/web/scrapestandard (1-5cr)metered
Search the web/v1/web/searchstandard (2-120cr)metered
List interactive web sessions/v1/web/sessionsstandard (0cr)
Get an interactive web session/v1/web/sessions/{session_id}standard (0cr)
Close an interactive web session/v1/web/sessions/{session_id}standard (0cr)
Execute an interaction in a web session/v1/web/sessions/{session_id}/executestandard (0cr)
Extract structured data from a web page/v1/web/extractadvanced (5cr)metered
Create an interactive web session/v1/web/sessionsadvanced (5cr)
Start an async web agent job/v1/web/agentpremium (25cr)

Platform notes

  • Method-aware surface. The sync data endpoints (scrape, search, map, extract) and all reads are GET. Async jobs, monitors, sessions, and parse use POST/PATCH/DELETE as listed above.
  • Variable-cost work is metered honestly. Crawl pages, PDF parsing, enhanced proxy, agent runs, and session time are held, settled to actual usage, and refunded when unused.
  • Some sites cannot be scraped. When /v1/web/scrape cannot fetch a site at all, it returns 404 RESOURCE_NOT_FOUND with error.details.reason: "site_not_supported" and retry_will_succeed: false, and charges nothing. Retrying returns the same answer. When the URL is on a platform SocialCrawl reads directly, error.details.suggestion names the endpoint to use instead, for example /v1/instagram/profile or /v1/reddit/post.
  • A markdown scrape can be served by a second extractor. When the primary scraper cannot fetch a page, a markdown-only request is retried on a second extractor at the same 1-credit price. That extractor returns the whole page, so only_main_content does not apply to it.
  • Job ids are job_…, monitor ids wm_…, session ids ws_…. All three are scoped to the API key that created them, so listing only ever shows your own.
  • POST /v1/web/crawl and POST /v1/web/batch-scrape honour an Idempotency-Key header. Resubmitting the same key returns the original job rather than starting a second one.
  • Content responses follow the unified WebPage schema. Job, monitor, and session lists return the top-level pagination block. Send its opaque next_cursor back verbatim as cursor; the public token preserves the database keyset, and a malformed or stale token safely restarts at the first page.

Next steps