Web Scraping Start Crawl API
Scrape Web Scraping Start Crawl data with one API call. Starts an async crawl job and returns a SocialCrawl job id (job_...). Poll or cancel it via GET/DELETE /v1/web/jobs/{job_id}, or track it with a webhook.
Last updated October 2026Maintained by the SocialCrawl team
Starts an async crawl job that walks a site and scrapes every page it finds, returning a job id to poll.
Use it when you need many pages from one site; check progress with GET jobs/{job_id}.
Searching 68 platforms in parallel
What can you do with the Start Crawl API?
The Start Crawl endpoint gives you structured Web Scraping data with computed fields in a single request. No scraping infrastructure to build or maintain.
Example Request
curl -X POST -H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com","limit":"10"}' \
"https://www.socialcrawl.dev/v1/web/crawl"import requests
response = requests.post(
"https://www.socialcrawl.dev/v1/web/crawl",
json={
'url': 'https://example.com',
'limit': '10',
},
headers={"x-api-key": "YOUR_API_KEY"},
)
data = response.json()const response = await fetch(
"https://www.socialcrawl.dev/v1/web/crawl",
{
method: "POST",
headers: {
"x-api-key": "YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
url: "https://example.com",
limit: "10",
}),
},
);
const data = await response.json();Parameters
| Parameter | Required | Description |
|---|---|---|
| url | Yes | Root URL to crawl. |
| limit | No | Maximum pages to crawl. |
| max_depth | No | How many link hops from the root URL to follow. 1 is the root page and the pages it links to. Leave unset to let `limit` alone bound the crawl. |
| allow_backward_links | No | Crawl the whole domain rather than only URLs under the starting path. Send it when the root URL is a deep page but you want the entire site. |
| allow_external_links | No | Follow links off the domain. Off by default, and easy to make expensive: `limit` is the only thing bounding it. |
| include_paths | No | CSV of URL path patterns to crawl, e.g. '/blog/.*,/docs/.*'. Everything else is skipped, which is the cheapest way to narrow a crawl. |
| exclude_paths | No | CSV of URL path patterns to skip, e.g. '/tag/.*'. Applied after include_paths. |
| webhook_url | No | Optional webhook URL for terminal job updates. |
| formats | No | CSV of output formats for each crawled page, e.g. 'markdown' (default) or 'markdown,html'. |
| goal | No | Optional. What you are looking for on the site, in your own words. SocialCrawl maps the site, scrapes the URLs that likely hold it (always at least two), and lists every skipped URL with p. Without this param the crawl follows links by depth only. |
What does the Web Scraping Start Crawl API return?
This example shows the endpoint's response shape and fields. Values are illustrative; this is not a live capture.
Example response
{
"success": true,
"platform": "web",
"endpoint": "/v1/web/crawl",
"data": {
"job_id": "job_9f3k2n8d1",
"kind": "crawl",
"resource": "crawl",
"status": "queued",
"progress": null,
"invalid_urls": null,
"result_ref": null,
"result": null,
"error": null,
"credits_hold": 10,
"credits_charged": 0,
"created_at": "2026-07-09T09:30:00.000Z",
"updated_at": "2026-07-09T09:30:00.000Z",
"expires_at": "2026-07-16T09:30:00.000Z"
},
"credits_used": 10,
"credits_remaining": 990,
"request_id": "req-a1b2c3d4e5f6",
"cached": false
}Illustrative fixture for the documented response shape. Values are examples, not a production capture.
How does the Web Scraping Start Crawl API work?
Send a GET request with your API key and get back clean, structured JSON in our unified schema. Supported computed fields are populated when the source provides the required inputs.
Method
POST
Response
JSON
How do you scrape social media data in seconds?
The fastest social media scraping API for developers. Scrape profiles, posts, comments, and analytics from 68 platforms covering 10B+ monthly active users.
One schema, every platform
Query 68 platforms with identical response structures. Write your integration once.
Computed fields, not just scraped
When an endpoint supports these metrics and the source provides the required inputs, the normalized record includes engagement_rate, estimated_reach, content_category, and language. Ready to use.
See your data before you code
Visual Data Explorer. Paste any URL, get rich result cards, sortable tables, CSV export.
import requests
response = requests.get(
'https://www.socialcrawl.dev/v1/tiktok/profile',
params={'handle': 'charlidamelio'},
headers={'x-api-key': 'sc_YOUR_API_KEY'}
)
data = response.json(){
"success": true,
"platform": "tiktok",
"data": {
"author": {
"username": "charlidamelio",
"followers": 152400000
},
"engagement": {
"likes": 12400000000,
"engagement_rate": 0.087
},
"metadata": {
"language": "en",
"content_category": "lifestyle"
}
}
}Ready to scrape Web Scraping Start Crawl data?
Get your API key and start pulling Web Scraping data in under 60 seconds.
