Tavily Crawl API
Scrape Tavily Crawl data with one API call. Multi-page crawl starting from a root URL. Returns each crawled page with its extracted content (unlike map, which returns only URLs). Use `instructions` to guide the crawler in natural language — Tavily uses an LLM to follow only the paths matching your intent. Use `select_*` / `exclude_*` filters (comma-separated regex patterns) to constrain scope.
Last updated August 2026Maintained by the SocialCrawl team
Returns pages found by following links from a starting URL, each with its extracted text content.
Use it when you want the content of many pages on a site. Plain-language instructions steer which paths get followed.
Searching 48 platforms in parallel
What can you do with the Crawl API?
The Crawl endpoint gives you structured Tavily data with computed fields in a single request. No scraping infrastructure to build or maintain.
Example Request
curl -H "x-api-key: YOUR_API_KEY" \
"https://www.socialcrawl.dev/v1/tavily/crawl?url=https%3A%2F%2Fdocs.tavily.com"import requests
response = requests.get(
"https://www.socialcrawl.dev/v1/tavily/crawl",
params={
'url': 'https://docs.tavily.com',
},
headers={"x-api-key": "YOUR_API_KEY"},
)
data = response.json()const response = await fetch(
"https://www.socialcrawl.dev/v1/tavily/crawl?url=https%3A%2F%2Fdocs.tavily.com",
{
headers: { "x-api-key": "YOUR_API_KEY" },
},
);
const data = await response.json();Parameters
| Parameter | Required | Description |
|---|---|---|
| url | Yes | Root URL to begin crawling. |
| max_depth | No | Maximum link depth from the root URL. Defaults to 1. |
| max_breadth | No | Maximum number of links followed per level (per page). Defaults to 20. |
| limit | No | Total number of pages the crawler will process before stopping. Defaults to 50. |
| instructions | No | Natural-language instructions for the crawler (e.g. 'Find all product pages with pricing'). |
| select_paths | No | Comma-separated regex patterns — only crawl URLs whose path matches. |
| select_domains | No | Comma-separated regex patterns — only crawl URLs whose domain matches. |
| exclude_paths | No | Comma-separated regex patterns — skip URLs whose path matches. |
| exclude_domains | No | Comma-separated regex patterns — skip URLs whose domain matches. |
| allow_external | No | Whether to follow links to external domains. Defaults to true. |
| extract_depth | No | Per-page extraction strategy. `basic` is faster; `advanced` handles harder pages. (basic | advanced) |
| format | No | Output format for extracted content. `markdown` (default) or `text`. (markdown | text) |
| categories | No | Comma-separated list of category hints to bias the crawl toward. |
What does the Tavily Crawl API return?
Every response follows one unified schema. Here is a real, unmodified response body, so you can see the exact fields you get back before spending a credit.
Example response
{
"success": true,
"platform": "instagram",
"endpoint": "/v1/instagram/engagement",
"data": {
"engagement_rate_percentages": 38.33,
"recent_posts": 12,
"followers": 87608035,
"comments": 528912,
"likes": 33049046,
"recent_posts_explanation": "Statistics based on the last 12 posts",
"id_user": "2278169415",
"username": "mrbeast",
"is_private": false,
"posts_details": [
{
"likes": 5636982,
"comments": 69484,
"taken_at": 1781457954,
"datetime": "2026-06-14 20:25:54",
"hours_since_post": 461,
"time_ago": "19 days ago",
"likes_per_hour": 12228,
"comments_per_hour": 151
},
{
"likes": 20000768,
"comments": 223510,
"taken_at": 1732824650,
"datetime": "2024-11-28 23:10:50",
"hours_since_post": 13971,
"time_ago": "2 years ago",
"likes_per_hour": 1432,
"comments_per_hour": 16
},
{
"likes": 929226,
"comments": 30633,
"taken_at": 1782232475,
"datetime": "2026-06-23 19:34:35",
"hours_since_post": 246,
"time_ago": "10 days ago",
"likes_per_hour": 3777,
"comments_per_hour": 125
},
{
"likes": 487761,
"comments": 22482,
"taken_at": 1781799425,
"datetime": "2026-06-18 19:17:05",
"hours_since_post": 366,
"time_ago": "15 days ago",
"likes_per_hour": 1333,
"comments_per_hour": 61
},
{
"likes": 712265,
"comments": 15716,
"taken_at": 1781366405,
"datetime": "2026-06-13 19:00:05",
"hours_since_post": 487,
"time_ago": "20 days ago",
"likes_per_hour": 1463,
"comments_per_hour": 32
},
{
"likes": 1475116,
"comments": 35386,
"taken_at": 1781277094,
"datetime": "2026-06-12 18:11:34",
"hours_since_post": 512,
"time_ago": "21 days ago",
"likes_per_hour": 2881,
"comments_per_hour": 69
},
{
"likes": 1108220,
"comments": 26632,
"taken_at": 1780160249,
"datetime": "2026-05-30 19:57:29",
"hours_since_post": 822,
"time_ago": "1 months ago",
"likes_per_hour": 1348,
"comments_per_hour": 32
},
{
"likes": 542948,
"comments": 28476,
"taken_at": 1779375582,
"datetime": "2026-05-21 17:59:42",
"hours_since_post": 1040,
"time_ago": "1 months ago",
"likes_per_hour": 522,
"comments_per_hour": 27
},
{
"likes": 698514,
"comments": 24401,
"taken_at": 1779120014,
"datetime": "2026-05-18 19:00:14",
"hours_since_post": 1111,
"time_ago": "2 months ago",
"likes_per_hour": 629,
"comments_per_hour": 22
},
{
"likes": 468000,
"comments": 13548,
"taken_at": 1778947209,
"datetime": "2026-05-16 19:00:09",
"hours_since_post": 1159,
"time_ago": "2 months ago",
"likes_per_hour": 404,
"comments_per_hour": 12
},
{
"likes": 526594,
"comments": 24411,
"taken_at": 1777737719,
"datetime": "2026-05-02 19:01:59",
"hours_since_post": 1495,
"time_ago": "2 months ago",
"likes_per_hour": 352,
"comments_per_hour": 16
},
{
"likes": 462652,
"comments": 14233,
"taken_at": 1777580305,
"datetime": "2026-04-30 23:18:25",
"hours_since_post": 1538,
"time_ago": "2 months ago",
"likes_per_hour": 301,
"comments_per_hour": 9
}
]
},
"credits_used": 5,
"credits_remaining": 9999,
"request_id": "req-8Kq2ZmR4vT9xLb3P",
"cached": false
}Example captured from the Instagram API. Every SocialCrawl endpoint returns this same unified schema, so your Tavily Crawl response has the same fields.
How does the Tavily Crawl API work?
Send a GET request with your API key and get back clean, structured JSON in our unified schema. Supported computed fields are populated when the source provides the required inputs.
Method
GET
Response
JSON
How do you scrape social media data in seconds?
The fastest social media scraping API for developers. Scrape profiles, posts, comments, and analytics from 48 platforms covering 10B+ monthly active users.
One schema, every platform
Query 48 platforms with identical response structures. Write your integration once.
Computed fields, not just scraped
When an endpoint supports these metrics and the source provides the required inputs, the normalized record includes engagement_rate, estimated_reach, content_category, and language — ready to use.
See your data before you code
Visual Data Explorer — paste any URL, get rich result cards, sortable tables, CSV export.
import requests
response = requests.get(
'https://www.socialcrawl.dev/v1/tiktok/profile',
params={'handle': 'charlidamelio'},
headers={'x-api-key': 'sc_YOUR_API_KEY'}
)
data = response.json(){
"success": true,
"platform": "tiktok",
"data": {
"author": {
"username": "charlidamelio",
"followers": 152400000
},
"engagement": {
"likes": 12400000000,
"engagement_rate": 0.087
},
"metadata": {
"language": "en",
"content_category": "lifestyle"
}
}
}Have a question? We got answers
Find answers to frequently asked questions about SocialCrawl's API, pricing, and capabilities.
Contact usHow do I crawl a website with the Tavily API?
How does LLM-driven path selection work?
What parameters control the crawl?
When should I use crawl instead of map or search?
How much does the Tavily Crawl API cost?
Ask AI about SocialCrawl
Ready to scrape Tavily Crawl data?
Get your API key and start pulling Tavily data in under 60 seconds.
