Web Scraping Parse Document API
Scrape Web Scraping Parse Document data with one API call. Uploads a document through multipart/form-data and returns parsed web-style content.
Last updated August 2026Maintained by the SocialCrawl team
Returns the text of a document you upload, such as a PDF, as clean markdown along with its page count.
Use it when you have the file itself; to read a normal web page use scrape instead.
Searching 50 platforms in parallel
What can you do with the Parse Document API?
The Parse Document endpoint gives you structured Web Scraping data with computed fields in a single request. No scraping infrastructure to build or maintain.
Example Request
curl -X POST -H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"file":"document.pdf"}' \
"https://www.socialcrawl.dev/v1/web/parse"import requests
response = requests.post(
"https://www.socialcrawl.dev/v1/web/parse",
json={
'file': 'document.pdf',
},
headers={"x-api-key": "YOUR_API_KEY"},
)
data = response.json()const response = await fetch(
"https://www.socialcrawl.dev/v1/web/parse",
{
method: "POST",
headers: {
"x-api-key": "YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
file: "document.pdf",
}),
},
);
const data = await response.json();Parameters
| Parameter | Required | Description |
|---|---|---|
| file | Yes | Multipart file field. |
| filename | No | |
| mime_type | No | Optional MIME type override. |
| url | No |
What does the Web Scraping Parse Document API return?
This example shows the endpoint's response shape and fields. Values are illustrative; this is not a live capture.
Example response
{
"success": true,
"platform": "web",
"endpoint": "/v1/web/parse",
"data": {
"page": {
"url": null,
"final_url": null,
"status_code": null,
"scrape_id": "sc_6k9d3f8b4",
"fetched_at": "2026-07-09T09:30:02.000Z",
"title": "document.pdf",
"description": null,
"source_type": null,
"content": {
"markdown": "# Quarterly report\n\nRevenue grew 12%...",
"html": null,
"raw_html": null,
"summary": null
},
"media": {
"screenshot_url": null,
"audio_url": null,
"video_url": null
},
"extraction": null,
"answer": null,
"highlights": null,
"change_tracking": null,
"page_count": 3,
"total_page_count": 3,
"fetch": {
"cache_state": null,
"cached_at": null,
"proxy_tier": null
}
}
},
"credits_used": 1,
"credits_remaining": 999,
"request_id": "req-a1b2c3d4e5f6",
"cached": false
}Illustrative fixture for the documented response shape. Values are examples, not a production capture.
How does the Web Scraping Parse Document API work?
Send a GET request with your API key and get back clean, structured JSON in our unified schema. Supported computed fields are populated when the source provides the required inputs.
Method
POST
Response
JSON
How do you scrape social media data in seconds?
The fastest social media scraping API for developers. Scrape profiles, posts, comments, and analytics from 50 platforms covering 10B+ monthly active users.
One schema, every platform
Query 50 platforms with identical response structures. Write your integration once.
Computed fields, not just scraped
When an endpoint supports these metrics and the source provides the required inputs, the normalized record includes engagement_rate, estimated_reach, content_category, and language — ready to use.
See your data before you code
Visual Data Explorer — paste any URL, get rich result cards, sortable tables, CSV export.
import requests
response = requests.get(
'https://www.socialcrawl.dev/v1/tiktok/profile',
params={'handle': 'charlidamelio'},
headers={'x-api-key': 'sc_YOUR_API_KEY'}
)
data = response.json(){
"success": true,
"platform": "tiktok",
"data": {
"author": {
"username": "charlidamelio",
"followers": 152400000
},
"engagement": {
"likes": 12400000000,
"engagement_rate": 0.087
},
"metadata": {
"language": "en",
"content_category": "lifestyle"
}
}
}Ready to scrape Web Scraping Parse Document data?
Get your API key and start pulling Web Scraping data in under 60 seconds.
