SocialCrawl

Response Schema

The unified SocialCrawl response envelope, headers, and partial-data warnings

Every SocialCrawl response uses the same envelope, on success and on error. Write one parser and it handles every platform.

Successful response

Response
{
  "success": true,
  "platform": "tiktok",
  "endpoint": "/v1/tiktok/profile",
  "data": {
    "author": { "username": "charlidamelio", "followers": 156800000 },
    "computed": {
      "engagement_rate": null,
      "language": null,
      "content_category": null,
      "estimated_reach": null
    }
  },
  "credits_used": 1,
  "credits_remaining": 4999,
  "request_id": "req-XXXXX",
  "cached": false
}

The computed leaves are null here because this abridged author carries no likes_count and no bio, and because estimated_reach is always null on a profile. Every computed field is a real value or an honest null, never a filled-in guess, so type your parser for T | null on all four. Computed fields states the exact rule per field.

Envelope fields

  • successbooleanrequired

    true for successful responses, false for errors.

  • platformstringrequired

    Platform identifier such as tiktok or instagram. Account endpoints like /v1/credits/balance report meta.

  • endpointstringrequired

    The full endpoint path that was called.

  • dataobjectrequired

    Platform-specific, normalised payload. Lists are returned as items plus optional next_cursor and total. The computed sub-object is attached to every Author and Post, and judged lists add labels to their rows and a report such as data.labels.

  • credits_usedintegerrequired

    Credits consumed by this request. 0 on cache hits and idempotent replays.

  • credits_remaininginteger | nullrequired

    Your current committed balance when known. null when a response cannot establish it without an auxiliary balance lookup.

  • request_idstringrequired

    Unique identifier in the form req-XXXXX, for debugging and log lookup.

  • cachedbooleanrequired

    true when served from cache. Cache hits cost 0 credits.

  • paginationobject

    List endpoints only. The uniform pagination contract next_cursor, has_more, page_size, identical on every platform. A creator post list called with since or stop_at_id also carries stopped_at.

  • metaobject

    Present only when the response carries a hint. meta.hint names one call that does what your recent calls did in fewer steps. Absent on almost every response.

  • idempotent_replayboolean

    Present and true only when this response is an idempotency replay. Absent on a normal response.

Two more fields live inside data rather than at the envelope root: data.dropped and data._warnings[].

data.dropped: Discarded-row counter

On list endpoints, data.dropped counts upstream items omitted because they could not be repaired to the endpoint schema. A healthy list response carries "dropped": 0. It is the only signal that a page came back short, so read it on every page of a drain rather than inferring completeness from items.length.

dropped sits next to items inside data, while pagination sits at the envelope root. Reading response.dropped returns undefined and silently looks like zero.

data._warnings[]: Partial-data channel

When the field-map or computed-field pipeline hits ambiguous upstream data, it attaches a human-readable notice to data._warnings: string[]:

Response
{
  "data": {
    "author": { "username": "..." },
    "computed": { "engagement_rate": 1.0 },
    "_warnings": [
      "computed.engagement_rate: value exceeded 1.0 (raw: 1.42); clamped"
    ]
  }
}

Treat entries as advisory. The response is still valid. Empty arrays are omitted, so _warnings is present only when there is something to report.

A query parameter the endpoint does not declare is ignored, never forwarded, and named in _warnings with the parameter the endpoint does support, for example "days was ignored: did you mean recent_days?", or "sentiment was ignored: this endpoint has no sentiment parameter; label=sentiment is free here.". The request is never rewritten for you. An alias we accept and apply for you (username for handle on tiktok/profile, video_id for url on youtube/video) is not ignored, so it gets no line.

Data-level reports

Lists that carry SocialCrawl labels add a report next to items, so you can tell a judged page from an unjudged one without reading every row. All of them are on by default, cost nothing, and disappear with judgments=off.

KeyOnWhat it reports
data.labelscomment, post and review listsThe labels on the rows: presets, mode (default or requested), status, rows, labelled, cached, pending, pending_ids, unjudged, dropped_ids, extra_credits. When you name labels with label=, data.labels.default carries the report of the default ones
data.relevancepost search lists, search/everywhere, search/forumsHow many rows were scored against your query and, with relevance=filter, which were dropped (dropped_ids). On the post search lists it also carries origin, status, pending and pending_ids
data.stance_splitsearch/everywherePositive, negative, neutral and no-opinion counts and shares, overall and per platform, with the result ids behind every count
data.stories, data.story_groupingsearch/news (JSON only)Articles about the same event grouped into stories across editions and languages, and how many articles were judged. data.stories is null when grouping did not finish
data.account_kindssearch/creatorsHow many creators were judged for account kind

status is one of three values: complete (every row that can be judged carries its judgment), partial (some rows are null, either still being judged, counted in pending, or not judged this time), or skipped (judging was switched off for this request because capacity was short, and the page is exactly what judgments=off returns plus the report). A pending row is filled on your next call for the same page, including a free cache hit. See Labels for the details, including the search endpoints, which do not fill late rows.

meta.hint: A shorter path for what you just did

When a run of your calls repeats what one other call already does, the response to the call that completes the run can carry a top-level meta.hint naming that call. It appears only on a successful GET:

Response
{
  "success": true,
  "platform": "youtube",
  "endpoint": "/v1/youtube/search",
  "data": {
    "items": [
      /* ... */
    ],
    "dropped": 0
  },
  "credits_used": 1,
  "credits_remaining": 4998,
  "request_id": "req-XXXXX",
  "cached": false,
  "meta": {
    "hint": {
      "code": "search_multi",
      "message": "You searched \"matcha latte\" on 3 platforms one call at a time; /v1/search/multi returns each platform's own search rows in one call at the same price per platform, and a platform with no rows is not charged.",
      "method": "GET",
      "path": "/v1/search/multi?query=matcha%20latte&platforms=tiktok,instagram,youtube"
    }
  }
}
codeWhen it appearsPoints to
creator_cardYou looked up the same handle on the profile endpoints of three platforms the creator card covers, within 30 minutesGET /v1/prism/creator-card
batch_profilesYou looked up 50 different handles on one platform's profile endpoint within 30 minutesPOST /v1/prism/profiles
search_multiYou searched the same query on the search endpoints of three platforms that search/multi covers, within 30 minutesGET /v1/search/multi
incremental_syncYou fetched page 1 of the same creator post list again within 24 hours, without since or stop_at_idThe same list with stop_at_id (Pagination)
comment_labelsAfter a search, you opened the comments on five or more posts from the same platform within 30 minutes, without label=Filtering on the computed.labels every comment row already carries at no extra credits, or GET /v1/prism/comments, with its price in the message
search_repeatYou ran page 1 of the same search on the same endpoint on an earlier visit (more than 30 minutes ago, within two days), without seen= or a date filterThe same search with seen= (or the endpoint's date filter)

method and path are the call to make, and message says why in one sentence. When the suggested call costs more than the calls it replaces, the message states both prices. A key receives each hint at most once per UTC day, and a response carries at most one. A hint never appears on an error or on an idempotent replay, it costs nothing, and the rest of the response is unchanged. Send the request header X-SocialCrawl-Hints: off to never receive one. If your parser rejects unknown keys, allow an optional meta object at the envelope root.

Error response

Response
{
  "success": false,
  "error": {
    "type": "INSUFFICIENT_CREDITS",
    "message": "Your account has 0 credits remaining. This endpoint requires 1 credits.",
    "status": 402,
    "doc_url": "https://www.socialcrawl.dev/docs/errors#insufficient-credits"
  },
  "credits_used": 0,
  "credits_remaining": 0,
  "request_id": "req-abc123"
}

credits_used is the net committed charge for the failed request. It is 0 when no deduction occurred or the refund committed. If an immediate refund fails, it remains non-zero while exact compensation is queued; keep the request_id for reconciliation.

doc_url is always the anchor form /docs/errors#<code-slug>, never a per-code sub-path. See Errors for the full error-code table and refund rules.

Response headers

HeaderValue
X-Request-IdMatches request_id in the body. Use it to correlate logs
X-Credits-UsedCredits consumed. 0 on cache hits, empty-upstream 404s, 405/409/422, and idempotent replays
X-Credits-RemainingCurrent balance when known. On an idempotency replay, omitted when the balance lookup fails; body credits_remaining is null instead
X-CacheHIT (served from cache, 0 credits) or MISS. Force a MISS with Cache-Control: no-cache
X-Idempotent-Replay"true" on idempotent replays. Only present when the response was replayed
X-RateLimit-Limit / -Remaining / -ResetRequest-window headroom, on every authenticated response. See Rate limits
X-Concurrency-Limit / -RemainingIn-flight headroom, on every authenticated response. Not an alias of X-RateLimit-*
X-Upstream-RetriesHow many upstream retries this call needed. Omitted when there were none
Retry-AfterSeconds to wait. Sent on both 429s (seconds until the window resets for RATE_LIMITED, a static 1 for CONCURRENCY_LIMIT), on a transient 503 (30), and on a 402 served from the credit cooldown (60). Absent otherwise
AllowAccepted method set on 405 responses: "GET, HEAD" for registry reads or "POST" for POST-only routes

List endpoints

List archetypes (PostList, CommentList, SearchResult) are normalised to a consistent shape regardless of upstream key names:

Response
{
  "data": {
    "items": [
      /* ... */
    ],
    "next_cursor": "eyJwYWdlIjoyfQ==",
    "total": 1247
  }
}

The data-level next_cursor and total are legacy per-payload fields, included only when upstream provides them.

Use the top-level pagination block for pagination. Every list response also carries a uniform pagination object at the envelope root. It is identical on all platforms and all three pagination styles (cursor, page, offset):

Response
{
  "data": {
    "items": [
      /* ... */
    ],
    "dropped": 0
  },
  "pagination": {
    "next_cursor": "sc.eyJ2IjoyLCJjIjoiMTc3NTM5NDk1MDAwMCJ9",
    "has_more": true,
    "page_size": 30
  }
}

The loop is the same everywhere. Read pagination.next_cursor, send it back as the cursor query parameter, and stop when pagination.has_more is false. You never handle a per-platform input-param name yourself. The API rewrites the sc. cursor to the endpoint's native param, max_cursor or after for example, on the way in. For the full drain recipe, see Pagination.

On a list in date order, a call that sends since or stop_at_id gets a fourth field, pagination.stopped_at. It is known_id or since when the walk reached that boundary, end when the list ran out first, and null while there is more to walk. A call that sends neither, or a list that is not in date order, has no stopped_at key. See Incremental sync and Date window on every list.

Idempotent requests

Any /v1/* request can be made safely retriable by sending an Idempotency-Key header. A UUIDv4 is recommended.

cURL
curl 'https://www.socialcrawl.dev/v1/tiktok/profile?handle=charlidamelio' \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Idempotency-Key: 7a5e1b4c-2d8f-4a3b-9c1e-6e8b4d2a1f3c"
TypeScript
const res = await fetch(
  "https://www.socialcrawl.dev/v1/tiktok/profile?handle=charlidamelio",
  {
    headers: {
      "x-api-key": process.env.SOCIALCRAWL_API_KEY!,
      "Idempotency-Key": crypto.randomUUID(),
    },
  },
);

const body = await res.json();
console.log(body.idempotent_replay, body.credits_used);
Python
import os
import uuid

import requests

res = requests.get(
    "https://www.socialcrawl.dev/v1/tiktok/profile",
    params={"handle": "charlidamelio"},
    headers={
        "x-api-key": os.environ["SOCIALCRAWL_API_KEY"],
        "Idempotency-Key": str(uuid.uuid4()),
    },
)

body = res.json()
print(body.get("idempotent_replay"), body["credits_used"])

Replayed calls keep the cached payload immutable except for billing metadata: credits_used becomes 0 and idempotent_replay becomes true. A known current balance refreshes body credits_remaining and the X-Credits-Remaining header; no balance row resolves to 0. On a transient balance lookup failure, body credits_remaining is null and X-Credits-Remaining is omitted.

GET endpoints cannot combine Idempotency-Key with Accept: text/event-stream, because an interrupted stream has no complete response body to replay. Request Accept: application/json when you need idempotent replay. The unsupported combination returns an unbilled 400 INVALID_REQUEST. POST batch streams keep their documented batch replay behavior.

Keys are scoped per account. A new request holds a 5-minute execution claim; settled outcomes remain replayable for 24 hours, subject to the 64 KB body cap. See Errors for contention, replay-limit, claim-store, and payload-mismatch outcomes.

Force a fresh fetch

Responses from cache-enabled endpoints are cached and the cache is shared, so a repeat call returns instantly and costs 0 credits (X-Cache: HIT). To bypass the cache for a single call, send the standard Cache-Control: no-cache request header:

cURL
curl 'https://www.socialcrawl.dev/v1/tiktok/profile?handle=charlidamelio' \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Cache-Control: no-cache"

The request fetches live from the source and is billed at the normal endpoint cost, a forced MISS, so X-Cache: MISS and a non-zero X-Credits-Used. The fresh result is written back, so your next plain call gets a free HIT. Only the no-cache directive triggers it. Cache-Control: no-store on its own does not. Registry-marked uncached endpoints skip this cache contract.

For the freshness windows per data type, the shared-cache model, and guidance on when to bypass, see Caching.

Next steps