# Sentiment Analysis API: Every Comment Scored, Same Price (https://www.socialcrawl.dev/blog/comment-sentiment-analysis-api)
> SocialCrawl's sentiment analysis API scores every TikTok, Instagram, YouTube and X comment for sentiment and intent, same credits, judgments=off to disable.
A `GET` call to a comment-list endpoint on TikTok, Instagram, YouTube, Reddit, X or Facebook now returns a `computed.labels` block on every comment (or `labels: null` when a comment couldn't be judged): sentiment, question, purchase intent and complaint. It's a sentiment analysis API built into the call you already make, not a second step. Two page-level summaries, `data.label_share` and `data.comment_language`, ride along on the same response at no extra credits. It costs exactly what the same call cost before the labels existed. `judgments=off` turns them back off if you'd rather not have them.
This covers the sentiment analysis applications teams actually run on comments: monitoring, triage and lead-detection workflows that label comments you're already fetching, on a call you were already making, not a string you paste into a separate scoring tool. Every example below is a real request against a live post, run on 2026-10-01, with the actual (trimmed) response next to it. Four platforms got measured live for this post: TikTok, Instagram, YouTube and X. Reddit and Facebook carry the same fields on their comment endpoints, named below, but weren't part of this run's live checks.
## What does the sentiment analysis API return on each comment?
The labels ride on the comment list you already fetch: no second call, no separate scoring endpoint to wire up.
```bash
curl "https://www.socialcrawl.dev/v1/tiktok/post/comments?url=https://www.tiktok.com/@stoolpresidente/video/7623818255903329566" \
-H "x-api-key: $SOCIALCRAWL_API_KEY"
```
One comment from that response, trimmed to the part that matters:
```json
{
"comment": { "text": "Where’s Miss Peaches" },
"computed": {
"language": "en",
"labels": {
"sentiment": { "level": 2, "score_0_1": 0.5, "confidence": 0.94 },
"question": { "p": 0.75 },
"purchase_intent": { "p": 0.03 },
"complaint": { "p": 0.02 }
}
}
}
```
Four scored fields per comment: a sentiment level, score and confidence, plus a probability each for question, purchase intent and complaint. `computed.language` comes along on the same response. That one comment reads as mostly a question (`p: 0.75`), which is right: it's asking where someone is.
## Does labelling cost extra credits?
Same TikTok video, same URL, with `judgments=off` added:
```bash
curl "https://www.socialcrawl.dev/v1/tiktok/post/comments?url=https://www.tiktok.com/@stoolpresidente/video/7623818255903329566&judgments=off" \
-H "x-api-key: $SOCIALCRAWL_API_KEY"
```
`computed` drops to `{"language": "en"}`, no `labels` key, no `data.label_share`. The bill doesn't move. Both calls cost 1 credit. We confirmed the same pattern across four platforms by billing both versions of the same call:
| Platform | Endpoint | Baseline credits | `judgments=off` credits | Same price? |
|---|---|---:|---:|---|
| TikTok | `tiktok/post/comments` | 1 | 1 | Yes |
| Instagram | `instagram/post/comments` | 5 | 5 | Yes |
| YouTube | `youtube/video/comments` | 1 | 1 | Yes |
| Twitter/X | `twitter/tweet/replies` | 1 | 1 | Yes |
Same price, no extra credits, on every platform we billed. Reddit's `reddit/post/comments` and Facebook's `facebook/post/comments` and `facebook/post/comment/replies` carry the same fields according to our endpoint notes, but this run didn't bill them live.
## What does label_share tell you about a whole comment page?
A single comment call also works as a Twitter sentiment analysis API: it returns a page-level stat, with its own confidence interval.
```bash
curl "https://www.socialcrawl.dev/v1/twitter/tweet/replies?url=https://x.com/News24/status/2093257170653483101" \
-H "x-api-key: $SOCIALCRAWL_API_KEY"
```
`data.label_share.presets.sentiment.negative` on 36 replies to a post about fuel-price pass-through:
```json
{ "n": 36, "counted": 27, "abstained": 6, "share": 0.9, "low": 0.7745, "high": 0.9593, "too_few": false }
```
90% of the replies the model could judge are negative, interval 0.77 to 0.96. `counted` (27) is the number of replies in the negative bucket, and the share is taken over the 30 replies the model could judge (36 minus 6 abstained), so 27 of 30 is 0.9. A comment the system can't judge counts as abstained, not as a silent "no". SocialCrawl surfaces verbatim comments as the receipts behind the number, such as *"@News24 News24 wants us to pay to read an article about how we will pay more for stuff"*.
One `GET` call turned a raw reply thread into "90% negative, here's the evidence, here's the margin of error." That's usually a separate NLP pipeline you build and maintain yourself. For teams running comment labels into a monitoring workflow, this is the exact number a [social media monitoring API](/blog/social-media-monitoring-api) call would feed into an alert.
## What happens when the sample is too small to trust?
```bash
curl "https://www.socialcrawl.dev/v1/instagram/post/comments?url=https://www.instagram.com/p/DXidPIVDU6M/" \
-H "x-api-key: $SOCIALCRAWL_API_KEY"
```
`data.label_share.presets.sentiment.positive` on this 15-comment Instagram page:
```json
{ "n": 15, "counted": 2, "abstained": 6, "share": null, "low": 0.0765, "high": 0.4964, "too_few": true }
```
The interval, 0.08 to 0.50, is wider than 0.25, so `share` comes back `null` and `too_few: true`. No headline percentage gets printed on a sample this thin. We doubled the sample with `scan_pages=2` on the same post, 30 comments instead of 15, and every sentiment bucket still came back `too_few: true`, because a third of the page (13 of 30 comments) couldn't be judged either way. The honesty holds under more data, not just small data. A naive count of 2 positive out of 15 would print 13% here; `too_few` is there so you don't ship that number.
## Can you re-scan a comment thread without paying twice?
Instagram's `instagram/post/comments` takes a `scan_pages` parameter (1 to 3, default 1) that reads multiple pages, dedupes, and re-sorts by like count:
```bash
curl "https://www.socialcrawl.dev/v1/instagram/post/comments?url=https://www.instagram.com/p/DXidPIVDU6M/&scan_pages=2" \
-H "x-api-key: $SOCIALCRAWL_API_KEY"
```
`data.scan` on that call:
```json
{ "pages_requested": 2, "pages_walked": 2, "pages_billed": 2, "comments_scanned": 30, "order": "likes", "complete": false, "stopped": "pages" }
```
Both pages added new comments (30 distinct against 15 on a single-page call), so both got billed, 10 credits total. The billing rule is conditional, not a fixed discount: a page that only repeats earlier comments isn't billed, and credits held for pages that were never read are refunded in the same call, but on this specific live run neither case fired, because the second page added new comments. Check your own call's `data.scan` block rather than assuming either outcome.
TikTok's `tiktok/post/comments` takes the same idea with a `sort` parameter, building on the pagination that [the TikTok comment-scraping tutorial](/blog/how-to-scrape-tiktok-comments-python) already covers. `sort=recent` is the nearest this endpoint gets to real-time sentiment analysis: it sorts the comments the call actually read by date, newest first, and `scan_pages` (up to 3) widens how much of the thread it reads. It does not turn the endpoint into a live stream:
```bash
curl "https://www.socialcrawl.dev/v1/tiktok/post/comments?url=https://www.tiktok.com/@stoolpresidente/video/7623818255903329566&sort=recent" \
-H "x-api-key: $SOCIALCRAWL_API_KEY"
```
`data.scan` came back `{"pages_requested": 1, "pages_walked": 1, "pages_billed": 1, "comments_scanned": 49, "order": "published_at", "complete": false, "stopped": "pages"}`, alongside this in `data._warnings`: "TikTok has no newest-first comment order, so these are the newest of the 49 comments on the first 1 page, sorted by date. data.scan.complete is false. Pass scan_pages (up to 3) to scan more of the thread." `data.comment_recency.recent_share_is_floor` is `true`. That warning and floor flag exist to stop a caller building a spike-detector from assuming they saw the whole thread when they only read the first page of it.
## Does comment_language tell you where your audience is?
No. `data.comment_language` reports a language mix, not a country or an audience location. On all four sampled pages it came back 100% English (`en`, share 1.0), which is a property of the posts we picked, not a platform-wide fact. On TikTok, `prism/audience-language` places the comment-language mix beside an audience sample and one page of follower locations, each with its own count. This post did not measure it live.
## What would you build yourself instead?
Before comment-level labels showed up on the call itself, running sentiment analysis NLP on social comments meant wiring two systems together: something that fetches the comments, and something that scores the text. The three hyperscaler NLP APIs are the usual pick for the scoring half, each one an AI sentiment analysis tool you prompt or call separately, and all three price by text volume, not by comment.
[Google Cloud Natural Language](https://cloud.google.com/products/natural-language/pricing) charges $1.00 per 1,000 "units" of 1,000 characters each after the first 5,000 units a month at no charge. [AWS Comprehend](https://aws.amazon.com/comprehend/pricing/) charges per 100-character unit with a 3-unit (300-character) minimum on every single request, so a two-word reply like "so true" still bills as three 100-character units (300 characters). [Azure AI Language](https://azure.microsoft.com/en-us/pricing/details/language/) charges $1.00 per 1,000 text records on its Standard pay-as-you-go tier for the first 500,000 records a month (US East), where a record is up to 1,000 characters, so a two-word comment still bills as a whole record. A newer option is routing comment text through a general-purpose LLM instead: [the cheapest 2026-generation models](https://www.spheron.network/blog/llm-api-pricing-comparison-gpt-claude-gemini-deepseek-2026/) start at $0.14 per million input tokens (DeepSeek V4-Flash), with OpenAI's budget GPT-5.6 Luna at $0.20 input and $1.20 output, but you're still building the fetch, the prompt and the output parsing yourself.
None of these four options fetch a comment for you. They're scoring functions waiting for text you already have, which, for social comments, is the harder half of the problem most roundups skip over. If you want a comparison of those standalone scoring tools instead of an API that fetches the comment for you, [8 sentiment analysis tools](/blog/sentiment-analysis-tools) covers that ground directly.
## How do you start?
Sign up for a SocialCrawl account to get an API key and starting credits. The first call is a raw `curl`, no SDK to install:
```bash
curl "https://www.socialcrawl.dev/v1/youtube/video/comments?url=https://www.youtube.com/watch?v=dQw4w9WgXcQ" \
-H "x-api-key: $SOCIALCRAWL_API_KEY"
```
That returns the same `computed.labels` block shown above, on the same comment call the [YouTube comment scraper guide](/blog/youtube-comment-scraper) walks through for pagination. From there: the [TikTok](/platforms/tiktok/post-comments) and [Instagram](/platforms/instagram/post-comments) comment endpoint docs cover the full parameter list for `scan_pages` and `sort`, the [visual explorer](/explorer) lets you see a labelled response before you write a line of code, and [`prism/comments`](/platforms/prism/comments) walks every comment page of one post, on any supported platform, and returns the same labels plus `data.label_share` and `data.comment_language` at its unchanged price.
## FAQ
### Does SocialCrawl's comment sentiment analysis cost extra?
No. The labels aren't a separate product or tier: a comment call bills the same with labels on as with them off. `judgments=off` disables `computed.labels`, `data.label_share` and `data.comment_language` on the response; no SDK or code change is required beyond the query parameter.
### What does label_share actually measure on a comment page?
The fraction of judged comments in a preset bucket (a sentiment level, question, purchase intent, or complaint), alongside the sample size, how many comments were left out as unjudgeable, and a 90% confidence interval. When that interval is wider than 0.25, `too_few: true` suppresses the headline percentage instead of printing a number the sample can't support.
### Does comment_language tell you where an audience is from?
No, it's a language mix, not a country. On TikTok, `prism/audience-language` places that mix beside an audience sample and one page of follower locations.
### Can you sort TikTok or Instagram comments by recency instead of relevance?
Yes. On TikTok, `sort=recent` sorts the comments the call read newest first, and `scan_pages` (1 to 3) reads more of the thread; `data.scan.complete` says whether it reached the end. On Instagram, `scan_pages` (1 to 3) with `sort=recent` returns the kept comments newest first instead of by likes. A page that only repeats earlier comments isn't billed, and credits held for pages that were never read are refunded in the same call. Check `data.scan` on your own call.
### Which platforms carry the comment labels today?
TikTok, Instagram, YouTube, Reddit, X (via tweet replies) and Facebook carry the fields by endpoint name. This post's live numbers cover TikTok, Instagram, YouTube and X specifically; Reddit and Facebook weren't part of this run's live price and response checks.