Sponsored Post Detection API: Flag Undisclosed Ads
Flag sponsored posts and undisclosed ads in TikTok, X, YouTube and Reddit search results. One giveaway query lost 26 of 29 rows to the engagement bait filter.
To find sponsored posts and undisclosed ads in bulk, add label=sponsored,intent,niche to a TikTok, X, YouTube or Reddit search. Each row comes back with computed.labels.sponsored, a score for paid or gifted promotion that also says whether a disclosure was found, plus an intent and a niche. Add label=quality and exclude=engagement_bait and the engagement bait rows are dropped, with their ids listed in data.labels.dropped_ids.
That is sponsored post detection with the SocialCrawl API: one key and the same fields on four platforms. sponsored is a probability with evidence flags, so you pick the cut-off. Every number below comes from live calls on 1 October 2026 and describes one page of results.
Why does a #ad hashtag filter miss undisclosed ads?
Searching #ad finds only the creators who disclosed the way you thought to look for. Regulators who sampled real posts found many who did not.
The European Commission and consumer authorities in 22 Member States, Norway and Iceland checked 576 influencers. Their 14 February 2024 press release reports that 97% posted commercial content and only 20% systematically disclosed it. 38% skipped platform labels such as Instagram's paid partnership label and wrote "collaboration", "partnership" or a generic thanks to the brand instead.
In May 2025 the UK's Advertising Standards Authority published an AI-led review of more than 50,000 Instagram and TikTok pieces. As reported by Lewis Silkin (we did not fetch the ASA report itself), 80% of the undisclosed ads made no attempt to label, and 20% used inadequate labels such as "gifted", "PR trip" or "affiliate". Those counts are per post, so do not set them beside the EU's per-influencer numbers.
The FTC's influencer guidelines, in its Endorsement Guides Q&A, say "just because a platform offers this feature is no guarantee that it's an effective way for influencers to disclose their material connection to a brand." They say a post that starts with #ad or "Ad:" would likely be effective and call #ambassador ambiguous.
So a detector has to read the caption and see which brand is being thanked. Modash's Collaborations API and HypeAuditor's Likely Sponsored help page also detect sponsored posts, though neither says it scores every row of a keyword search (public pages, accessed 2026-10-02). For influencer marketing, these labels sit on the post rows you already pull.
How do you flag sponsored posts in a TikTok search?
Name every label you want in the search. This query was built to pull disclosures, so it stress-tests the field and says nothing about how much of TikTok is sponsored.
curl -G -H "x-api-key: $SOCIALCRAWL_API_KEY" \
"https://www.socialcrawl.dev/v1/tiktok/search" \
--data-urlencode "query=#ad skincare serum gifted" \
--data-urlencode "label=sponsored,intent,niche,quality" \
--data-urlencode "label_evidence=1"HTTP 200, 30 rows, 3 credits. The trimmed data.labels block:
{
"data": {
"labels": {
"presets": ["sponsored", "intent", "niche", "quality"],
"rows": 30,
"labelled": 30,
"cached": 4,
"skipped": 0,
"unjudged": 0,
"dropped": 0,
"dropped_ids": [],
"extra_credits": 2,
"mode": "requested",
"status": "complete",
"pending": 0
}
},
"credits_used": 3
}All 30 rows carried all four label groups on the first read. The 3 credits are 1 for the search plus 2 for quality.
A post with a disclosure marker:
{
"post": {
"content": { "text": "⚠️ Warning: Once you see the glow, there’s no going back. ✨ #syntheticperformer #ad #bomierepartner #skincare #BeautyTok" }
},
"computed": {
"labels": {
"sponsored": { "p": 0.85, "disclosed": true, "undisclosed": false, "brand": null },
"niche": { "label": "beauty_skincare", "confidence": 1, "taxonomy": "sc-niche-v1" }
},
"labels_evidence": {
"sponsored": { "quote": "✨ #syntheticperformer #ad #bomierepartner #skincare #BeautyTok", "sentence_index": 1 }
}
}
}No marker at all, the case a hashtag filter cannot see. A creator filming for a brand names it and never writes #ad:
{
"post": {
"content": { "text": "Your favorite series — how I film UGC videos 🎥✨ Today I’m shooting for iUNIK and their Beta-Glucan Power Moisture Serum ... @IUNIK US #iunik #iunikbetaglucan ... #ugccreator" }
},
"computed": {
"labels": {
"sponsored": { "p": 0.85, "disclosed": false, "undisclosed": true, "brand": "@iunik" },
"niche": { "label": "beauty_skincare", "confidence": 1, "taxonomy": "sc-niche-v1" }
}
}
}Its evidence quote is the whole caption, so it is left out. Plain advice:
{
"post": {
"content": { "text": "Save the guesswork for something else ✨ Here’s how often to use these popular skincare ingredients ..." }
},
"computed": {
"labels": {
"sponsored": { "p": 0.04, "disclosed": false, "undisclosed": false, "brand": null },
"niche": { "label": "beauty_skincare", "confidence": 1, "taxonomy": "sc-niche-v1" }
}
}
}With a cut-off of p at 0.5 or above (ours, not the API's), 23 of the 30 rows qualified. 11 had disclosed: true. 10 had undisclosed: true, which the API sets when p is at least 0.7 and no marker is found. 2 sat at 0.52 and 0.56 with neither flag, and 18 of the 23 named a brand. The other 7 ran from 0.04 to 0.46.
The request names its labels on purpose. In the earlier run, a TikTok search with no label parameter had 0 of 30 rows labeled on the first read (data.labels.status: "partial") and an X search 13 of 20. label=quality alone returned all four groups on 28 of 30 rows, since the default labels run in a separate background block. Naming all four returned 30 of 30. The TikTok video search endpoint post covers the base search.
What do disclosed and undisclosed get wrong?
Both flags come from a text rule that looks for a marker, while p carries the model's judgment, and the two can disagree.
A UGC creator thanks Olay and tags the post #GiftedbyOlay. The page returned:
{
"post": {
"content": { "text": "Hey @Olay Skin Care ! Thanks a million for making a serum that’s skips the layering & packs in 5 benefits in 1! Definitely a new favorite to my skincare lineup. #GiftedbyOlay #ugccreator ..." }
},
"computed": {
"labels": {
"sponsored": { "p": 0.96, "disclosed": false, "undisclosed": true, "brand": "@olay" }
},
"labels_evidence": {
"sponsored": { "quote": "#GiftedbyOlay #ugccreator #Skincare #ugcjourney", "sentence_index": 3 }
}
}
}The post is caught, with p 0.96 and the right brand. The flag is wrong, because a caption that says "gifted" reads undisclosed: true. The marker rule matches #gifted or "gifted" only as a whole word, so a hashtag with the brand glued on does not count, and it has no form for "ad" without a hash sign. The same page read #giftedbyskinbetter the same way, and four of the five captions with a bare "ad" read undisclosed: true (the fifth also said "Gifted"). undisclosed: true means the rule found no marker it recognizes, so counting it as "the creator hid the sponsorship" reports creators who wrote "gifted" as undisclosed ads.
The reverse happens too, on a YouTube search from the earlier run (type=videos, label=sponsored,intent,niche, query sponsored skincare review gifted by brand):
{
"post": {
"content": { "text": "how to get paid brand deals & gifted PR as a small creator 💰 (tips & tricks) *NO GATEKEEPING*" }
},
"computed": {
"labels": {
"sponsored": { "p": 0.05, "disclosed": true, "undisclosed": false, "brand": null },
"niche": { "label": "business_career", "confidence": 0.99, "taxonomy": "sc-niche-v1" }
},
"labels_evidence": null
}
}A how-to about getting gifted PR reads disclosed: true because the title contains "gifted", and the p of 0.05 is the right answer. disclosed: true with a low p means a keyword matched in a post that is probably not an ad.
brand is the account the post mentions, and in several rows it is cut short at the first space or accented letter of the mention: on the TikTok page above, @lanc for Lancôme, @dr for Dr Barbara Sturm and @derma for Derma Factory. Check it before you treat it as a clean brand name.
The labels read the caption (the title, for YouTube), not the video. A creator who discloses only by voice, or only through the platform's paid partnership toggle, will not show disclosed: true. And given the FTC and ASA lines above, disclosed: true means a marker was found, not that it is adequate.
label_evidence=1 is the review step. It costs no extra credit and returns the sentence behind a score as computed.labels_evidence, a sibling of labels. It is null when no single sentence shows the signal, which is common outside sponsored.
What niche and intent does each post carry?
computed.labels.niche is one of 33 topic niches, such as beauty_skincare and finance_investing, with personal_no_niche and other as the two ways out (taxonomy sc-niche-v1, 35 values in total). label is null when the caption is too thin or confidence is under 0.6. Two of the 30 TikTok rows came back null, and both captions were little more than hashtags and mentions.
computed.labels.intent.label is one of asking_for_recommendation, comparing_options, switching_away, complaining, promoting, news_or_discussion or other. It is null when confidence is under 0.5, which happened on 4 of the 30 TikTok rows (confidences 0.42, 0.46, 0.42 and 0.34). switching_away did not appear in any of our samples. buyer flags a buying intent and seller is a 0 to 1 score, as in this organic X post from the earlier run, trimmed:
{
"post": {
"content": { "text": "I'm sorry but I need the skincare routine. Whatever it is, idc, ..." }
},
"computed": {
"labels": {
"sponsored": { "p": 0.03 },
"intent": { "label": "asking_for_recommendation", "confidence": 0.99, "buyer": true, "seller": 0.02 },
"quality": { "fact_density": 0 }
}
}
}A sponsorship-tracking job can keep niche.label == "beauty_skincare" and sponsored.p >= 0.5. A lead-finding job can keep intent.buyer == true. For comment sentiment, a sibling label family, see the sentiment analysis tools post.
How do you filter engagement bait out of a search result?
Meta's transparency center defines engagement bait as posts that "explicitly request engagement (such as votes, shares, comments, tags, likes, or other reactions)". We found no published prevalence figure.
label=quality adds five fields under computed.labels.quality: fact_density (an integer from 0 to 3), engagement_bait, rage_bait and secondhand (each 0 to 1), and post_aim (inform, opinion, sell, entertain, provoke or other). One row from the page above:
{
"post": {
"content": { "text": "fall skincare is officially back in rotation ... adding the Cordiale Camellia Revitalizing Boosting Serum to my routine ... @Cordiale #Cordiale #Gifted #Skincare #fallskincare" }
},
"computed": {
"labels": {
"sponsored": { "p": 0.96, "disclosed": true, "undisclosed": false, "brand": "@cordiale" },
"quality": { "fact_density": 2, "engagement_bait": 0.02, "rage_bait": 0.01, "secondhand": 0.04, "post_aim": "sell" }
}
}
}exclude=engagement_bait drops rows with engagement_bait at or above 0.8, the API's own cut-off. It needs label=quality in the same call. Sent with no label, it returns HTTP 400 and 0 credits, with a message that starts "Parameter 'exclude' has no effect without 'label'" and did_you_mean: "label". Two giveaway queries, one on TikTok and one on X:
curl -G -H "x-api-key: $SOCIALCRAWL_API_KEY" \
"https://www.socialcrawl.dev/v1/tiktok/search" \
--data-urlencode "query=giveaway follow like tag 3 friends to win" \
--data-urlencode "label=sponsored,intent,niche,quality" \
--data-urlencode "exclude=engagement_bait"
curl -G -H "x-api-key: $SOCIALCRAWL_API_KEY" \
"https://www.socialcrawl.dev/v1/twitter/search/tweets" \
--data-urlencode "query=giveaway RT follow tag a friend to win" \
--data-urlencode "sort=top" \
--data-urlencode "label=sponsored,intent,niche,quality" \
--data-urlencode "exclude=engagement_bait"The TikTok response, trimmed to the counters and the first two dropped ids:
{
"data": {
"dropped": 0,
"labels": {
"presets": ["sponsored", "intent", "niche", "quality"],
"rows": 29,
"labelled": 29,
"cached": 11,
"dropped": 26,
"dropped_ids": ["768983…", "769163…", "..."],
"extra_credits": 1,
"mode": "requested",
"status": "complete"
}
},
"credits_used": 2
}| Query | Rows on the page | Rows kept | Rows dropped | Credits |
|---|---|---|---|---|
TikTok giveaway follow like tag 3 friends to win | 29 | 3 | 26 | 2 |
X giveaway RT follow tag a friend to win, sort=top | 20 | 2 | 18 | 2 |
The three TikTok survivors scored 0.09, 0.43 and 0.26, and one is a post-giveaway note: "i'll contact the winner on friday!! and next giveaway at 40k". Among the dropped ids is a post the earlier run scored 0.97: "RULES- must be over 18, live in Great Britain, share,like,save the post. Follow me and tag three friends."
The two X survivors scored 0.71 and 0.73. One was a product giveaway that asks for a like, an RT and a follow, kept because 0.71 is under the line. If that row is bait for you, apply your own cut-off to quality.engagement_bait.
Two calls to the same query fetch the page separately. In the earlier run's with and without pair, 20 of 30 TikTok ids repeated and 13 of 20 on X, so a before-and-after pair compares different rows, and scores near the line moved (one TikTok post 0.87 then 0.77, an X post 0.81 then 0.75).
Top-level data.dropped stayed 0 when 26 rows were dropped, so read data.labels.dropped and data.labels.dropped_ids.
Giveaway rows ask for likes, comments and shares, so they can shift the counts behind engagement rate benchmarks. In a social media monitoring API job, exclude=engagement_bait keeps them out of the dataset you store.
Can the API flag rage bait?
Yes, as a score. quality.rage_bait runs from 0 to 1, and post_aim: "provoke" marks posts written to start a fight. Oxford University Press named "rage bait" its 2025 Word of the Year (announced 1 December 2025) and defines it as "online content deliberately designed to elicit anger or outrage ... typically posted in order to increase traffic to or engagement with a particular web page or social media account." Oxford reports usage tripled in 12 months.
In the earlier run's Reddit sample (/v1/reddit/search, 25 rows, sort=top, timeframe=month), two rows had post_aim: "provoke" and one scored rage_bait 0.70, the only row of 25 at or above 0.5. The highest rage_bait across the 30 TikTok skincare rows above was 0.11.
There is no exclude=rage_bait, so filter on the score yourself and pick your own cut-off (we used 0.5):
import os
import requests
r = requests.get(
"https://www.socialcrawl.dev/v1/reddit/search",
params={
"query": "skincare serum review",
"sort": "top",
"timeframe": "month",
"label": "sponsored,intent,niche,quality",
},
headers={"x-api-key": os.environ["SOCIALCRAWL_API_KEY"]},
timeout=60,
)
items = r.json()["data"]["items"]
def rage_bait(row):
labels = (row.get("computed") or {}).get("labels") or {}
return (labels.get("quality") or {}).get("rage_bait") or 0
calm = [row for row in items if rage_bait(row) < 0.5]Rows with no quality block count as 0, so unlabeled rows stay in.
Which platforms return these labels?
We ran four, each on one hand-picked skincare query and one page. The TikTok row is the call above, the others come from the earlier run, and "three" labels means quality was left out.
| Platform and endpoint | Query | Labels | Rows | Credits | p of 0.5 or more |
|---|---|---|---|---|---|
TikTok /v1/tiktok/search | #ad skincare serum gifted | four | 30 | 3 | 23 of 30 |
X /v1/twitter/search/tweets | skincare (#ad OR #gifted OR #sponsored) | three | 20 | 1 | 20 of 20, all disclosed: true |
X /v1/twitter/search/tweets | my skincare routine -#ad -#gifted -#sponsored -filter:links | four | 20 | 2 | 2 of 20 |
YouTube /v1/youtube/search, type=videos | sponsored skincare review gifted by brand | three | 20 | 1 | 2 of 20, both undisclosed: true |
Reddit /v1/reddit/search, sort=top, timeframe=month | skincare serum review | four | 25 | 2 | 0 of 25 (highest 0.16) |
Every row on every page carried labels. The counts describe those pages only.
Instagram, Threads, LinkedIn, Facebook and Hacker News post lists are listed as labeled in the docs, but we did not run them for this post. The Reddit data APIs comparison covers the Reddit side of the market.
How much does sponsored post detection cost per call?
Preview it first. dry_run=1 prices a labeled call for 0 credits and fetches nothing:
curl -G -H "x-api-key: $SOCIALCRAWL_API_KEY" \
"https://www.socialcrawl.dev/v1/tiktok/search" \
--data-urlencode "query=#ad skincare serum gifted" \
--data-urlencode "label=sponsored,intent,niche,quality" \
--data-urlencode "label_evidence=1" \
--data-urlencode "dry_run=1"{
"data": {
"estimate": {
"rows_expected": 120,
"rows_cached": 0,
"label_credits_min": 0,
"label_credits_max": 5,
"base_credits": 1
}
},
"credits_used": 0
}The quote is 1 base credit plus 0 to 5 label credits. rows_expected: 120 is the TikTok page cap, and the live call returned 30 rows. sponsored, intent and niche add nothing, even when named. In the earlier run they cost 1 credit on 30 TikTok rows, the same as the bare search. quality adds one credit per started block of 25 rows it labels fresh, and 120 rows is five blocks.
What the live calls charged:
- The 30-row TikTok call with all four labels: 3 credits, 1 base plus
extra_credits: 2. Four rows were already in the label cache, which leaves 26 to label, or two blocks. - The TikTok giveaway call: 2, with 11 of 29 rows cached and
extra_credits: 1. The earlier run of that query, sent withCache-Control: no-cacheandqualityalone, charged 3. That fits the header switching label reuse off, but two observations are not a controlled test. - The X giveaway call: 2, with 8 of its 20 rows cached.
exclude=engagement_baitadds nothing. In the earlier run both giveaway queries cost the same with and without it (3 on TikTok, 2 on X).- A whole page served from cache is free. Repeating an X call without
Cache-Control: no-cachereturnedx-cache: HITand 0 credits.
Read data.labels.cached and extra_credits on each response rather than budgeting a fixed figure. The four curl calls in this post cost us 7 credits.
How do you run your first labeled call?
- Create a SocialCrawl API key (the pricing page lists the plans). Auth is the
x-api-keyheader, one key for every platform above. - Run the first curl above, read
computed.labels.sponsoredon each row and pick the cut-off for the undisclosed ads you will review. We counted at 0.5 and make no claim about its accuracy. - Add
label_evidence=1for the sentence behind a score, andlabel=qualitywithexclude=engagement_baitwhen you need bait filtered. Rundry_run=1first on large pages. - Check
brandagainst your own brand list, and readpanddisclosedtogether.
Your rows will differ from ours (checked 2026-10-01). The field names will not. The explorer shows the payload before you write a line, and the TikTok endpoint reference lists every parameter.
Frequently asked questions
How do I find sponsored posts in bulk on TikTok, X, YouTube or Reddit?
Send the search call with label=sponsored,intent,niche, read computed.labels.sponsored.p on each row and keep the rows above your cut-off (we used 0.5). disclosed and undisclosed say whether a marker was found. We did not run Instagram or the other listed platforms for this post.
How can I tell if a post is a paid or gifted promotion?
The API has no field that separates paid from gifted. It returns one score, p, with disclosed, undisclosed and brand, and a gifted-product post in our TikTok sample scored 0.96. Read the label_evidence quote and the caption to decide which kind of deal it was.
What is engagement bait and how do I filter it from search results?
Meta defines engagement bait as posts that explicitly request votes, shares, comments, tags, likes or other reactions. Ask for label=quality and exclude=engagement_bait in the same call. Rows scoring 0.8 or more are dropped, and their ids appear in data.labels.dropped_ids.
Can I drop rage-bait posts from a keyword search?
Not with a parameter, since exclude takes engagement_bait only. Request label=quality, read computed.labels.quality.rage_bait on each row and filter in your own code, as the Python snippet above does at 0.5.
Does the API tell me which brand a post promotes?
It returns brand, the promoted account taken from the post's own @mentions, or null when none fits. Sometimes the mention is cut short (@lanc for Lancôme), so check it against your brand list.
How much does it cost to label sponsored posts?
sponsored, intent and niche add no credits to the search. quality adds one credit per started block of 25 rows it labels fresh, so a 30-row TikTok search with all four labels charged 3 credits. A dry_run=1 call quotes a min and max for 0 credits.
Are undisclosed ads illegal?
It depends on the country and the post. The FTC's Endorsement Guides expect a clear disclosure of a material connection to a brand. This post is not legal advice, and a label from this API is evidence for a reviewer to check.
Related posts
Sentiment Analysis API: Every Comment Scored, Same Price
SocialCrawl's sentiment analysis API scores every TikTok, Instagram, YouTube and X comment for sentiment and intent, same credits, judgments=off to disable.
Brand Switching: 4 API Calls to Find Who Left Evernote
Brand switching from public posts: find people who say they left a competitor, in their own words, with the reason and a link. Four live API calls on Evernote.
TikTok Account Finder: 28-30 Lookalikes for 5 Credits
One call to SocialCrawl's TikTok similar-accounts endpoint returns 28 to 30 seed-specific creators for 5 credits. Real requests, real responses, zero overlap.
