{"id":"44bdc746-6c44-4442-8841-a33926fb9424","arxiv_id":"2509.05838","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"An audit of TikTok in Italy finds that accounts set to age 13 and 18+ receive similar levels of harmful content, especially when actively searching.","lead":"This preprint reports an automated audit of TikTok's youth safety, using 20 fake accounts aged 13 and 18+ in Italy to measure exposure to harmful videos. The authors find little difference between the two age groups, casting doubt on TikTok's age-based content moderation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central effect could be be a setup artifact: paper never verifies TikTok treats self-reported age-13 accounts as minors. Test age enforcement before interpreting null result.","rationale":"The reader's weakest_assumption is exactly the load-bearing concern I identify. The paper's central claim—minimal difference in harmful content exposure between adult and youth accounts—depends on TikTok actually recognizing and acting on the self-reported age of 13. If that premise fails, the comparison compares two adult-like accounts and the null result is uninformative about age-based moderation. This is more fundamental than the classifier-precision issue because a perfect classifier cannot rescue an invalid manipulation. The paper is transparent about being preliminary and lists limitations, but it omits this validation. The proposed test is feasible and would settle the concern: inspect account settings or API payloads for age-enforcement signals, or attempt an age-restricted action. If enforcement is confirmed, the null result stands as a real (if preliminary) finding; if not, the central claim should be re-framed or re-run. Since the reader already conditionalized the verdict on exactly this concern, I do not propose changing the verdict; CONDITIONAL remains appropriate, and my verdict_should_be is UNCHANGED relative to the reader's assessment.","tokens_in":856,"tokens_out":712,"duration_ms":77064,"concrete_test":"Before or during data collection, validate the manipulation per account: after each Youth account is created, log in and record (a) the Privacy/Safety settings (DM availability, comment defaults, Live access) and (b) the raw JSON responses from the TikTok web API for an FYF request, checking for viewer age/minor metadata and any restricted-content flags. Pre-register a pass criterion: if any of the 10 Youth accounts has unrestricted DMs, no minor-age field, or returns the same FYF payload shape as Adult accounts, the manipulation is not confirmed; if all 10 show enforced minor status, re-run the Section 4.3 comparison and see whether the 0.034 vs 0.023 medians persist.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 creates 10 accounts 'set with an age of 13' and 10 'slightly above 18', but the paper never checks whether TikTok actually treats these accounts as minors. The entire adult-vs-youth comparison is a platform-condition comparison: if the age-13 accounts do not receive the under-18 experience (age-restricted features, Youth Safety Mode, age-grouped FYF policy, restricted search results), then the low and similar harm rates in Section 4.3 are an artifact of a failed manipulation rather than evidence that age-based moderation is absent. This is especially sharp in the active-search protocol: the Youth account searched harm-adjacent terms and received results, and the paper even notes that the Italian keyword 'Alcol' returned results for youth accounts—if under-age search filtering were in effect, such unfiltered results would not appear. The paper reports no validation of account state (e.g., API fields, settings, or feature availability), so the null result cannot be interpreted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an observational audit of TikTok's age-based content moderation. The authors created 10 youth (age 13) and 10 adult (age >18) sockpuppet accounts in Italy, collected over 7,000 videos over four days via passive \"For You Feed\" scrolling and active searches with harm-adjacent keywords, and used GPT-4o on video descriptions to estimate harmful content prevalence, Detoxify to measure comment toxicity, and VideoLLaMA3 plus manual annotation to evaluate automated classifiers. The main finding is that harmful video prevalence during passive scrolling is low and similar across age groups (median 3.4% for adults vs 2.3% for youth; fewer than 10% overall), while active searching raises prevalence to roughly 28% in both groups. The authors conclude that TikTok's age-based moderation may not meaningfully differentiate exposure in practice.","tokens_in":6256,"tokens_out":4372,"duration_ms":51812,"significance":"If the central claim were supported, the paper would be a useful contribution to DSA-relevant platform accountability research: it uses a multi-day, multi-account sockpuppet design; covers both passive and active interaction modes; grounds harmful-content categories in TikTok's official community guidelines; and includes manual annotation for evaluating automated classifiers. The paper is also transparent about several limitations. However, the central comparison is not currently established because the age manipulation is unvalidated, the main classifier has low precision, the two youth-account outliers are ignored, and no statistical tests or confidence intervals are provided. These are fixable, but they are load-bearing for the headline conclusion.","major_comments":[{"comment":"The central adult-vs-youth comparison rests on the assumption that setting an account's registered age to 13 causes TikTok to treat it as a minor. The paper never validates this manipulation: there is no check of Youth Safety Mode, age-restricted features, search filtering, or any platform-side marker. The observation that the Italian keyword \"Alcol\" still returned results for youth accounts is consistent both with absent age-based moderation and with the account not being recognized as a minor; the design cannot distinguish these. The null result in Figure 3 may therefore be a setup artifact rather than evidence about TikTok's moderation.","section":"Section 3.1 / Section 4.3"},{"comment":"The harmful-prevalence estimates in Section 4.3 are produced by GPT-4o annotations of video descriptions, yet Section 4.4 reports that GPT-4o has precision of only 59% against manual labels on the 100-video validation set, and no recall is reported. When the estimated positive rate is only 2.3-3.4%, a classifier with 59% precision can substantially overstate true prevalence; the paper reports no correction, no confidence intervals, and no sensitivity analysis. The claim that fewer than 10% of videos are harmful is therefore not quantitatively supported by the evidence presented.","section":"Section 4.4 / Section 4.3"},{"comment":"Figure 3 shows two youth accounts with approximately 14% and 25% harmful proportions, several times the median for both groups. The text summarizes the result as \"fewer than 10% of the videos were predicted to be harmful for both adult and youth accounts\" and uses this to support the \"minimal differences\" conclusion, but it never analyzes or even discusses these outliers. Moreover, no statistical test is reported for the adult-vs-youth comparison. With N=10 accounts per group and high apparent variance, the conclusion requires at least an account-level test, effect size, or explicit treatment of outliers.","section":"Section 4.3 / Figure 3"},{"comment":"The active-search protocol uses only one Adult and one Youth account, but Section 4.3 reports aggregate prevalence rates of 28.44% (adult) and 27.91% (youth) and describes the increase over passive exposure as \"significantly higher.\" No confidence intervals, per-account variability, or keyword-level breakdowns are given. With a single account per condition, the active-search findings cannot support RQ2 or the comparison between age groups.","section":"Section 3.1 / Section 4.3"}],"minor_comments":[{"comment":"The keyword table lists only 7 of the 13 harmful-content categories used in the classifier prompt (e.g., suicide/self-harm, sexual abuse, hate speech, and harassment are absent). Please clarify whether active searches actually covered all framework categories or only a subset, since this affects interpretation of the active-search results.","section":"Appendix B, Table 2"},{"comment":"The manuscript says VideoLLaMA3 was tested on \"a random selection of 100 videos\" and then describes the manual sample as 50 per age group with 25 per group from keyword searches. Please clarify the exact sampling procedure and report the overlap between the classifier-evaluation set and the manual-evaluation set.","section":"Section 4.4"},{"comment":"Precision alone is reported for GPT-4o and VideoLLaMA3. Please report recall, F1, and the confusion matrix against manual labels; without recall, the trade-off between missed harmful content and false alarms cannot be assessed.","section":"Section 4.4"},{"comment":"Several claims would be easier to evaluate with explicit sample sizes: the number of videos collected per account/condition, the number of search videos per keyword, and the total N used for each figure. The current \"over 7,000 videos\" is too coarse.","section":"Section 3.1 / Results"},{"comment":"Minor typographical and formatting issues include \"FORYOUFEED(FYF)\" (missing space), incomplete reference entries (e.g., WSJ Staff 2021), and the unresolved difference between the 13 categories in the prompt and the shorter keyword table.","section":"Text and references"}],"recommendation":"major_revision","confidential_remarks":"This is a preliminary arXiv report whose headline conclusion is plausible but currently underevidenced. The main fixes are within the scope of a revision: validate the youth-account condition, report classifier performance with confidence intervals and a corrected prevalence estimate, analyze account-level outliers, and run statistical tests on the age-group comparison. I therefore recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a genuinely useful pilot study, not a decisive finding. It does something the literature needs—a systematic Italian TikTok audit with 20 fresh accounts, 7,000+ videos, passive and active protocols, and a manual annotation check on the classifiers. The result that adult and youth accounts see nearly identical harmful-content rates (medians about 3.4% vs 2.3%) is a new empirical datapoint that contradicts Eltaher et al.'s report for similar platforms. That contrast is worth knowing about.\n\nWhat's good: the pipeline is clearly described and mostly reproducible; the authors are honest about their limitations; and they evaluate both text-only GPT-4o and video-based VideoLLaMA3 against human labels. The fact that VideoLLaMA3 doesn't beat text-only GPT-4o is a useful negative result for automated auditing.\n\nThe soft spots are load-bearing, not cosmetic. First, the adult-vs-youth comparison assumes TikTok treats an account registered as age 13 as a minor and applies youth-specific moderation. The paper never systematically verifies that—no checking of account settings, available features, or API fields. That said, the 'alcohol' vs 'Alcol' anecdote (English keyword censored for youth, Italian equivalent returning results) is some evidence that age-aware filtering is active, so the stress-test concern about a total setup artifact is partially mitigated. Still, a few minutes checking account state would have closed the loop.\n\nSecond, the classifier precision of 59% means the prevalence estimates are noisy; the near-identical distributions could reflect classifier bias rather than platform behavior. No confidence intervals or statistical tests accompany the central comparison, and the two youth accounts with 14% and 25% harmful rates are mentioned but never dissected—they might be outliers that materially change the story. The active-search component uses only one account per group, so the 28% vs 27.9% similarity is basically anecdotal.\n\nWho this is for: researchers and regulators working on platform auditing and youth protection. It deserves a serious referee, but the verdict you'd expect is \"revise\"—validate the age treatment, add stats, and report per-account variation properly. If those hold up, it becomes a credible check on age-based moderation claims.\n\nRecommendation: yes, engage. Send to peer review rather than desk reject, but referee it with real scrutiny on the account-state verification and the classifier evaluation.\n\nBest.","headline":"Useful preliminary audit of TikTok age-based moderation, but the headline null result is not yet firmly supported: the age-13 accounts are never verified as being treated as minors, and the main classifier has 59% precision.","tokens_in":6711,"tokens_out":1745,"would_cite":false,"duration_ms":23066,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An audit using scripted fake accounts on TikTok found that accounts declared as 13-year-olds and adults were shown nearly identical shares of videos that violate the platform's own youth-safety categories; active keyword searches raised the","keywords":["TikTok audit","youth safety","content moderation","harmful content","sockpuppet accounts","algorithmic auditing","LLM classification","For You feed"],"falsifier":"Register fresh 13-year-old and adult accounts on TikTok mobile and web in Italy, and before collecting any feed data check whether the youth accounts exhibit known minor restrictions such as no direct messages, safety prompts, or limited search suggestions. If they do not, the near-identical feed exposure cannot be attributed to ineffective age moderation because the accounts may not be treated as minors at all.","tokens_in":5965,"feed_emoji":"📱","tokens_out":7413,"duration_ms":81381,"temperature":0.7,"pith_summary":"This paper tries to establish whether TikTok's age-based safety measures actually change what minors are exposed to, using sockpuppet accounts registered as 13-year-olds and as adults in Italy. Over four days and more than 7,000 videos, passive For You scrolling produced nearly identical harmful-content estimates for both age groups, with median proportions around 3.4% for adults and 2.3% for youth. Active keyword searches raised harmful exposure to about 28% for both groups, suggesting that the platform's safeguards are weaker for search than for the recommendation feed. The paper also measures how well LLM and vision-language classifiers detect harmful content, reporting precision near 59%, and frames the pipeline as a repeatable way to audit a very large online platform. A reader should care because the result directly bears on whether age-based content moderation protects children in practice.","feed_headline":"TikTok audit finds 13-year-olds get same harmful feed as adults","feed_subtitle":"7,000 videos scrolled across fresh youth and adult accounts: passive exposure looks alike, active searches raise both to ~28%.","key_machinery":"The load-bearing mechanism is the paired sockpuppet audit: ten accounts created with a registered age of 13 and ten registered just above 18, run on TikTok's web interface in Italy, each with scripted sessions of passive scrolling and active keyword searching. The account pair is what lets the authors attribute any difference, or lack of difference, in the feed to age. The second mechanism is the label model: GPT-4o classifies each video as harmful or not based on the video description, using a category list taken from TikTok's own Community Guidelines, with Detoxify scoring comment toxicity and a 100-video manual annotation set used to check classifier precision.","core_discovery":"The paper's central claim is empirical: when fresh accounts are registered as 13-year-olds and as adults, and then scrolled through TikTok's For You Feed under identical automated sessions in Italy, the estimated share of videos that violate TikTok's own youth-safety categories is nearly the same. In passive scrolling, fewer than 10% of videos were classified as harmful for either group, with median proportions of 0.034 for adults and 0.023 for youth. Two youth accounts stood apart, with about 14% and 25% harmful videos, but the overall distributions overlapped heavily. When accounts actively searched harm-adjacent keywords, the harmful share jumped to roughly 28% for both adults and youth,","pith_inferences":["Editorial extension: The near-null difference may be an artifact of self-reported age. If TikTok's web registration does not apply minor-specific restrictions to fresh accounts, the comparison tests account creation, not the platform's youth enforcement.","Editorial extension: The authors observed that English 'alcohol' was censored for youth accounts while Italian 'Alcol' still returned results. This suggests translation-blind moderation; an audit could map which harmful keywords are filtered across languages and find systematic gaps.","Editorial extension: Because the large-sample harm estimates rely on GPT-4o text-only labels, visuals that violate guidelines without textual clues are systematically undercounted; a multimodal relabeling could change the estimated prevalence and possibly the adult/youth difference.","Editorial extension: The two youth outlier accounts with 14% and 25% harmful shares suggest per-account variance may be driven by recommendation dynamics rather than age; a larger account sample could separate personalization noise from age effects."],"forward_implications":["If the central result holds, a self-reported age of 13 is not enough to change the mix of For You content on TikTok's web version in Italy; age-based filtering is not visibly differentiating the recommendation feed.","Active search for harm-related keywords is a stronger exposure pathway than passive scrolling, and it raises harmful exposure for adults and youth alike, so moderation policy and audits should focus on search.","Detecting harmful videos from the description alone yields precision around 59%, so automated audits need better multimodal classifiers before they can replace human annotation at scale.","Because the audit uses only public data and platform-visible accounts, the same pipeline can be rerun by external researchers or regulators to check compliance with the Digital Services Act's minor-protection requirements."],"supporting_citations":[{"why":"Provides context on what can and cannot be learned about TikTok through its research API, motivating the scraping-based audit.","marker":"(Corso et al., 2024)"},{"why":"Documents FYF personalization factors used to justify controlling for account age and behavior in the audit design.","marker":"(Boeker and Urman, 2022)"},{"why":"Prior experimental audit finding that 13-year-old accounts encountered harmful content more frequently and rapidly, the baseline this study extends and complicates.","marker":"(Eltaher et al., 2025)"},{"why":"Supplies the session-length and watch-time statistics used to calibrate the scripted passive scrolling sessions.","marker":"(Yang et al., 2025)"},{"why":"Provides the daily video-consumption estimate of about 90 videos used to set the 88 videos per account per day protocol.","marker":"(Zannettou et al., 2024)"},{"why":"Supplies the 0.5–0.7 threshold commonly used to classify a comment as toxic with Detoxify scores.","marker":"(Hua et al., 2020)"},{"why":"Provides the Fleiss Kappa interpretation used to report inter-annotator agreement on the manual harmful-content labels.","marker":"(Landis and Koch, 1977)"}],"fun_headline_variants":["TikTok youth and adult feeds show similar harm levels","Fresh 13yo and adult TikTok accounts get similar harmful feeds","Active searches spike harmful TikTok content for all ages","TikTok age filters fail in 7,000-video audit","Audit: TikTok youth accounts see no safer feed than adults"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The comparison hinges on TikTok actually applying youth-specific moderation to accounts that self-report age 13, but the paper does not verify this, for example by checking for restricted features or age verification on those accounts.","fun_headline_variants_meta":{"raw":{"variants":["TikTok youth and adult feeds show similar harm levels","Fresh 13yo and adult TikTok accounts get similar harmful feeds","Active searches spike harmful TikTok content for all ages","TikTok age filters fail in 7,000-video audit","Audit: TikTok youth accounts see no safer feed than adults"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":996,"prompt_tokens":642,"completion_tokens":354,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":386,"completion_tokens_details":{"reasoning_tokens":284}},"tokens_in":386,"tokens_out":354,"duration_ms":4233,"temperature":1.0,"reasoning_tokens":284,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T04:57:31.512859+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Register fresh 13-year-old and adult accounts on TikTok mobile and web in Italy, and before collecting any feed data check whether the youth accounts exhibit known minor restrictions such as no direct messages, safety prompts, or limited search suggestions. If they do not, the near-identical feed exposure cannot be attributed to ineffective age moderation because the accounts may not be treated as minors at all.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides context on what can and cannot be learned about TikTok through its research API, motivating the scraping-based audit."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents FYF personalization factors used to justify controlling for account age and behavior in the audit design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the session-length and watch-time statistics used to calibrate the scripted passive scrolling sessions."},{"cited_title":"Analyzing User Engagement with TikTok's Short Format Video Recommendations using Data Donations","cited_arxiv_id":"2301.04945","evidence_quote":"Provides the daily video-consumption estimate of about 90 videos used to set the 88 videos per account per day protocol."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 0.5–0.7 threshold commonly used to classify a comment as toxic with Detoxify scores."}],"review_version":1}