{"id":"b8d0e9fd-bb9c-45fd-ac4b-b6da168f88d6","arxiv_id":"2505.11160","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Simulated 13-year-old accounts on TikTok, YouTube, and Instagram encountered harmful videos more often and sooner than simulated 18-year-old accounts in a 3,000-video audit.","lead":"Researchers created age 13 and age 18 accounts on TikTok, YouTube, and Instagram, then scrolled and searched through 3,000 videos to measure how often harmful content reached minors. They report that 13-year-old accounts encountered harmful videos more often and faster than 18-year-old accounts, suggesting weak age-based moderation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 6 cannot reproduce the paper's headline percentages (e.g., YouTube 13 passive sums to about 18.51%, not 15%; YouTube 18 passive to about 9.01%, not 8.17%), and its Privacy & Sec. entries contradict Section 3.5's exclusion, so the central quantitative claim is unverifiable as reported.","rationale":"The paper addresses a timely and policy-relevant question, and the direction of its main finding is consistent with prior work on algorithmic exposure of minors, so the central claim is not implausible. However, the evidence presented for the precise quantitative version fails an internal consistency check. My hand sums of Table 6 do not match the abstract and Section 4.2 percentages: YouTube 13 passive totals about 18.51% rather than 15%; YouTube 18 passive totals about 9.01% rather than 8.17%; TikTok 13 passive totals about 8.67% rather than 7.83%. In addition, Section 3.5 says Privacy and Security was excluded from annotation, but Table 6 reports nonzero Privacy & Sec. percentages in many cells. These contradictions mean the published numbers cannot be independently checked. I also flag the single-account design in Section 3.4: one account per age-by-mode cell provides no estimate of account-to-account variability, and the time-to-first-harmful values are single observations. The reader's weakest_assumption, adult-annotator validity, is a genuine threat and is self-acknowledged in Section 3.6, but I treat the numerical mismatch as the more immediate blocker: even perfect adult labels would not make the reported percentages reproducible. The legal-safeguards section is largely a legislative review rather than an evaluation, but that is secondary. These issues are correctable through data release, reanalysis, and additional accounts, so I do not move the reader's CONDITIONAL verdict; I would make raw-data release and the corrected aggregation a hard condition, and would move to REJECT if reanalysis does not recover the reported pattern.","tokens_in":17704,"tokens_out":15680,"duration_ms":132846,"concrete_test":"Recompute the row totals for every Platform/Age/Method block in Table 6 from the published entries and compare them with the percentages in the abstract, Section 4.1, and Section 4.2; in the same pass, check whether any nonzero Privacy & Sec. entries remain, given Section 3.5's exclusion statement. If the totals differ, request the raw per-video annotation data and the aggregation script, then regenerate all summary percentages and time-to-first-harmful values from the raw data. This settles whether the headline numbers are supported by the paper's own data or are artifacts of an undefined denominator or an excluded-category inconsistency.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is internal inconsistency in the supporting data. The abstract and Section 4.2 report specific rates, such as YouTube 13 passive 15%, YouTube 18 passive 8.17%, and TikTok 13 passive 7.83%; Section 4.6 summarizes minors at 7.83-15%. These cannot be reproduced from Table 6, the only disaggregated results table. Summing its Low/Medium/High rows by block gives, for example, YouTube 13 passive about 18.51% (Low 18.01 + Medium 0.50 + High 0.00), YouTube 18 passive about 9.01%, TikTok 13 passive about 8.67%, and Instagram 13 passive about 16.50%. The discrepancies are 0.5-3.5 percentage points, far beyond rounding. Additionally, Section 3.5 states that Privacy and Security and Enforcement Actions were excluded from annotation, yet Table 6 contains a Privacy & Sec. column with nonzero values in nearly every block. Either the table uses a different denominator or scope, or the aggregation is wrong; in neither case can the reader verify the headline percentages. A further structural limitation is Section 3.4's single account per age-by-mode cell, so the 13-versus-18 differences and time-to-first-harmful values (e.g., YouTube 3:06) are single draws with no variance estimate or statistical test. The Section 3.6 adult-label caveat is valid but secondary to these numerical problems.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an experimental audit of content moderation on TikTok, YouTube, and Instagram using sock-puppet accounts aged 13 and 18. For each platform, the authors created accounts for each age and, for TikTok and YouTube, for both passive scrolling and search-based scrolling, while Instagram was limited to passive scrolling. They collected approximately 3,000 recommended videos and manually labeled each as harmful or not using a unified taxonomy derived from the platforms' community guidelines, also assigning severity levels. The central finding is that 13-year-old accounts encounter harmful videos more frequently and more quickly than 18-year-old accounts; for example, 15% of YouTube recommendations to 13-year-old accounts during passive scrolling were harmful versus 8.17% for 18-year-old accounts, and the first harmful video appeared at 3:06 minutes for the younger group. The paper concludes that platform moderation is materially weaker for minors and that stronger age verification and enforcement are needed.","tokens_in":17934,"tokens_out":8662,"duration_ms":65440,"significance":"If the results were verifiable, this would be a useful empirical contribution to an active policy debate. The study covers three major platforms, uses a unified harm taxonomy, and attempts to compare both passive and active exposure modes, addressing gaps in prior work that focused on single platforms or single harm categories. The authors are transparent about the subjective nature of harm and explicitly acknowledge the adult-labeling limitation. However, the manuscript's quantitative claims are currently undermined by internal inconsistencies in the only disaggregated results table and by the absence of any statistical inference, so the headline conclusions cannot be relied upon as reported.","major_comments":[{"comment":"The reported headline percentages cannot be reproduced from the paper's only disaggregated results table. Summing the Low, Medium, and High rows for YouTube 13 Passive in Table 6 gives 18.51% (Low 18.01% + Medium 0.50% + High 0.00%), not the 15% stated in the abstract and Section 4.2. Similarly, YouTube 18 Passive sums to 9.01% rather than 8.17%, and TikTok 13 Passive sums to 8.67% rather than 7.83%. These discrepancies are far beyond rounding error and mean that the central quantitative claim of the paper cannot be verified from the presented data.","section":"Section 4.2 / Table 6"},{"comment":"Section 3.5 explicitly states that Privacy and Security and Enforcement Actions were excluded from manual annotation, yet Table 6 contains a 'Privacy & Sec.' column with nonzero values in several blocks, including TikTok 13 Passive Low (0.17%), TikTok 18 Passive Low (0.17%), and Instagram 13 Passive Low (0.17%). This contradicts the described labeling scope and suggests either a different annotation procedure or an undisclosed change in category handling, further undermining the table's validity.","section":"Section 3.5 / Table 6"},{"comment":"The experimental design uses a single account per age-by-mode cell (one 13-year-old passive, one 13-year-old search-based, one 18-year-old passive, one 18-year-old search-based per platform). Consequently, all comparisons between age groups and interaction modes, including the time-to-first-harmful-video values (e.g., YouTube 3:06 vs 1:28), rest on single draws with no variance estimate or statistical test. Without replication or error bars, the paper's comparative claims (e.g., 'children commonly encounter harmful content in under five minutes, compared to roughly nine minutes for adults') are not supported.","section":"Section 3.4 / Sections 4.1-4.4"},{"comment":"The study's outcome measure relies on adult annotators applying a taxonomy derived from platform policies, yet the paper's central claim concerns what is harmful for 13-year-olds. Section 3.6 acknowledges this disconnect, but no evidence is provided that adult severity ratings correspond to adolescent perceptions. Because every quantitative result (the 15% vs 8.17% rates, time-to-first-harm figures, and severity distributions) depends on these adult labels, the validity of the measure is load-bearing and is not established in the manuscript.","section":"Section 3.6 / Section 5"}],"minor_comments":[{"comment":"The phrase 'videos moderation' should be 'video moderation'.","section":"Section 4.6"},{"comment":"The abbreviation 'RSD' is used without definition; the manuscript should spell out 'Regulation (EU) 2022/2065 (Digital Services Act)' at first use.","section":"Section 2.2.2"},{"comment":"The reference to 'the Methodology Section' is vague; the authors should cite Section 3.2 specifically when describing the harmful content framework.","section":"Section 3.5"},{"comment":"The paper states that inter-annotator agreement was verified by two additional experts, but it reports no agreement statistics (e.g., Cohen's kappa) or counts of disagreements, making the reliability of the labeling process difficult to assess.","section":"Section 3.5"},{"comment":"The title and abstract emphasize 'legal safeguards,' but the empirical study does not evaluate any legal instrument; the legal discussion is a literature review only. The authors should either narrow the scope or adjust the framing to accurately reflect the absence of a legal intervention.","section":"Title / Abstract"},{"comment":"The heatmap colors described in the text are not visible in a black-and-white printout; the table should be interpretable without color, for example by adding numeric labels or distinct patterns.","section":"Table 6"}],"recommendation":"major_revision","confidential_remarks":"The table inconsistency is the most serious issue and must be resolved before the paper can be considered further. The n=1 per cell design is a fundamental limitation that limits the strength of the claims; even if the numbers are corrected, the authors should explicitly frame the results as descriptive rather than inferential. The paper would benefit from making the raw data and annotation code available for reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful study to have in the conversation, but the numbers as reported don't add up. The direction—13-year-old accounts get harmful content faster and more often than 18-year-old accounts—matches prior single-platform work and is probably right. But the abstract and Section 4.2 cite rates like YouTube 13 passive 15% and YouTube 18 passive 8.17%, and summing the categories in Table 6 gives roughly 18.5% and 9.0%. That is not a rounding issue. Table 6 also includes nonzero Privacy & Security entries even though Section 3.5 explicitly excludes those categories from annotation. So a reader cannot verify the central claim from the paper's own data.\n\nWhat's genuinely good: the cross-platform design, the unified harm taxonomy, the 13-versus-18 comparison, the inclusion of Instagram, the passive-versus-search mode comparison, and the time-to-first-harmful measure. These are real additions over much of the literature, which tends to be single-platform or single-category. The annotation protocol with multi-labeller adjudication is thoughtful, and they openly flag the adult-label caveat, the non-English exclusion, and the minimal-engagement design.\n\nSoft spots beyond the arithmetic: there is only one account per age-by-mode cell, so the 13-versus-18 differences and the 3:06 time-to-harm are single draws with no variance estimate or statistical test. The 'legal safeguards' part of the title overreaches—the paper reviews DSA, OSA, and KOSA but does not evaluate their effectiveness empirically. And the 'speed' result depends on labels that stream through adult annotators, a gap they acknowledge. All of this is fixable with reanalysis, data release, and more careful wording. The arithmetic inconsistency is the real blocker; without a corrected Table 6 or raw data, the precise rates in the abstract cannot be taken at face value.\n\nWho this is for: regulators, platform-safety researchers, and anyone tracking the DSA/OSA/KOSA debate who wants a current cross-platform snapshot. The paper deserves a serious referee—the topic is important and the design is worth engaging with—but the current version should not be published as-is. I would ask for the raw data, corrected tables, and an analysis that treats each account as one observation rather than hundreds of videos. I would not cite the headline numbers until they are reproduced.","headline":"A policy-relevant cross-platform audit with a plausible direction of effect, but the paper's own data table cannot reproduce its headline percentages, so the quantitative claims are not yet usable.","tokens_in":18566,"tokens_out":1942,"would_cite":false,"duration_ms":19516,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Simulated 13-year-old accounts are served harmful videos more often and sooner than accounts declared as 18.","keywords":["Social Media","Content Moderation","Online Harm","Algorithmic Transparency","Child Safety","Age-Restricted Content","Platform Policies","recommendation algorithms"],"falsifier":"Re-run the audit with the same 3,000 videos rated by panels of adolescents, with annotators blind to which account age each video came from: if the YouTube 15% versus 8.17% passive-scroll gap shrinks or reverses under teen harm ratings, the paper's conclusion that platforms fail younger users through recommendation feeds is an artifact of adult labelling. A second check would be platform-side transparency data showing 13-year-old accounts receiving equal or lower harmful-recommendation rates than 18-year-old accounts under fresh-account conditions.","tokens_in":17443,"feed_emoji":"📱","tokens_out":6636,"duration_ms":62457,"temperature":0.7,"pith_summary":"The paper sets out to test whether age-based content moderation on the three largest short-video platforms actually protects younger users. It creates fresh accounts declared as 13 and 18 years old, scrolls them through 3,000 videos under passive and search-based conditions, and labels every video against a unified harm framework drawn from the platforms' own community guidelines. The central claim is that the 13-year-old accounts consistently encounter videos judged harmful more frequently and more quickly than the 18-year-old accounts: on YouTube alone, 15% of passive recommendations to the younger accounts were rated harmful versus 8.17% for adults, with first harmful exposure at about three minutes. If this holds, it matters because video is now the main way minors use these platforms and because regulation in the EU and UK assumes platforms are already mitigating minors' exposure to harmful recommendations.","feed_headline":"Audit: simulated teens meet harmful videos faster than adult accounts","feed_subtitle":"A 3,000-video test across YouTube, TikTok and Instagram finds minors' feeds riskier within minutes.","key_machinery":"The load-bearing instrument is a Unified Harmful Content Framework, a taxonomy assembled by merging the community guidelines of YouTube, Instagram, and TikTok into one list of harm categories with severity levels (low, medium, high). Around it the authors build a controlled audit: two fresh accounts per platform per age (13 and 18), two interaction modes (pure passive scrolling and search-based scrolling with normal then low-risk keywords), fixed 20-second view durations, and a four-annotator adjudication pipeline for labelling. The framework converts each platform's policy language into comparable measurement, so that a percentage of recommended videos rated harmful, a time-to-first-harmful-video, and a category-by-severity distribution can be computed across platforms. The mechanism's key move is comparing exactly matched account types whose only deliberate difference is declared age and scrolling mode.","core_discovery":"On the paper's own terms, the discovery is that accounts declared as 13 do not receive safer recommendation feeds than accounts declared as 18; they receive riskier ones. Across platforms and interaction modes, 13-year-old accounts saw harmful content in 7.83% to 15% of videos, while 18-year-old accounts saw 4.67% to 8.33%, and the gap appeared without any user searches, likes, or follows. YouTube was the clearest case: passive scrolling recommended harmful videos to 15% of the minor feed versus 8.17% of the adult feed, and the first harmful video appeared after an average of 3:06 minutes for minors versus roughly nine minutes for adults. The paper also finds that low-severity harm dominates the harmful content served to both ages, with Sensitive and Mature Themes the most common category, and it interprets the overall pattern as evidence that recommendation systems amplify rather than suppress harmful material for minors.","pith_inferences":["Our inference: if the age gap is real, strengthening age verification alone will not fix it, because the accounts in this study truthfully declared age 13 and still received more harmful recommendations; the feed composition itself would have to change.","Our inference: the result suggests an engagement-based explanation rather than a content-policy one: recommendation algorithms may optimise watch time over safety, and if 13-year-olds' viewing patterns differ, the same core algorithm will surface different content even under identical moderation rules.","Our inference: the adult-labelling caveat could be turned into a direct test by having adolescents rate the same videos; such participatory labelling would likely shift severity boundaries and could reveal whether the 15% versus 8.17% gap is experienced by teens as a harm gap or as a tone gap.","Our inference: the method is repeatable as a lightweight regulatory audit; a small number of seeded accounts with fixed protocols can benchmark whether a platform's protections for minors improve after policy changes."],"forward_implications":["If the finding is correct, YouTube, TikTok, and Instagram's age-based filtering is not delivering the protection their community guidelines promise for the youngest permitted users, at least for fresh accounts with minimal interaction.","Regulators auditing under the EU Digital Services Act or the UK Online Safety Act would have a concrete reason to demand recommender-level transparency: today's published enforcement reports count removals, while this study measures what the default feed actually serves within minutes.","Low-severity harmful content, not just extreme material, would need to be treated as a moderation target, since it dominates what minors receive and repeated exposure may normalize harm.","The time-to-first-harm measure implies that moderation quality cannot be judged by aggregate take-down volumes; a child's first minutes on a platform are the decisive exposure window.","Search-based behavior is not uniformly riskier: on YouTube searching reduced minors' harmful exposure sharply from 15% to 8%, while TikTok stayed flat at 7.83%, so platform-specific algorithm audits are needed rather than one-size-fits-all conclusions."],"supporting_citations":[{"why":"Supplies the finding that age verification on social media is easily bypassed, and supports the study's account-creation approach and its call for robust verification.","marker":"[29]"},{"why":"Establishes the prior baseline that inappropriate content can be reached from safe YouTube content within ten recommendations, motivating the exposure-rate measurement.","marker":"[31]"},{"why":"Prior demonstration that simulated 13-year-old TikTok accounts encounter mental-health related harmful content within hours of use.","marker":"[33]"},{"why":"Prior experiment using simulated 13-year-old TikTok accounts to measure stereotype exposure, supporting the seeded-account methodology.","marker":"[36]"},{"why":"Prior timing result showing suicide-related content on TikTok within 2.6 minutes, which motivates the study's speed-to-harm measure.","marker":"[38]"},{"why":"Prior audit of a 13-year-old TikTok account showing algorithmic promotion of self-harm and suicide content, directly comparable to this study's minor-account design.","marker":"[40]"},{"why":"One of the three platform policy documents merged into the Unified Harmful Content Framework.","marker":"[41]"},{"why":"TikTok's community guidelines are used both as input to the harm taxonomy and as the policy baseline against which enforcement is compared.","marker":"[42]"},{"why":"YouTube's community guidelines are used both as input to the harm taxonomy and as the policy baseline against which enforcement is compared.","marker":"[43]"}],"fun_headline_variants":["Simulated 13-year-olds see more harmful videos than 18-year-olds","YouTube serves minors harmful content faster, audit finds","13-year-old sims get harmful content faster and more often","Younger sim accounts face harmful videos sooner, 3,000-video test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results rest on adult researchers' manual ratings of what is harmful and how severe it is; if adults and 13-year-olds systematically disagree about those judgments, the age-group gap could reflect adult perceptions rather than actual risk to minors.","fun_headline_variants_meta":{"raw":{"variants":["Simulated 13-year-olds see more harmful videos than 18-year-olds","YouTube serves minors harmful content faster, audit finds","13-year-old sims get harmful content faster and more often","Younger sim accounts face harmful videos sooner, 3,000-video test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001326,"raw_usage":{"total_tokens":5447,"prompt_tokens":1047,"completion_tokens":4400,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":663,"completion_tokens_details":{"reasoning_tokens":4325}},"tokens_in":663,"tokens_out":4400,"duration_ms":29766,"temperature":1.0,"reasoning_tokens":4325,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:56:09.669165+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the audit with the same 3,000 videos rated by panels of adolescents, with annotators blind to which account age each video came from: if the YouTube 15% versus 8.17% passive-scroll gap shrinks or reverses under teen harm ratings, the paper's conclusion that platforms fail younger users through recommendation feeds is an artifact of adult labelling. A second check would be platform-side transparency data showing 13-year-old accounts receiving equal or lower harmful-recommendation rates than 18-year-old accounts under fresh-account conditions.","supporting_citations":[{"cited_title":"The digital loophole: Evaluating the effectiveness of child age verification methods on social media","cited_arxiv_id":null,"evidence_quote":"Supplies the finding that age verification on social media is easily bypassed, and supports the study's account-creation approach and its call for robust verification."},{"cited_title":"Disturbed youtube for kids: Characterizing and detecting inappropriate videos targeting young children","cited_arxiv_id":null,"evidence_quote":"Establishes the prior baseline that inappropriate content can be reached from safe YouTube content within ten recommendations, motivating the exposure-rate measurement."},{"cited_title":"Driven into the darkness: How tiktok encourages self-harm and suicidal ideation, 2023","cited_arxiv_id":null,"evidence_quote":"Prior demonstration that simulated 13-year-old TikTok accounts encounter mental-health related harmful content within hours of use."},{"cited_title":"Surveilling young people online: An investigation into tiktok’s data processing practices, 2021","cited_arxiv_id":null,"evidence_quote":"Prior experiment using simulated 13-year-old TikTok accounts to measure stereotype exposure, supporting the seeded-account methodology."},{"cited_title":"Deadly by design: Tiktok pushes harmful content promoting eating disorders and self-harm into young users’ feeds, 2022","cited_arxiv_id":null,"evidence_quote":"Prior timing result showing suicide-related content on TikTok within 2.6 minutes, which motivates the study's speed-to-harm measure."},{"cited_title":"Suicide, incels, and drugs: How tiktok’s deadly algorithm harms kids, 2023","cited_arxiv_id":null,"evidence_quote":"Prior audit of a 13-year-old TikTok account showing algorithmic promotion of self-harm and suicide content, directly comparable to this study's minor-account design."},{"cited_title":"Community standards, 2024","cited_arxiv_id":null,"evidence_quote":"One of the three platform policy documents merged into the Unified Harmful Content Framework."},{"cited_title":"Community guidelines overview, 2024","cited_arxiv_id":null,"evidence_quote":"TikTok's community guidelines are used both as input to the harm taxonomy and as the policy baseline against which enforcement is compared."},{"cited_title":"Community guidelines, 2024","cited_arxiv_id":null,"evidence_quote":"YouTube's community guidelines are used both as input to the harm taxonomy and as the policy baseline against which enforcement is compared."}],"review_version":1}