{"id":"2de9bd35-bb5c-476a-a6d8-40ceb7c6ef20","arxiv_id":"2508.01042","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Roughly 25% of top TikTok search results in Spain, Germany, and Poland contained synthetic AI imagery in June 2025, mostly posted by accounts specialized in AI content.","lead":"This report measured how much AI-generated imagery appears in top search results on TikTok and Instagram across three EU countries, and found about a quarter of TikTok results contained it. It also introduces a taxonomy of what it calls Agentic AI Accounts, which produce most of this content.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fresh proxy-created accounts may get cold-start TikTok rankings; the 25% AI-prevalence figure could be an artifact of collection setup rather than typical user experience.","rationale":"The paper's central quantitative claim is a prevalence estimate, so the validity of the estimate hinges entirely on whether the search results captured are the search results users would see. The methodology creates new accounts behind residential proxies precisely to obtain logged-in results, but this setup introduces a confounding variable: account freshness. Since TikTok personalizes search rankings and has no prior data on a new account, it may fall back to a default or engagement-optimized ranking that is not representative of established users; such rankings are plausibly enriched with exactly the high-volume AI-specialist content the paper measures. This threat is not speculative: the paper itself notes web/app differences, yet does not test account-age effects, and the 80% AAA claim is sensitive to the same bias because a cold-start ranking may favor accounts that post frequently and uniformly. I am not arguing the finding is false; the manual annotation protocol, disclosed disagreement rates, and cross-country consistency are genuine strengths that make the measurement credible. But a prevalence claim without ruling out a systematic collection artifact should remain conditional. The concrete A/B test is feasible and would decisively show whether the proxy-fresh-account pipeline changes the result. Other concerns (hashtag selection, broad definition of synthetic content) are real but secondary: the paper scopes claims to the chosen hashtags and defines its terms clearly. Thus verdict_should_be is UNCHANGED: the reader's CONDITIONAL verdict already reflects this concern, and the proposed test is the natural condition for acceptance.","tokens_in":32049,"tokens_out":6567,"duration_ms":83597,"concrete_test":"Repeat the TikTok data collection for the same 13 hashtags in the same countries using two conditions: (A) the original fresh-proxy protocol, and (B) accounts warmed up for 5-7 days with realistic browsing, follows, and likes before searching. Compare the top-30 result sets (overlap, rank correlation) and the manually annotated AI prevalence between conditions. If the prevalence in condition B is outside the 95% confidence interval of condition A, or if rank overlap is below ~50%, the cold-start artifact is confirmed and the 25% figure should be re-estimated. A lower-cost corroboration is a panel of 10-20 local users per country searching on their personal devices and screenshotting top results for one hashtag, to benchmark against the proxy results.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central 25% claim is measured from TikTok search results fetched with newly created accounts logged in via residential proxies (Methodology, 'Data collection'). New accounts have no watch, like, or follow history, and TikTok's search ranking is personalized and known to exhibit cold-start behavior; the platform may serve a different, often lower-quality or engagement-bait-heavy result set to unpersonalized accounts. The paper acknowledges possible web/app differences but does not address account-age or account-reputation effects, and it does not report how many accounts were created per country or whether all queries for a country were run from the same account. If fresh accounts are systematically shown more AI-specialist content (e.g., high-volume Agentic AI Accounts), then both the ~25% prevalence and the ~86% AAA share would be inflated, and the headline claim would not describe what ordinary logged-in users see.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a cross-platform, cross-country measurement of AI-generated visual content in TikTok and Instagram search results. Over two collection dates in June 2025, the authors annotated the top 30 results for 13 hashtags (politics, health, history, pope-related) in Spain, Germany, and Poland, coding each item for synthetic AI imagery (full/partial, photorealistic) and for the presence of platform or user AI labels. They find that roughly 25% of TikTok search results contain synthetic AI imagery, with lower prevalence on Instagram, and that more than 80% of TikTok AI-tagged content was posted by accounts whose last 10 posts were exclusively AI, which they term Agentic AI Accounts. They propose a taxonomy of such accounts, discuss examples of AI slop, and argue that platform labeling is insufficient and often hidden. The paper includes a detailed codebook, qualitative examples, and regulatory implications.","tokens_in":32175,"tokens_out":7194,"duration_ms":82875,"significance":"If the reported prevalence is accurate, the finding that one in four top TikTok search results contains synthetic AI imagery is important and policy-relevant, and the cross-national design is a strength. The manual annotation methodology is systematic, with dual coding and reported disagreement rates, and the public availability of the codebook and dataset supports reproducibility. However, the main quantitative claims rest on a convenience sample collected through newly created accounts, and the 'Agentic AI Account' category is defined by a threshold that makes the 80% share partly tautological. The paper is valuable as a descriptive snapshot and a taxonomy, but the headline numbers should be treated as provisional until the account-selection issue is addressed.","major_comments":[{"comment":"The headline prevalence figures (Table 8 and Executive Summary) are measured from search results obtained with newly created accounts behind residential proxies. TikTok search ranking is personalized, and fresh accounts have no watch, like, or follow history, so the set of top results may differ systematically from what established logged-in users see. The paper acknowledges web/app differences (Methodology) but does not address account-age or personalization effects, nor does it report how many accounts were created per country or whether all queries for a country used the same account. Please provide a robustness analysis (e.g., compare against a small number of long-standing accounts, or against non-logged-in top-6 results) or explicitly restrict the claim to 'search results as seen by a newly created account.'","section":"Methodology – Data collection"},{"comment":"The operational definition of an Agentic AI Account as an account whose most recent 10 posts consist exclusively of synthetic AI imagery means that the statement 'over 80% of synthetic AI content on TikTok was posted by Agentic AI Accounts' is in large part a restatement of the classification rule rather than an empirical discovery about automation. The paper further asserts that these accounts automate content creation and posting through pipelines, but no direct evidence (posting frequency, timing patterns, API signatures, or otherwise) is presented to distinguish automated from human-run content farms. Please provide such evidence, or rename the category to a more neutral term (e.g., 'AI-specialist accounts') and mark the automation inference as a hypothesis.","section":"Taxonomy of Agentic AI Accounts"},{"comment":"The paragraph following Figure 13 reports that synthetic AI content is 'statistically significantly' shared approximately 1.15 times more often than other content, but it does not report the test used, the effect size, a p-value, or a confidence interval, and no multiple-comparison correction is mentioned. Because this is the only inferential statistic in the manuscript, please provide the full test details or remove the significance claim.","section":"Generative AI content on TikTok"}],"minor_comments":[{"comment":"The text refers to 'Table 9' when reporting labeling percentages, but Table 9 is an earlier examples table; the labeling percentages actually appear in Table 16.","section":"AI Label on TikTok and Instagram"},{"comment":"The entries in Table 16 are inconsistent in formatting (e.g., '22.45' without a percent sign, '6.45 %' with a space); please standardize the notation.","section":"Table 16"},{"comment":"The text reporting the highest single-hashtag prevalence for #pope in Poland as 'over 53%' contradicts Table 15, which reports 35.71% and 37.84% for the aggregate #pope/translation in Poland; please reconcile the numbers.","section":"Generative AI content on TikTok"},{"comment":"The prevalence estimates in Table 8 are point percentages without confidence intervals; given the small number of queries, countries, and collection dates, the authors should state the uncertainty, for example by treating each hashtag-country-date combination as a cluster.","section":"Results of our research"},{"comment":"The term 'Agentic AI Accounts' is used in the executive summary before its operational definition appears later in the Methodology and Taxonomy sections; a forward reference would help readers who skip the executive summary.","section":"Executive summary"}],"recommendation":"major_revision","confidential_remarks":"This is a credible civil-society investigation with good transparency, but it is not a conventional hypothesis-testing paper. The main reviewer concern is the cold-start collection artifact; if the authors can add a robustness check or temper the executive summary, the paper could be publishable as a descriptive study. The journal should weigh whether the lack of inferential statistics and the small convenience sample meet its standards."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious look. It is one of the first cross-country, cross-platform measurements of AI-generated content prevalence in search results, and it introduces a sensible taxonomy of what it calls Agentic AI Accounts. The headline finding—about 25% of top TikTok search results for ordinary hashtags contain synthetic AI imagery—is new data, and it is presented with an uncommon level of methodological honesty: two coders, a published codebook, a reported ~4% disagreement rate, and an explicit statement that datasets and code are attached. That is real evidence and it earns credit.\n\nWhat it does well is scope itself. It does not claim to measure feeds or personalized recommendations; it measures search results for a small set of hashtags on three dates in three countries. The manual annotation is careful, and the qualitative taxonomy of slop types and AAA subtypes is a genuine conceptual contribution beyond the existing media reporting.\n\nThe soft spots are real but not fatal. The sample is narrow: 13 hashtags, top 30 results, three countries, two TikTok snapshots, one Instagram snapshot. Percentages are reported without confidence intervals, and the “growing” language in the conclusion is not supported by two time points. On Instagram the AI-content count is so small (13 posts) that any percentage is fragile.\n\nThe stress-test concern about fresh proxy-created accounts is legitimate and the paper does not fully address it. New accounts have no watch history, and TikTok search ranking is personalized and known to have cold-start behavior. If unpersonalized accounts are served more engagement-bait or AI-specialist content, the 25% figure could overstate what a typical logged-in user sees. This does not sink the paper, but it should be discussed explicitly and ideally tested by comparing a few queries from established accounts with different histories. The paper acknowledges web/app differences but not account-age effects, and it does not report how many accounts were created or whether the same account ran all queries in a country. That is the main methodological gap.\n\nThe AAA definition (last 10 posts exclusively AI) is a clean operationalization, and the claim that over 80% of AI content comes from such accounts is not circular in a harmful way—it is a structural finding about the supply of this content. The circularity burden is low.\n\nBottom line: this is a useful descriptive study for platform governance and DSA enforcement discussions. It deserves peer review, not a desk reject. A good referee should push for more detail on account creation, robustness to cold-start, and a more careful hedged conclusion about growth. I would bring it to a reading group and would cite it if I worked on AI content moderation or platform transparency.","headline":"A transparent, useful first measurement of AI slop prevalence in search results, with a plausible headline figure that reviewers should stress-test on account cold-start effects.","tokens_in":32741,"tokens_out":1868,"would_cite":true,"duration_ms":27438,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One in four top TikTok search results is AI-generated, report finds.","keywords":["synthetic AI imagery","AI slop","TikTok search results","Instagram","Agentic AI Accounts","AI labeling","algorithmic virality"],"falsifier":"A replication study using a diverse panel of real user accounts with varied browsing histories and devices, on both the web and app versions of TikTok, would settle the claim. If the share of AI content in the top 30 search results across the same hashtags and countries falls well below 25% when using such accounts, the prevalence claim would be refuted.","tokens_in":31825,"feed_emoji":"🤖","tokens_out":5209,"duration_ms":58064,"temperature":0.7,"pith_summary":"This paper attempts to establish that synthetic AI imagery is a common and mostly unlabelled presence in TikTok search results, and that most of it is produced by a small number of high-volume automated accounts. The authors manually annotated the top 30 search results for 13 hashtags across Spain, Germany, and Poland on TikTok and Instagram in June 2025. They find roughly one in four TikTok search results contains AI-made visuals, while Instagram shows far less. They also propose a taxonomy of 'Agentic AI Accounts' that post exclusively or mainly AI content via automated pipelines, and argue these accounts exploit platform algorithms. If true, this would mean recommendation systems are being gamed at scale by automated AI content, outpacing current labelling and moderation.","feed_headline":"One in four top TikTok search results is AI-generated","feed_subtitle":"Most comes from automated 'Agentic AI' accounts; fewer than half of AI videos are labelled.","key_machinery":"The central object is the 'Agentic AI Account' (AAA), defined as an account that posts synthetic AI imagery exclusively or primarily, using partial or fully automated pipelines for research, creation, and distribution. The paper classifies AAAs into Mono-Topic (one format or subject), Poly-Topic (varied trends), and Hybrid (synthetic and regular imagery with AI-generated narratives). The other key mechanism is the manual annotation codebook for detecting synthetic AI imagery, applied by two independent coders to avoid confirmation bias. Together, the taxonomy and the codebook carry the argument: the taxonomy explains who produces the AI content, and the codebook provides the evidence that the content is synthetic.","core_discovery":"On TikTok, approximately 25% of the top 30 search results for common hashtags such as #health, #history, and #trump contain synthetic AI imagery, with similar shares across Spain, Germany, and Poland. On Instagram, the share is about 4.9%. More than 80% of the AI imagery on TikTok and Instagram is photorealistic, and over 80% of AI content on TikTok is posted by what the authors call Agentic AI Accounts, defined as accounts whose recent posts consist exclusively of synthetic AI imagery produced through partial or fully automated pipelines. Only about half of AI content on TikTok is labelled as AI, and on Instagram only 3 of 13 AI posts were labelled; labels are often hidden behind clicks or absent on the web version. The paper concludes that current AI labelling practices are insufficient and that Agentic AI Accounts represent a new form of automated content production that platforms are not adequately addressing.","pith_inferences":["The prevalence numbers are likely time-sensitive: as generative video tools become widespread and watermarking remains restricted, the share of AI content in search results could grow beyond 25% in the near term.","The taxonomy of Agentic AI Accounts could be extended into a longitudinal detection method: monitoring posting rates and content homogeneity might allow platforms or researchers to flag AAAs before content goes viral, but this would require operationalizing the 'recent 10 posts exclusively synthetic' criterion into an automated classifier.","A testable extension: if platform algorithms truly favor engagement over provenance, then disabling sharing or demoting AI-posted content would disproportionately reduce the reach of Agentic AI Accounts; the paper's sharing-rate finding suggests a natural experiment for platform-level interventions."],"forward_implications":["If the 25% prevalence is representative, TikTok search results are a significant vector for AI-generated imagery on ordinary topics, not just niche AI tags.","The dominance of Agentic AI Accounts implies that a relatively small number of automated producers shape the visible AI content in search results, making moderation more tractable if platforms target such accounts.","The low labelling rate (roughly half on TikTok, 23% on Instagram) suggests current platform policies and self-disclosure are not effective, meaning DSA obligations for prominent marking are not being met in practice.","The finding that AI content is shared 1.15 times more often than non-AI content on TikTok suggests synthetic content may have a viral advantage, increasing the incentive for automated production.","Because Instagram's web version shows no AI labels, users on desktop are systematically deprived of AI disclosure information."],"supporting_citations":[{"why":"The media investigation that frames 'AI slop' as a brute-force attack on recommendation algorithms, motivating the paper's central question.","marker":"[1]"},{"why":"The human detection guidebook for synthetic AI imagery that underpins the manual annotation codebook.","marker":"[4]"},{"why":"The investigative reporting on AI slop production economies that informs the definition of Agentic AI Accounts.","marker":"[7]"}],"fun_headline_variants":["One in four top TikTok search results is AI-generated","TikTok search: 25% AI, 80% from automated accounts","AI slop makes up 25% of TikTok's top search hits","Agentic AI accounts flood TikTok search with fakes","Study: TikTok search results are 25% AI-generated"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The prevalence numbers assume that search results collected from newly created, logged-in accounts behind residential proxies in each country approximate the top results ordinary users see; if platform search results vary by account history, device, or IP reputation, the 25% figure could be an artifact of the collection setup.","fun_headline_variants_meta":{"raw":{"variants":["One in four top TikTok search results is AI-generated","TikTok search: 25% AI, 80% from automated accounts","AI slop makes up 25% of TikTok's top search hits","Agentic AI accounts flood TikTok search with fakes","Study: TikTok search results are 25% AI-generated"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000263,"raw_usage":{"total_tokens":1590,"prompt_tokens":928,"completion_tokens":662,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":575}},"tokens_in":544,"tokens_out":662,"duration_ms":8381,"temperature":1.0,"reasoning_tokens":575,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:51:58.881466+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication study using a diverse panel of real user accounts with varied browsing histories and devices, on both the web and app versions of TikTok, would settle the claim. If the share of AI content in the top 30 search results across the same hashtags and countries falls well below 25% when using such accounts, the prevalence claim would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The media investigation that frames 'AI slop' as a brute-force attack on recommendation algorithms, motivating the paper's central question."},{"cited_title":"AIF Guidebook: A Human Guide to Detecting Synthetic AI Imagery","cited_arxiv_id":null,"evidence_quote":"The human detection guidebook for synthetic AI imagery that underpins the manual annotation codebook."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The investigative reporting on AI slop production economies that informs the definition of Agentic AI Accounts."}],"review_version":1}