{"id":"b680eb46-a204-4389-8125-9fb9bc2738eb","arxiv_id":"2504.18140","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A replication of prior TikTok recommender audits finds poor reproducibility, short-lived findings, and strong dependence on evaluation metrics.","lead":"This paper tried to rerun earlier studies that test TikTok's recommendation algorithm with automated fake accounts, and found those studies are hard to reproduce and their conclusions go stale quickly. It matters because regulators such as the EU rely on such audits, and this work shows why audit results need standardized, repeatable methods.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'watch is strongest' result rests on an unverified GDPR inference: absent explicit actions in the data export does not prove the recommender never received them.","rationale":"The reader's weakest assumption is exactly the GDPR-based inference about explicit actions, and I agree that it is the most load-bearing weakness in the paper's specific RQ2 findings. The paper is an honest replication study with released code and data, and the central qualitative message, that one-shot algorithmic audits of TikTok are poorly reproducible and quickly outdated, is supported by the documented obstacles: missing code, inaccessible repositories, platform changes, bot bans, and metric sensitivity. However, the concrete claim that watch is the strongest personalisation factor relative to like and follow depends on the premise that explicit actions were actually transmitted to or processed by the recommender in the web sessions. The GDPR export check is the sole evidence for this premise, and it is not conclusive because GDPR export contents, action reversal, and separate interaction pipelines can all produce the same observation. The paper partially acknowledges this in Section 5, but the abstract and Section 7 present the finding without the caveat. I therefore do not change the reader's conditional verdict: the overall contribution is valuable and likely correct at the qualitative level, but the specific factor ranking needs a direct experimental check before it can be considered established.","tokens_in":16285,"tokens_out":4666,"duration_ms":54136,"concrete_test":"Run a controlled web-interface experiment with fresh accounts: (a) like a fixed set of topic-defined videos, (b) follow creators who post that topic, and (c) a control with no explicit actions. Keep all actions active and do not reverse them; record the hashtag similarity of the next 200 recommended videos to the liked or followed topics. If (a) and (b) show no detectable shift relative to (c), the GDPR-based inference is supported. If they do shift, the 'watch is strongest' conclusion and the explanation for missing GDPR records are both wrong. As a second check, request a GDPR export from a human web-only account with active, unreversed likes and follows; if explicit actions appear in that export, the paper's claim that web explicit actions are not recorded is refuted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most concrete new finding, that the watch action provides the strongest personalisation impact, depends on the claim that explicit like and follow actions are weak or ineffective in the web interface. The only evidence for this is the GDPR export check described in Section 4, where explicit actions were missing for most accounts while watch history matched almost exactly. The authors infer that the recommender may not have taken these explicit actions into consideration. But absence from a GDPR export is not equivalent to non-use by the recommender: TikTok may store explicit interaction signals in a separate pipeline or export category, the bot's actions may have been reversed before the data request (the ethics section states that reversible actions were undone after the study), or the web interface may log actions for anti-abuse purposes without including them in the user-facing export. The paper itself hedges in Section 5, noting that the explicit-versus-implicit comparison 'may be biased by the recommender not taking explicit actions into consideration.' Because the abstract and Section 7 state the watch-action finding without that caveat, the ranking of personalisation factors is not established. This does not threaten the broader reproducibility argument, which is well supported by the documented process, but it does mean the specific RQ2 finding should not be cited as a stable property of TikTok's recommender until the GDPR inference is independently confirmed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper re-runs and extends prior TikTok sockpuppeting audits (Boeker and Urman 2022; Vombatkere et al. 2024; partly Mousavi et al. 2024) in January-February 2025. For RQ1 it documents concrete reproducibility barriers: missing or unusable released code, incomplete methodology descriptions, content and platform change, and bot bans or proxy failures. For RQ2 it compares location, watch duration, liking, and following via bot scenarios and two hashtag-similarity metrics plus video popularity, concluding that watch action now provides the strongest personalization signal, with like and follow showing an initial exploration phase. The authors release code and data and argue for longitudinal, more authentic, reproducible audits.","tokens_in":16432,"tokens_out":3172,"duration_ms":35370,"significance":"If the RQ1 findings are taken at face value, the paper makes a valuable contribution to the algorithmic-auditing literature by showing, in a documented multi-month replication effort, that current one-shot sockpuppeting audits are hard to reproduce and that conclusions are time-sensitive and metric-sensitive. The public release of code and data is a clear strength and should be credited. The RQ2 ranking, however, is the weaker part of the paper: several scenarios ran only once, no uncertainty quantification is reported, the main 'watch is strongest' claim depends on an unverified GDPR-export inference, and the authors themselves show that metric choice flips conclusions. The paper is honest in its Limitations section, but the abstract and Section 7 state the ranking without the same caveats. As a replication study, the central reproducibility argument is sound; as a claim about the current causal effect of user actions on TikTok personalization, it needs substantially more support.","major_comments":[{"comment":"The claim that explicit actions (like and follow) are not taken into consideration by the recommender rests on an inference from GDPR data exports: Section 4 states that 'while the watch history had an almost exact match, the explicit actions (like and follow) were missing for most of the accounts.' Absence from a GDPR export does not establish that the actions never reached the recommender; the signals could be stored in a separate pipeline, logged for anti-abuse purposes but excluded from the user-facing export, or affected by the study's own reversibility procedure described in Section 6. The paper itself hedges in Section 5 ('may be biased by the recommender not taking explicit actions into consideration'), but the abstract and Section 7 present 'the watch action provides the strongest personalisation impact' without that caveat. Because this finding is load-bearing for RQ2 and is cited in the contributions, the authors should either verify the inference (e.g., by running a mobile-app audit or otherwise confirming that explicit web actions are absent from the recommender's inputs) or reframe the ranking as an observation conditional on incomplete knowledge of how web actions are processed.","section":"Section 4 and Section 5"},{"comment":"The quantitative personalization comparisons lack error bars and significance tests, while several scenarios were run only once (Table 1, 'Rep.' column). This is especially problematic because the noise level is high: Section 5 reports that feed similarity between two control users ranges from 2% to 28% (average 11%). Against this baseline, reported differences such as the -2.06% watch effect and the +9.06% random-like effect in Table 3 are not convincingly distinguishable from noise. Bot bans, proxy misconfiguration, and incomplete sessions (acknowledged in Section 6) further reduce the effective sample for some scenarios. Without confidence intervals, significance tests, or at least per-run variability, the claim that watch action is 'the strongest' personalisation factor is not established by the reported measurements.","section":"Section 5, Table 1, Table 3"},{"comment":"The paper demonstrates a strong dependence on the evaluation metric: Table 2 (video popularity) shows no consistent personalization advantage for the personalised user (e.g., random implicit feedback gives -2.11% for the personalised user versus -37.60% for the control), while Table 3 (hashtag basic match similarity) is used to conclude that watch is strongest. The authors acknowledge in the text and in Section 6 that changing the metric or its strictness can lead to 'completely different findings.' Given this acknowledged metric sensitivity, choosing one metric as the basis for the headline ranking is not robust. The authors should either report the ranking across all metrics and discuss disagreement explicitly in the conclusions, or restrict the headline claim to the specific metric used.","section":"Section 5, Table 2 vs. Table 3"}],"minor_comments":[{"comment":"In the watch-duration scenario list, the items are numbered '1) 50% [S9]; 2) 200% [S10]; or 4) 400% [S11]'; the numbering skips '3)'. Please fix the enumeration.","section":"Section 3, Table 1"},{"comment":"The statement that 'some of the bots (in 5 cases) were banned' is followed by '15 more accounts were banned after finishing the audit study.' It would be clearer to state the total number of accounts, the scenarios in which bans occurred, and the timing relative to data collection.","section":"Section 4"},{"comment":"The claim that 'there are no audits on Instagram or YouTube Shorts' is too absolute unless it is explicitly scoped to the systematic review's search date and method; please add the cutoff or qualify the statement.","section":"Section 4"},{"comment":"The y-axis label 'ratio' is not defined in the figure captions; please state that it is the fraction of videos containing at least one predefined hashtag or substring, as described in the text.","section":"Figures 2 and 3"},{"comment":"The Limitations paragraph lists three important threats (hashtag-based metrics, generic hashtags, and flagged bots) but does not explicitly connect them to the specific claims in Section 5; adding one or two sentences mapping each limitation to the affected finding would improve transparency.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The RQ1 reproducibility argument is well supported by the documented process and by the authors' decision to release code and data. The RQ2 ranking, however, is the main point of risk: it depends on an unverified GDPR inference, has no uncertainty quantification, and is metric-dependent. I would not advise rejecting the paper, because the reproducibility findings are valuable and the limitations are partially acknowledged; but the headline claims in the abstract and Section 7 need to be either strengthened with additional evidence or explicitly conditioned on the methodological caveats."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a replication of two TikTok sockpuppeting audits (Boeker & Urman, Vombatkere et al.), run in early 2025 with released code and data. The real contribution is the documented failure mode: one-shot audits of this kind are hard to reproduce, and findings are fragile with respect to metric choice and time. That message comes through clearly and is supported by the authors' own process.\n\nWhat it does well: it doesn't just say \"we couldn't reproduce\"; it catalogs the specific barriers—missing or incomplete code from reference studies, platform changes (livestreams, ads, HTML), bot bans, and the need to reverse-engineer details. The GDPR post-audit check is a genuinely useful idea for detecting whether bot actions were actually registered. The public release of code and data is commendable and makes the replication effort checkable. The paper also shows, convincingly, that changing the hashtag similarity metric flips the conclusion—that alone is a valuable methodological warning.\n\nThe soft spots are real but mostly in the quantitative RQ2 layer. The claim that \"watch is the strongest personalisation factor\" rests on a shaky inference: the absence of like/follow actions in the GDPR export is taken as evidence that the recommender never received them. That is plausible but unverified—separate pipelines, reversible actions undone after the study, or anti-abuse logging could explain the gap. The paper hedges in Section 5 but the abstract and contributions state the finding without the hedge. Also, most scenarios ran once, there are no variance estimates or significance tests, and several sessions were incomplete due to bans. The paper acknowledges these limitations honestly, which is to its credit, but it means the specific ranking of personalisation factors should not be cited as an established property of TikTok's recommender.\n\nThe central reproducibility argument, however, holds up. I would accept this for peer review—a serious referee can push for the caveats and sensitivity analysis, but the core message deserves to be out there. The paper is useful for anyone doing algorithmic audits, especially those planning sockpuppeting studies on TikTok or other short-video platforms.","headline":"A valuable, honest replication study whose central reproducibility argument holds up; the specific 'watch is strongest' finding needs a caveat until the GDPR inference is independently confirmed.","tokens_in":17054,"tokens_out":1408,"would_cite":true,"duration_ms":14236,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sockpuppeting audits of TikTok fail to reproduce, and their findings on what drives personalisation hold only in the short term.","keywords":["algorithmic audit","sockpuppeting","TikTok","reproducibility","personalisation","recommender systems","GDPR data access","short-term validity"],"falsifier":"Run the same watch-duration versus explicit-action scenarios through TikTok's mobile app while intercepting network traffic to record every request. If likes and follows sent from the app change the feed as much as or more than watch duration, the paper's claim that explicit actions have little effect on the web and that watch is the strongest signal is contradicted. Alternatively, recompute the paper's comparisons using human-annotated video topics instead of hashtag similarity; if the ranking of personalisation factors changes, the watch-dominance finding is a metric artifact.","tokens_in":16009,"feed_emoji":"🎭","tokens_out":8779,"duration_ms":83417,"temperature":0.7,"pith_summary":"This paper tries to rerun earlier sockpuppeting audits of TikTok's For You page—audits that create automated fake users to see what steers recommendations—and asks whether their results survive contact with a changed platform. It reports that they do not: the earlier audits are poorly reproducible because code, data, and procedural details were missing, and because TikTok's content, interface, and bot detection moved on. Rerunning the scenarios in early 2025, the paper finds that the watch action now provides the strongest personalisation signal, stronger than liking or following, and that the effect grows with watch duration. It also finds that conclusions flip depending on the evaluation metric chosen, so the one-shot audit approach itself is the problem. If the paper is right, regulators cannot rely on a single audit snapshot; they need reproducible, longitudinal audit methods.","feed_headline":"TikTok audit findings don't reproduce and go stale fast","feed_subtitle":"Re-running earlier audits three years later flips which signals matter: watching videos now beats liking or following.","key_machinery":"The load-bearing machinery is the paired-control sockpuppet audit. Two fresh TikTok accounts, identical except for one manipulated personalisation factor (location, watch duration, liking, or following), scroll 250 For You videos per session for four sessions; the feeds are compared through video popularity trends and hashtag similarity (strict Jaccard and a lenient substring-based 'basic match'). A post-audit data request under European privacy law is then used to check whether the bot's actions were actually logged, which is what exposes the missing explicit actions on the web interface. This machinery lets the paper separate what the recommender did from what the audit recorded, and it is the basis for both the reproducibility critique and the metric-sensitivity claim.","core_discovery":"The paper's central discovery is that the standard sockpuppet audit, applied to TikTok's For You page, fails a basic reproducibility check: trying to rerun two earlier audits after a gap of roughly three years required about nine person-months, extensive reverse-engineering, methodological fixes, and still produced different conclusions. In the new data, the implicit watch action has the strongest personalisation impact, with longer or repeated watching strengthening the effect, whereas likes and follows show an exploration phase followed by exploitation after around 1,000 videos. The previous hierarchy—follow first, watch about as strong as like—no longer holds. A post-audit check via user-requestable privacy data shows that explicit like and follow actions are missing from the recorded histories for most web-interface accounts, suggesting these actions may not reach the recommender at all on the web, which would bias any audit run through the browser. Finally, changing the evaluation metric (strict Jaccard vs. lenient basic match, or video popularity) is enough to reverse the apparent conclusions, so the paper frames its findings as evidence that algorithmic-audit findings are short-lived and method-dependent.","pith_inferences":["An implication the paper leaves open: if explicit actions really are dropped on the web pipeline, the earlier 'follow is strongest' finding may have been an artifact of the web interface, and a mobile-app replication could restore the old ranking.","Relatedly, the paper's 'watch is strongest' result may describe the web pipeline more than the recommender itself; on mobile, where likes and follows are recorded, the ranking of personalisation factors could differ.","The observed switch from exploration to exploitation for likes at roughly 1,000 interactions suggests a practical audit threshold: runs shorter than that may systematically understate explicit-action influence, so run length should be reported and treated as a boundary condition.","The interest-ratio landing at 36%, the low end of the earlier 30–50% band, could be reused as a moving benchmark: future audits that fall outside the band would trigger suspicion of an algorithm change before any deeper analysis."],"forward_implications":["Any single-run audit that releases no code, no data, and no precise scenario details cannot be independently verified, so regulators cannot tell whether a changed result reflects a changed algorithm or a changed audit.","Findings about which user actions drive TikTok personalisation carry an expiration date: the same scenarios that put follow first in earlier audits put watch first in early 2025, so one-shot studies should be labelled with the date and platform state they captured.","If explicit actions are indeed missing from web-interface signal paths, future TikTok audits should run through the mobile app or verify recorded actions with user-requestable data before drawing conclusions about likes and follows.","Audit conclusions are metric-dependent: strict hashtag Jaccard, lenient basic match, and video-popularity trends can point in opposite directions, so an audit should report several evaluation measures and make its choice explicit.","Because feed diversity has risen sharply since earlier audits (similarity between two bots doing the same thing fell from about 35% to about 10–11%), personalisation effects are now harder to detect and audits need larger sample sizes or stronger scenarios."],"supporting_citations":[{"why":"Defines the original personalisation scenarios and the earlier finding that following is the strongest signal, which this paper reruns and overturns.","marker":"[4]"},{"why":"Supplies the exploration-versus-exploitation framing and the 30–50 percent personalised-feed baseline used for comparison.","marker":"[28]"},{"why":"Provides the hashtag-similarity evaluation measures (strict Jaccard and lenient basic match) reused in the analysis.","marker":"[16]"},{"why":"Shows how user-requestable platform data can be used to inspect real traces, the approach adapted here to verify whether bot actions were recorded.","marker":"[30]"},{"why":"Demonstrates that audit methodology choices change YouTube audit findings, motivating the paper's metric-sensitivity analysis.","marker":"[5]"},{"why":"Argues for continuous and automated social media audits, the direction the paper's reproducibility findings support.","marker":"[22]"},{"why":"Documents the practical obstacles of auditing TikTok through the mobile app, the alternative the paper points to for web-interface signal loss.","marker":"[13]"}],"fun_headline_variants":["TikTok audit findings don't reproduce or last","TikTok audits: poor reproducibility, short-term validity","Re-running TikTok audits flips the findings","TikTok audit results fade and fail to replicate","Watching TikTok now trumps likes, but only briefly"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The explanation for the like and follow results assumes the user-requestable data file is a complete record of every action the account performed, so an action missing from it was never received by the recommender rather than logged elsewhere or filtered out for bot accounts.","fun_headline_variants_meta":{"raw":{"variants":["TikTok audit findings don't reproduce or last","TikTok audits: poor reproducibility, short-term validity","Re-running TikTok audits flips the findings","TikTok audit results fade and fail to replicate","Watching TikTok now trumps likes, but only briefly"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000697,"raw_usage":{"total_tokens":3170,"prompt_tokens":985,"completion_tokens":2185,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":2109}},"tokens_in":601,"tokens_out":2185,"duration_ms":17730,"temperature":1.0,"reasoning_tokens":2109,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:22:54.872241+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same watch-duration versus explicit-action scenarios through TikTok's mobile app while intercepting network traffic to record every request. If likes and follows sent from the app change the feed as much as or more than watch duration, the paper's claim that explicit actions have little effect on the web and that watch is the strongest signal is contradicted. Alternatively, recompute the paper's comparisons using human-annotated video topics instead of hashtag similarity; if the ranking of personalisation factors changes, the watch-dominance finding is a metric artifact.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the hashtag-similarity evaluation measures (strict Jaccard and lenient basic match) reused in the analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates that audit methodology choices change YouTube audit findings, motivating the paper's metric-sensitivity analysis."}],"review_version":1}