{"id":"975cdb70-ac3d-4fc7-b200-70717e3c5f03","arxiv_id":"2501.17831","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"TikTok's recommendation algorithm served Republican-seeded test accounts more co-partisan content than Democratic-seeded accounts during the 2024 U.S. presidential race.","lead":"A controlled audit of 323 TikTok accounts found that accounts seeded with Republican content received about 11.8% more party-aligned recommendations than accounts seeded with Democratic content, while Democratic accounts saw more opposing-party content. The result adds evidence to debates about platform neutrality and TikTok's influence on elections.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Transcript-only outcome measurement: the representativeness check only compares marginal partisan distributions, not the bot-condition difference, so the headline asymmetry may not generalize to the 77% of videos without transcripts.","rationale":"The reader's weakest assumption identified both the conditioning-channel representativeness and the transcript-only labeling as central concerns. I focus on the transcript issue because it affects the measurement of the outcome itself, which is the foundation of the entire paper. The authors' representativeness check is a step in the right direction, but it is insufficient: it only compares marginal partisan distributions between transcript and non-transcript videos, whereas the central claim concerns a difference between experimental conditions. If the algorithm's partisan skew differs for non-transcript content (e.g., silent memes, visual-only videos), the headline 11.8% and 7.5% figures could be biased in either direction. The proposed interaction test is feasible with the existing human-annotated data and would directly settle whether the transcript-only restriction changes the conclusion. This does not change the reader's CONDITIONAL verdict: the internal validity is strong, but external validity remains conditional on this untested assumption. I therefore recommend UNCHANGED, with the specific interaction test as a required revision before the claim can be treated as generalizing to all TikTok recommendations.","tokens_in":47030,"tokens_out":3809,"duration_ms":42798,"concrete_test":"Recompute the ideological skew and the between-condition difference (Democrat-conditioned vs Republican-conditioned bots) using the human-annotated non-transcript videos (Data Representativeness section; the 4,000-video sample, or a larger sampled set if resources allow). Specifically, among political videos, estimate the proportion of Republican-aligned minus Democratic-aligned videos separately for transcript and non-transcript videos within each bot condition, then test the interaction between transcript availability and bot condition in a logistic or linear model. If the interaction is not significant and the non-transcript point estimate of the asymmetry falls within the confidence interval of the transcript-based estimate, the concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central asymmetry (11.8% more co-partisan, 7.5% less cross-partisan) is estimated exclusively from the 40,264 unique videos with English transcripts (22.8% of unique recommendations; Section 'Measuring Political Content'). The Data Representativeness check (Methods, Figure 7) compares transcript and non-transcript videos on marginal partisan distributions using chi-square tests and finds no significant difference. However, the paper's claim is about the difference between Republican- and Democratic-conditioned bots. The authors do not test whether this between-condition difference is the same in the transcript subsample as in the non-transcript sample. If transcript availability is correlated with bot condition or with partisanship conditional on condition, the measured asymmetry could be biased. For example, if Republican-conditioned bots receive a higher share of silent/visual Republican-aligned videos (which lack transcripts), the transcript-only analysis would understate Republican co-partisan recommendations; the reverse bias is also possible. The 4,000-video human-annotated non-transcript sample (2,000 per bot condition) may be adequate for marginal comparisons but has limited power to detect an interaction, and no interaction test is reported. Thus the headline numbers rest on an untested homogeneity assumption that is central to external validity.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a large-scale sock-puppet audit of TikTok's recommendation algorithm during the 2024 U.S. presidential race. The authors created 323 experimental runs across New York, Texas, and Georgia, conditioning fresh accounts with Democratic, Republican, or no partisan videos, and collected roughly 394,000 videos from April 30 to November 11, 2024. Political and partisan labels were assigned via a three-LLM ensemble with human validation on transcripts. The central finding is that Republican-conditioned bots received approximately 11.8% more co-partisan recommendations than Democratic-conditioned bots, while Democratic-conditioned bots received approximately 7.5% more cross-partisan recommendations. The authors argue that this asymmetry exists across all three states, persists after controlling for video- and channel-level engagement metrics, is driven mainly by negative partisanship content, and is not explained by observable engagement or asymmetric homophily. The paper also includes channel-level mismatch analyses, topic analyses, a misinformation exploration, and a survey of TikTok users' perceptions.","tokens_in":47179,"tokens_out":7530,"duration_ms":80521,"significance":"If the central result holds, the paper is a valuable empirical contribution to algorithmic auditing and to the public debate about platform neutrality. Its strengths include a longitudinal design spanning the election period, a large number of independent experimental runs, preregistration, a human-validated classification pipeline, and a serious set of robustness checks: counterfactual sampling under 48 engagement-based models, an asymmetric-homophily test, and a sensitivity analysis for latent engagement metrics. The replication materials and dataset are promised on GitHub, which is a further strength. However, the headline numbers are estimated only from videos with English transcripts, and the representativeness check does not test the between-condition difference that is the paper's main claim. This, together with the hand-picked conditioning channel set and unclustered video-level regressions, means the paper needs additional work before its central claim can be fully accepted.","major_comments":[{"comment":"The representativeness check in the Data Representativeness section tests whether transcript and non-transcript videos have the same marginal partisan distribution within each bot condition, but the paper's headline claim is a difference between conditions. The 11.8% and 7.5% estimates are computed only from the 40,264 transcript videos (22.8% of unique recommendations), and Figure 7 does not report an interaction test between bot condition and transcript availability. For example, if Republican-conditioned bots are disproportionately recommended Republican-aligned videos without transcripts, the transcript-only sample would understate or overstate the between-condition asymmetry. Please report the between-condition difference in partisan composition for the 4,000 human-annotated non-transcript videos, with a formal interaction test or a confidence interval, and state whether the headline asymmetry holds in that sample.","section":"Methods, Data Representativeness; Measuring Political Content; Figure 7"},{"comment":"The conditioning treatment consists of 12 Democrat- and 12 Republican-aligned channels selected by searching politically charged keywords and matched on follower count and cumulative likes. This leaves open the possibility that unmeasured channel characteristics (content style, posting frequency, network position) are correlated with partisanship and influence downstream recommendations. A robustness analysis that reruns the conditioning with alternative or leave-one-out channel sets, or a model that includes channel-level random effects, would strengthen the claim that the asymmetry is attributable to partisan content rather than to the particular channels used. At minimum, the scope of the conclusion should be stated as conditional on the selected conditioning channels.","section":"Methods, Conditioning Stage; Table 4"},{"comment":"The logistic regressions are estimated at the bot-video level, but the text does not report clustering or a multilevel structure. Since videos are nested within bots and within channels, and each bot contributes up to 1,200 recommendations, the independence assumption is questionable. I request cluster-robust standard errors (by bot and/or by channel) or a mixed-effects specification, and a statement of whether the key coefficients (e.g., the Democrat bot odds ratio of 2.81 in Supplementary Table 13) remain significant after accounting for this dependence.","section":"Figure 2H; Supplementary Tables 13-15"}],"minor_comments":[{"comment":"The introduction states that the bots analyzed \"over 340,000 videos encountered through 381 experimental runs,\" while the abstract reports roughly 394,000 videos and 323 experiments; these numbers should be reconciled.","section":"Introduction"},{"comment":"The abstract reports 11.8% more co-partisan and 7.5% less cross-partisan content, while the introduction reports 11.5% and 8.0%; please clarify which definition corresponds to the headline estimate and use consistent numbers throughout.","section":"Abstract and Introduction"},{"comment":"The sensitivity analysis refers to \"Table 1 below\" when it means Table 11, and the number of counterfactual models is inconsistent: the main text says 48 scenarios, the Methods says 39 models, and Supplementary Table 5 lists 54 rows; please correct the counts.","section":"Methods, Sensitivity Analysis"},{"comment":"The text says a latent metric would need to be \"98 times bigger than the gap in likes\" and also \"9.4 times as strongly\" as the combined metric; both statements are internally defensible, but they should be placed side by side or cross-referenced so readers do not mistake them for conflicting numbers.","section":"Robustness Checks, Sensitivity Analysis"}],"recommendation":"major_revision","confidential_remarks":"I found the empirical core credible and the robustness checks thoughtful. The main obstacle is transcript-sample generalizability: the missing interaction test is fixable using the already-collected human-annotated non-transcript data, and if the authors provide that analysis I would be willing to accept the revised version. The conditioning-channel selection is a secondary external-validity concern that can be addressed by careful framing rather than by additional data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. This is the first controlled longitudinal audit of TikTok's political recommendations, and it's a good one. The headline result—Republican-conditioned bots received ~11.8% more co-partisan recommendations and ~7.5% less cross-partisan recommendations than Democratic-conditioned bots—is supported by a careful experimental design and multiple robustness checks. I'd trust it as a measurement of what these particular bots saw.\n\nWhat's new: prior work on TikTok politics focused on demographics or content supply; this is the first sock-puppet audit of the recommendation algorithm itself. The design is solid: 323 runs over 27 weeks across NY, TX, GA; conditioning stage followed by recommendation stage; pair-matching to equalize exposure; human-validated LLM ensemble classification with high agreement; counterfactual sampling by engagement metrics; logistic regressions; sensitivity analysis showing a latent metric would need to be ~9x larger than the combined likes/plays/shares gap to explain the skew. The negative partisanship decomposition is a nice touch. Data and code are public.\n\nSoft spots, in proportion. The outcome is measured only on videos with English transcripts (22.8% of unique videos, 26.2% of all recommendations). The paper's representativeness check compares the partisan distribution of transcript vs non-transcript videos within each bot condition and finds no significant difference. That's reassuring, but it doesn't directly test whether the between-condition difference is the same in the two subsamples. The 4,000-video human-annotated non-transcript sample has limited power for an interaction test, and none is reported. I don't think this sinks the result, but it's the right thing for a referee to ask about. Second, the conditioning channels are 24 hand-picked accounts matched on followers and cumulative likes only. If Republican channels happen to be more algorithmically promotable for style or network reasons, the asymmetry could be an artifact of the specific channel set. The engagement controls mitigate this, but don't fully eliminate it. Third, the preregistration was filed after nearly all data had been collected; it's a registered analysis plan, not a prospective preregistration. Minor editorial issues: a missing citation placeholder in the asymmetric homophily section and some inconsistent numbers across the abstract and intro.\n\nWho benefits: platform audit researchers, political communication people, and anyone interested in algorithmic accountability. It deserves serious refereeing. My recommendation: send it to review, and ask the authors to address the transcript interaction and conditioning channel representativeness, and clean up the editorial gaps.","headline":"First controlled TikTok audit with a credible Republican skew, weakened but not sunk by transcript-only outcome measurement and narrow conditioning channel selection.","tokens_in":47766,"tokens_out":5522,"would_cite":true,"duration_ms":49144,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"During the 2024 U.S. presidential campaign, TikTok's recommendation algorithm served Republican-conditioned accounts about 11.8% more co-partisan videos than Democratic-conditioned accounts, and served Democratic-conditioned accounts…","keywords":["TikTok","recommendation algorithm","algorithmic audit","partisan bias","2024 U.S. presidential election","sock puppet accounts","negative partisanship"],"falsifier":"A direct test would be an adversarial replication: fresh conditioning channels chosen by an independent panel, matched on content style and posting frequency, with all recommended videos (including those without transcripts) labeled by human annotators; if the Republican-conditioned co-partisan advantage disappears or reverses, the central claim fails. A complementary check would test whether any latent engagement feature available to TikTok has a Republican-Democrat gap larger than roughly 9.4 times the best observed composite gap.","tokens_in":46789,"feed_emoji":"📱","tokens_out":9105,"duration_ms":78738,"temperature":0.7,"pith_summary":"This paper sets out to test whether TikTok's recommendation algorithm treats the two major U.S. parties symmetrically. Using 323 controlled 'sock puppet' accounts in New York, Texas, and Georgia, seeded with Democratic, Republican, or no partisan videos, the authors logged roughly 394,000 recommended videos from April to November 2024 and labeled them with an LLM ensemble. They find that Republican-seeded accounts received about 11.8% more party-aligned recommendations than Democratic-seeded accounts, while Democratic-seeded accounts received about 7.5% more opposite-party recommendations. The asymmetry appears in all three states and survives controls for video- and channel-level engagement, and it is driven mainly by negative-partisanship content, particularly anti-Democratic videos. If the finding holds, it would mean TikTok's opaque feed, on which users have only indirect control, systematically shapes political exposure differently for the two parties during a presidential election.","feed_headline":"TikTok skewed Republican in 2024 race, 323-bot audit finds","feed_subtitle":"Seeded Democratic accounts got 7.5% more opposing-party videos; engagement metrics don't explain the gap.","key_machinery":"The load-bearing mechanism is the sock-puppet algorithmic audit: freshly created TikTok accounts, geo-spoofed via VPN and GPS mocking, that first watch up to 400 researcher-selected partisan videos (the conditioning stage) and then log the videos appearing on the 'For You' page (the recommendation stage). The outcome measure is an ideological-content score ranging from -1 (all Democratic-aligned) to +1 (all Republican-aligned), computed as the proportion of Republican-aligned minus Democratic-aligned political videos. To rule out engagement-based explanations, the authors build 48 counterfactual models that sample videos proportional to observed video- and channel-level metrics, plus a sensitivity analysis showing that an unobserved engagement metric would need a Republican-Democrat gap roughly 9.4 times the best observed composite gap to explain the skew. A logistic-regression model of ideological mismatch, controlling for state, week, engagement, and conditioning, isolates the partisan conditioning effect.","core_discovery":"The central claim is that TikTok's recommendation algorithm displayed a measurable Republican skew during the 2024 U.S. presidential race. In paired weekly comparisons, Republican-conditioned bots were served approximately 11.8% more co-partisan videos than Democratic-conditioned bots, and Democratic-conditioned bots were served approximately 7.5% more cross-partisan videos. The authors operationalize this as an ideological-content score, the share of recommended political videos aligned with Republicans minus the share aligned with Democrats, and show the skew grew over the campaign and was present in a projected Democratic state, a projected Republican state, and a swing state. They argue the asymmetry is not a byproduct of engagement: counterfactual models that resample videos weighted by likes, comments, shares, plays, channel followers, or verification status predict little or no Republican advantage, and a logistic regression shows Democratic bots remain roughly 2.8 times as likely to receive ideologically mismatched videos even after controlling for comment partisanship. The mismatch is concentrated in negative-partisanship content and in top Republican channels such as the official Trump account and Fox News being recommended to Democratic bots.","pith_inferences":["If the skew is driven by engagement optimization around negative partisanship, then the same mechanism should produce out-party animosity biases in non-U.S. elections and in non-political domains where negativity drives watch time; this is a testable prediction the paper does not make.","A replication that matches conditioning channels on content format and posting frequency, rather than only on followers and likes, would reveal whether the asymmetry is about partisanship or about stylistic differences between the two channel sets.","Because stance labels come only from transcript-bearing videos (22.8% of unique videos), a multimodal replication that labels visual-only political content could shift the measured asymmetry; this is an open empirical question.","The survey of 1,000 users shows Republicans perceiving more co-partisan and positive content, consistent with the experimental skew, but linking those perceptions to the actual recommendation logs would be a stronger test of whether real users notice the asymmetry."],"forward_implications":["Democratic-leaning users can expect to see more opposing-party content than Republican-leaning users see, making the partisan experience asymmetric.","The Republican skew persists across New York, Texas, and Georgia, so it is not an artifact of one state's political environment.","Engagement metrics (likes, shares, comments, plays, followers, verification) do not explain the skew; counterfactual sampling by these metrics predicts little or no Republican advantage.","Negative-partisanship videos, especially Anti Democrat content, drive the mismatch; positive pro-party content contributes much less to the skew.","Top Republican channels, including Trump's and Fox News', were recommended to Democratic bots more often than top Democratic channels were recommended to Republican bots."],"supporting_citations":[{"why":"Provides the general algorithmic-audit methodology that the sock-puppet design operationalizes.","marker":"[66]"},{"why":"Prior audit of YouTube's recommendation algorithm whose experimental design and engagement controls are adapted to TikTok.","marker":"[29]"},{"why":"Precedent for auditing a recommendation system for ideologically congenial and problematic content.","marker":"[67]"},{"why":"Earlier TikTok election audit whose hashtag-based party classification this paper replaces with transcript-based stance classification.","marker":"[56]"},{"why":"API used to retrieve extra pre-election videos for channels with few labeled videos, supporting channel-level partisan classification.","marker":"[57]"},{"why":"Supplies the validated topic-classification prompt used to identify topics of political videos.","marker":"[59]"},{"why":"Used as best practice for manually validating channel-level ideological classification.","marker":"[30]"}],"fun_headline_variants":["323-bot audit: TikTok's 2024 feed skewed Republican","GOP-seeded TikTok bots got 11.8% more aligned videos in 2024","TikTok's Republican skew in 2024 not explained by engagement metrics","Negative partisanship fueled TikTok's 2024 GOP recommendation bias","TikTok recommended 7.5% more cross-party videos to Dem-seeded bots"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that the 24 conditioning channels, matched only on follower count and cumulative likes, represent partisan political content on TikTok, and that stance labels inferred from transcripts (available for 22.8% of unique videos) faithfully represent the partisanship of all recommended videos.","fun_headline_variants_meta":{"raw":{"variants":["323-bot audit: TikTok's 2024 feed skewed Republican","GOP-seeded TikTok bots got 11.8% more aligned videos in 2024","TikTok's Republican skew in 2024 not explained by engagement metrics","Negative partisanship fueled TikTok's 2024 GOP recommendation bias","TikTok recommended 7.5% more cross-party videos to Dem-seeded bots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2433,"prompt_tokens":1057,"completion_tokens":1376,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":673,"completion_tokens_details":{"reasoning_tokens":1272}},"tokens_in":673,"tokens_out":1376,"duration_ms":10844,"temperature":1.0,"reasoning_tokens":1272,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:31:25.728514+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be an adversarial replication: fresh conditioning channels chosen by an independent panel, matched on content style and posting frequency, with all recommended videos (including those without transcripts) labeled by human annotators; if the Republican-conditioned co-partisan advantage disappears or reverses, the central claim fails. A complementary check would test whether any latent engagement feature available to TikTok has a Republican-Democrat gap larger than roughly 9.4 times the best observed composite gap.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the general algorithmic-audit methodology that the sock-puppet design operationalizes."},{"cited_title":"Asymmetric ideological segregation in exposure to political news on 1009 facebook","cited_arxiv_id":null,"evidence_quote":"Prior audit of YouTube's recommendation algorithm whose experimental design and engagement controls are adapted to TikTok."},{"cited_title":"Tikapi: Tiktok api for developers (2025)","cited_arxiv_id":null,"evidence_quote":"Precedent for auditing a recommendation system for ideologically congenial and problematic content."},{"cited_title":"The use of tiktok for political campaigning in canada: the case of jagmeet singh","cited_arxiv_id":null,"evidence_quote":"Earlier TikTok election audit whose hashtag-based party classification this paper replaces with transcript-based stance classification."},{"cited_title":"& Votta, F","cited_arxiv_id":null,"evidence_quote":"API used to retrieve extra pre-election videos for channels with few labeled videos, supporting channel-level partisan classification."},{"cited_title":"& Abidin, C","cited_arxiv_id":null,"evidence_quote":"Supplies the validated topic-classification prompt used to identify topics of political videos."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Used as best practice for manually validating channel-level ideological classification."}],"review_version":1}