REVIEW 4 major objections 6 minor 40 references
Algorithmic Audit of Personalisation Drift in Polarising Topics on TikTok
T0 review · 4 major / 6 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read TikTok personalises strongly, but steers misinformation topics toward safe neutral content while sustaining US politics and reinforcing stance.
desk verdict Solid multi-topic TikTok sockpuppet audit that cleanly separates preference, topic, and stance drift; politics behaves unlike climate/vaccines, with one taxonomy caveat on the stance result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Three measured drifts over time bins—preference-aligned, polarisation-topic, and polarisation-stance—produced by sockpuppet accounts whose watch/like/bookmark decisions are driven by an LLM user-interaction predictor (with audio transcript) on the For You feed after a controlled seeding phase.
What would settle it
Repeat the same seed and interaction protocol with long-lived real-user accounts (or donated feeds) over the same topics: if misinformation topics no longer neutralise toward cooking, or mixed US-politics accounts no longer prefer oppose, the reported drift patterns fail to transfer.
Extended reading notes
Core claim
TikTok’s recommendation trajectories differ markedly by topic. Preference-aligned personalisation is strong. Misinformation-themed topics show a strong neutralising polarisation-topic drift toward neutral and safe content, while US politics shows high sustained topic share without that neutralising drift. Polarisation-stance behaviour generally reinforces the seeded stance; mixed-polarity US-politics accounts show a significant preference and small drift toward the oppose stance.
Load-bearing premise
That new bot accounts on US proxies, for about 10–16 days, interacting only through LLM-chosen strong feedback signals, are treated enough like real users that measured feed ratios equal real personalisation drift.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a sockpuppet algorithmic audit of TikTok personalisation on four polarising topics (flat earth, vaccines, climate change, US politics) plus a neutral cooking baseline. Using 68 controlled accounts, LLM-driven relevance/stance decisions (GPT-4.1 + Whisper), and multi-day For You interactions, it measures three constructed drifts: preference-aligned (interest vs unrelated), polarisation-topic (polarising vs neutral interest), and polarisation-stance (support vs oppose). Main claims: strong preference-aligned personalisation; a neutralising polarisation-topic trajectory for misinformation-themed topics (especially climate and vaccines) versus sustained high US-politics topic share without that neutralising drift; and, for stance, general reinforcement of the seeded stance, with mixed-polarity US-politics accounts showing a significant preference and small drift toward the oppose stance.
Significance. If the measured trajectories transfer beyond the sockpuppet setup, the work is a timely multi-topic audit that goes beyond election-only studies and cleanly separates overall personalisation, topic-level polarisation share, and stance share. Strengths include a relatively large account set, reported LLM topic/stance accuracies on a 350-video set, a Whisper ablation, Mann–Whitney tests, a hashtag-popularity check arguing popularity alone does not explain topic differences, explicit ethics process, and planned data/code release. The topic-dependent contrast (neutralising misinfo pathways vs politics equilibrium) is the most policy-relevant contribution for platform governance and DSA-style transparency.
major comments (4)
- [§5 / Appendix A Table 2 / Fig. 5] §5 and Appendix A Table 2: the polarisation-stance claim for US politics (stance reinforcement under single-seed; significant oppose preference and small drift under mixed polarity, Fig. 5, p≈2.1e-22) rests on a Trump-centric taxonomy (support = Trump/Republicans/conservatives/right; oppose = Biden/Harris/Democrats/liberals/left). That axis is not validated against human left–right, party-only, or anti-incumbent labels on audit videos. If high-engagement anti-Trump content is systematically coded oppose without being left-liberal, the oppose excess and fitted drift can be inflated. Please re-label a stratified sample of recommended political videos with an independent human scheme (or coarser party/left–right codes) and show whether the oppose preference survives.
- [§3.2–§3.4] §3.2–§3.4: outcome topic/stance ratios are produced by the same LLM pipeline that chooses watch/like/bookmark actions. Validation is on a separately searched 350-video set (topic ~95–98%, stance ~90–100%), not on a random sample of For You recommendations from the audit itself, where mixed, satirical, or event-driven content is more likely. Residual label error is therefore load-bearing for all three drifts. Report human agreement on a stratified sample of recommended videos (by topic, day, and user group) and sensitivity of the fitted drifts under label noise or alternative prompts.
- [§3.1 / Fig. 5] §3.1 and Fig. 5: the mixed-polarity US-politics group that underpins the oppose-preference/drift claim uses only four accounts and a shorter window (Jan 2026). With event-driven content acknowledged in §5, n=4 is thin for a central stance conclusion. Either expand this arm, report account-level trajectories and uncertainty, or demote the mixed-polarity oppose drift from a primary finding to an exploratory result.
- [§4 / Appendix D] §4 and Appendix D: the neutralising polarisation-topic drift for climate/vaccines/flat earth is interpreted as recommender behaviour relative to cooking. Hashtag view totals are a weak inventory proxy and do not establish that polarising videos remain available in the ranking pool after seeding. Without a content-availability or search-side control (e.g., parallel unseeded or search-only baselines for the same queries), suppression cannot be cleanly separated from sparse supply, safety filtering, or cooking’s extreme popularity. Strengthen the causal language or add such a control.
minor comments (6)
- [Abstract / §1] Abstract and §1 over-claim slightly relative to the body: climate change yields almost no polarising videos, so stance conclusions for that topic should be explicitly scoped out in the abstract.
- [§3.1] §3.1: clarify why the mixed-polarity arm was run months later (Jan 2026 vs Sep 2025) and how temporal confounds are handled when comparing to other groups.
- [Figures 2–5] Figures 2–5: axis labels, bin definitions (30-minute), and exact regression specification for “drift” lines should appear in captions so the plots are self-contained.
- [§3.4] §3.4: state whether Mann–Whitney tests are applied to binned ratios, raw counts, or user-level aggregates, and whether multiple-comparison correction is used across topics/stances.
- [Appendix C] Appendix C raw video totals are useful; consider a compact table in the main text summarising topic-relevant / cooking / unrelated counts per user group.
- [§3.1 / §4–§5] Minor language issues: “left-learning” (should be left-leaning), “flatearth” consistency, and occasional tense shifts in §4–§5.
Circularity Check
Empirical sockpuppet audit; mild dual-use of the same LLM for interaction and labels, not a derivation that redefines its target.
-
other
[§3.2 User Interaction Predictor; §3.4 Evaluation methodology; §4 RQ1 results (Fig. 2–3)]
"When the video is related to the topic and stance of interest for the user, the action returned by the user interaction predictor is to watch the video in full, like it and bookmark it. In any other case, the video is skipped. ... As with all the phases of the study, the topic and stance are assigned by the user interaction predictor. ... we observe a strong preference-aligned drift, where the number of videos of interest quickly increases to around 70%."
The same LLM both selects positive feedback (training the feed toward its class) and defines the numerator of preference-aligned and stance ratios. High measured personalisation is therefore partly the closed loop of reinforcing and counting that class, not fully independent external labeling. Mitigated by separate validation accuracy; does not force between-topic contrasts.
full rationale
This paper is an observational algorithmic audit, not a first-principles derivation. Drift quantities are defined as empirical ratios of recommended videos labeled by topic/stance over time bins; they are not claimed as closed-form predictions from fitted parameters or uniqueness theorems. The only residual circularity is methodological: the same GPT-4.1 user-interaction predictor both decides which videos receive watch/like/bookmark feedback (thus shaping the recommender trajectory) and supplies the topic/stance labels used to compute preference-aligned, polarisation-topic, and polarisation-stance ratios. That dual use makes measured preference-aligned personalisation partly a closed loop of reinforcing and counting the same LLM-defined class. It is mitigated by a separately collected 350-video human-annotated validation set with high reported accuracy, and it does not force the paper’s distinctive between-topic contrasts (neutralising drift for climate/vaccines vs equilibrium for US politics). Self-citations to prior audits by overlapping authors are background, not load-bearing uniqueness claims. No fitted-parameter-as-prediction, ansatz smuggling, or renaming of a known result as a derivation. Score 2 for one minor dual-use step that is not central to the strongest comparative claims.
Assumptions & free parameters
free parameters (4)
- seed_video_count
- daily_interaction_duration
- aggregation_bin_width
- skip_delay_seconds
assumptions (5)
- domain assumption Newly created sockpuppet accounts on residential US proxies receive recommendation behaviour comparable enough to ordinary users for drift conclusions to generalise.
- domain assumption GPT-4.1 topic/stance labels (with Whisper transcripts) are accurate enough that ratio trends reflect true content composition, not systematic misclassification.
- domain assumption Simultaneous full watch + like + bookmark is a valid strong interest signal for studying personalisation drift.
- ad hoc to paper Cooking is a sufficiently representative neutral/safe baseline against which polarising-topic share can be compared.
- domain assumption Short-window US-centric observations (Sep 2025 / Jan 2026) support claims about how TikTok treats these topics in general.
invented entities (3)
-
preference-aligned drift
-
polarisation-topic drift
-
polarisation-stance drift
Cite this review
Pith. "Pith review of Algorithmic Audit of Personalisation Drift in Polarising Topics on TikTok." pith.science (2026). https://pith.science/paper/JEW5COK2
@misc{pith2026260320723,
author = {Pith},
title = {Pith review of: Algorithmic Audit of Personalisation Drift in Polarising Topics on TikTok},
year = {2026},
howpublished = {\url{https://pith.science/paper/JEW5COK2}},
note = {Machine review of arXiv:2603.20723}
}
read the original abstract
Social media platforms have become an integral part of everyday life, serving as a primary source of news and information for many users. These platforms increasingly rely on personalised recommendation systems that shape what users see and engage with. While these systems are optimised for engagement, concerns have emerged that they may also drive users toward more polarised perspectives, particularly in contested domains such as politics, climate change, vaccines, and conspiracy theories. In this paper, we present an algorithmic audit of personalisation drift on TikTok in these polarising topics. Using controlled accounts designed to simulate users with interests aligned with or opposed to different polarising topics, we systematically measure the extent to which TikTok steers content exposure toward specific topics and polarities over time. Specifically, we investigated: 1) a preference-aligned drift (showing a strong personalisation towards user interests), 2) a polarisation-topic drift (showing a strong neutralising effect for misinformation-themed topics, and a high preference and reinforcement of interest of US politic topic); and 3) a polarisation-stance drift (showing a preference of oppose stance towards US politics topic and a general reinforcement of users' stance by recommending items aligned with their stance towards polarising topics). Overall, our findings provide evidence that recommendation trajectories differ markedly across topics, with some pathways amplifying polarised viewpoints more strongly than others and offer insights for platform governance, transparency and user awareness.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Guy Aridor, Duarte Goncalves, and Shan Sikdar. 2020. Deconstructing the Filter Bubble: User Decision-Making and Recommender Systems. InProceedings of the 14th ACM Conference on Recommender Systems(Virtual Event, Brazil) (RecSys ’20). Association for Computing Machinery, New York, NY, USA, 82–91. doi:10.1145/3383313.3412246
-
[2]
Cameron Ballard, Ian Goldstein, Pulak Mehta, Genesis Smothers, Kejsi Take, Victoria Zhong, Rachel Greenstadt, Tobias Lauinger, and Damon McCoy. 2022. Conspiracy Brokers: Understanding the Monetization of YouTube Conspiracy Theories. InProceedings of the ACM Web Conference 2022 (WWW ’22). Association for Computing Machinery, New York, NY, USA, 2707–2718. d...
doi:10.1145/3485447 2022
-
[3]
Jack Bandy. 2021. Problematic Machine Behavior: A Systematic Literature Review of Algorithm Audits.Proc. ACM Hum.-Comput. Interact.5, CSCW1, Article 74 (April 2021), 34 pages. doi:10.1145/3449148
doi:10.1145/3449148 2021
-
[4]
Maximilian Boeker and Aleksandra Urman. 2022. An Empirical Investigation of Personalization Factors on TikTok. InProceedings of the ACM Web Conference 2022 (Virtual Event, Lyon, France)(WWW ’22). Association for Computing Machinery, New York, NY, USA, 2298–2309. doi:10.1145/3485447.3512102
-
[5]
Shea Brown, Jovana Davidovic, and Ali Hasan. 2021. The algorithm audit: Scoring the algorithms that score us.Big Data & Society8, 1 (2021), 2053951720983865. arXiv:https://doi.org/10.1177/2053951720983865 doi:10.1177/2053951720983865
-
[6]
Mert Can Cakmak, Nitin Agarwal, and Diwash Poudel. 2025. Investigating Algorithmic Bias in YouTube Shorts.arXiv preprint arXiv:2507.04605(2025)
arXiv 2025
-
[7]
Ryan Evans, Daniel Jackson, and Jaron Murphy. 2023. Google News and machine gatekeepers: Algorithmic personalisation and news diversity in online news Pecher et al. search.Digital Journalism11, 9 (2023), 1682–1700
2023
-
[8]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schel- ten, Alex Vaughan, et al . 2024. The llama 3 herd of models.arXiv preprint arXiv:2407.21783(2024)
arXiv 2024
Show all 40 references
-
[9]
Muhammad Haroon, Anshuman Chhabra, Xin Liu, Prasant Mohapatra, Zubair Shafiq, and Magdalena Wojcieszak. 2022. YouTube, The Great Radicalizer? Auditing and Mitigating Ideological Biases in YouTube Recommendations. arXiv:2203.10666 [cs](March 2022). arXiv:2203.10666 [cs]
2022 arXiv
-
[10]
Eslam Hussein, Prerna Juneja, and Tanushree Mitra. 2020. Measuring Misinfor- mation in Video Search Platforms: An Audit Study on YouTube.Proc. ACM Hum.- Comput. Interact.4, CSCW1, Article 048 (May 2020), 27 pages. doi:10.1145/3392854
2020 doi
-
[11]
Hazem Ibrahim, HyunSeok Daniel Jang, Nouar Aldahoul, Aaron R Kaufman, Talal Rahwan, and Yasir Zaki. 2025. TikTok’s recommendations skewed to- wards Republican content during the 2024 US presidential race.arXiv preprint arXiv:2501.17831(2025)
2025 arXiv
-
[12]
Prerna Juneja, Md Momen Bhuiyan, and Tanushree Mitra. 2023. Assessing enactment of content regulation policies: A post hoc crowd-sourced audit of election misinformation on YouTube. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems(Hamburg, Germany...
2023
-
[13]
Jonas Kaiser and Adrian Rauchfleisch. 2020. Birds of a feather get recommended together: Algorithmic homophily in YouTube’s channel recommendations in the United States and Germany.Social Media+ Society6, 4 (2020), 2056305120969914
2020
-
[14]
Levi Kaplan and Piotr Sapiezynski. 2024. Comprehensively Auditing the TikTok Mobile App. InCompanion Proceedings of the ACM Web Conference 2024(Singa- pore, Singapore)(WWW ’24). Association for Computing Machinery, New York, NY, USA, 1198–1201. doi:10.1145/3589335.3651260
2024 doi
-
[15]
Daniel Klug, Yiluo Qin, Morgan Evans, and Geoff Kaufman. 2021. Trick and Please. A Mixed-Method Study On User Assumptions About the TikTok Algorithm. In Proceedings of the 13th ACM Web Science Conference 2021(Virtual Event, United Kingdom)(WebSci ’21). Association for Computin...
2021 doi
-
[16]
Mark Ledwich, Anna Zaitsev, and Anton Laukemper. 2022. Radical bubbles on YouTube? Revisiting algorithmic extremism with personalised recommendations. First Monday(2022)
2022
-
[17]
Matej Mosnar, Adam Skurla, Branislav Pecher, Matus Tibensky, Jan Jakubcik, Adrian Bindas, Peter Sakalik, and Ivan Srba. 2025. Revisiting Algorithmic Audits of TikTok: Poor Reproducibility and Short-term Validity of Findings. InProceed- ings of the 48th International ACM SIGIR ...
2025 doi
-
[18]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. InInternational conference on machine learning. PMLR, 28492–28518
2023
-
[19]
Manoel Horta Ribeiro, Raphael Ottoni, Robert West, Virgílio A. F. Almeida, and Wagner Meira. 2020. Auditing Radicalization Pathways on YouTube. InProc. of the 2020 Conference on Fairness, Accountability, and Transparency. ACM, New York, NY, USA, 131–141. doi:10.1145/3351095.3372879
2020 doi
-
[20]
Robertson, David Lazer, and Christo Wilson
Ronald E. Robertson, David Lazer, and Christo Wilson. 2018. Auditing the Personalization and Composition of Politically-Related Search Engine Results Pages. InProceedings of the 2018 World Wide Web Conference(Lyon, France) (WWW ’18). International World Wide Web Conferences St...
2018 doi
-
[21]
Christian Sandvig, Kevin Hamilton, Karrie Karahalios, and Cedric Langbort. 2014. Auditing algorithms: Research methods for detecting discrimination on internet platforms.Data and discrimination: converting critical concerns into productive inquiry22, 2014 (2014), 4349–4357
2014
-
[22]
Donghee Shin and Kulsawasd Jitkajornwanich. 2024. How algorithms promote self-radicalization: audit of TikTok’s algorithm using a reverse engineering method.Social Science Computer Review42, 4 (2024), 1020–1040
2024
-
[23]
Kasia Söderlund, Emma Engström, Kashyap Haresamudram, Stefan Larsson, and Pontus Strimling. 2024. Regulating high-reach AI: On transparency directions in the Digital Services Act.Internet policy review13, 1 (2024), 1–31
2024
-
[24]
Sara Solarova, Matúš Mesarčík, Branislav Pecher, and Ivan Srba. 2026. Beyond the Checkbox: Strengthening DSA Compliance Through Social Media Algorithmic Auditing.arXiv:2601.18405 [cs](Jan. 2026). arXiv:2601.18405 [cs]
2026
-
[25]
Sara Solarova, Matej Mosnar, Matus Tibensky, Jan Jakubcik, Adrian Bindas, Simon Liska, Filip Hossner, Matúš Mesarčík, and Ivan Srba. 2026. The DSA’s Blind Spot: Algorithmic Audit of Advertising and Minor Profiling on TikTok. arXiv:2603.05653 [cs.CY] https://arxiv.org/abs/2603.05653
2026 arXiv
-
[26]
Kirill Solovev, Chiara Drolsbach, Emma Demirel, and Nicolas Pröllochs. 2025. TikTok Rewards Divisive Political Messaging During the 2025 German Federal Election.arXiv preprint arXiv:2509.10336(2025)
2025 arXiv
-
[27]
Larissa Spinelli and Mark Crovella. 2020. How YouTube Leads Privacy-Seeking Users Away from Reliable Information. InAdjunct Publication of the 28th ACM Conference on User Modeling, Adaptation and Personalization. ACM, New York, NY, USA, 244–251. doi:10.1145/3386392.3399566
2020 doi
-
[28]
Ivan Srba, Robert Moro, Matus Tomlein, Branislav Pecher, Jakub Simko, Elena Stefancova, Michal Kompan, Andrea Hrckova, Juraj Podrouzek, Adrian Gavornik, and Maria Bielikova. 2023. Auditing YouTube’s Recommendation Algorithm for Misinformation Filter Bubbles.ACM Trans. Recomm. ...
2023 doi
-
[29]
Gemma Team. 2024. Gemma. (2024). doi:10.34740/KAGGLE/M/3301
2024 doi
-
[30]
Qwen Team. 2024. Qwen2.5: A Party of Foundation Models. https://qwenlm. github.io/blog/qwen2.5/
2024
-
[31]
Kjerstin Thorson, Kelley Cotter, Mel Medeiros, and Chankyung Pak. 2021. Al- gorithmic inference, political interest, and exposure to news and politics on Facebook.Information, Communication & Society24, 2 (2021), 183–200
2021
-
[32]
Matus Tomlein, Branislav Pecher, Jakub Simko, Ivan Srba, Robert Moro, Elena Ste- fancova, Michal Kompan, Andrea Hrckova, Juraj Podrouzek, and Maria Bielikova
-
[33]
InFifteenth ACM Conference on Recommender Systems
An Audit of Misinformation Filter Bubbles on YouTube: Bubble Bursting and Recent Behavior Changes. InFifteenth ACM Conference on Recommender Systems. ACM, New York, NY, USA, 1–11. doi:10.1145/3460231.3474241
-
[34]
Aleksandra Urman, Mykola Makhortykh, and Aniko Hannak. 2024. Mapping the Field of Algorithm Auditing: A Systematic Literature Review Identifying Research Trends, Linguistic and Geographical Disparities.arXiv preprint arXiv:2401.11194 (2024)
2024 arXiv
-
[35]
Karan Vombatkere, Sepehr Mousavi, Savvas Zannettou, Franziska Roesner, and Krishna P. Gummadi. 2024. TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds. InProceedings of the ACM Web Conference 2024(Singapore, Singapore)(WWW ...
2024 doi
-
[36]
Linda Xue, Francesco Corso, Nicolo’ Fontana, Geng Liu, Stefano Ceri, and Francesco Pierri. 2025. Towards an Automated Framework to Audit Youth Safety on TikTok.arXiv preprint arXiv:2509.05838(2025)
2025 arXiv
-
[37]
Can Yang, Xinyuan Xu, Bernardo Pereira Nunes, and Sean Wolfgand Matsui Siqueira. 2023. Bubbles bursting: Investigating and measuring the personalisation of social media searches.Telematics and Informatics82 (2023), 101999
2023
-
[38]
Jinyi Ye, Luca Luceri, and Emilio Ferrara. 2025. Auditing Political Exposure Bias: Algorithmic Amplification on Twitter/X During the 2024 U.S. Presidential Election. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25). Association...
2025 doi
-
[39]
Gummadi, Elissa M
Savvas Zannettou, Olivia Nemes-Nemeth, Oshrat Ayalon, Angelica Goetzen, Krishna P. Gummadi, Elissa M. Redmiles, and Franziska Roesner. 2024. Analyzing User Engagement with TikTok’s Short Format Video Recommendations using Data Donations. InProceedings of the 2024 CHI Conferenc...
2024 doi
-
[40]
earth is flat
Lisa Zieringer and Diana Rieger. 2023. Algorithmic recommendations’ role for the interrelatedness of counter-messages and polluted content on YouTube–a network analysis.Computational Communication Research5, 1 (2023), 109. A Additional Methodology Details In this section, we p...
2023
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.