REVIEW 5 major objections 5 minor 1 cited by
Investigating Algorithmic Bias in YouTube Shorts
T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims YouTube Shorts' recommendation algorithm systematically replaces politically sensitive content with entertainment within the first recommendation step, regardless of watch time.
desk verdict Solid data effort, but missing control seeds and inferential statistics make the 'political bias' headline an over-reach. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The core object is a recommendation chain: each seed Short is opened in a fresh, logged-out browser session and the next 50 recommended Shorts are recorded, with 3-, 15-, or 60-second dwell times simulating watch-time. Relevance, topic category (politics, non-entertainment, entertainment), and emotion (joy, sadness, anger, neutral, fear) are assigned by GPT-4o from each video's title and transcript, after validation against benchmark datasets for emotion, topic, and relevance. This setup turns an opaque proprietary recommender into measured drift curves of relevance, topic, emotion, and engagement as functions of recommendation depth.
What would settle it
Run a controlled, logged-in session seeded with a South China Sea or Taiwan-election Short, record the first ten recommendations while the user actually watches for 15 or 60 seconds, and compute semantic relevance to the seed topic; if mean relevance remains substantially above near-zero at depth 5 in this setting, the claimed immediate, watch-time-independent drift is contradicted for real user conditions.
Extended reading notes
Core claim
The central claim is that YouTube Shorts' recommendation chain diverges from politically sensitive seed content almost immediately and does so consistently across watch-time conditions. For both the South China Sea and Taiwan election datasets, seed videos are highly relevant and predominantly political, but by the first recommended video relevance scores approach zero and entertainment content becomes the majority, with joy/happiness rising and anger or fear declining. Engagement metrics jump after the first recommendation, indicating that highly viewed and liked videos are disproportionately promoted, and in the 60-second condition periodic sponsored content appears beyond depth 10. The authors conclude that watch time does not reduce algorithmic bias, and that the platform systematically prefers emotionally positive, engagement-optimized content over politically sensitive material.
Load-bearing premise
The watch-time results assume that pausing for 3, 15, or 60 seconds in an automated browser is perceived by YouTube as a genuine watch-time signal, but the study offers no evidence that the platform registers these pauses as engagement.
Editorial extensions
If this is right
- A user who starts with a politically sensitive Short will, within one recommendation step, mostly see entertainment content rather than related political coverage.
- Longer viewing does not counteract the drift, so the common assumption that engagement signals keep a feed on-topic is not supported for Shorts.
- High-engagement videos are systematically promoted, which can reinforce popularity bias and crowd out niche or serious topics.
- In longer viewing sessions, sponsored or promotional content is inserted periodically after about the tenth recommendation, suggesting monetized content is timed to sustained attention.
Reading between the lines
- Beyond the paper, the near-zero relevance at depth 1 may reflect Shorts' feed design, which could be dominated by global or trending content rather than topic affinity; one test would be whether entertainment seeds also drift topically.
- Beyond the paper, the watch-time conclusion depends on the untested assumption that an automated browser pause is registered by YouTube as genuine watch-time, so a logged-in real-user study is needed before generalizing the 'watch time does not reduce bias' claim.
- Beyond the paper, emotion labels come from titles and transcripts, so visually or musically affective Shorts with little text may be misclassified; a multimodal annotation could change the measured emotional drift.
- Beyond the paper, because sessions are logged-out and started from clean browsers, the results describe cold-start recommendations, not the personalized feeds most users experience.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an algorithm audit of YouTube Shorts recommendations. The authors seed fresh, logged-out browsing sessions with videos on the 2024 Taiwan presidential election, the South China Sea dispute, and a general content category; collect recommendation chains up to depth 50 under three simulated watch-time pauses (3s, 15s, 60s); classify 685,842 videos using GPT-4o on relevance, topic, and emotion; and report drift from politically sensitive seeds toward entertainment and positive affect, a preference for high-engagement videos, and no reduction in bias with longer watch time. The paper claims to be the first comprehensive analysis of algorithmic drift in YouTube Shorts along these dimensions.
Significance. If the central claims were established, this would be a useful contribution to the algorithm-auditing literature: the study is large in scale, measures observed recommendations rather than hypothetical model outputs, compares three watch-time conditions, and uses multi-dimensional labeling with an explicit benchmark-based model selection. The controlled fresh-session collection design is a strength, as is the authors' transparency about the validation benchmarks. However, the current evidence does not support the headline political-bias conclusion: the design lacks a matched non-political control, the statistical analysis is almost entirely descriptive, the classification is validated only on domain-mismatched benchmarks, and the watch-time manipulation is not validated as a genuine engagement signal. The significance is therefore potential rather than realized.
major comments (5)
- [Section 3.1, Table 1; Section 4.2; Abstract] The central claim that YouTube Shorts 'systematically' deprioritizes politically sensitive content cannot be separated from generic popularity bias because the only comparison condition is not a matched control. The General Content set is an aggregate of 15 categories and is already entertainment-dominated at depth 0 (Figures 2g-2i), so the observed drift from political seeds to entertainment is equally consistent with cold-start regression toward the dominant Shorts distribution in fresh, logged-out sessions. A matched control of non-political, non-entertainment seeds (e.g., science or education) with similar length, recency, and initial engagement is required to attribute the drift to topic sensitivity. Section 4.2's own concession that the effect 'likely reflects engagement optimization and popularity bias' admits that the mechanism has not been identified as specifically political. This is a load-bearing gap for the abstract's conclusion about politically sensitive content.
- [Sections 4.1-4.4, Figures 1-4] All drift claims rest on visual inspection of mean depth curves. The paper reports no confidence intervals, significance tests, effect sizes, or multiple-comparison corrections; standard deviations are shown as additional curves rather than as uncertainty bands around the means. In particular, the claim that 'watch time does not meaningfully impact topic relevance' (Section 4.1) and the detailed statements about differences among the 3s, 15s, and 60s conditions (Sections 4.2 and 4.4) are asserted from plotted means without any statistical comparison. The strength of the headline conclusions--'consistent drift,' 'systematic preference,' 'does not reduce algorithmic bias'--is not matched by the reported evidence.
- [Section 3.2, Table 4; Section 4.1] The classifier used for all downstream analyses is validated only on generic benchmark datasets that do not match the YouTube Shorts distribution. Emotion accuracy is 66.95% and relevance accuracy is 78.73%, and relevance is precisely the measure supporting the 'approach[ing] near-zero values at depth 1' claim. Because relevance is judged against a broadly formulated topic ('the South China Sea dispute' or 'the Taiwan election') rather than against the specific seed video, and because no domain-specific validation or error analysis on Shorts titles/transcripts is provided, classifier noise could plausibly inflate or even generate the apparent sharp relevance drop. This is a measurement-validity concern for the paper's primary drift metric.
- [Section 3.1, Section 5 (RQ3)] The watch-time manipulation is not validated as a genuine engagement signal. Pausing a Selenium browser for 3, 15, or 60 seconds is assumed to be interpreted by YouTube as watch time, but no evidence is offered that the platform registers these pauses as dwell time, and no post-hoc check verifies that the recommendation outputs differ in a way attributable to the pause rather than to elapsed time or other confounds. Section 3.1 only says the delays were incorporated 'to better simulate natural user behavior,' which does not establish signal validity. Since RQ3 and the conclusion that 'watch time does not reduce algorithmic bias' depend entirely on this assumption, the watch-time dimension of the study is unsupported.
- [Section 4.4, Figure 4] The engagement analysis conflates the properties of recommended videos with the algorithm's promotion decisions. Showing that mean logged views, likes, and comments rise from depth 0 to depth 1 does not demonstrate that the algorithm 'disproportionately promotes' high-engagement content without comparing against a suitable base rate of Shorts engagement or otherwise adjusting for confounds. Seed videos were collected through keyword searches and may be systematically less viral than the general Shorts population, making the observed depth-0 to depth-1 rise partly a selection artifact. The claim of popularity bias therefore needs a baseline comparison or an appropriate control, not just a within-chain trend.
minor comments (5)
- [Table 2 and Section 3.1] Table 2's Total column sums to 681,300, while the text reports 685,842 videos 'including both root and recommended videos'; the difference is exactly the number of root videos, so the table appears to omit root videos from the Total entries. The column header and the text should be reconciled.
- [Figure captions, Figures 1-4] The captions describe standard deviations as plotted curves, which is visually confusing; showing shaded confidence bands or reporting per-depth summary statistics in tables would make the uncertainty clearer and would also strengthen the paper's inferential claims.
- [Sections 4.2 and 4.4] The paper repeatedly attributes observed patterns to 'ads or sponsored content' (e.g., Figures 2c, 3c, 4c), but no detection method for ads or sponsored videos is described. The authors should either operationalize this category or avoid making specific claims about promotional insertions.
- [Section 3.1] The general content condition includes 'News & Politics' as one of its 15 categories, so it is not a clean non-political baseline; the near-zero political share at depth 0 in Figure 2 should be interpreted and reported with this composition in mind.
- [References] A few references are missing publication venues or complete author information (e.g., [11], [17], [25]); the bibliography should be brought into a consistent format.
Circularity Check
No significant circularity: the drift measurements are externally labeled observations, not fitted or definitional outputs.
full rationale
The paper's central claims (drift from political to entertainment content, emotion shift toward joy/neutral, engagement amplification, and watch-time insensitivity) are direct empirical measurements of recommendation outputs collected from fresh, logged-out Selenium sessions. The labels are produced by GPT-4o, which was selected by validation on external benchmark datasets (GoEmotions, DailyDialog, BBC News, News Category, MS MARCO, WikiQA) reported in Tables 3-4; no classifier weight or threshold is fitted to the YouTube Shorts data whose drift is then reported, so the 'prediction' is not forced by construction. The relevance, topic, and emotion categories are operational definitions, and the finding that depth-0 political seeds lose relevance at depth 1 is an observed property of the recommendation chain, not an analytic consequence of the definition. The paper's own Section 4.2 notes that the drift 'likely reflects engagement optimization and popularity bias,' which is a substantive interpretation rather than a circular reduction to the seed keywords. Self-citations (refs [8]-[10], [12], [20]-[21]) appear in the literature review and methodology for prior findings and transcript tooling, but the load-bearing conclusion does not depend on accepting those prior results: the current study's data collection and external model validation stand independently. The absence of a matched non-political, non-entertainment control is a validity and generalizability concern about confounding with base-rate popularity bias, not a circularity in the derivation chain. No equation or fitted parameter is reused as an output, so score 0 is appropriate.
Assumptions & free parameters
free parameters (2)
- watch-time thresholds =
3s, 15s, 60s
- recommendation depth limit =
50
assumptions (4)
- domain assumption GPT-4o classifications are valid for YouTube Shorts content
- domain assumption Pausing 3s/15s/60s in a headless browser is interpreted by YouTube as genuine watch-time signals
- domain assumption Recommendation chain depth ordering reflects YouTube's recommendation priority
- domain assumption A fresh WebDriver with no cookies and no logged-in user is a neutral environment
Cite this review
Pith. "Pith review of Investigating Algorithmic Bias in YouTube Shorts." pith.science (2026). https://pith.science/paper/D2XSLZED
@misc{pith2026250704605,
author = {Pith},
title = {Pith review of: Investigating Algorithmic Bias in YouTube Shorts},
year = {2026},
howpublished = {\url{https://pith.science/paper/D2XSLZED}},
note = {Machine review of arXiv:2507.04605}
}
read the original abstract
The rapid growth of YouTube Shorts, now serving over 2 billion monthly users, reflects a global shift toward short-form video as a dominant mode of online content consumption. This study investigates algorithmic bias in YouTube Shorts' recommendation system by analyzing how watch-time duration, topic sensitivity, and engagement metrics influence content visibility and drift. We focus on three content domains: the South China Sea dispute, the 2024 Taiwan presidential election, and general YouTube Shorts content. Using generative AI models, we classified 685,842 videos across relevance, topic category, and emotional tone. Our results reveal a consistent drift away from politically sensitive content toward entertainment-focused videos. Emotion analysis shows a systematic preference for joyful or neutral content, while engagement patterns indicate that highly viewed and liked videos are disproportionately promoted, reinforcing popularity bias. This work provides the first comprehensive analysis of algorithmic drift in YouTube Shorts based on textual content, emotional tone, topic categorization, and varying watch-time conditions. These findings offer new insights into how algorithmic design shapes content exposure, with implications for platform transparency and information diversity.
Forward citations
Cited by 1 Pith paper
-
Algorithmic Audit of Personalisation Drift in Polarising Topics on TikTok
TikTok personalises strongly, neutralises climate/vaccine/flat-earth content toward safe neutral topics, but sustains and often stance-reinforces US politics—with a tilt toward the oppose stance when both sides are seeded.
Reference graph
Works this paper leans on
-
[10]
Social Network Analysis and Mining 14(1), 1–42 (2024)
Cakmak, M.C., Agarwal, N., Oni, R.: The bias beneath: analyzing drift in youtube’s algorithmic recommendations. Social Network Analysis and Mining 14(1), 1–42 (2024)
work page 2024
-
[12]
Cakmak, M.C., Agarwal, N., Dagtas, S., Poudel, D.: Unveiling bias in youtube shorts: Analyzing thumbnail recommendations and topic dynamics. In: International Confer- ence on Social Computing, Behavioral-Cultural Modeling and Prediction and Behavior Representation in Modeling and Simulation, pp. 205–215 (2024). Springer
work page 2024
-
[1]
Shorts on the Rise: Assessing the Effects of YouTube Shorts on Long-Form Video Content
Rajendran, P.T., Creusy, K., Garnes, V.: Shorts on the rise: Assessing the effects of youtube shorts on long-form video content. arXiv preprint arXiv:2402.18208 (2024)
work page Pith review arXiv 2024
-
[2]
Aubin, C., Liedke, J.: Social Media and News Fact Sheet
St. Aubin, C., Liedke, J.: Social Media and News Fact Sheet. Pew Research Center. Accessed: October 3, 2024 (2024). https://www.pewresearch.org/journalism/fact-sheet/ social-media-and-news-fact-sheet/
work page 2024
-
[3]
Violot, C., Elmas, T., Bilogrevic, I., Humbert, M.: Shorts vs. regular videos on youtube: A comparative analysis of user engagement and content creation trends. arXiv preprint arXiv:2403.00454 (2024)
work page Pith review arXiv 2024
-
[4]
Michigan Journal of International Law 45(1), 93–154 (2024)
Seo, Y.: Power shift, the south china sea dispute, and the role of international law. Michigan Journal of International Law 45(1), 93–154 (2024)
work page 2024
-
[5]
The New York Times: How Taiwan’s Election Fits Into the Island’s Past, and Its Future. The New York Times. Accessed: 2024-06-30 (2024). https://www.nytimes.com/2024/01/ 14/world/asia/taiwan-election-china-lai-ching-te.html
work page 2024
-
[6]
In: Boratto, L., Faralli, S., Marras, M., Stilo, G
Kirdemir, B., Kready, J., Mead, E., Hussain, M.N., Agarwal, N.: Examining video rec- ommendation bias on youtube. In: Boratto, L., Faralli, S., Marras, M., Stilo, G. (eds.) Advances in Bias and Fairness in Information Retrieval, pp. 106–116. Springer, Cham (2021)
work page 2021
Show all 28 references
-
[7]
Applied Network Science 10(1), 1–38 (2025)
Cakmak, M.C., Agarwal, N.: Influence of symbolic content on recommendation bias: ana- lyzing youtube’s algorithm during taiwan’s 2024 election. Applied Network Science 10(1), 1–38 (2025)
2025
-
[8]
In: Proceedings of the International Conference on Advances in Social Networks Analysis and Mining, pp
Cakmak, M.C., Okeke, O., Onyepunuka, U., Spann, B., Agarwal, N.: Analyzing bias in rec- ommender systems: A comprehensive evaluation of youtube’s recommendation algorithm. In: Proceedings of the International Conference on Advances in Social Networks Analysis and Mining, pp. 7...
2024
-
[9]
In: Cherifi, H., Rocha, L.M., Cherifi, C., Donduran, M
Cakmak, M.C., Okeke, O., Onyepunuka, U., Spann, B., Agarwal, N.: Investigating bias in youtube recommendations: Emotion, morality, and network dynamics in china-uyghur content. In: Cherifi, H., Rocha, L.M., Cherifi, C., Donduran, M. (eds.) Complex Networks & Their Applications...
2024
-
[11]
Yang, C.: Bias in short-video recommender systems: user-centric evaluation on tiktok (2022) 14
2022
-
[13]
Hunnego, M.: Exploring the consistency and variability of algorithmic filter bubbles: A comparative analysis of instagram reels and tiktok. B.S. thesis, University of Twente (2024)
2024
-
[14]
In: Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, pp
Zhang, Y., Bai, Y., Chang, J., Zang, X., Lu, S., Lu, J., Feng, F., Niu, Y., Song, Y.: Lever- aging watch-time feedback for short-video recommendations: A causal labeling framework. In: Proceedings of the 32nd ACM International Conference on Information and Knowledge Management...
2023
-
[15]
PBS News
Klepper, D., Wu, H.: How Taiwan preserved election integrity by fighting back against disinformation. PBS News. Accessed: 2024-06-30 (2024). https://tinyurl.com/46n27r5w
2024
-
[16]
https://www.atlanticcouncil.org/ programs/digital-forensic-research-lab/
Atlantic Council: Digital Forensic Research Lab. https://www.atlanticcouncil.org/ programs/digital-forensic-research-lab/. Accessed: 2025-02-25 (2025)
2025
-
[17]
https://cil.nus.edu.sg/research/ocean-law-policy/south-china-sea/
Centre for International Law, National University of Singapore: Ocean Law and Pol- icy – South China Sea. https://cil.nus.edu.sg/research/ocean-law-policy/south-china-sea/. Accessed: 2025-02-25 (2025)
2025
-
[18]
https://entreresource.com/ youtube-video-categories-full-list-explained-and-which-you-should-use/
EntreResource: YouTube Video Categories – Full List Explained and Which You Should Use in 2023. https://entreresource.com/ youtube-video-categories-full-list-explained-and-which-you-should-use/. Accessed: 2025-02-25 (2023)
2023
-
[19]
Streamers: Youtube Scraper. APIFY. Accessed: 2024-01-10. https://apify.com/streamers/ youtube-scraper
2024
-
[20]
In: 2023 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp
Cakmak, M.C., Okeke, O., Spann, B., Agarwal, N.: Adopting parallel processing for rapid generation of transcripts in multimedia-rich online information environment. In: 2023 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp. 832–837 (2023)...
2023
-
[21]
In: 2024 IEEE Interna- tional Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp
Cakmak, M.C., Agarwal, N.: High-speed transcript collection on multimedia platforms: Advancing social media research through parallel processing. In: 2024 IEEE Interna- tional Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp. 857–860 (2024). https://doi.org...
2024
-
[22]
IEEE Engineering Management Review 52(2), 153– 164 (2024) https://doi.org/10.1109/EMR.2024.3353338
Joosten, J., Bilgram, V., Hahn, A., Totzek, D.: Comparing the ideation quality of humans with generative artificial intelligence. IEEE Engineering Management Review 52(2), 153– 164 (2024) https://doi.org/10.1109/EMR.2024.3353338
2024
-
[23]
https://arxiv.org/abs/2005.00547
Demszky, D., Movshovitz-Attias, D., Ko, J., Cowen, A., Nemade, G., Ravi, S.: GoEmotions: A Dataset of Fine-Grained Emotions (2020). https://arxiv.org/abs/2005.00547
2020 arXiv
-
[24]
https://arxiv.org/abs/1710.03957
Li, Y., Su, H., Shen, X., Li, W., Cao, Z., Niu, S.: DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset (2017). https://arxiv.org/abs/1710.03957
2017 arXiv
-
[25]
https://www.kaggle.com/datasets/bhavikjikadara/ bbc-news-articles
Jikadara, B.: BBC News Articles. https://www.kaggle.com/datasets/bhavikjikadara/ bbc-news-articles. Kaggle (2023)
2023
-
[26]
https://arxiv.org/abs/2209.11429
Misra, R.: News Category Dataset (2022). https://arxiv.org/abs/2209.11429
2022 arXiv
-
[27]
Nguyen, T., Rosenberg, M., Song, X., Gao, J., Tiwary, S., Majumder, R., Deng, L.: Ms marco: A human-generated machine reading comprehension dataset (2016)
2016
-
[28]
In: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp
Yang, Y., Yih, W.-t., Meek, C.: Wikiqa: A challenge dataset for open-domain question answering. In: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 2013–2018 (2015) 15
2015
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.