REVIEW 3 major objections 5 minor 74 references
Auditing News Curation Systems: A Case Study Examining Algorithmic and Editorial Logic in Apple News
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Apple News's algorithmic Trending Stories show no personalization or localization and concentrate on fewer and softer-news sources than the human-curated Top Stories.
desk verdict A worthwhile audit of Apple News with a genuine simulator-equivalence caveat that deserves peer review but needs tighter validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The personalization test asked dozens of workers to screenshot the app at the same times. Nearly every screenshot showed the same headlines as a control device, and no story appeared for one user only. A simulated-location test found no differences across all 50 US state capitals. So Trending Stories appear to be a single national list for everyone. Over two months, the algorithmically selected stories came from a more concentrated set of sources: the top three sources made up 45% of all stories, versus 24% for the human-edited section. The algorithm also featured more celebrity and entertainment stories, while the editors chose more policy and international news.
The paper is an example of algorithm auditing, treating a proprietary platform as a black box and observing it from the outside. It also offers a reusable framework and open-source code for future audits of similar news curators.
Extended reading notes
Core claim
Results showed that the human-curated Top Stories section features fewer stories per day and exhibits greater source diversity and greater source evenness compared to the algorithmically-curated Trending Stories section.
Load-bearing premise
The audit assumes that the Appium-controlled iPhone simulator presents the same content that real iPhones would show in the United States. If Apple serves different content to simulators (for example, test modes, different device identifiers, or time-dependent A/B variants), the findings on personalization, localization, and content distributions would not generalize to actual users. This assumption enters in the 'Sock Puppet Auditing via Appium' section and is used for the extended data collection in Experiment 4, which collected all content from a single simulated device.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an audit of Apple News, focusing on the algorithmically curated Trending Stories section and the human-editorially curated Top Stories section. It introduces a three-part audit framework (mechanism, content, consumption), then applies it through four experiments: update frequency of Trending Stories, a GPS-based localization test using Appium simulators, a crowdsourced personalization test with synchronized Mechanical Turk screenshots, and a 62-day automated collection from a single simulator comparing the two sections. The headline findings are that Trending Stories show no meaningful personalization or localization in the tested settings, that the algorithmic section has higher source concentration and lower Shannon equitability than the editorial section, and that the algorithmic section skews toward soft news while the editorial section skews toward policy and international news. The paper interprets these differences as manifestations of algorithmic versus editorial logic.
Significance. If the findings hold, this is a valuable early empirical characterization of Apple News, a platform with substantial gatekeeping power but little public data. The study is methodologically constructive: the code is released, the measurements are computed directly from observed data without fitted parameters, and the structure of the personalization experiment (synchronized crowd screenshots compared with a control simulator) is a useful template for auditing closed platforms. The diversity comparison is supported by an appropriate significance test (Hutcheson's t, t=11.17, p<0.001). The main limitation is that the extended content audit rests on a single simulated device with only limited real-device validation, so the generalizability of the central source-diversity claim is not yet fully established.
major comments (3)
- [Experiment 4 / Results: Content / Table 2] The central content comparison (Table 2, Figures 3–4) is based on a single Appium-controlled simulator over 62 days, but the only real-device validation reported is Experiment 3, which covers Trending Stories at three synchronized time points and does not validate Top Stories or the full collection period. If the simulator receives different content than production iPhones (due to device/profile signals, IP-based location, or A/B variants), the source diversity and evenness differences could be artifacts of the measurement channel. I ask the authors to add validation (for example, periodic real-device screenshots for both sections during the 62-day window) or to explicitly constrain the claims to what the single-device data can support.
- [Experiment 2: Testing for location-based adaptation] The localization test manipulates only the simulator's GPS coordinates; the manuscript does not report varying the host IP address or Apple ID across conditions. Because Apple News may infer location from IP or account information rather than GPS, the null result does not rule out location-based adaptation. The paper should soften 'no evidence of localization' to 'no evidence of GPS-based localization' and discuss this residual confound in the limitations.
- [Discussion: Algorithmic vs. Editorial Logic] The comparison is observational: Top Stories and Trending Stories differ not only in human versus algorithmic curation but also in section purpose (top stories versus trending content), number of slots, and update frequency. The abstract and conclusion attribute the diversity difference to curation type ('human curation outperformed algorithmic curation'), but the design does not isolate curation logic from these other factors. I recommend adding an explicit caveat that the observed differences are consistent with, but not uniquely caused by, the editorial-versus-algorithmic distinction.
minor comments (5)
- [Table 2] The text reports 1,268 Top Stories collected, while Table 2 lists 1,267; the discrepancy should be reconciled.
- [Table 1] The row 'meghan markle 33*' in the Trending Stories n-gram table is confusing: the asterisk footnote says the n-gram appeared twice in the other section, but the row is placed among Trending n-grams. Please clarify the intended placement and annotation.
- [Abstract] The claim of being 'the first data-backed characterization of Apple News in the United States' is strong; given prior work by Brown (2018a, 2018b) on Apple News, consider softening the novelty claim or specifying more precisely what aspect is new.
- [Experiment 1] The distinction between platform-wide and user-specific update frequency is clear in principle, but the finding that the app must be closed to see updates is important and should be highlighted explicitly in the main results rather than only in the narrative.
- [Audit Methods] The paper does not state the iOS version or Apple News version used in the simulator; since platform behavior can change between releases, including this information would improve reproducibility and comparability.
Assumptions & free parameters
assumptions (4)
- domain assumption The Appium iPhone simulator presents content identical to real iPhones for Apple News.
- domain assumption Trending Stories is algorithmically curated and Top Stories is editorially curated.
- domain assumption The n-gram log-ratio procedure identifies topic salience, and the resulting categories reflect 'soft news' and 'hard news'.
- standard math The Hutcheson t-test assumptions are appropriate for comparing Shannon diversity indices between the two sections.
Cite this review
Pith. "Pith review of Auditing News Curation Systems: A Case Study Examining Algorithmic and Editorial Logic in Apple News." pith.science (2026). https://pith.science/paper/WPNE6OLW
@misc{pith2026190800456,
author = {Pith},
title = {Pith review of: Auditing News Curation Systems: A Case Study Examining Algorithmic and Editorial Logic in Apple News},
year = {2026},
howpublished = {\url{https://pith.science/paper/WPNE6OLW}},
note = {Machine review of arXiv:1908.00456}
}
read the original abstract
This work presents an audit study of Apple News as a sociotechnical news curation system that exercises gatekeeping power in the media. We examine the mechanisms behind Apple News as well as the content presented in the app, outlining the social, political, and economic implications of both aspects. We focus on the Trending Stories section, which is algorithmically curated, and the Top Stories section, which is human-curated. Results from a crowdsourced audit showed minimal content personalization in the Trending Stories section, and a sock-puppet audit showed no location-based content adaptation. Finally, we perform an extended two-month data collection to compare the human-curated Top Stories section with the algorithmically curated Trending Stories section. Within these two sections, human curation outperformed algorithmic curation in several measures of source diversity, concentration, and evenness. Furthermore, algorithmic curation featured more "soft news" about celebrities and entertainment, while editorial curation featured more news about policy and international events. To our knowledge, this study provides the first data-backed characterization of Apple News in the United States.
Figures
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence 'output.state := if if FUNCTION not #0 #1 if FUNCTION and 'skip pop #0 if FUNCTIO...
- [2]
-
[3]
Bagdikian, B. H. 1983. The media monopoly . Beacon Press
work page 1983
-
[4]
Baker, P., and Potts, A. 2013. 'Why do white people have thin lips?' Google and the perpetuation of stereotypes via auto-complete search forms . Critical Discourse Studies 10(2):187--204
work page 2013
-
[5]
Bakshy, E.; Messing, S.; and Adamic, L. A. 2015. Exposure to ideologically diverse news and opinion on Facebook (supplementary materials) . Science 348(6239):1130--1132
work page 2015
-
[6]
Bandy, J., and Diakopoulos, N. 2019. Getting to the Core of Algorithmic News Aggregators: Applying a crowdsourced audit to the trending stories section of Apple News . In Computation + Journalism Symposium
work page 2019
-
[7]
Bawden, D., and Robinson, L. 2009. The dark side of information: Overload, anxiety and other paradoxes and pathologies . Journal of Information Science 35(2):180--191
work page 2009
-
[8]
Bernstein, M. S.; Karger, D. R.; Miller, R. C.; and Brandt, J. 2012. Analytic Methods for Optimizing Realtime Crowdsourcing . In Proceedings of Collective Intelligence 2012
work page 2012
Show all 74 references
-
[9]
Bozdag, E., and van den Hoven, J. 2015. Breaking the filter bubble: democracy and design . Ethics and Information Technology 17(4):249--265
2015
-
[10]
Brown, P. 2018a. Apple News UK editors rely on six outlets for 75 percent of Top Stories . Columbia Journalism Review
-
[11]
Brown, P. 2018b. Study: Apple News's human editors prefer a few major newsrooms . Columbia Journalism Review
-
[12]
Buolamwini, J. 2018. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification . In Proceedings of Machine Learning Research , volume 81, 1--15
2018
-
[13]
Burch, S. 2018. Tim Cook Explains Why Apple News Needs Human Editors . The Wrap
2018
-
[14]
J.; and Narayanan, A
Caliskan, A.; Bryson, J. J.; and Narayanan, A. 2017. Semantics derived automatically from language corpora contain human-like biases . Science 356(6334):183--186
2017
-
[15]
Chakraborty, A.; Ghosh, S.; Ganguly, N.; and Gummadi, K. P. 2015. Can Trending News Stories Create Coverage Bias? On the Impact of High Content Churn in Online News Media . In Computation + Journalism Symposium
2015
-
[16]
Chaney, A. J. B.; Stewart, B. M.; and Engelhardt, B. E. 2019. How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility . In RecSys
2019
-
[17]
Chen, L.; Mislove, A.; and Wilson, C. 2016. An Empirical Analysis of Algorithmic Pricing on Amazon Marketplace . In Proceedings of the 25th International Conference on World Wide Web - WWW '16 , 1339--1349
2016
-
[18]
L.; Nematzadeh, A.; Menczer, F.; and Flammini, A
Ciampaglia, G. L.; Nematzadeh, A.; Menczer, F.; and Flammini, A. 2018. How algorithmic popularity bias hinders or promotes quality . Scientific Reports 8(1):15951
2018
-
[19]
Davies, J. 2017. The Guardian pulls out of Facebook's Instant Articles and Apple News . Digiday
2017
-
[20]
Diakopoulos, N.; Trielli, D.; Stark, J.; and Mussenden, S. 2018. I Vote For-How Search Informs Our Choice of Candidate . Digital Dominance: The Power of Google, Amazon, Facebook, and Apple
2018
-
[21]
Diakopoulos, N. 2015. Algorithmic Accountability: Journalistic investigation of computational power structures . Digital Journalism 3(3):398--415
2015
-
[22]
Diakopoulos, N. 2019. Automating the News: How Algorithms Are Rewriting the Media . Harvard University Press
2019
-
[23]
Dotan, T. 2018. Inside Apple's Courtship of News Publishers . The Information
2018
-
[24]
Entman, R. 1993. Framing - Toward Clarification of a Fractured Paradigm . Journal of Communication 43(4):51--58
1993
-
[25]
Epstein, R., and Robertson, R. E. 2017. A method for detecting bias in search rankings, with evidence of systematic bias related to the 2016 presidential election . Technical report, Technical Report White Paper no. WP-17-02. American Institute for Behavioral Research and Technology
2017
-
[26]
Feiner, L. 2019. Apple reports services margins along weaker iPhone sales for Q1 2019
2019
-
[27]
R.; Goel, S.; and Rao, J
Flaxman, S. R.; Goel, S.; and Rao, J. M. 2016. Filter Bubbles, Echo Chambers, and Online News Consumption . Public Opinion Quarterly 80(S1):298----320
2016
-
[28]
Friedman, B., and Nissenbaum, H. 1996. Bias in computer systems . ACM Transactions on Information Systems 14(3):330--347
1996
-
[29]
Gaither, C. 2005. Web Giants Go With Different Angles in Competition for News Audience . Los Angeles Times
2005
-
[30]
S.; and Smith, J
Garfinkel, S.; Matthews, J.; Shapiro, S. S.; and Smith, J. M. 2017. Toward algorithmic transparency and accountability . Communications of the ACM 60(9):5--5
2017
-
[31]
Gillespie, T. 2014. The Relevance of Algorithms . Media technologies: Essays on communication, materiality, and society 167:167--194
2014
-
[32]
Gillespie, T. 2018. Custodians of the internet: Platforms, content moderation, and the hidden decisions that shape social media . Yale University Press
2018
-
[33]
Haim, M.; Graefe, A.; and Brosius, H. B. 2018. Burst of the Filter Bubble?: Effects of personalization on the diversity of Google News . Digital Journalism 6(3):330--343
2018
-
[34]
z y \' n ski, P.; Khaki, A
Hann \' a k, A.; Sapie \. z y \' n ski, P.; Khaki, A. M.; Lazer, D.; Mislove, A.; and Wilson, C. 2013. Measuring Personalization of Web Search . In Proceedings of the 22nd international conference on World Wide Web , 527----538. ACM
2013
-
[35]
Hannak, A.; Soeller, G.; Lazer, D.; Mislove, A.; and Wilson, C. 2014. Measuring Price Discrimination and Steering on E-commerce Web Sites . In Proceedings of the 2014 Conference on Internet Measurement Conference - IMC '14 , 305--318
2014
-
[36]
Hara, K.; Adams, A.; Milland, K.; Savage, S.; Callison-Burch, C.; and Bigham, J. P. 2018. A Data-Driven Analysis of Workers' Earnings on Amazon Mechanical Turk . In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems - CHI '18 , 1--14
2018
-
[37]
Hardie, A. 2014. Log Ratio: An informal introduction . ESRC Centre for Corpus Approaches to Social Science (CASS)
2014
-
[38]
Helberger, N.; Karppinen, K.; and D'Acunto, L. 2018. Exposure diversity as a design principle for recommender systems . Information Communication and Society 21(2):191--207
2018
-
[39]
Himma, K. E. 2007. The concept of information overload: A preliminary step in understanding the nature of a harmful information-related condition . Ethics and Information Technology 9(4):259--272
2007
-
[40]
Hutcheson, K. 1970. A test for comparing diversities based on the shannon formula . Journal of Theoretical Biology 29(1)
1970
-
[41]
D., and Nissenbaum, H
Introna, L. D., and Nissenbaum, H. 2000. Shaping the web: Why the politics of search engines matters . Information Society 16(3):169--185
2000
-
[42]
Kay, M.; Matuszek, C.; and Munson, S. A. 2015. Unequal Representation and Gender Stereotypes in Image Search Results for Occupations . In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems - CHI '15 , 3819--3828
2015
-
[43]
Kliman-Silver, C.; Hannak, A.; Lazer, D.; Wilson, C.; and Mislove, A. 2015. Location, Location, Location: The Impact of Geolocation on Web Search Personalization . In Proceedings of the 2015 Internet Measurement Conference , 121----127. ACM
2015
-
[44]
Kwak, H.; Lee, C.; Park, H.; and Moon, S. 2010. What is Twitter, a Social Network or a News Media? In Proceedings of the 19th international conference on World wide web , 591--600. ACM
2010
-
[45]
E., and Shaw, D
McCombs, M. E., and Shaw, D. L. 1972. The Agenda-Setting Function of Mass Media . Public Opinion Quarterly 36(2):176
1972
-
[46]
McCombs, M. 2005. A Look at Agenda-setting: past, present and future . Journalism Studies 6(4):543--557
2005
-
[47]
M \" o ller, J.; Trilling, D.; Helberger, N.; and van Es, B. 2018. Do not blame it on the algorithm: an empirical assessment of multiple recommender systems and their impact on content diversity . Information Communication and Society 21(7):959--977
2018
-
[48]
Napoli, P. 2011. Exposure Diversity Reconsidered . Journal of Information Policy 1(4):246--259
2011
-
[49]
Nechushtai, E., and Lewis, S. C. 2019. What kind of news gatekeepers do we want machines to be? Filter bubbles, fragmentation, and the normative dimensions of algorithmic recommendations . Computers in Human Behavior 90:298--307
2019
-
[50]
Nicas, J. 2018. Apple News's Radical Approach: Humans Over Machines . New York Times
2018
-
[51]
O'Neil, C. 2017. Weapons of math destruction : how big data increases inequality and threatens democracy . Broadway Books
2017
-
[52]
Oremus, W. 2018. The Temptation of Apple News . Slate
2018
-
[53]
Pariser, E. 2011. The filter bubble: how the new personalized Web is changing what we read and how we think . Penguin
2011
-
[54]
Pielou, E. C. 1966. The Measurement of Diversity in Different Types of Biological Collections . Journal of Theoretical Biology 13(C):131--144
1966
-
[55]
Puschmann, C. 2018. Beyond the bubble: Assessing the diversity of political search results . Digital Journalism 1--20
2018
-
[56]
Reinemann, C.; Stanyer, J.; Scherr, S.; and Legnante, G. 2012. Hard and soft news: A review of concepts, operationalizations and key findings
2012
-
[57]
E.; Jiang, S.; Joseph, K.; Friedland, L.; Lazer, D.; and Wilson, C
Robertson, R. E.; Jiang, S.; Joseph, K.; Friedland, L.; Lazer, D.; and Wilson, C. 2018. Auditing Partisan Audience Bias within Google Search . Proceedings of the ACM on Human-Computer Interaction 2(CSCW):1--22
2018
-
[58]
E.; Lazer, D.; and Wilson, C
Robertson, R. E.; Lazer, D.; and Wilson, C. 2018. Auditing the Personalization and Composition of Politically-Related Search Engine Results Pages . In WWW 2018: The 2018 Web Conference , 955--965. ACM
2018
-
[59]
J.; Dodds, P
Salganik, M. J.; Dodds, P. S.; and Watts, D. J. 2006. Experimental study of inequality and unpredictability in an artificial cultural market . Science 311(5762):854--856
2006
-
[60]
Sandvig, C.; Hamilton, K.; Karahalios, K.; and Langbort, C. 2014. Auditing Algorithms: Research Methods for Detecting Discrimination on Internet Platforms . In Data and discrimination: converting critical concerns into productive inquiry , 1----23
2014
-
[61]
Schroeder, R., and Kralemann, M. 2005. Journalism Ex Machina--Google News Germany and its news selection processes . Journalism Studies 6(2):245--247
2005
-
[62]
Schudson, M. 1995. The Power of News . Harvard University Press
1995
-
[63]
Seaver, N. 2017. Algorithms as culture: Some tactics for the ethnography of algorithmic systems . Big Data & Society 4(2)
2017
-
[64]
Shannon, C. E. 1948. A Mathematical Theory of Communication . The Bell System Technical Journal Vol.27(1948):379--423
1948
-
[65]
J., and Vos, T
Shoemaker, P. J., and Vos, T. P. 2009. Gatekeeping theory . Routledge
2009
-
[66]
Tran, K. 2018. Apple News reaches 90 million readers . Business Insider
2018
-
[67]
Trielli, D., and Diakopoulos, N. 2019. Search as News Curator: The Role of Google in Shaping Attention to News Information . In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . ACM
2019
-
[68]
Ulken, E. 2005. A Question of Balance: Are Google News search results politically biased? Ph.D. Dissertation, USC Annenberg School for Communication
2005
-
[69]
Vijaymeena, M., and Kavitha, K. 2016. A Survey on Similarity Measures in Text Mining . Machine Learning and Applications: An International Journal 3(1):19--28
2016
-
[70]
Vincent, N.; Johnson, I.; Sheehan, P.; and Hecht, B. 2019. Measuring the Importance of User-Generated Content to Search Engines . In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM)
2019
-
[71]
Wang, S. 2018. Google News gets an update, with more AI-driven curation, more labeling, more reader controls (fingers crossed for no AI flops)
2018
-
[72]
Webber, W.; Moffat, A.; and Zobel, J. 2010. A similarity measure for indefinite rankings . ACM Transactions on Information Systems 28(4)
2010
-
[73]
Weiss, M. 2018. Digiday Research poll: Apple News dominates publisher focus after Facebook's news feed changes - Digiday . Digiday
2018
-
[74]
Weizenbaum, J. 1976. Computer power and human reason: from judgment to calculation . WH Freeman & Co
1976
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.