Pith. sign in

REVIEW 3 major objections 5 minor 74 references

Auditing News Curation Systems: A Case Study Examining Algorithmic and Editorial Logic in Apple News

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Apple News's algorithmic Trending Stories show no personalization or localization and concentrate on fewer and softer-news sources than the human-curated Top Stories.

desk verdict A worthwhile audit of Apple News with a genuine simulator-equivalence caveat that deserves peer review but needs tighter validation. read the letter →

arxiv 1908.00456 v1 pith:WPNE6OLW submitted 2019-08-01 cs.CY cs.HC

classification cs.CYcs.HC
keywords newscurationsectionstoriesapplealgorithmicauditcontent
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Apple News is a news app on iPhones. It has two prominent sections: Top Stories, chosen by human editors, and Trending Stories, chosen by an algorithm. This study audited both sections to learn how they work and what they show. The researchers used three methods: crowdsourced screenshots from real users, automated 'sock puppet' simulators, and a two-month continuous scrape.

The personalization test asked dozens of workers to screenshot the app at the same times. Nearly every screenshot showed the same headlines as a control device, and no story appeared for one user only. A simulated-location test found no differences across all 50 US state capitals. So Trending Stories appear to be a single national list for everyone. Over two months, the algorithmically selected stories came from a more concentrated set of sources: the top three sources made up 45% of all stories, versus 24% for the human-edited section. The algorithm also featured more celebrity and entertainment stories, while the editors chose more policy and international news.

The paper is an example of algorithm auditing, treating a proprietary platform as a black box and observing it from the outside. It also offers a reusable framework and open-source code for future audits of similar news curators.

Extended reading notes

Core claim

Results showed that the human-curated Top Stories section features fewer stories per day and exhibits greater source diversity and greater source evenness compared to the algorithmically-curated Trending Stories section.

Load-bearing premise

The audit assumes that the Appium-controlled iPhone simulator presents the same content that real iPhones would show in the United States. If Apple serves different content to simulators (for example, test modes, different device identifiers, or time-dependent A/B variants), the findings on personalization, localization, and content distributions would not generalize to actual users. This assumption enters in the 'Sock Puppet Auditing via Appium' section and is used for the extended data collection in Experiment 4, which collected all content from a single simulated device.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents an audit of Apple News, focusing on the algorithmically curated Trending Stories section and the human-editorially curated Top Stories section. It introduces a three-part audit framework (mechanism, content, consumption), then applies it through four experiments: update frequency of Trending Stories, a GPS-based localization test using Appium simulators, a crowdsourced personalization test with synchronized Mechanical Turk screenshots, and a 62-day automated collection from a single simulator comparing the two sections. The headline findings are that Trending Stories show no meaningful personalization or localization in the tested settings, that the algorithmic section has higher source concentration and lower Shannon equitability than the editorial section, and that the algorithmic section skews toward soft news while the editorial section skews toward policy and international news. The paper interprets these differences as manifestations of algorithmic versus editorial logic.

Significance. If the findings hold, this is a valuable early empirical characterization of Apple News, a platform with substantial gatekeeping power but little public data. The study is methodologically constructive: the code is released, the measurements are computed directly from observed data without fitted parameters, and the structure of the personalization experiment (synchronized crowd screenshots compared with a control simulator) is a useful template for auditing closed platforms. The diversity comparison is supported by an appropriate significance test (Hutcheson's t, t=11.17, p<0.001). The main limitation is that the extended content audit rests on a single simulated device with only limited real-device validation, so the generalizability of the central source-diversity claim is not yet fully established.

major comments (3)
  1. [Experiment 4 / Results: Content / Table 2] The central content comparison (Table 2, Figures 3–4) is based on a single Appium-controlled simulator over 62 days, but the only real-device validation reported is Experiment 3, which covers Trending Stories at three synchronized time points and does not validate Top Stories or the full collection period. If the simulator receives different content than production iPhones (due to device/profile signals, IP-based location, or A/B variants), the source diversity and evenness differences could be artifacts of the measurement channel. I ask the authors to add validation (for example, periodic real-device screenshots for both sections during the 62-day window) or to explicitly constrain the claims to what the single-device data can support.
  2. [Experiment 2: Testing for location-based adaptation] The localization test manipulates only the simulator's GPS coordinates; the manuscript does not report varying the host IP address or Apple ID across conditions. Because Apple News may infer location from IP or account information rather than GPS, the null result does not rule out location-based adaptation. The paper should soften 'no evidence of localization' to 'no evidence of GPS-based localization' and discuss this residual confound in the limitations.
  3. [Discussion: Algorithmic vs. Editorial Logic] The comparison is observational: Top Stories and Trending Stories differ not only in human versus algorithmic curation but also in section purpose (top stories versus trending content), number of slots, and update frequency. The abstract and conclusion attribute the diversity difference to curation type ('human curation outperformed algorithmic curation'), but the design does not isolate curation logic from these other factors. I recommend adding an explicit caveat that the observed differences are consistent with, but not uniquely caused by, the editorial-versus-algorithmic distinction.
minor comments (5)
  1. [Table 2] The text reports 1,268 Top Stories collected, while Table 2 lists 1,267; the discrepancy should be reconciled.
  2. [Table 1] The row 'meghan markle 33*' in the Trending Stories n-gram table is confusing: the asterisk footnote says the n-gram appeared twice in the other section, but the row is placed among Trending n-grams. Please clarify the intended placement and annotation.
  3. [Abstract] The claim of being 'the first data-backed characterization of Apple News in the United States' is strong; given prior work by Brown (2018a, 2018b) on Apple News, consider softening the novelty claim or specifying more precisely what aspect is new.
  4. [Experiment 1] The distinction between platform-wide and user-specific update frequency is clear in principle, but the finding that the app must be closed to see updates is important and should be highlighted explicitly in the main results rather than only in the narrative.
  5. [Audit Methods] The paper does not state the iOS version or Apple News version used in the simulator; since platform behavior can change between releases, including this information would improve reproducibility and comparability.
Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The audit does not introduce free parameters or invented entities. It relies on several domain assumptions about the platform and the validity of the measurement tools, the most important being that a simulator is a faithful proxy for real devices and that the section labels (algorithmic vs editorial) are accurate. These are external assumptions, not fitted from the data.

assumptions (4)
  • domain assumption The Appium iPhone simulator presents content identical to real iPhones for Apple News.
    All content-level results (localization, source diversity, topic n-grams) come from simulator data. The personalization experiment offers partial validation because 89.2% of real user screenshots matched the simulator control, but this does not establish full equivalence across all versions, accounts, or times.
  • domain assumption Trending Stories is algorithmically curated and Top Stories is editorially curated.
    Attributed to a New York Times report (Nicas, 2018) and Apple's public statements; the paper does not independently verify the provenance of the lists, and this attribution is the basis for comparing 'algorithmic logic' and 'editorial logic'.
  • domain assumption The n-gram log-ratio procedure identifies topic salience, and the resulting categories reflect 'soft news' and 'hard news'.
    The paper does not use an independent, validated classification of soft/hard news; the conclusion is based on inspecting the most salient bigrams and trigrams from each section (Table 1).
  • standard math The Hutcheson t-test assumptions are appropriate for comparing Shannon diversity indices between the two sections.
    The test is a standard method for comparing sample diversity. However, the stories are not independent random draws; they are time-series observations from a single channel, so the p-value should be interpreted cautiously.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Auditing News Curation Systems: A Case Study Examining Algorithmic and Editorial Logic in Apple News." pith.science (2026). https://pith.science/paper/WPNE6OLW

@misc{pith2026190800456,
  author       = {Pith},
  title        = {Pith review of: Auditing News Curation Systems: A Case Study Examining Algorithmic and Editorial Logic in Apple News},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WPNE6OLW}},
  note         = {Machine review of arXiv:1908.00456}
}
read the original abstract

This work presents an audit study of Apple News as a sociotechnical news curation system that exercises gatekeeping power in the media. We examine the mechanisms behind Apple News as well as the content presented in the app, outlining the social, political, and economic implications of both aspects. We focus on the Trending Stories section, which is algorithmically curated, and the Top Stories section, which is human-curated. Results from a crowdsourced audit showed minimal content personalization in the Trending Stories section, and a sock-puppet audit showed no location-based content adaptation. Finally, we perform an extended two-month data collection to compare the human-curated Top Stories section with the algorithmically curated Trending Stories section. Within these two sections, human curation outperformed algorithmic curation in several measures of source diversity, concentration, and evenness. Furthermore, algorithmic curation featured more "soft news" about celebrities and entertainment, while editorial curation featured more news about policy and international events. To our knowledge, this study provides the first data-backed characterization of Apple News in the United States.

Figures

Figures reproduced from arXiv: 1908.00456 by the authors.

Figure 1
Figure 1. Two screenshots from the Apple News app taken [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. A depiction of our two audit methods: crowd [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 5
Figure 5. Relative distribution of Trending Stories over the [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figures from the paper (2 more)
Figure 6
Figure 6. Figure 6: Relative distribution of Top Stories over the hours [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 4
Figure 4. Figure 4: Relative distribution of Top Stories across top ten [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

74 extracted references · 73 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence 'output.state := if if FUNCTION not #0 #1 if FUNCTION and 'skip pop #0 if FUNCTIO...

  2. [2]

    Getting Featured Stories Approved

    Apple Inc. Getting Featured Stories Approved

  3. [3]

    Bagdikian, B. H. 1983. The media monopoly . Beacon Press

  4. [4]

    Baker, P., and Potts, A. 2013. 'Why do white people have thin lips?' Google and the perpetuation of stereotypes via auto-complete search forms . Critical Discourse Studies 10(2):187--204

  5. [5]

    Bakshy, E.; Messing, S.; and Adamic, L. A. 2015. Exposure to ideologically diverse news and opinion on Facebook (supplementary materials) . Science 348(6239):1130--1132

  6. [6]

    Bandy, J., and Diakopoulos, N. 2019. Getting to the Core of Algorithmic News Aggregators: Applying a crowdsourced audit to the trending stories section of Apple News . In Computation + Journalism Symposium

  7. [7]

    Bawden, D., and Robinson, L. 2009. The dark side of information: Overload, anxiety and other paradoxes and pathologies . Journal of Information Science 35(2):180--191

  8. [8]

    S.; Karger, D

    Bernstein, M. S.; Karger, D. R.; Miller, R. C.; and Brandt, J. 2012. Analytic Methods for Optimizing Realtime Crowdsourcing . In Proceedings of Collective Intelligence 2012

Show all 74 references
  1. [9]

    Bozdag, E., and van den Hoven, J. 2015. Breaking the filter bubble: democracy and design . Ethics and Information Technology 17(4):249--265

  2. [10]

    Brown, P. 2018a. Apple News UK editors rely on six outlets for 75 percent of Top Stories . Columbia Journalism Review

  3. [11]

    Brown, P. 2018b. Study: Apple News's human editors prefer a few major newsrooms . Columbia Journalism Review

  4. [12]

    Buolamwini, J. 2018. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification . In Proceedings of Machine Learning Research , volume 81, 1--15

  5. [13]

    Burch, S. 2018. Tim Cook Explains Why Apple News Needs Human Editors . The Wrap

  6. [14]

    J.; and Narayanan, A

    Caliskan, A.; Bryson, J. J.; and Narayanan, A. 2017. Semantics derived automatically from language corpora contain human-like biases . Science 356(6334):183--186

  7. [15]

    Chakraborty, A.; Ghosh, S.; Ganguly, N.; and Gummadi, K. P. 2015. Can Trending News Stories Create Coverage Bias? On the Impact of High Content Churn in Online News Media . In Computation + Journalism Symposium

  8. [16]

    Chaney, A. J. B.; Stewart, B. M.; and Engelhardt, B. E. 2019. How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility . In RecSys

  9. [17]

    Chen, L.; Mislove, A.; and Wilson, C. 2016. An Empirical Analysis of Algorithmic Pricing on Amazon Marketplace . In Proceedings of the 25th International Conference on World Wide Web - WWW '16 , 1339--1349

  10. [18]

    L.; Nematzadeh, A.; Menczer, F.; and Flammini, A

    Ciampaglia, G. L.; Nematzadeh, A.; Menczer, F.; and Flammini, A. 2018. How algorithmic popularity bias hinders or promotes quality . Scientific Reports 8(1):15951

  11. [19]

    Davies, J. 2017. The Guardian pulls out of Facebook's Instant Articles and Apple News . Digiday

  12. [20]

    Diakopoulos, N.; Trielli, D.; Stark, J.; and Mussenden, S. 2018. I Vote For-How Search Informs Our Choice of Candidate . Digital Dominance: The Power of Google, Amazon, Facebook, and Apple

  13. [21]

    Diakopoulos, N. 2015. Algorithmic Accountability: Journalistic investigation of computational power structures . Digital Journalism 3(3):398--415

  14. [22]

    Diakopoulos, N. 2019. Automating the News: How Algorithms Are Rewriting the Media . Harvard University Press

  15. [23]

    Dotan, T. 2018. Inside Apple's Courtship of News Publishers . The Information

  16. [24]

    Entman, R. 1993. Framing - Toward Clarification of a Fractured Paradigm . Journal of Communication 43(4):51--58

  17. [25]

    Epstein, R., and Robertson, R. E. 2017. A method for detecting bias in search rankings, with evidence of systematic bias related to the 2016 presidential election . Technical report, Technical Report White Paper no. WP-17-02. American Institute for Behavioral Research and Technology

  18. [26]

    Feiner, L. 2019. Apple reports services margins along weaker iPhone sales for Q1 2019

  19. [27]

    R.; Goel, S.; and Rao, J

    Flaxman, S. R.; Goel, S.; and Rao, J. M. 2016. Filter Bubbles, Echo Chambers, and Online News Consumption . Public Opinion Quarterly 80(S1):298----320

  20. [28]

    Friedman, B., and Nissenbaum, H. 1996. Bias in computer systems . ACM Transactions on Information Systems 14(3):330--347

  21. [29]

    Gaither, C. 2005. Web Giants Go With Different Angles in Competition for News Audience . Los Angeles Times

  22. [30]

    S.; and Smith, J

    Garfinkel, S.; Matthews, J.; Shapiro, S. S.; and Smith, J. M. 2017. Toward algorithmic transparency and accountability . Communications of the ACM 60(9):5--5

  23. [31]

    Gillespie, T. 2014. The Relevance of Algorithms . Media technologies: Essays on communication, materiality, and society 167:167--194

  24. [32]

    Gillespie, T. 2018. Custodians of the internet: Platforms, content moderation, and the hidden decisions that shape social media . Yale University Press

  25. [33]

    Haim, M.; Graefe, A.; and Brosius, H. B. 2018. Burst of the Filter Bubble?: Effects of personalization on the diversity of Google News . Digital Journalism 6(3):330--343

  26. [34]

    z y \' n ski, P.; Khaki, A

    Hann \' a k, A.; Sapie \. z y \' n ski, P.; Khaki, A. M.; Lazer, D.; Mislove, A.; and Wilson, C. 2013. Measuring Personalization of Web Search . In Proceedings of the 22nd international conference on World Wide Web , 527----538. ACM

  27. [35]

    Hannak, A.; Soeller, G.; Lazer, D.; Mislove, A.; and Wilson, C. 2014. Measuring Price Discrimination and Steering on E-commerce Web Sites . In Proceedings of the 2014 Conference on Internet Measurement Conference - IMC '14 , 305--318

  28. [36]

    Hara, K.; Adams, A.; Milland, K.; Savage, S.; Callison-Burch, C.; and Bigham, J. P. 2018. A Data-Driven Analysis of Workers' Earnings on Amazon Mechanical Turk . In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems - CHI '18 , 1--14

  29. [37]

    Hardie, A. 2014. Log Ratio: An informal introduction . ESRC Centre for Corpus Approaches to Social Science (CASS)

  30. [38]

    Helberger, N.; Karppinen, K.; and D'Acunto, L. 2018. Exposure diversity as a design principle for recommender systems . Information Communication and Society 21(2):191--207

  31. [39]

    Himma, K. E. 2007. The concept of information overload: A preliminary step in understanding the nature of a harmful information-related condition . Ethics and Information Technology 9(4):259--272

  32. [40]

    Hutcheson, K. 1970. A test for comparing diversities based on the shannon formula . Journal of Theoretical Biology 29(1)

  33. [41]

    D., and Nissenbaum, H

    Introna, L. D., and Nissenbaum, H. 2000. Shaping the web: Why the politics of search engines matters . Information Society 16(3):169--185

  34. [42]

    Kay, M.; Matuszek, C.; and Munson, S. A. 2015. Unequal Representation and Gender Stereotypes in Image Search Results for Occupations . In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems - CHI '15 , 3819--3828

  35. [43]

    Kliman-Silver, C.; Hannak, A.; Lazer, D.; Wilson, C.; and Mislove, A. 2015. Location, Location, Location: The Impact of Geolocation on Web Search Personalization . In Proceedings of the 2015 Internet Measurement Conference , 121----127. ACM

  36. [44]

    Kwak, H.; Lee, C.; Park, H.; and Moon, S. 2010. What is Twitter, a Social Network or a News Media? In Proceedings of the 19th international conference on World wide web , 591--600. ACM

  37. [45]

    E., and Shaw, D

    McCombs, M. E., and Shaw, D. L. 1972. The Agenda-Setting Function of Mass Media . Public Opinion Quarterly 36(2):176

  38. [46]

    McCombs, M. 2005. A Look at Agenda-setting: past, present and future . Journalism Studies 6(4):543--557

  39. [47]

    M \" o ller, J.; Trilling, D.; Helberger, N.; and van Es, B. 2018. Do not blame it on the algorithm: an empirical assessment of multiple recommender systems and their impact on content diversity . Information Communication and Society 21(7):959--977

  40. [48]

    Napoli, P. 2011. Exposure Diversity Reconsidered . Journal of Information Policy 1(4):246--259

  41. [49]

    Nechushtai, E., and Lewis, S. C. 2019. What kind of news gatekeepers do we want machines to be? Filter bubbles, fragmentation, and the normative dimensions of algorithmic recommendations . Computers in Human Behavior 90:298--307

  42. [50]

    Nicas, J. 2018. Apple News's Radical Approach: Humans Over Machines . New York Times

  43. [51]

    O'Neil, C. 2017. Weapons of math destruction : how big data increases inequality and threatens democracy . Broadway Books

  44. [52]

    Oremus, W. 2018. The Temptation of Apple News . Slate

  45. [53]

    Pariser, E. 2011. The filter bubble: how the new personalized Web is changing what we read and how we think . Penguin

  46. [54]

    Pielou, E. C. 1966. The Measurement of Diversity in Different Types of Biological Collections . Journal of Theoretical Biology 13(C):131--144

  47. [55]

    Puschmann, C. 2018. Beyond the bubble: Assessing the diversity of political search results . Digital Journalism 1--20

  48. [56]

    Reinemann, C.; Stanyer, J.; Scherr, S.; and Legnante, G. 2012. Hard and soft news: A review of concepts, operationalizations and key findings

  49. [57]

    E.; Jiang, S.; Joseph, K.; Friedland, L.; Lazer, D.; and Wilson, C

    Robertson, R. E.; Jiang, S.; Joseph, K.; Friedland, L.; Lazer, D.; and Wilson, C. 2018. Auditing Partisan Audience Bias within Google Search . Proceedings of the ACM on Human-Computer Interaction 2(CSCW):1--22

  50. [58]

    E.; Lazer, D.; and Wilson, C

    Robertson, R. E.; Lazer, D.; and Wilson, C. 2018. Auditing the Personalization and Composition of Politically-Related Search Engine Results Pages . In WWW 2018: The 2018 Web Conference , 955--965. ACM

  51. [59]

    J.; Dodds, P

    Salganik, M. J.; Dodds, P. S.; and Watts, D. J. 2006. Experimental study of inequality and unpredictability in an artificial cultural market . Science 311(5762):854--856

  52. [60]

    Sandvig, C.; Hamilton, K.; Karahalios, K.; and Langbort, C. 2014. Auditing Algorithms: Research Methods for Detecting Discrimination on Internet Platforms . In Data and discrimination: converting critical concerns into productive inquiry , 1----23

  53. [61]

    Schroeder, R., and Kralemann, M. 2005. Journalism Ex Machina--Google News Germany and its news selection processes . Journalism Studies 6(2):245--247

  54. [62]

    Schudson, M. 1995. The Power of News . Harvard University Press

  55. [63]

    Seaver, N. 2017. Algorithms as culture: Some tactics for the ethnography of algorithmic systems . Big Data & Society 4(2)

  56. [64]

    Shannon, C. E. 1948. A Mathematical Theory of Communication . The Bell System Technical Journal Vol.27(1948):379--423

  57. [65]

    J., and Vos, T

    Shoemaker, P. J., and Vos, T. P. 2009. Gatekeeping theory . Routledge

  58. [66]

    Tran, K. 2018. Apple News reaches 90 million readers . Business Insider

  59. [67]

    Trielli, D., and Diakopoulos, N. 2019. Search as News Curator: The Role of Google in Shaping Attention to News Information . In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . ACM

  60. [68]

    Ulken, E. 2005. A Question of Balance: Are Google News search results politically biased? Ph.D. Dissertation, USC Annenberg School for Communication

  61. [69]

    Vijaymeena, M., and Kavitha, K. 2016. A Survey on Similarity Measures in Text Mining . Machine Learning and Applications: An International Journal 3(1):19--28

  62. [70]

    Vincent, N.; Johnson, I.; Sheehan, P.; and Hecht, B. 2019. Measuring the Importance of User-Generated Content to Search Engines . In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM)

  63. [71]

    Wang, S. 2018. Google News gets an update, with more AI-driven curation, more labeling, more reader controls (fingers crossed for no AI flops)

  64. [72]

    Webber, W.; Moffat, A.; and Zobel, J. 2010. A similarity measure for indefinite rankings . ACM Transactions on Information Systems 28(4)

  65. [73]

    Weiss, M. 2018. Digiday Research poll: Apple News dominates publisher focus after Facebook's news feed changes - Digiday . Digiday

  66. [74]

    Weizenbaum, J. 1976. Computer power and human reason: from judgment to calculation . WH Freeman & Co

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.