{"id":"c02f50aa-5b92-4a8c-a0b3-2e3ce5f5dd6b","arxiv_id":"1908.00456","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Apple News's algorithmic Trending Stories show no personalization or localization and concentrate on fewer and softer-news sources than the human-curated Top Stories.","lead":"This paper audited Apple News and found that its algorithmically selected Trending Stories are nearly identical for all users, with no location-based adaptation, and are less diverse in sources and heavier on soft news than the human-edited Top Stories. The findings matter because Apple News reaches over 85 million people, so the contrast between algorithmic and editorial curation shapes what news the public sees.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulator equivalence is the weak link: the two-month diversity audit relies on one Appium simulator, validated against real iPhones only for Trending Stories at three time points.","rationale":"The reader's CONDITIONAL verdict is appropriate. The paper's strongest claim is descriptive and the data do support it under the stated collection method, but that method's generalizability rests on the simulator being a faithful proxy for real iPhones. The crowd experiment (Experiment 3) is genuine independent support: 83 synchronized screenshots with near-total overlap show that for Trending Stories at those moments, real users and the simulator saw essentially the same set. That is a real strength. It does not, however, cover the full 62-day window, the Top Stories section, or location-dependent serving behavior. The localization experiment is a separate weak point because simulated GPS coordinates need not affect IP-based personalization. The proposed validation panel is a direct, low-cost check: if simulator and real-device feeds converge, the diversity result stands; if they diverge, the headline claim needs to be qualified to a single simulated device or re-estimated from a device panel. I therefore keep the existing CONDITIONAL verdict rather than moving it.","tokens_in":17375,"tokens_out":6724,"duration_ms":76143,"concrete_test":"Recruit a small panel of 10-20 iPhone users in at least five U.S. states and have them run the paper's published apple-news-scraper for Top Stories and Trending Stories twice daily for two weeks, while the same scraper runs on the Appium simulator. Compare, per day and overall, the overlap coefficient, Shannon equitability, and top-source share between simulator and real devices. Also run the simulator behind VPN endpoints in several states while keeping GPS fixed, and with two different simulated GPS locations while IP stays fixed, to disentangle GPS-based from IP-based localization. If real-device feeds contain location-specific or device-specific stories, or if source concentration differs materially, Experiment 4's diversity comparison must be recomputed or conditioned on device class.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central diversity comparison (Table 2, Figures 3 and 4) is built on Experiment 4, which collected 62 days of Top Stories and Trending Stories from a single Appium-controlled iPhone simulator. The only real-device validation is Experiment 3, which matched 83 crowd-worker screenshots against the simulator for Trending Stories at three synchronized moments; it does not cover Top Stories and does not cover the full collection period. If Apple News serves simulator-specific content, whether because of different device identifiers, missing carrier or account signals, or host-IP location rather than simulated GPS location, the source distributions in Experiment 4 may not represent what real U.S. iPhone users see. The localization test is especially exposed: Appium can inject GPS coordinates into the simulator, but Apple News may infer location from the host Mac's IP address or from the Apple ID, so the null result in Experiment 2 may only show that GPS injection does not change the feed. Because the paper explicitly uses the absence of personalization and localization to justify collecting from a single simulator, a failure of either validation would put the headline source-diversity result, and the editorial-versus-algorithmic logic interpretation, on unverified ground.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an audit of Apple News, focusing on the algorithmically curated Trending Stories section and the human-editorially curated Top Stories section. It introduces a three-part audit framework (mechanism, content, consumption), then applies it through four experiments: update frequency of Trending Stories, a GPS-based localization test using Appium simulators, a crowdsourced personalization test with synchronized Mechanical Turk screenshots, and a 62-day automated collection from a single simulator comparing the two sections. The headline findings are that Trending Stories show no meaningful personalization or localization in the tested settings, that the algorithmic section has higher source concentration and lower Shannon equitability than the editorial section, and that the algorithmic section skews toward soft news while the editorial section skews toward policy and international news. The paper interprets these differences as manifestations of algorithmic versus editorial logic.","tokens_in":17514,"tokens_out":5419,"duration_ms":59872,"significance":"If the findings hold, this is a valuable early empirical characterization of Apple News, a platform with substantial gatekeeping power but little public data. The study is methodologically constructive: the code is released, the measurements are computed directly from observed data without fitted parameters, and the structure of the personalization experiment (synchronized crowd screenshots compared with a control simulator) is a useful template for auditing closed platforms. The diversity comparison is supported by an appropriate significance test (Hutcheson's t, t=11.17, p<0.001). The main limitation is that the extended content audit rests on a single simulated device with only limited real-device validation, so the generalizability of the central source-diversity claim is not yet fully established.","major_comments":[{"comment":"The central content comparison (Table 2, Figures 3–4) is based on a single Appium-controlled simulator over 62 days, but the only real-device validation reported is Experiment 3, which covers Trending Stories at three synchronized time points and does not validate Top Stories or the full collection period. If the simulator receives different content than production iPhones (due to device/profile signals, IP-based location, or A/B variants), the source diversity and evenness differences could be artifacts of the measurement channel. I ask the authors to add validation (for example, periodic real-device screenshots for both sections during the 62-day window) or to explicitly constrain the claims to what the single-device data can support.","section":"Experiment 4 / Results: Content / Table 2"},{"comment":"The localization test manipulates only the simulator's GPS coordinates; the manuscript does not report varying the host IP address or Apple ID across conditions. Because Apple News may infer location from IP or account information rather than GPS, the null result does not rule out location-based adaptation. The paper should soften 'no evidence of localization' to 'no evidence of GPS-based localization' and discuss this residual confound in the limitations.","section":"Experiment 2: Testing for location-based adaptation"},{"comment":"The comparison is observational: Top Stories and Trending Stories differ not only in human versus algorithmic curation but also in section purpose (top stories versus trending content), number of slots, and update frequency. The abstract and conclusion attribute the diversity difference to curation type ('human curation outperformed algorithmic curation'), but the design does not isolate curation logic from these other factors. I recommend adding an explicit caveat that the observed differences are consistent with, but not uniquely caused by, the editorial-versus-algorithmic distinction.","section":"Discussion: Algorithmic vs. Editorial Logic"}],"minor_comments":[{"comment":"The text reports 1,268 Top Stories collected, while Table 2 lists 1,267; the discrepancy should be reconciled.","section":"Table 2"},{"comment":"The row 'meghan markle 33*' in the Trending Stories n-gram table is confusing: the asterisk footnote says the n-gram appeared twice in the other section, but the row is placed among Trending n-grams. Please clarify the intended placement and annotation.","section":"Table 1"},{"comment":"The claim of being 'the first data-backed characterization of Apple News in the United States' is strong; given prior work by Brown (2018a, 2018b) on Apple News, consider softening the novelty claim or specifying more precisely what aspect is new.","section":"Abstract"},{"comment":"The distinction between platform-wide and user-specific update frequency is clear in principle, but the finding that the app must be closed to see updates is important and should be highlighted explicitly in the main results rather than only in the narrative.","section":"Experiment 1"},{"comment":"The paper does not state the iOS version or Apple News version used in the simulator; since platform behavior can change between releases, including this information would improve reproducibility and comparability.","section":"Audit Methods"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for ICWSM and contributes a reproducible audit method for a closed news platform. My main concern is that the central diversity claim depends on single-simulator data with limited real-device validation; this is fixable with additional validation or appropriately weakened claims, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a genuinely useful audit of a major news aggregator, and the main directional claims probably hold, but the paper's central two-month comparison leans on a single simulated iPhone whose equivalence to real devices is only partly validated.\n\nWhat's new: first data-backed look at Apple News in the U.S., combining crowdsourced screenshots, a sock-puppet simulator, and long-running scraping to characterize both mechanism and content. The framework (mechanism, content, consumption) is a nice contribution in itself. The empirical work is mostly careful: the personalization test uses synchronized screenshots and an overlap coefficient, and the null result (0.97 average overlap, rising to 1.00 when the previous control set is included) is convincing as far as it goes. The localization test is weaker because GPS injection may not be how Apple determines location, but the two runs are suggestive. The diversity comparison is statistically solid: Hutcheson's t=11.17, p<0.001, and the descriptive stats (top-3 source share 45.2% vs 23.7%) are stark. The code is on GitHub.\n\nWhere it gets soft: the extended collection in Experiment 4 uses one Appium simulator, and the validation against real iPhones only covers Trending Stories at three synchronized moments. Top Stories—the editorial section that drives the headline comparison—never gets that check. The paper itself says it used a single account after finding no personalization or localization, but the personalization test didn't examine Top Stories, and the localization test may have tested GPS injection rather than real-world location signals. So there is a real, though bounded, risk that simulator-specific content could distort the source distributions. A second soft spot: raw data aren't shared, so the diversity numbers can't be independently rechecked. The soft/hard news classification is qualitative n-gram inspection, not a validated scheme—minor.\n\nNone of this breaks the paper. The editorial-vs-algorithmic story is plausible and the directional findings are likely robust. A referee should push on the simulator equivalence and ask for data release.\n\nRecommendation: this deserves serious peer review. It's a solid contribution to computational journalism and algorithm auditing. I'd bring it to reading group and would cite it for the method and the personalization null result. Send it to review with a request to address the simulator validation gap, ideally by adding a short real-device check for Top Stories or by discussing the limitation explicitly.","headline":"A worthwhile audit of Apple News with a genuine simulator-equivalence caveat that deserves peer review but needs tighter validation.","tokens_in":18077,"tokens_out":4027,"would_cite":true,"duration_ms":37585,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-14T15:55:18.370459+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}