{"id":"4a283a23-2ab6-4444-927f-638cbb1c3241","arxiv_id":"2510.24354","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"In a human experiment, personalized, engagement-rewarding rankings increased consumption of extreme and like-minded news relative to a neutral ranking.","lead":"This paper shows that ranking algorithms which personalize results and reward likes or shares push users toward more extreme and like-minded news. The effect appears in both simulations and a controlled online experiment, suggesting that common engagement-optimizing design choices can amplify polarization.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Experiment 2's click-level Mann–Whitney tests ignore that clicks within a run share a dynamically updated ranking; with only 3 runs per topic-condition, the reported p<0.001 values are not statistically justified, weakening the central causal evidence.","rationale":"The central empirical claim rests on Experiment 2, which is the only direct test with human participants under dynamically updated rankings. The statistical analysis in Sections 3.5 and 4.3 uses click-level Mann–Whitney tests, but clicks within a run are not independent because each click changes the ranking shown to subsequent users via Eq. (5). A run is the true experimental unit, and there are only 3 runs per topic-condition. Consequently, the reported p<0.001 values are not supported by the design; the smallest possible one-tailed p for a 3-vs-3 Mann–Whitney test is 0.05. This concern is load-bearing because the paper's causal language in Section 5 is justified by these significance tests. The reader's invariance concern is real but secondary: it affects the mechanistic interpretation of the simulations, not the direct experimental comparison between the two algorithm configurations. A run-level reanalysis would settle whether the significance claims survive. If they do not, the experimental contribution is weakened to directional qualitative evidence, though the model and static results remain suggestive. This does not require rejecting the paper, but it does require tempering the statistical claims, so the conditional verdict is unchanged.","tokens_in":19290,"tokens_out":10518,"duration_ms":110713,"concrete_test":"Reanalyze Experiment 2 at the run level: for each of the 24 runs, compute the mean Consumption Extremism and Consumption Polarization over the final window (w=200, discarding the first ~50 periods), so each topic-condition has n=3 run-level observations. Compare conditions with a permutation test or a mixed-effects model with random intercepts for participant and run, reporting two-tailed p-values. If the differences remain significant at run level, the experimental claim survives; if not, the paper should present the result as directional evidence only.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the unit of analysis in the dynamic ranking experiment (Section 3.5, Fig. 6). The paper applies one-tailed Mann–Whitney U tests 'considering each individual click,' but clicks in the same run are not independent: Eq. (5) updates item popularities after each interaction, so the ranking and the click distribution evolve along a single trajectory within each run. The design contains only 3 independent runs per topic-condition (24 runs total), so the effective replication for each comparison is 3 versus 3 runs, not roughly 600 versus 600 clicks. With n=3 per group, the smallest achievable one-tailed Mann–Whitney p is 0.05, so the reported p<0.001 values in Fig. 6 cannot be obtained from a valid run-level test. This directly undermines the statistical significance claims and the 'causal link' stated in Section 5, even though the direction of the effect may be correct. The click-distribution chi-square tests in Fig. 7 have the same nesting problem.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes and tests a mechanism by which popularity-based ranking algorithms amplify political extremism and polarization when they reward active engagement and personalize results. The authors specify a dynamical model in which users with a five-point political stance click on ranked news items with position-biased, stance-conditional probabilities, and sometimes 'highlight' items with a probability that depends on both stances; item popularities are updated per stance group via Eq. (5) with parameters λ (personalization) and η (active-engagement reward). The model inputs (user stance distribution, β, C, H) are estimated by maximum likelihood from a static ranking experiment with 432 participants. Simulations across a grid of (λ, η) predict that both consumption extremism and polarization increase with λ and η. A dynamic experiment with 1,534 participants compares the corner conditions (λ=0, η=0) and (λ=1, η=100), reporting increases in extremism and polarization under the personalized, engagement-rewarding condition, and the paper concludes with claims of statistical significance and of a 'causal link' between algorithmic design and amplification.","tokens_in":19523,"tokens_out":14607,"duration_ms":122280,"significance":"If the statistical concerns below are addressed, this is a useful contribution: it provides a transparent, minimal model of user–algorithm feedback whose parameters are estimated from human data and whose qualitative predictions are validated in an independent human-in-the-loop experiment. The calibrated-then-confirmed design is a genuine strength, as is the honest Limitations section, which concedes that only the two corner configurations were tested, that the simulations are not designed for quantitative predictions, and that the outcome measures concern consumption rather than attitude change. The paper's headline claims — 'inevitably driven' in the Abstract and the 'causal link' statement in Section 5 — currently exceed what the evidence supports, mainly because the reported significance rests on a click-level analysis that inflates the effective sample size. With a run-level re-analysis and appropriately hedged wording, the qualitative finding is credible and relevant to debates about platform design and regulation.","major_comments":[{"comment":"The unit-of-analysis concern raised in review is confirmed by the text of §3.5. The one-tailed Mann–Whitney U tests reported in Fig. 6 are applied 'considering each individual click,' but clicks within a single run are not independent: Eq. (5) updates group-specific popularities after every interaction, so each run is one trajectory of a shared ranking, and there are only 3 runs per topic-condition. Treating roughly 600 clicks as independent observations overstates the effective replication by about two orders of magnitude. With n1 = n2 = 3 runs, the smallest achievable one-tailed Mann–Whitney p is 0.05, so the reported per-topic p < 0.001 values cannot be obtained from any valid run-level test. The chi-square tests in Fig. 7 have the same nesting problem, and the 20%/13% effect sizes stated in §4.3.2 and §5 inherit it. I ask the authors to re-analyze at the run level (e.g., exact permutation tests on the 3 vs. 3 runs per topic, mixed models with run as a random effect, or a null distribution generated by simulating the ranking dynamics), and to restate the significance claims in Sections 4.3 and 5 accordingly.","section":"§3.5, Fig. 6; §4.3; §5"},{"comment":"The measurement window is inconsistent with the paper's own convergence analysis. Appendix D states that steady state is reached 'between 200 and 300 interactions,' yet the dynamic experiment has about 253 interactions per run, discards only the first 50 periods, and computes metrics over a window of size w = 200 — i.e., roughly clicks 53–253, most of which lie in the transient regime identified by the authors. The claim in §3.1.3 that metrics are computed 'when the probability distribution in Eq. (4) is approximately stationary' is therefore not satisfied, and the reported effect magnitudes are averages over non-stationary trajectories. Averaging over the transient will attenuate rather than inflate the between-condition differences, so this may not invalidate the qualitative direction, but it does undermine the quantitative claims. Please report a sensitivity analysis using only interactions after the convergence time identified in Fig. D5, and reconcile the design description in §3.5 with Appendix D.","section":"§3.5 and Appendix D"},{"comment":"The claims that the mechanism makes amplification 'inevitable' (Abstract) and that 'any increase in personalization or highlight reward can produce a shift' (§4.2) go beyond what is tested. Experimentally, only the two extreme corners (λ=0, η=0) and (λ=1, η=100) are compared — a limitation the authors explicitly acknowledge in the Limitations section. The monotonic grid in Fig. 4 is reported as cell means over 1,000 simulations without error bars or statistical comparison between adjacent cells, so the 'any increase' claim is not quantified, and the simulation is a projection of the fitted parameters (β, C, H) and the update rule rather than an independent derivation. I recommend adding uncertainty quantification to Fig. 4, softening 'inevitably' and 'any increase' to claims that match the two-corner experiment plus simulation, and stating explicitly in the Conclusions which parts of the claim rest on experiment and which rest on simulation.","section":"Abstract, §4.2, Fig. 4, Limitations"}],"minor_comments":[{"comment":"The assignment of participants to the two parameterizations, and to the three repetitions within each topic-condition, is not described; please state the randomization procedure and confirm that the two conditions were balanced on participant characteristics.","section":"§3.5"},{"comment":"The caption calls the increases 'small,' but the η axis runs from at most 1.0 to 100.0 in logarithmic scale, so the smallest displayed reward is a doubling of popularity for a highlight; clarify whether η=0 is included in the grid (a log axis cannot display zero) and re-word 'small increases' accordingly.","section":"Fig. 4 caption"},{"comment":"The text states that the U-shaped engagement pattern is 'statistically significant' without reporting a test; please provide the test used and its p-value.","section":"§4.1, Fig. 3(d)"},{"comment":"'The simulations reproduce the main experimental effects with surprising accuracy' is a visual judgment; please report a quantitative agreement measure (e.g., mean absolute deviation between simulated and experimental metric values per topic-condition).","section":"§4.3.1"},{"comment":"Reference [7] contains a corrupted author name ('Michaundefined'), and Appendix F contains a typo ('finaly'); both should be corrected.","section":"Reference list, Appendix F"},{"comment":"Table 1 lists only D_stu as estimated from data, but the text of §3.1 states that the news-stance distribution D_sn is uniform by assumption; please make this asymmetry explicit in the table or its footnote for clarity.","section":"Table 1 and §3.1"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the unit-of-analysis problem in Experiment 2 is the main barrier to acceptance. It is, however, a fixable statistical issue rather than a fatal design flaw: the experiment has a plausible causal structure, the qualitative direction of the effects is consistent with the simulation, and the authors' own Limitations section shows appropriate caution about scope. I would expect a revised version to re-run the tests at run level, add the requested convergence sensitivity analysis, and temper the headline claims. If the run-level analysis collapses the significance entirely, the paper would then need to be re-positioned as an existence proof with a credible direction but weak inferential support. Given the paper's interdisciplinary placement, I would also ask the editor to weigh the strictness of the statistical standard applied here against what is typical for the venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper earns its keep with the dynamic ranking experiment, not the model. The model is mostly a repackaging of known ingredients (position bias, homophily, U-shaped engagement) and its simulated amplification is a projection of parameters fitted in the static experiment. That part is fine as a formalization, but it would be circular if it were the only evidence. It isn't. Experiment 2 compares 1,534 participants interacting with dynamically updated rankings under two algorithm settings, with 24 runs (3 per topic per condition). Running a controlled human-in-the-loop ranking study at this scale is real work, and the qualitative pattern—more extreme and same-stance clicks under personalization plus highlight rewards—is credible.\n\nThe soft spot is the statistics. They report one-tailed Mann–Whitney U tests 'considering each individual click.' Clicks in the same run share a dynamically updated ranking, so they are not independent observations. With only three independent runs per topic-condition, the effective n is 3 vs 3, for which the minimum one-tailed Mann–Whitney p is 0.05. The reported p<0.001 values in Fig. 6 are not obtainable from a valid run-level test. The chi-square tests in Fig. 7 have the same nesting problem. This doesn't mean the effect is absent—the direction is consistent across topics and the simulation matches—but the significance claims as written are not justified. A run-level analysis (e.g., comparing run means, or a mixed model with run as random effect) may well confirm the pattern with n=3, but the stars need to be recalculated.\n\nMinor issues: the simulation heatmap in Fig. 4 has no error bars, and no code or data are included, which makes the bootstrap and run-level checks impossible to reproduce. The paper honestly lists its own limitations, including the joint manipulation of eta and lambda, so I won't repeat those.\n\nBottom line: the mechanism is plausible and the experimental design is valuable, but the inferential statistics are overstated. This is fixable. I'd send it to peer review with the expectation of a major revision focused on the unit-of-analysis problem and reproducibility.","headline":"A genuinely new human-in-the-loop experiment supports a plausible mechanism, but the click-level tests overstate significance because clicks within a run are not independent.","tokens_in":20057,"tokens_out":1904,"would_cite":true,"duration_ms":16942,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Rewarding active engagement and personalizing popularity rankings amplifies consumption extremism and polarization, via a feedback loop the paper models and tests with human participants.","keywords":["extremism","polarization","ranking algorithms","personalization","active engagement","position bias","agent-based model","human-subject experiment"],"falsifier":"Collect the dynamic-condition data and test whether the observed click and highlight rates by rank and stance still match the static estimates: if a goodness-of-fit test rejects the model's predicted stationary click distribution under ($\\lambda=1,\\eta=100$), or if a replication of Experiment 2 fails to find higher extremism and polarization in the personalized engagement-reward condition than in the click-only condition, the central claim is falsified.","tokens_in":19113,"feed_emoji":"📈","tokens_out":6543,"duration_ms":57343,"temperature":0.7,"pith_summary":"This paper isolates a mechanism by which ranking algorithms can amplify extremism and polarization. It argues that once users click higher-ranked items, prefer like-minded news, and engage more actively at the ideological extremes, a popularity-based ranking that rewards engagement and personalizes by user stance becomes a self-reinforcing loop: extreme users' highlights push extreme content to the top of like-minded users' lists, and position bias then drives clicks to it. The authors formalize the loop in a dynamical model, calibrate it with a static ranking experiment, and test it in a dynamic ranking experiment with hundreds of human participants. If the mechanism holds, small changes in two algorithmic parameters — the weight given to active engagement and the degree of personalization — predictably shift news consumption toward extreme and same-stance content.","feed_headline":"Engagement rewards and personalization push clicks toward extreme news","feed_subtitle":"Dynamic-ranking experiments with ~2,000 users show a 20% rise in extreme, same-stance clicks.","key_machinery":"The carrying object is the group-specific popularity score $p^g_n(t)$ updated by the rule in Eq. (5): a click from a user in group $g$ adds 1 to item $n$'s popularity in that group, or $1+\\eta$ if the item was highlighted, while clicks from users outside $g$ contribute only $(1-\\lambda)$ as much. The ranking shown to group $g$ orders items by $p^g_n$, so $\\lambda$ is the personalization dial (how little other groups' behavior matters) and $\\eta$ is the active-engagement reward. Click probabilities are the product of a position-bias factor $R(r_n)=\\beta^{N-r_n}$ and a stance-conditioned click matrix $C_{s_n,s_u}$, normalized over the list. What this machinery does is convert individual behavioral tendencies into a visibility feedback loop: engagement-weighted, group-specific popularity makes extreme, like-minded content rise for the users most likely to click it.","core_discovery":"The paper's central claim is that a specific feedback loop, not user preferences alone, drives consumption toward extremism and polarization. It starts from four empirical regularities: users click items higher in a ranking more often; users prefer news aligned with their own stance; users at the ideological extremes engage (like/share) more than moderates; and platforms rank by popularity. The paper formalizes the loop with a discrete-time model in which each user group (left, center, right) maintains its own popularity score per news item; the personalization parameter $\\lambda$ controls how much a group's ranking ignores other groups' clicks, and $\\eta$ multiplies the popularity boost of highlighted items. Simulations on parameters estimated from a static ranking experiment with 432 participants predict that both metrics rise monotonically with $\\lambda$ and $\\eta$. A dynamic ranking experiment with 1,534 participants then compares click-based, non-personalized rankings ($\\lambda=0,\\eta=0$) with personalized rankings that strongly reward highlights ($\\lambda=1,\\eta=100$). Clicks on extreme same-stance content rose by about 20% and same-stance clicks by 13%, with significant increases in consumption extremism for three of four topics and polarization for three of four, matching the simulated direction.","pith_inferences":["The paper's evidence is about consumption; the paper explicitly stops short of claiming durable attitude change or radicalization, so the natural next test is whether exposure shifts of this size move beliefs over longer horizons.","If the mechanism is right, an A/B test that lowers $\\eta$ or $\\lambda$ on a real feed should measurably reduce extreme-content consumption without any change in user beliefs — a testable design intervention the paper does not run.","The model predicts weaker amplification on platforms or topics where extreme users are not the most active; comparing the same ranking algorithm across such contexts could separate the engagement-profile assumption from the algorithm parameters.","Because $\\lambda=1$ creates three parallel popularity contests, a natural extension is to ask whether even a small amount of cross-group mixing ($0<\\lambda<1$) breaks the loop; the paper's monotone simulations suggest it does not, but this was not tested experimentally."],"forward_implications":["Every increase in personalization ($\\lambda$) or active-engagement reward ($\\eta$) in the simulated model raises consumption extremism and polarization, with personalization the stronger driver of polarization.","In the human dynamic experiment, the personalized, engagement-rewarding ranking increased consumption extremism in all four topics (significant except Climate Change) and polarization in all but Gender, with clicks on extreme same-stance content up about 20% and same-stance clicks up 13%.","Under the engagement-rewarding personalized condition, centrist items are demoted while same-stance extreme items rise, so exposure diversity shrinks even for users whose own preferences did not change.","Because the two experimental conditions differ only in the ranking algorithm's parameters, the consumption shift is attributable to algorithmic design rather than to a change in user preferences.","The mechanism offers a possible explanation for engagement-driven changes on real platforms, such as Facebook's 2018 update weighting 'meaningful social interactions' more heavily."],"supporting_citations":[{"why":"Supplies the empirical basis for position bias and like-minded clicking on Facebook, grounding hypotheses H1 and H2.","marker":"[6]"},{"why":"Documents the U-shaped relationship between ideological extremity and active engagement on the Facebook URL dataset, grounding hypothesis H3 and the real-world reference for high engagement weight.","marker":"[20]"},{"why":"Provides the 'few-get-richer' result showing how popularity-based rankings can generate self-reinforcing visibility advantages.","marker":"[21]"},{"why":"Introduces the ranking-for-engagement model that this paper's discrete formulation most directly extends.","marker":"[22]"},{"why":"Models opinion dynamics through algorithmic gatekeepers, an antecedent for the user-algorithm feedback loop formalized here.","marker":"[23]"},{"why":"Field-experiment evidence that feed-ranking choices shape exposure and engagement, the empirical context the mechanism must speak to.","marker":"[26]"},{"why":"Field evidence that active engagement signals such as reshares amplify political news, motivating the engagement-reward parameter $\\eta$.","marker":"[27]"},{"why":"Journalistic documentation of Facebook's 'Meaningful Social Interactions' update, the concrete real-world case mapped to the high-$\\eta$ scenario.","marker":"[28]"},{"why":"Contrasting field result that like-minded sources are prevalent but not attitudinally polarizing, framing the paper's focus on consumption amplification rather than attitude change.","marker":"[41]"}],"fun_headline_variants":["Rewarding engagement in rankings amplifies extremism and polarization","Personalized rankings reward clicks on extreme same-stance content","20% rise in extreme same-stance clicks from engagement-based rankings","Feedback loop in rankings drives extremism and polarization","How personalization and engagement rewards fuel extreme content"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The click and highlight probabilities estimated from a static, randomized ranking are assumed to remain valid when rankings become personalized and engagement-weighted; if users' response to rank or content changes under the dynamic condition, the mechanism the model describes is not necessarily the one that produced the experimental outcomes.","fun_headline_variants_meta":{"raw":{"variants":["Rewarding engagement in rankings amplifies extremism and polarization","Personalized rankings reward clicks on extreme same-stance content","20% rise in extreme same-stance clicks from engagement-based rankings","Feedback loop in rankings drives extremism and polarization","How personalization and engagement rewards fuel extreme content"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00046,"raw_usage":{"total_tokens":2315,"prompt_tokens":966,"completion_tokens":1349,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":1270}},"tokens_in":582,"tokens_out":1349,"duration_ms":10570,"temperature":1.0,"reasoning_tokens":1270,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:41:20.780026+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect the dynamic-condition data and test whether the observed click and highlight rates by rank and stance still match the static estimates: if a goodness-of-fit test rejects the model's predicted stationary click distribution under ($\\lambda=1,\\eta=100$), or if a replication of Experiment 2 fails to find higher extremism and polarization in the personalized engagement-reward condition than in the click-only condition, the central claim is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the U-shaped relationship between ideological extremity and active engagement on the Facebook URL dataset, grounding hypothesis H3 and the real-world reference for high engagement weight."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the 'few-get-richer' result showing how popularity-based rankings can generate self-reinforcing visibility advantages."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the ranking-for-engagement model that this paper's discrete formulation most directly extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Models opinion dynamics through algorithmic gatekeepers, an antecedent for the user-algorithm feedback loop formalized here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Field evidence that active engagement signals such as reshares amplify political news, motivating the engagement-reward parameter $\\eta$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Journalistic documentation of Facebook's 'Meaningful Social Interactions' update, the concrete real-world case mapped to the high-$\\eta$ scenario."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contrasting field result that like-minded sources are prevalent but not attitudinally polarizing, framing the paper's focus on consumption amplification rather than attitude change."}],"review_version":2}