{"id":"6b427783-e42f-4a24-9d44-f659bcb59d0e","arxiv_id":"2607.15284","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Users who could adjust a news recommender's stance and topic sliders changed how extreme their feed became depending on their starting point, but did not consistently increase political diversity.","lead":"A 102-person study tested a news app that let people see and adjust the political slant and topics it recommended. The adjustable interface changed how extreme and diverse the news became compared to thumbs up/down, but effects varied by user; showing people what they actually read made them realize their feed was less diverse than they thought.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Subgroup findings likely inflated by regression to the mean: users are binned on the baseline of the same outcome measure, so the extremeness/stance/up-vote 'sweet spot' pattern needs an independent-baseline robustness check.","rationale":"I read the paper as a genuine user study with released code/data and a useful comparison of two interaction designs; the conditional verdict is appropriate. The most load-bearing risk is exactly the within-group baseline binning: the paper's headline 'sweet spot' and heterogeneous-effect language comes from low/medium/high subgroups defined by the same outcome measured at baseline. Random assignment gives the between-group RQ2 tests some protection, so I would not reject on this basis, but the within-group directional claims are not secure without an independent-baseline robustness check. The authors' own limitations section discusses power and external validity but not this conditioning artifact. A single re-analysis with an independent baseline would settle whether the concern lands. If it fails, the paper should be reworded from 'users steer the system' to 'users directly manipulate the sliders to change the displayed top-K distribution.'","tokens_in":19626,"tokens_out":12565,"duration_ms":152966,"concrete_test":"Re-run the Section 4 subgroup analysis (Table 3, Figure 2) with low/medium/high bins defined by an independent baseline—e.g., the user's pre-questionnaire self-reported ideology and topic interest, or a held-out half of the initial top-K ranking—instead of the same m_b used as the outcome. Check whether extremeness, stance, and up-vote delta signs and significance (low subgroup increasing, medium/high decreasing) survive. If they disappear or become non-significant, the 'sweet spot' and heterogeneous-effect claims are artifacts of outcome-conditioned binning; if they survive, the regression-to-the-mean concern is answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 bins all 102 users into low/medium/high by the begin value m_b of the very measure being analyzed. Because m_b and m_e are repeated, noisy measurements of the same recommender state (top-K aggregate of a bootstrap-seeded model; up-vote ratio based on 10 votes), a user with an unusually low begin value will tend to show an increase at end even without any preference change. This regression to the mean mechanically produces the reported pattern of high subgroups decreasing and low subgroups increasing. The between-group RQ2 comparisons are partly protected by randomization, but the paper's 'extremeness sweet spot' and the abstract's 'mild-content users may move to extremes' are within-group claims, and with only ~34 users per subgroup (unevenly split across arms) plus no baseline-balance check, the directionality of the subgroup changes is not secure. Section 5 notes sample-size/power limitations but does not address this conditioning artifact. For treatment users, the coupling is tighter still: the sliders directly recalibrate u_t and u_s (Eq. 1), so the end top-K distribution used to compute extremeness/stance/diversity is partly user-set rather than an emergent trajectory.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a user study of 102 Mechanical Turk participants comparing a traditional political news recommender (up-vote/down-vote only) with an enhanced interface that displays the system's inferred political-stance and topic-interest profile and lets users adjust sliders to modify future recommendations. The authors define four outcome measures — average extremeness, political stance, diversity (normalized entropy), and up-vote ratio — computed over the top-K recommended articles before and after the session. They claim that the treatment group steered the system toward less extreme content when initially shown extreme content, toward more extreme content when initially shown mild content, and toward the political center; that diversity decreased in both groups; and that the transparent interface increased user awareness of filter bubbles. The analysis is primarily based on within-group begin-to-end changes and treatment-control differences within subgroups defined by the initial value of each outcome measure.","tokens_in":19941,"tokens_out":4882,"duration_ms":56996,"significance":"If the findings hold, the paper makes a worthwhile contribution to the user-control and filter-bubble literature by providing behavioral evidence from a deployed system rather than simulation. The study design includes attention checks, bootstrapped confidence intervals, and a publicly available code/data repository, which are concrete strengths. However, the central subgroup-specific conclusions are vulnerable to regression to the mean because subgroups are formed from the baseline of the same outcome measure, and the statistical evidence for many specific effects is weaker than the abstract suggests once the full multiple-testing family is considered. The awareness claim also rests on a comparison that was only run in the control arm. With appropriate robustness analyses and revised interpretations, the paper could be a solid empirical contribution.","major_comments":[{"comment":"The low/medium/high subgroups are formed by ranking all 102 users on the begin value m_b of the very measure being analyzed. Because m_b and m_e are repeated, noisy measurements of the same recommender state (top-K aggregate from a bootstrap-seeded model; up-vote ratio from 10 votes), a user with an extreme m_b will tend to move toward the mean at m_e even with no change in preferences. This mechanically produces the observed pattern of high subgroups decreasing and low subgroups increasing. This directly threatens the 'extremeness sweet spot' interpretation and the associated abstract claims. Please provide an independent baseline for subgrouping (e.g., pre-questionnaire ideology or initial model parameters), or model the change as a function of m_b with an explicit regression-to-the-mean correction, or compare against a no-interaction repeated-measures control.","section":"Section 4, subgroup construction and Figure 2"},{"comment":"Table 3 reports 48 t-tests (4 measures × 4 groups × 3 comparisons: C begin/end, T begin/end, C vs T change), but the caption describes a Bonferroni correction for 'four hypotheses per measure.' Within each measure there are actually 12 tests, so a proper Bonferroni threshold would be approximately 0.004, not 0.0125. Many reported significant effects (e.g., extremeness low Δ=.28, p=.034; political stance strong liberal Δ=.35, p=.064; diversity medium Δ=.15, p=.047; up-vote medium Δ=.14, p=.033) would not survive an FDR or full-family correction. Please report multiplicity-adjusted p-values or clearly pre-specify a smaller confirmatory family and label the remaining results as exploratory.","section":"Table 3 and statistical significance"},{"comment":"The claim that the enhanced interface helped users realize they were in a filter bubble is not directly supported by the Qb/Qc comparison. The pre/post diversity question was administered only to the control group, with a histogram shown between the two responses; the treatment group did not receive the same manipulation. Thus the observed drop in Qc shows that transparency about the article distribution changes self-reports in the control group, but it does not establish that the treatment UI increased awareness relative to the control group. The Qd self-report is a single Likert item and is susceptible to demand characteristics. Please add the corresponding Qb–Qc comparison in the treatment arm or soften the causal wording in the abstract and conclusions.","section":"Section 4, Post-questionnaire analysis"},{"comment":"In the treatment arm, the sliders directly update u_t and u_s in Eq. (1), so the begin-to-end changes in extremeness, stance, and diversity are partly a consequence of the pre-specified mapping from user input to ranking, rather than an emergent trajectory from user preferences alone. This is by design, but the RQ1 interpretation should be qualified as measuring the joint effect of direct control and subsequent system updates. In addition, λ=0.4 is selected on the basis of undisclosed 'preliminary analysis' and no sensitivity analysis is reported; since the slider-based recalibration recomputes the top-K ranking through this weighted sum, the magnitude and direction of treatment effects could depend on λ and on K=200. Please report sensitivity analyses for these parameters.","section":"Section 3, Eq. (1) and sensitivity of λ and K"}],"minor_comments":[{"comment":"The paper contains several typos (e.g., 'bu it' instead of 'but it' in the post-questionnaire paragraph) and some undefined notation early (K in Eqs. 2–4 is not defined until later). A careful proofread is needed.","section":"Throughout"},{"comment":"Because subgroups are based on the full 102-user ranking, the number of control and treatment users in each low/medium/high bin may be unbalanced. Please report the arm sizes per subgroup, particularly for the political-stance subgroups labeled 'strong liberal,' 'liberal,' and 'conservative.'","section":"Section 4, subgroup samples"},{"comment":"The caption does not specify the direction of the one-tailed tests used for the 'Change' columns, nor does it justify why RQ2 tests are one-tailed while RQ1 tests are two-tailed. Please state the alternative hypotheses explicitly.","section":"Table 3 caption"},{"comment":"Many bootstrapped 95% confidence intervals are visually invisible at the plotted scale; consider numeric error bars or separate panels with annotations. Reporting effect sizes (e.g., Cohen's d) alongside p-values would also aid interpretation given the modest sample size.","section":"Figure 2"},{"comment":"The abstract states that 'the transparent approach helped users realize that they were in a filter bubble,' but the only direct evidence is the within-control Qb→Qc change. Please align the abstract with the actual between-group evidence, or add the missing treatment-arm comparison.","section":"Abstract and Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a useful empirical study with publicly released code and data, and the topic is appropriate for the venue. The regression-to-the-mean concern is genuine and central to the subgroup-specific claims; I would like to see a robustness analysis using an independent baseline or a continuous model with explicit RTM correction. The multiple-testing issue is also significant but easily addressed by reporting FDR-adjusted p-values or a clearly pre-specified confirmatory analysis. I do not see grounds for rejection, but the current abstract overstates what the design can establish."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my take on arXiv:2607.15284. It's a genuine contribution: a live political news recommender with a transparency-and-control UI (sliders for political stance and topic interest) run with 102 users, half randomized to a control group with only up/down-vote buttons, and the code and data are released. That combination goes beyond the mostly simulation-based prior work. The central between-group finding — that users with the enhanced UI change the system's output more strongly than up/down-voting alone — is credible because of the randomization, and the t-tests and bootstrap CIs are standard practice.\n\nThe soft spot is real and load-bearing: the subgroup analysis bins users into low/medium/high by the begin value of the very measure being studied. Begin and end are repeated measurements of the same recommender state, so part of the observed pattern (high subgroups decreasing, low subgroups increasing) is mechanically expected regression to the mean. That threatens the paper's 'extremeness sweet spot' interpretation and the abstract's claim that mild-content users may move to extremes. The between-group RQ2 comparisons are protected by randomization, but the subgroup-specific directional claims are not secure. Multiple comparisons are only weakly addressed (a per-measure Bonferroni correction across four rows), and several key effects sit at p between .02 and .05, so they would not survive a stricter correction. There's also a subtle issue that in the treatment arm the end top-K values are partly user-set through the sliders, so the 'trajectory' is not a purely emergent system response.\n\nA few things are genuinely good. The awareness finding in the post-questionnaire — control users reported lower diversity after seeing their actual stance distribution — is a clean demonstration of the transparency gap, even though the earlier Qb question showed no group difference. And the authors are honest about sample size and power limitations, though they don't actually address the regression-to-mean artifact.\n\nWho is this for? HCI and recommender-systems researchers, and anyone who wants a concrete example of how transparency tools interact with user behavior. It deserves a serious referee: the empirical surface area is real, the data are available, and the core findings can likely be rescued with an independent baseline or a different grouping strategy. I'd expect a heavy revision before it's solid.\n\nRecommendation: send it to peer review.","headline":"A real empirical contribution undermined at the subgroup level by regression to the mean; the between-group effects and the awareness finding are the credible core.","tokens_in":20421,"tokens_out":4330,"would_cite":true,"duration_ms":41347,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports a 102-user experiment showing that letting readers adjust inferred political stance and topic sliders changes what a news recommender feeds them — reducing extremeness for most while cutting political diversity overall.","keywords":["filter bubbles","news recommendation","user control","transparency","political diversity","user study","recommender systems","echo chambers"],"falsifier":"A researcher could re-run the analysis with a change-score model that conditions on the begin value (e.g., regress end−begin on begin and treatment, or use a regression-discontinuity at the low/medium/high cutpoints). If the apparent 'sweet spot' — high values falling and low values rising — disappears once begin values are controlled nonlinearly, then the central steering claim is a regression artifact. Alternatively, a randomized experiment that assigns the same initial extremeness level to users independent of their own profile would settle it.","tokens_in":19521,"feed_emoji":"📰","tokens_out":4440,"duration_ms":43273,"temperature":0.7,"pith_summary":"The paper tries to show that giving people direct, transparent control over a political news recommender — sliders for political stance and topic interest — genuinely changes the recommender's trajectory compared with up/down voting alone. It finds heterogeneous effects: users initially shown extreme or moderately extreme articles steer toward less extreme content, while users initially shown mild content may steer toward more extreme content. In both groups, political diversity of the recommendations decreases. The transparency also makes users aware that they were in a filter bubble and more informed about how the recommender works. If these findings hold, they imply that user-control interfaces are a double-edged tool: they can moderate extremeness for some users but do not by themselves restore diverse news diets.","feed_headline":"User sliders cut extreme news but not filter bubbles","feed_subtitle":"In a 102-user test, transparency helps readers moderate news feeds, but political diversity still falls.","key_machinery":"The mechanism is a recalibrated scoring function s = λ s_r + (1−λ)(Sim(u_t,a_t)+Sim(u_s,a_s))/2, where s_r is a content-based classifier's up-vote probability, u_s and u_t are user political-stance and topic-interest vectors initialized from a pre-questionnaire, and a_s/a_t are article stance/topic vectors. The enhanced interface lets users adjust u_t and u_s by moving sliders, reranking the top articles by binary search. The paper's measures — average political stance, extremeness (absolute stance), diversity (normalized entropy over five stances) and up-vote ratio — are computed on the top 200 ranked articles at the beginning and end of each session, with users split into low/medium/high s","core_discovery":"The central discovery is that an interactive transparency-and-control interface changes the recommendation trajectory in a user-dependent way. In the treatment group, users whose recommender began with medium or high extremeness significantly reduced extremeness, while those who began with low extremeness increased it, producing a kind of extremeness 'sweet spot'. Users also moved political stance toward the center, yet normalized entropy of political stances in the top-ranked articles decreased for both control and treatment, meaning diversity fell. The up-vote ratio rose most for users who initially up-voted few articles, and the post-questionnaire showed that transparency made users reali","pith_inferences":["The subgroup analysis bins users by the beginning value of the very measure being studied (extremeness, diversity, up-vote ratio), so part of the high-decrease/low-increase pattern may be regression to the mean rather than user steering; a design that randomizes initial positions or matches users on begin values would disentangle these.","If the diversity decrease is real, a combined interface that lets users express 'diversity' as a goal (not just stance and topic) might recover the lost diversity; the paper's own discussion suggests more complex per-topic preference expressions.","The awareness effect (control users revising their diversity estimate after seeing the histogram) suggests a low-cost policy lever: platforms could display a stance histogram of the user's recommendation feed without any interactive sliders.","Because the study covers one hour and 102 users, the extremeness sweet spot could be a short-horizon effect; longer-term measurement (weeks) could test whether users keep pulling toward the center or drift back."],"forward_implications":["If replicated, the results imply that transparency-and-slider interfaces are a practical way to let willing users pull a recommender away from extreme partisan content.","The 'sweet spot' pattern suggests users do not seek maximal extremity; there is a comfort zone of moderate partisanship.","Because diversity fell even when extremeness fell, moderating extremeness and preserving viewpoint diversity are distinct goals that may need separate design levers.","The post-questionnaire result — control users changed their diversity rating after seeing the actual stance histogram — implies users cannot accurately perceive a filter bubble without external transparency.","Heterogeneous effects in the treatment group warn that the same control tool can move different users in opposite directions."],"fun_headline_variants":["Transparent news controls expose bubbles, not eliminate them","Slider control shifts news extremes — both ways","Awareness of filter bubbles rises, but diversity falls","User power over news: a double-edged slider","Control over news recommender: users see, but diversity drops"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's subgroup-specific conclusions assume that users binned by their starting value of the same outcome measure show changes that reflect genuine steering, not statistical regression to the mean.","fun_headline_variants_meta":{"raw":{"variants":["Transparent news controls expose bubbles, not eliminate them","Slider control shifts news extremes — both ways","Awareness of filter bubbles rises, but diversity falls","User power over news: a double-edged slider","Control over news recommender: users see, but diversity drops"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000243,"raw_usage":{"total_tokens":1350,"prompt_tokens":711,"completion_tokens":639,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":455,"completion_tokens_details":{"reasoning_tokens":563}},"tokens_in":455,"tokens_out":639,"duration_ms":7574,"temperature":1.0,"reasoning_tokens":563,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T12:10:41.944025+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A researcher could re-run the analysis with a change-score model that conditions on the begin value (e.g., regress end−begin on begin and treatment, or use a regression-discontinuity at the low/medium/high cutpoints). If the apparent 'sweet spot' — high values falling and low values rising — disappears once begin values are controlled nonlinearly, then the central steering claim is a regression artifact. Alternatively, a randomized experiment that assigns the same initial extremeness level to users independent of their own profile would settle it.","supporting_citations":[],"review_version":1}