{"id":"cf33c81e-ee75-4eaa-b000-5d935d86accc","arxiv_id":"2509.00199","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Recommendation-system A/B tests can systematically underestimate a new algorithm's true deployment effect because a small treatment group cannot trigger the full ecosystem's feedback loops, a bias the paper formalizes and illustrates.","lead":"A Roblox research paper argues that A/B tests in recommendation systems can understate the value of new algorithms because small test groups cannot activate the ecosystem-wide feedback loops (virality, creator adaptation, retraining) that a full rollout would create. The authors formalize this as 'algorithm adaptation bias' and illustrate it with two internal case studies.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. 4 does not constrain the sign of Bias(ρ); the central claim that partial-rollout experiments favor production rests on an unproven assertion, not on the formalism or data.","rationale":"The reader's CONDITIONAL verdict is appropriate, but the decisive issue is narrower than the reader states. Even setting aside the unquantified case studies, the theoretical core does not imply the paper's headline conclusion: Eq. 4 is an identity and, absent structural assumptions, says nothing about the sign of Bias(ρ). The 'in practice' assertion is not a derivation. I would not reject the paper—it is a short workshop paper that explicitly aims to raise awareness, and the definition of Bias(ρ) is a useful conceptual contribution—but the conditional status is justified. The proposed simulation test would settle the sign question. I partially agree with the reader's weakest_assumption: the reader flags both missing empirical quantification and the unproven sign; I see the sign as the single load-bearing issue, with the case studies secondary. For this reason I recommend keeping the verdict unchanged at CONDITIONAL.","tokens_in":5203,"tokens_out":7291,"duration_ms":83004,"concrete_test":"Build a minimal performative model with state S_ρ=(1−ρ)S_0+ρS_1, equilibria S_0,S_1 defined as fixed points, and outcome functions y_1(s), y_0(s). Sweep parameters controlling (i) how strongly y_1 depends on state and (ii) how much S_ρ shifts with ρ. Compute Bias(ρ) via Eq. 3 for ρ=0.01, 0.1, 0.5. If any plausible parameter set yields Bias(ρ)>0 while S_ρ is close to S_0, the 'often negative' assertion is disproved; if all plausible sets yield Bias(ρ)<0, the paper's intuition is corroborated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central practical claim—that partial-rollout A/B tests systematically favor the production policy and cause missed launches—requires Bias(ρ) < 0 in the settings of interest. The paper's formalization does not establish this. Equation (4) writes Bias(ρ) as a difference of two adaptation gaps: [E_{D_ρ}Y(π1)−E_{D(π1)}Y(π1)] − [E_{D_ρ}Y(π0)−E_{D(π0)}Y(π0)]. The text (Section 3.1) asserts that for ρ≪1, D_ρ is 'typically much closer to D(π0)', making the second gap small and the first gap negative. But no theorem, model, or calibration bounds these gaps. A policy can be better under the control distribution than under its own equilibrium (e.g., a novelty policy that excites users initially but fatigues at scale), making the first gap positive; the second gap is nonzero unless D_ρ = D(π0) and could cancel or reverse the sign. The listed mechanisms are plausible but one-directional; none is formalized. Thus the sign of Bias(ρ) is load-bearing and currently unsubstantiated. The two case studies cannot repair this: they are aggregate pre/post comparisons without counterfactuals or uncertainty quantification.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper formalizes a bias that arises when recommender-system A/B tests evaluate a candidate policy on a small treatment share while the rest of the traffic is governed by the incumbent policy. It defines the full-deployment policy-level estimand τ★ and the partial-rollout experimental estimand τ_exp(ρ), defines algorithm adaptation bias as their difference, and decomposes it into two adaptation gaps (Eq. 4). The paper argues that in typical partial rollouts the treatment arm is evaluated under a control-dominated distribution, so the bias is often negative, i.e., experiments favor the production policy and lead to missed launches. It supports this claim with two qualitative real-world case studies (a thumbnail UI redesign and ranking-objective changes) and proposes mitigation strategies: model-data separation, ramped rollouts with 50/50 phases, and post-hoc diagnostics.","tokens_in":5573,"tokens_out":4776,"duration_ms":53677,"significance":"The formalization is a useful step: it clearly distinguishes the estimand of a partial-rollout A/B test from the estimand of a full deployment, and connects the problem to performative prediction and interference literature. If the sign and magnitude of the bias were established, the practical consequences for launch decisions would be substantial. The proposed mitigation ideas are plausible and worth discussion. However, the paper is a position/awareness paper: the central directional claim (that bias 'often favors the production variant') is not proven by the formalism or the empirical section. The two case studies are anecdotal and lack the quantitative rigor needed to support the claim.","major_comments":[{"comment":"The central practical assertion requires Bias(ρ) < 0 in the settings of interest. Equation (4) writes Bias(ρ) as [E_{D_ρ}Y(π1)−E_{D(π1)}Y(π1)] − [E_{D_ρ}Y(π0)−E_{D(π0)}Y(π0)]. The text claims that for ρ≪1, D_ρ is 'typically much closer' to D(π0), making the first gap negative and the second small. No theorem, model, or calibration is provided to bound these gaps. A candidate policy can be better under the control distribution than under its own equilibrium (e.g., novelty effects that fatigue at scale), making the first gap positive; the second gap is nonzero unless D_ρ=D(π0) and can cancel or reverse the sign. The listed mechanisms are plausible but one-directional. The sign of Bias(ρ) is load-bearing for the paper's conclusion and needs support from a formal model, enough to identify regimes where the sign is guaranteed, or quantitative evidence.","section":"3.1, Eq. (4)"},{"comment":"The empirical evidence consists of two examples reported only qualitatively: 'neutral or non-significant impact' pre-launch and 'consistently revealed a positive lift' post-launch. There are no effect sizes, confidence intervals, sample sizes, experiment durations, or descriptions of the pre/post designs. Post-launch comparisons are not randomized and are subject to time-varying confounds, concurrent product changes, and regression to the mean. These cases cannot establish that the divergence is due to algorithm adaptation bias rather than other mechanisms. At minimum, one case should be documented with quantitative results, details on the exact experiment/rollout phases, and a discussion of alternative explanations.","section":"4, Sections 4.1–4.2"},{"comment":"The formalism assumes the existence (and uniqueness, at least conceptually) of stationary/equilibrium distributions D(π0), D(π1), and the mixture distribution D_ρ, and assumes that D_ρ is 'closer' to D(π0) for small ρ. No conditions are given for these distributions to exist or for the mixture to converge to the control-only distribution as ρ→0. The decomposition in Eq. (4) is definitionally valid, but it cannot yield quantitative conclusions without a concrete data-generating model for the feedback loop (users, creators, retraining). A simple worked example or a reference to a framework (e.g., performative prediction) that guarantees the ordering would strengthen the paper.","section":"3.1, D(π) and D_ρ definitions"}],"minor_comments":[{"comment":"The phrase 'often favor the production variant' is stated as a fact in the abstract and introduction, but the body of the paper does not establish this with data or theory. Consider softening to 'can favor' or 'may favor' in summary sections, reserving stronger claims for settings explicitly supported.","section":"Abstract and Introduction"},{"comment":"The DOI 'https://doi.org/10.1145/nnnnnnn.nnnnnnn' is a placeholder and must be replaced before publication.","section":"ACM Reference Format"},{"comment":"Reference [11] ('Andrew Redgate and Jane Smith') appears with generic author names and a venue 'AI Ethics and Society' that is not clearly established; please verify this citation. Several other entries are missing page numbers or full titles (e.g., [3], [4]).","section":"References"},{"comment":"The appendix (Section A) is empty; either remove it or add content.","section":"Appendix"},{"comment":"The definition mixes position bias (rank effects) with presentation bias (UI-related). Clarify or cite distinct definitions for these two types.","section":"Section 2, Position/Presentation Bias"},{"comment":"The conclusion appropriately notes that 'additional research is required' for bias estimation and adjustment. This limitation should also be acknowledged earlier, in the abstract or introduction, to match the paper's actual evidentiary strength.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The paper is a short workshop-style position paper. For a journal-level venue, the evidence base for the central directional claim would need to be substantially strengthened, either with a formal model or with rigorous empirical documentation. I also note that the reference list contains several entries that may be difficult to verify (e.g., the 'AI Ethics and Society' article); the editor may wish to check this. The formalization is a genuine contribution and the topic is timely; with a more careful framing of the sign claim and a detailed case study, the paper could become publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a short, honest position paper that names a real problem: when you A/B test a recommender in a partial rollout, the measured effect may not be the effect you'd get after full deployment, because the ecosystem (users, creators, data pipelines) hasn't adapted to the new policy. The paper formalizes this as algorithm adaptation bias, defines it as a difference between two estimands, and decomposes it into two adaptation gaps. That's useful. It connects the idea to performative prediction and causal interference, and it offers practical mechanisms (virality, creator response, training feedback) that explain why small treatment arms often look worse than they would at scale.\n\nThe two case studies are illustrative but thin—no numbers, no counterfactuals, no controls. The paper would be stronger if it at least characterized the magnitude of the observed lift or shared aggregated metrics.\n\nThe main soft spot is the sign claim. Equation (4) is a tautology. It doesn't tell you whether the bias is negative or positive. The paper asserts that for small ρ, D_ρ is close to D(π0), so the second gap is small and the first gap is negative. But that's an assumption about the dynamic system, not a consequence of the formalism. A policy could look better in the short term and fatigue at scale, making the first gap positive. The mechanisms listed are all plausible and mostly one-directional, but they're not backed by a model or calibrated data. The paper itself acknowledges uncertainty in places, but the abstract's claim that results 'often favor the production variant' overstates what's shown.\n\nThat said, this is a workshop paper whose stated goal is to raise awareness and spark discussion. On that terms, it succeeds. The framework gives practitioners a way to talk about a phenomenon they've likely seen, and the proposed diagnostics (impression shift, engagement trajectories, popularity distribution shift) are reasonable first steps.\n\nI'd send it to peer review, but the referee should ask for either real data with error bars or a more careful statement about when the sign flips. It's not ready to be cited as evidence that the bias is systematic, but it's a useful reference for the concept.\n\nBest.","headline":"A useful, honest position paper that names and formalizes a real RecSys evaluation bias, but the central sign claim is asserted rather than demonstrated.","tokens_in":5935,"tokens_out":1882,"would_cite":true,"duration_ms":21787,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A/B tests run on a small slice of traffic can systematically understate what a recommender change would do if fully deployed, because the measured effect is taken in a world still shaped by the incumbent system.","keywords":["recommender systems","online experiments","algorithm adaptation bias","A/B testing","causal inference","performative prediction","interference","partial rollout"],"falsifier":"A controlled staged launch of the same algorithm change at, say, 1%, 10%, and 100% traffic with the same model version and no concurrent product changes would settle it: if estimated lifts do not grow monotonically toward the post-launch value, the claimed systematic negative bias is not universal. The paper's two case studies do not include such a control.","tokens_in":5160,"feed_emoji":"📉","tokens_out":5705,"duration_ms":62553,"temperature":0.7,"pith_summary":"This paper argues that standard A/B tests in recommender systems can be systematically wrong in a specific way: when a new algorithm is tested on a small share of traffic, the experiment measures the candidate in a world still dominated by the incumbent system, so the candidate never gets the ecosystem-level amplification (virality, creator adaptation, training feedback) it would receive at full deployment. The paper formalizes this as a gap between the experimental estimand and the full-deployment causal effect, and defines that gap as algorithm adaptation bias. Real-world examples show pre-launch tests reading neutral or small, while post-launch analyses show larger positive effects, implying missed launches. A careful reader should care because this bias is not a statistical artifact of low power; it is structural to partial-rollout evaluation in adaptive systems.","feed_headline":"Small-rollout A/B tests understate real recommender gains","feed_subtitle":"Formal bias term shows treatment traffic too small to trigger ecosystem feedback, so winners can look like losers.","key_machinery":"The load-bearing object is the bias identity Bias(ρ) = τ_exp(ρ) − τ★, where τ_exp(ρ) is the difference-in-means estimand under the partial-rollout mixture distribution and τ★ is the full-deployment policy-level effect. This decomposition turns the intuition about feedback loops into a measurable target: the bias is the sum of an adaptation gap for the treatment policy and an adaptation gap for the control policy, and the paper contends that for small rollout shares the treatment gap dominates.","core_discovery":"The central discovery is that the standard difference-in-means estimator in a partial-rollout A/B test targets an experimental estimand evaluated in the mixture distribution induced by the test, whereas the decision-relevant quantity is the effect under platform-wide replacement of the production policy by the candidate. The paper defines the gap between these as algorithm adaptation bias and decomposes it into two adaptation gaps, one for each arm; for small treatment share, the mixture is expected to be close to the incumbent system’s distribution, making the gap systematically negative, meaning the experiment favors the incumbent. The paper catalogs the mechanisms that produce this gap an","pith_inferences":["Beyond the paper: the formal estimand is policy-agnostic, so the same adaptation bias should appear in any adaptive system where the tested policy shapes the data it is trained on; recommender systems are just the clearest example.","Beyond the paper: because the mixture distribution depends on re-training cadence, the bias may not be constant over the experiment; comparing effect estimates from early versus late portions of a single partial rollout is a cheap, testable diagnostic the paper does not explicitly propose.","Beyond the paper: the sign of Bias(ρ) is an empirical curve, not a law; at higher treatment shares cross-variant interference could reverse the bias, so measuring the bias at several rollout shares would let teams calibrate a correction for launch decisions."],"forward_implications":["If the bias is systematic, existing pre-launch A/B results overstate the evidence against launching: true winning variants will be shelved or delayed.","Post-launch “surprise lift” becomes an expected outcome under adaptive systems; platforms should treat pre-launch neutral results as insufficient to reject a candidate when adaptation mechanisms are plausible.","Staged ramp-ups with a 50/50 traffic phase can detect the bias by comparing effect estimates across traffic levels; the growth trajectory of the effect itself becomes a signal.","Separating each variant’s training data from production traffic can remove part of the feedback bias, at an infrastructure cost.","UI and presentation changes are also affected, so the bias is not limited to model objective changes; any change that reshapes content exposure can suffer it."],"supporting_citations":[{"why":"Supplies the SUTVA framework that the bias violates.","marker":"[6]"},{"why":"Provides formal causal inference with interference, grounding the mixture estimands.","marker":"[5]"},{"why":"Offers design and analysis methods for reducing interference bias in experiments.","marker":"[3]"},{"why":"Supports the social and network spillover mechanisms behind the bias.","marker":"[7]"},{"why":"Formalizes how a policy shapes the distribution it is evaluated on.","marker":"[10]"},{"why":"Contextualizes biased estimation in adaptive experiments.","marker":"[4]"},{"why":"Represents the standard A/B testing practice whose estimand is called into question.","marker":"[9]"},{"why":"Supports the multiplayer and dynamic-interference mechanism in partial rollouts.","marker":"[14]"}],"fun_headline_variants":["Small treatment groups hide true recommender wins","A/B tests favor incumbent, penalize small rollout","Algorithm adaptation bias skews small-scale A/B tests","Why small-rollout tests underestimate top models","Partial rollout hides real impact of new recommenders"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The paper assumes, without quantifying it, that a small treatment share leaves the experiment world mostly like the incumbent system; if that is not true, the claimed direction of the bias does not follow from the math.","fun_headline_variants_meta":{"raw":{"variants":["Small treatment groups hide true recommender wins","A/B tests favor incumbent, penalize small rollout","Algorithm adaptation bias skews small-scale A/B tests","Why small-rollout tests underestimate top models","Partial rollout hides real impact of new recommenders"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1142,"prompt_tokens":726,"completion_tokens":416,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":358}},"tokens_in":470,"tokens_out":416,"duration_ms":4642,"temperature":1.0,"reasoning_tokens":358,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:49:41.008594+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled staged launch of the same algorithm change at, say, 1%, 10%, and 100% traffic with the same model version and no concurrent product changes would settle it: if estimated lifts do not grow monotonically toward the post-launch value, the claimed systematic negative bias is not universal. The paper's two case studies do not include such a control.","supporting_citations":[{"cited_title":"Imbens and Donald B","cited_arxiv_id":null,"evidence_quote":"Supplies the SUTVA framework that the bias violates."},{"cited_title":"Hudgens and M","cited_arxiv_id":null,"evidence_quote":"Provides formal causal inference with interference, grounding the mixture estimands."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the social and network spillover mechanisms behind the bias."},{"cited_title":"Perdomo, Christoph Mendler-Dünner, and Moritz Hardt","cited_arxiv_id":null,"evidence_quote":"Formalizes how a policy shapes the distribution it is evaluated on."},{"cited_title":"Gaps in the Thue--Morse word","cited_arxiv_id":"2102.01018","evidence_quote":"Contextualizes biased estimation in adaptive experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents the standard A/B testing practice whose estimand is called into question."},{"cited_title":"Treatment Effect Estimation Amidst Dynamic Network Interference in Online Gaming Experiments","cited_arxiv_id":"2402.05336","evidence_quote":"Supports the multiplayer and dynamic-interference mechanism in partial rollouts."}],"review_version":1}