{"id":"78dfaa48-abb7-4516-8897-229aed909f1f","arxiv_id":"2506.02881","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Simulation with optimism resimulates an adaptive experiment under the null with positively biased nuisance means, yielding asymptotically valid tests and narrower confidence intervals after bandit designs.","lead":"Researchers propose a simulation-based method for hypothesis tests and confidence intervals after adaptive experiments, such as multi-arm bandits. They add a small positive bias to estimated arm means (called simulation with optimism) and resimulate the experiment under the null, providing asymptotic type I error control for designs like explore-then-commit that standard methods cannot handle.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central guarantee depends on an unproven 'optimism widens quantiles' monotonicity that is verified only for three designs; Section E admits there is no unifying theory, so the 'wide variety' claim is unsupported.","rationale":"The reader's weakest_assumption is the same as the concern here: the paper verifies the 'optimism widens quantiles' principle only for the three designs in Theorem 1, and its own Limitations section concedes there is no unifying theory. I read the paper as a case-by-case proof effort rather than a general theorem, so the central risk is scope: a user applying the method to any other common adaptive design has no error-control guarantee. This is a correctness risk, not an internal inconsistency, and it does not undercut the three proved examples. The reader's conditional verdict already captures this risk, so no change to the verdict is needed.","tokens_in":25223,"tokens_out":24795,"duration_ms":287603,"concrete_test":"Implement Algorithm 2 with epsilon_a = loglogN / sqrt(N) for a two-arm design with equal means, T = 10,000, B = 10,000, over 10,000 replications, using two allocation rules: (i) clipped epsilon-greedy (Appendix C) and (ii) a 'pessimistic commit' rule that pulls the arm with the smaller sample mean in the second half of the experiment. If either rule's rejection rate at alpha = 0.05 exceeds 0.06, then the quantile-widening principle is violated and the 'wide variety' guarantee in the abstract is false; if both remain at or below 0.05, the concern is weakened but the missing general theorem still needs to be supplied.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is Lemma 7's condition (i): the simulated null distribution of the target-arm sample mean must have quantiles at least as extreme as the observed distribution for every sample path, for every alpha. The paper proves this only by separate limiting-distribution calculations for ETC, UCB, and the clipped reward-maximizing scheme (Examples 1-3). No general condition on the allocation rule is stated. Section E (Limitations) says explicitly: 'there exists minimal unifying theory on what/which designs this approach provides valid type I error.' Consequently, the abstract's claim of guarantees 'over a wide variety of common bandit designs' is not backed by a theorem. More importantly, the monotonicity itself is not obviously automatic: if a design increases the probability of pulling the target arm when other arms look better (e.g., an allocation rule targeting under-explored arms or a fairness-constrained rule), then positively biasing non-target means could shrink the simulated target-arm sample size, making the simulated quantiles less extreme and inflating type I error. The epsilon-greedy design used in the paper's own Appendix C is not covered by Theorem 1, so the paper's empirical demonstration there does not repair the theoretical gap. The claim is therefore conditionally correct for the three named designs but not for the stated scope.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a simulation-based inference method for adaptive experiments. Given an observed trajectory and a point null for a target arm, Algorithm 1 resimulates trajectories under the null using Gaussian outcomes, setting non-target arm means to positively biased estimates plus a vanishing optimism term and using sample variances. Algorithm 2 rejects the null when the observed sample mean falls outside the empirical quantiles of the simulated statistics, and Algorithm 3 inverts these tests to build confidence intervals and a heuristic point estimate. Theorem 1 claims asymptotic type I error control for an explore-then-commit design (Example 1), a UCB design (Example 2), and a clipped reward-maximizing design (Example 3) as the number of simulations B grows. Lemmas 1 and 2 provide power-one consistency and convergence of the confidence set/point estimate. Empirical comparisons on synthetic data and an MTurk adaptive experiment report tighter intervals than baseline anytime-valid and reweighting approaches.","tokens_in":25492,"tokens_out":15270,"duration_ms":155684,"significance":"Conditional on Theorem 1, the paper offers a genuinely different route to post-adaptive inference: it avoids conditional positivity, handles non-normal limiting distributions, and is computationally more attractive than nuisance-grid scanning. The case-by-case proofs for ETC, UCB, and the clipped reward-maximizing design are detailed, and the empirical study includes both synthetic and real-world adaptive data, with runtime scaling reported. The main weakness is that the announced scope substantially exceeds what is proved: the 'optimism widens quantiles' principle is verified only for three designs, and the paper's own limitations section concedes that no unifying theory is provided. This scope mismatch, rather than the core construction, is the principal reason the manuscript needs revision.","major_comments":[{"comment":"The abstract and Introduction claim guarantees 'over a wide variety of common bandit designs' and over a 'wide class of commonly used designs,' but Theorem 1 proves type I error control only for Examples 1, 2, and 3. Section E explicitly states that there is no unifying theory identifying which designs enjoy the guarantee, and the epsilon-greedy experiments in Appendix C fall outside Theorem 1 and therefore cannot repair the theoretical gap. The stated scope should be narrowed to the verified designs, or a general sufficient condition on the allocation rule should be proved.","section":"Abstract, §1, §E"},{"comment":"Lemma 7's condition (i) is the load-bearing 'optimism widens quantiles' property: for every sample path, the simulated null quantiles must be at least as extreme as the observed quantiles. The manuscript verifies this only through separate limiting calculations in Sections D.2.1-D.2.3 and gives no design-level condition that implies it. The property is not automatic: for an allocation rule that increases target-arm pulls when non-target arms look better, adding positive bias to non-target means could shrink the simulated target-arm sample size and make simulated quantiles less extreme, inflating type I error. In addition, condition (i) is stated with a sup/inf over the entire sample space of random quantiles, and the proof's passage to F(sup_omega ...) needs a uniformity or measurability argument that is not supplied. The theorem would be on solid ground if the monotonicity property were stated as a formal condition on the design and verified, rather than checked case by case.","section":"Lemma 7 and §D.2"},{"comment":"The proof of Lemma 2 in Section D.3 invokes Remark 5, which says that Algorithm 3 should be modified so that the confidence set always contains the empirical mean estimate. Algorithm 3 as printed already initializes the set with rho(H_T), so the gap is repairable, but the statement, proof, and remark are internally inconsistent, and Remark 5 is phrased as a future edit rather than as a property of the presented algorithm. Since Lemma 2 is the formal basis for the paper's confidence-interval consistency claim, this part of the manuscript needs to be brought into alignment before the result can be accepted as stated.","section":"Lemma 2, §D.3, Remark 5"}],"minor_comments":[{"comment":"The sentence 'Using these simulations, we characterize the distribution potentially non-normal sample mean test statistic to conduct inference' is grammatically incomplete and should be rewritten.","section":"Abstract"},{"comment":"The displayed nuisance vector ends with sigma_hat_2^2 where sigma_hat_K^2 is clearly intended.","section":"Algorithm 1, line 2"},{"comment":"Both conditions in Lemma 5 are labelled (i); the second should be labelled (ii).","section":"Lemma 5"},{"comment":"The text says that as G doubles the runtime 'doubles exactly,' but the reported values 11.50 to 20.21 and 20.66 to 39.17 are only approximately proportional; the wording should be softened.","section":"Table 1"},{"comment":"The phrase 'there exists minimal unifying theory' in the Limitations section should read 'there is little or no unifying theory,' and Remark 5 should be removed or converted into a formal part of the algorithm statement rather than a note about a future edit.","section":"§E and Remark 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about its main theoretical gap, and the core construction appears plausible for the three named designs. The mismatch between the abstract's 'wide variety' claim and the case-by-case proofs is the key issue; a careful revision that either narrows the claims or supplies a general sufficient condition would make this a solid contribution. I do not see a reproducibility concern: the algorithms are clear and the empirical details are sufficient."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Kevin— This one's worth a serious look, but read the limitations section before the abstract. The core trick—simulating the experiment under the null with positively biased nuisance means—is genuinely new, and for the three designs the paper actually proves (two-armed ETC, UCB, clipped reward-maximizing) the type I error argument is detailed and mostly convincing. It also handles nonparametric arm distributions and multiple arms, and the empirical gains on under-targeted arms are real: up to 50% narrower intervals in their setups with coverage matching. That's a practical advance over reweighting estimators that fail when conditional positivity fails.\n\nThe soft spot is scope. Theorem 1 is proven case-by-case for Examples 1-3, and Appendix E says plainly there is 'minimal unifying theory' on which designs the approach covers. The abstract's 'wide variety of common bandit designs' is not backed by a theorem. The load-bearing monotonicity—positive bias to non-target arms widens the simulated null quantiles—is not automatic. If the allocation rule responds to optimistic means by pulling the target arm more (for instance a rule that favors under-explored arms), the quantile widening could reverse. The paper doesn't state a general condition that rules this out. That's a real gap, but it's an honest one: they flag it themselves.\n\nThe other issues are minor. Lemma 2 quietly assumes the confidence set is non-empty a.s.; Remark 5 acknowledges and patches by including the sample mean. The bias term ε is a free parameter with a lower-but-no-upper-bound condition; they give sensible heuristics and show constant bias hurts power. The Appendix C epsilon-greedy experiments are not covered by the theorem, but they don't claim coverage there beyond empirics.\n\nBottom line: I'd send it to a good referee. The technique is new, the proofs for the named designs are substantive, and the limitation is disclosed rather than hidden. But the authors should either prove a general condition or walk the abstract back to 'for several common designs.' A revised version that fixes Lemma 2's assumption and releases code would be a solid paper.","headline":"Genuinely new simulation-with-optimism method with solid proofs for three designs, but the 'wide variety' claim overreaches and the paper admits it.","tokens_in":26000,"tokens_out":2711,"would_cite":true,"duration_ms":26610,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62L05","62F40","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"Simulating adaptive experiments with positively biased nuisance means yields hypothesis tests and confidence intervals with asymptotic type I error control for designs where standard reweighting fails.","keywords":["simulation-based inference","adaptive experiments","multi-armed bandits","type I error control","simulation with optimism","confidence intervals","post-experiment inference","explore-then-commit"],"falsifier":"Take an adaptive design not covered by Examples 1-3, for example a two-armed rule that commits to the target arm when the optimistic estimate of the other arm is high, so that raising the nuisance mean increases target-arm sampling. Resimulate under the null with the Theorem 1 bias and compare the empirical rejection rate at $\\theta^*$ over many independent replications; if the rate exceeds $\\alpha$ as $T$ and $B$ grow, the optimism-widens-quantiles principle fails for that design.","tokens_in":25013,"feed_emoji":"📊","tokens_out":8713,"duration_ms":79878,"temperature":0.7,"pith_summary":"After an adaptive experiment, the sample mean of an arm is a natural test statistic, but its distribution is shaped by the data-dependent sampling rule and is generally not normal; standard Wald intervals and reweighting schemes either require conditional positivity or are conservative. This paper proposes to learn the null distribution by resimulating the whole experiment many times, and proves that adding a carefully sized positive bias to the estimated means of all non-target arms—'simulation with optimism'—makes the simulated distribution conservative enough for valid tests. Theorem 1 establishes asymptotic type I error control for explore-then-commit, UCB, and clipped reward-maximizing designs, with no conditional-positivity assumption. Inverting the test gives confidence intervals and consistent point estimates, and the experiments show intervals up to 50% narrower than existing approaches, with the largest gains for arms the design samples least. If the guarantees hold, adaptive experiments can be analyzed with a simple simulation procedure instead of specialized asymptotics.","feed_headline":"Add optimism to resimulations and bandit tests become valid","feed_subtitle":"Resimulating adaptive trials with inflated nuisance means controls error rates and tightens confidence intervals by up to 50%.","key_machinery":"The mechanism is 'simulation with optimism': Algorithm 1 resimulates the experiment under the null by following the known adaptive policy, drawing each arm's outcomes as Gaussian with the target arm's mean fixed at $\\theta_0$ and every other arm's mean shifted upward by $\\epsilon_a$, and with variances estimated from the observed data. The bias $\\epsilon_a$ is chosen to dominate the law-of-iterated-logarithm scale $\\sqrt{\\log\\log N_T(a)/N_T(a)}$, which guarantees the optimistic nuisance eventually lies above the true mean almost surely. In the designs considered, that upward shift makes the simulated sample-mean distribution of the target arm wider than the true distribution, so comparing the observed statistic to simulated quantiles errs on the side of not rejecting. The proofs verify this quantile-widening case by case, using stability results for UCB, almost-sure convergence lemmas for random pull counts, and Glivenko-Cantelli convergence as the number of simulations $B$ grows.","core_discovery":"The paper's central claim is Theorem 1. In Algorithm 1, set the nuisance mean of each non-target arm $a$ to $\\hat\\mu_a = \\hat\\mu_T(a) + \\epsilon_a$, where $\\epsilon_a > 0$ and $\\sqrt{\\log\\log N_T(a)/N_T(a)}/\\epsilon_a \\to 0$, and set $\\hat\\sigma_a^2$ to the sample variance. Then the two-sided resimulation test in Algorithm 2 satisfies $\\limsup_{T\\to\\infty}\\lim_{B\\to\\infty} P(\\xi(\\theta^*, \\alpha, H_T) = 1) \\le \\alpha$ for the ETC, UCB, and clipped reward-maximizing designs, for every $\\alpha \\in [0,1]$. The force of the result is that it covers designs where the probability of selecting an arm can vanish, so asymptotic-normality reweighting is unavailable, and it bypasses the plug-in failure documented in Remark 2, where even $\\sqrt{T}$-consistent nuisance estimates do not make the simulated and observed statistics share a limiting distribution. Lemma 2 adds that the inverted confidence set collapses almost surely to $\\{\\theta^*\\}$ and the point estimate is strongly consistent.","pith_inferences":["The optimism principle suggests a general recipe: any adaptive design for which raising the other arms' means weakly increases the spread of the target arm's sample-mean distribution should inherit error control, and a unifying monotonicity condition would extend the theorem well beyond the three verified designs.","Because the bias term only needs to dominate the law-of-iterated-logarithm scale, one could tune $\\epsilon_a$ arm-by-arm, using a smaller bias for well-sampled arms to recover power while preserving the theorem's rate condition; the paper's appendix already shows that smaller bias improves power.","The same resimulation idea could apply to other test statistics or off-policy estimates whenever the adaptive policy is known, but the paper proves guarantees only for sample means and their differences, so such extensions would need fresh proofs."],"forward_implications":["Tests and confidence intervals for arm means and their differences become valid after explore-then-commit, UCB, and clipped reward-maximizing designs, even when the design violates conditional positivity.","In the paper's experiments, confidence intervals are up to 50% narrower than the best baseline, with the largest gains for arms the adaptive design samples least.","The sample-mean test has asymptotic power 1: any false null is rejected almost surely as the horizon grows, so the confidence set collapses to the true mean.","The procedure is computationally practical: constructing a confidence interval costs $O(G B)$ simulations of length $T$, where $G$ is the grid of nulls and $B$ the number of trajectories per null."],"supporting_citations":[{"why":"Baseline reweighting approach whose conditional-positivity requirement motivates the need for a different method; the paper's ETC example shows plug-in nuisances fail.","marker":"[11]"},{"why":"Supplies the stability theorem used in the proof for the UCB design in Example 2.","marker":"[15]"},{"why":"Supports Assumption 1 by showing regret-optimal schemes sample all arms at least on the order of log T.","marker":"[16]"},{"why":"Provides the real-world adaptive experiment data used in the empirical comparison.","marker":"[20]"},{"why":"Supplies the almost-sure convergence fact used to transfer limits through random pull counts.","marker":"[24]"},{"why":"Supplies the Glivenko-Cantelli theorem used to pass from a finite number B of simulations to the limiting simulated CDF.","marker":"[26]"},{"why":"Source of the law of iterated logarithm that sets the bias rate, and of anytime-valid bounds used for unbounded parameter grids.","marker":"[29]"}],"fun_headline_variants":["Optimistic resimulation rescues bandit inference","Positive bias in resimulations unlocks valid adaptive tests","Simulation with optimism shrinks bandit confidence intervals","Inflated nuisances in simulation make bandit tests valid","Resimulating with optimism tightens adaptive experiment CIs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The guarantee hinges on the principle that adding positive bias to the other arms' means makes the simulated null distribution of the target arm's sample mean at least as wide as the true distribution; the paper proves this only for its three example designs and states that no unifying theory for a broader class is known.","fun_headline_variants_meta":{"raw":{"variants":["Optimistic resimulation rescues bandit inference","Positive bias in resimulations unlocks valid adaptive tests","Simulation with optimism shrinks bandit confidence intervals","Inflated nuisances in simulation make bandit tests valid","Resimulating with optimism tightens adaptive experiment CIs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000906,"raw_usage":{"total_tokens":3920,"prompt_tokens":991,"completion_tokens":2929,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":2851}},"tokens_in":607,"tokens_out":2929,"duration_ms":24054,"temperature":1.0,"reasoning_tokens":2851,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:13:21.057646+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take an adaptive design not covered by Examples 1-3, for example a two-armed rule that commits to the target arm when the optimistic estimate of the other arm is high, so that raising the nuisance mean increases target-arm sampling. Resimulate under the null with the Theorem 1 bias and compare the empirical rejection rate at $\\theta^*$ over many independent replications; if the rate exceeds $\\alpha$ as $T$ and $B$ grow, the optimism-widens-quantiles principle fails for that design.","supporting_citations":[{"cited_title":"Offer-Westort, A","cited_arxiv_id":null,"evidence_quote":"Provides the real-world adaptive experiment data used in the empirical comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Glivenko-Cantelli theorem used to pass from a finite number B of simulations to the limiting simulated CDF."}],"review_version":1}