Pith. sign in

REVIEW 2 cited by

An Experimental Design for Anytime-Valid Causal Inference on Multi-Armed Bandits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.05794 v4 pith:FPGOFZEZ submitted 2023-11-09 stat.ME cs.LG

classification stat.MEcs.LG
keywords designsequenceanytime-validbanditexperimentationinferencemanagerszero
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Experimentation is crucial for managers to rigorously quantify the value of a change and determine if it leads to a statistically significant improvement over the status quo. As companies increasingly mandate that all changes undergo experimentation before widespread release, two challenges arise: (1) minimizing the proportion of customers assigned to the inferior treatment and (2) increasing experimentation velocity by enabling data-dependent stopping. This paper addresses both challenges by introducing the Mixture Adaptive Design (MAD), a new experimental design for multi-armed bandit (MAB) algorithms that enables anytime-valid inference on the Average Treatment Effect (ATE) for \emph{any} MAB algorithm. Intuitively, MAD "mixes" any bandit algorithm with a Bernoulli design, where at each time step, the probability of assigning a unit via the Bernoulli design is determined by a user-specified deterministic sequence that can converge to zero. This sequence lets managers directly control the trade-off between regret minimization and inferential precision. Under mild conditions on the rate the sequence converges to zero, we provide a confidence sequence that is asymptotically anytime-valid and guaranteed to shrink around the true ATE. Hence, when the true ATE converges to a non-zero value, the MAD confidence sequence is guaranteed to exclude zero in finite time. Therefore, the MAD enables managers to stop experiments early while ensuring valid inference, enhancing both the efficiency and reliability of adaptive experiments. Empirically, we demonstrate that the MAD achieves finite-sample anytime-validity while accurately and precisely estimating the ATE, all without incurring significant losses in reward compared to standard bandit designs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Confidence Horizons

    stat.ME 2026-08 conditional novelty 8.0 of 10

    A new family of 'asymptotic confidence horizons' provides large-sample anytime-valid coverage on bounded time windows, with closed-form boundary quantiles and connections to group sequential methods.

  2. Demonstration Experiments

    math.ST 2026-03 conditional novelty 7.0 of 10

    Valid global-null tests for threshold bandits under adaptive sampling, backed by a moderate-deviations principle for sequential t-statistics and a log-regret SNR allocation rule.

Pith tools