A cross-validation estimator that splits each A/B test's units eliminates the winner's-curse bias in evaluating decision rules across many weak experiments, with theory, simulations, and a Netflix case study reporting a 33% estimated gain.
Optimizing Returns from Experimentation Programs
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Experimentation in online digital platforms is used to inform decision making. Specifically, the goal of many experiments is to optimize a metric of interest. Null hypothesis statistical testing can be ill-suited to this task, as it is indifferent to the magnitude of effect sizes and opportunity costs. Given access to a pool of related past experiments, we discuss how experimentation practice should change when the goal is optimization. We survey the literature on empirical Bayes analyses of A/B test portfolios, and single out the A/B Testing Problem (Azevedo et al., 2020) as a starting point, which treats experimentation as a constrained optimization problem. We show that the framework can be solved with dynamic programming and implemented by appropriately tuning $p$-value thresholds. Furthermore, we develop several extensions of the A/B Testing Problem and discuss the implications of these results on experimentation programs in industry. For example, under no-cost assumptions, firms should be testing many more ideas, reducing test allocation sizes, and relaxing $p$-value thresholds away from $p = 0.05$.
citation-role summary
citation-polarity summary
fields
stat.ME 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Evaluating Decision Rules Across Many Weak Experiments
A cross-validation estimator that splits each A/B test's units eliminates the winner's-curse bias in evaluating decision rules across many weak experiments, with theory, simulations, and a Netflix case study reporting a 33% estimated gain.