Pith. sign in

Optimizing Returns from Experimentation Programs

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Experimentation in online digital platforms is used to inform decision making. Specifically, the goal of many experiments is to optimize a metric of interest. Null hypothesis statistical testing can be ill-suited to this task, as it is indifferent to the magnitude of effect sizes and opportunity costs. Given access to a pool of related past experiments, we discuss how experimentation practice should change when the goal is optimization. We survey the literature on empirical Bayes analyses of A/B test portfolios, and single out the A/B Testing Problem (Azevedo et al., 2020) as a starting point, which treats experimentation as a constrained optimization problem. We show that the framework can be solved with dynamic programming and implemented by appropriately tuning $p$-value thresholds. Furthermore, we develop several extensions of the A/B Testing Problem and discuss the implications of these results on experimentation programs in industry. For example, under no-cost assumptions, firms should be testing many more ideas, reducing test allocation sizes, and relaxing $p$-value thresholds away from $p = 0.05$.

citation-role summary

background 1

citation-polarity summary

fields

stat.ME 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Evaluating Decision Rules Across Many Weak Experiments

stat.ME · 2025-02-12 · conditional · novelty 6.0

A cross-validation estimator that splits each A/B test's units eliminates the winner's-curse bias in evaluating decision rules across many weak experiments, with theory, simulations, and a Netflix case study reporting a 33% estimated gain.

citing papers explorer

Showing 1 of 1 citing paper.

  • Evaluating Decision Rules Across Many Weak Experiments stat.ME · 2025-02-12 · conditional · none · ref 19 · internal anchor

    A cross-validation estimator that splits each A/B test's units eliminates the winner's-curse bias in evaluating decision rules across many weak experiments, with theory, simulations, and a Netflix case study reporting a 33% estimated gain.