Pith. sign in

REVIEW 1 cited by

Optimizing Returns from Experimentation Programs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.05508 v1 pith:5DT623AF submitted 2024-12-07 stat.ME econ.EMstat.AP

classification stat.MEecon.EMstat.AP
keywords experimentationtestingproblemdiscussexperimentsgoalmanyoptimization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Experimentation in online digital platforms is used to inform decision making. Specifically, the goal of many experiments is to optimize a metric of interest. Null hypothesis statistical testing can be ill-suited to this task, as it is indifferent to the magnitude of effect sizes and opportunity costs. Given access to a pool of related past experiments, we discuss how experimentation practice should change when the goal is optimization. We survey the literature on empirical Bayes analyses of A/B test portfolios, and single out the A/B Testing Problem (Azevedo et al., 2020) as a starting point, which treats experimentation as a constrained optimization problem. We show that the framework can be solved with dynamic programming and implemented by appropriately tuning $p$-value thresholds. Furthermore, we develop several extensions of the A/B Testing Problem and discuss the implications of these results on experimentation programs in industry. For example, under no-cost assumptions, firms should be testing many more ideas, reducing test allocation sizes, and relaxing $p$-value thresholds away from $p = 0.05$.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Decision Rules Across Many Weak Experiments

    stat.ME 2025-02 conditional novelty 6.0 of 10

    A cross-validation estimator that splits each A/B test's units eliminates the winner's-curse bias in evaluating decision rules across many weak experiments, with theory, simulations, and a Netflix case study reporting...

Pith tools