Pith. sign in

REVIEW 3 cited by

Improving the Estimation of Lifetime Effects in A/B Testing via Treatment Locality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.19618 v3 pith:7ELBTXX6 submitted 2024-07-29 stat.ME cs.LGecon.EMstat.APstat.ML

classification stat.MEcs.LGecon.EMstat.APstat.ML
keywords lifetimelocalitylong-termoutcomesshort-termsystemstestingtreatments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Utilizing randomized experiments to evaluate the effect of short-term treatments on the short-term outcomes has been well understood and become the golden standard in industrial practice. However, as service systems become increasingly dynamical and personalized, much focus is shifting toward maximizing long-term outcomes, such as customer lifetime value, through lifetime exposure to interventions. Our goal is to assess the impact of treatment and control policies on long-term outcomes from relatively short-term observations, such as those generated by A/B testing. A key managerial observation is that many practical treatments are local, affecting only targeted states while leaving other parts of the policy unchanged. This paper rigorously investigates whether and how such locality can be exploited to improve estimation of long-term effects in Markov Decision Processes (MDPs), a fundamental model of dynamic systems. We first develop optimal inference techniques for general A/B testing in MDPs and establish corresponding efficiency bounds. We then propose methods to harness the localized structure by sharing information on the non-targeted states. Our new estimator can achieve a linear reduction with the number of test arms for a major part of the variance without sacrificing unbiasedness. It also matches a tighter variance lower bound that accounts for locality. Furthermore, we extend our framework to a broad class of differentiable estimators, which encompasses many widely used approaches in practice. We show that all such estimators can benefit from variance reduction through information sharing without increasing their bias. Together, these results provide both theoretical foundations and practical tools for conducting efficient experiments in dynamic service systems with local treatments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploiting Similarities in A/B Testing with Off-Policy Estimation

    stat.ML 2025-06 reject novelty 6.0 of 10

    The authors propose a family of unbiased importance-weighted A/B testing estimators with a fallback to difference-in-means, but the claimed surrogate-optimal and misspecification-aware transforms are derived from an i...

  2. Experimental Designs for Multi-Item Multi-Period Inventory Control

    stat.ME 2025-01 conditional novelty 6.0 of 10

    Switchback experiments underestimate the global treatment effect in shared-capacity inventory systems, item-level randomization overestimates it, and a pairwise item-time design has intermediate bias.

  3. A Two-armed Bandit Framework for A/B Testing

    stat.ML 2025-07 conditional novelty 4.0 of 10

    A two-armed bandit based test statistic with permutation aggregation improves power for A/B testing in both i.i.d. and dynamic settings.

Pith tools