Pith. sign in

REVIEW 2 cited by

Large-scale Validation of Counterfactual Learning Methods: A Test-Bed

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1612.00367 v2 pith:S6GTTHSG submitted 2016-12-01 cs.LG cs.AIstat.ML

Large-scale Validation of Counterfactual Learning Methods: A Test-Bed

classification cs.LG cs.AIstat.ML
keywords learningoff-policydatamethodstest-bedadvertisinglarge-scalereal-world
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The ability to perform effective off-policy learning would revolutionize the process of building better interactive systems, such as search engines and recommendation systems for e-commerce, computational advertising and news. Recent approaches for off-policy evaluation and learning in these settings appear promising. With this paper, we provide real-world data and a standardized test-bed to systematically investigate these algorithms using data from display advertising. In particular, we consider the problem of filling a banner ad with an aggregate of multiple products the user may want to purchase. This paper presents our test-bed, the sanity checks we ran to ensure its validity, and shows results comparing state-of-the-art off-policy learning methods like doubly robust optimization, POEM, and reductions to supervised learning using regression baselines. Our results show experimental evidence that recent off-policy learning methods can improve upon state-of-the-art supervised learning techniques on a large-scale real-world data set.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Estimating Causal Effects from Data Generated by Stochastic Algorithms

    stat.ME 2026-07 accept novelty 8.0

    Logging the features and relative probability of one unexposed item alongside the exposed item identifies causal effects of content features from stochastic algorithms even with unobserved confounders.

  2. Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap

    cs.LG 2026-07 conditional novelty 6.0

    Reframing A/B assignment as a mixture policy and applying Δ-off-policy estimators yields an unbiased ATE estimator with variance provably no larger than difference-in-means whenever the tested policies overlap.