REVIEW 10 cited by
Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Off-policy evaluation (OPE) aims to estimate the performance of hypothetical policies using data generated by a different policy. Because of its huge potential impact in practice, there has been growing research interest in this field. There is, however, no real-world public dataset that enables the evaluation of OPE, making its experimental studies unrealistic and irreproducible. With the goal of enabling realistic and reproducible OPE research, we present Open Bandit Dataset, a public logged bandit dataset collected on a large-scale fashion e-commerce platform, ZOZOTOWN. Our dataset is unique in that it contains a set of multiple logged bandit datasets collected by running different policies on the same platform. This enables experimental comparisons of different OPE estimators for the first time. We also develop Python software called Open Bandit Pipeline to streamline and standardize the implementation of batch bandit algorithms and OPE. Our open data and software will contribute to fair and transparent OPE research and help the community identify fruitful research directions. We provide extensive benchmark experiments of existing OPE estimators using our dataset and software. The results open up essential challenges and new avenues for future OPE research.
Forward citations
Cited by 10 Pith papers
-
Estimating Causal Effects from Data Generated by Stochastic Algorithms
Logging the features and relative probability of one unexposed item alongside the exposed item identifies causal effects of content features from stochastic algorithms even with unobserved confounders.
-
Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing
Randomly intervening on binary predictions and testing outcome-distribution differences detects outcome performativity, with explicit sample-size formulas and regions where detection is infeasible.
-
Uncertainty Quantification and Causal Considerations for Off-Policy Decision Making
Three methods for off-policy evaluation: marginal ratio variance reduction, conformal predictive intervals, and causal bounds that falsify digital twins under unmeasured confounding.
-
When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits
An ordered diagnostic protocol screens proxy rewards and contextual-bandit policies for alignment and learnability before deployment, and shows offline batch estimates can mislead under delayed feedback.
-
From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation
A causal targeting system that optimizes treatment effects under global constraints and uses bandit exploration beat the incumbent prediction-based stack by a statistically significant 7.20% on LinkedIn Feed marketing...
-
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
A benchmark and small-scale evaluation suggesting LLM agents can modify off-policy evaluation code and sometimes improve the measured metrics, with the authors' two-agent framework the most reliable of those tested.
-
Off-Policy Evaluation of Ranking Policies via Embedding-Space User Behavior Modeling
The paper claims a new family of embedding-based off-policy estimators for ranking policies, but its central unbiasedness theorem is false under the stated assumptions.
-
Quick-Draw Bandits: Quickly Optimizing in Nonstationary Environments with Extremely Many Arms
Quick-Draw is a fast kernel-interpolation UCB bandit for large-arm nonstationary problems, but the claimed O*(sqrt T) regret bound does not follow from the paper's theorem.
-
Off-Policy Evaluation for Recommendations with Missing-Not-At-Random Rewards
Using both action-selection and reward-observation propensities gives an unbiased off-policy value estimator under position-based missingness and the no-direct-effect assumption.
-
Off-Policy Evaluation and Counterfactual Methods in Dynamic Auction Environments
Continuous off-policy estimators give directionally correct predictions of payment-policy performance in dynamic auctions, enabling counterfactual comparison and in-sample policy optimization from logged data.
Discussion (0). Continue with ORCID to comment.