An ordered diagnostic protocol screens proxy rewards and contextual-bandit policies for alignment and learnability before deployment, and shows offline batch estimates can mislead under delayed feedback.
Imbens, and Hyunseung Kang
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits
An ordered diagnostic protocol screens proxy rewards and contextual-bandit policies for alignment and learnability before deployment, and shows offline batch estimates can mislead under delayed feedback.