An ordered diagnostic protocol screens proxy rewards and contextual-bandit policies for alignment and learnability before deployment, and shows offline batch estimates can mislead under delayed feedback.
Hitsch, Sanjog Misra, and Walter W
1 Pith paper cite this work, alongside 32 external citations. Polarity classification is still indexing.
1
Pith paper citing it
32
external citations · OpenAlex
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits
An ordered diagnostic protocol screens proxy rewards and contextual-bandit policies for alignment and learnability before deployment, and shows offline batch estimates can mislead under delayed feedback.