Hackable confidence rewards let LLMs selectively answer wrong to maximize reward; non-hackable schemes form a spectrum that trades accuracy for calibration and should be tuned as a hyperparameter.
A statistical theory of target detection by pulsed radar,
1 Pith paper cite this work, alongside 791 external citations. Polarity classification is still indexing.
1
Pith paper citing it
791
external citations · external index
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models
Hackable confidence rewards let LLMs selectively answer wrong to maximize reward; non-hackable schemes form a spectrum that trades accuracy for calibration and should be tuned as a hyperparameter.