The authors derive a tractable max-min policy optimization over all r-correlated proxy rewards and show it yields more robust policies than ORPO, with an extension for linear rewards that also produces interpretable worst-case proxies.
We should note that the reason these frameworks are potentially applicable is that our formulation admits a closed-form solution for the inner minimization
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
The authors derive a tractable max-min policy optimization over all r-correlated proxy rewards and show it yields more robust policies than ORPO, with an extension for linear rewards that also produces interpretable worst-case proxies.