Pith. sign in

Rrm: Robust reward model training mitigates reward hacking.arXiv preprint arXiv:2409.13156

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

years

2026 5 2025 1

representative citing papers

Uncertainty-Aware Reward Modeling for Stable RLHF

cs.LG · 2026-06-18 · unverdicted · novelty 6.0

UARM equips reward models with quantile-based conformal prediction uncertainty and reweights GRPO advantages via heteroscedastic variance decomposition to improve calibration and reduce reward hacking in RLHF.

Optimal Transport for LLM Reward Modeling from Noisy Preference

cs.LG · 2026-05-07 · unverdicted · novelty 6.0

SelectiveRM applies optimal transport with a joint consistency discrepancy and partial mass relaxation to produce reward models that optimize a tighter upper bound on clean risk while autonomously dropping noisy preference samples.

citing papers explorer

Showing 6 of 6 citing papers.