Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2026 1

verdicts

UNVERDICTED 1

representative citing papers

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling

cs.LG · 2026-04-24 · unverdicted · novelty 7.0

Temporally Coherent Reward Modeling adds Monte Carlo and TD regularizers to Bradley-Terry training so reward model outputs at every token become conditional expectations of the final reward, improving token-level interpretability, process supervision from outcome data, and PPO efficiency.

citing papers explorer

Showing 1 of 1 citing paper.

  • Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling cs.LG · 2026-04-24 · unverdicted · none · ref 6

    Temporally Coherent Reward Modeling adds Monte Carlo and TD regularizers to Bradley-Terry training so reward model outputs at every token become conditional expectations of the final reward, improving token-level interpretability, process supervision from outcome data, and PPO efficiency.