Pith. sign in

Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

On the Robustness of Reward Models for Language Model Alignment

cs.CL · 2025-05-12 · conditional · novelty 5.0

Adding a batch-wise sum-to-zero penalty to Bradley-Terry reward modeling makes reward models more robust to unseen prompts and responses, according to experiments across multiple model families and benchmarks.

citing papers explorer

Showing 1 of 1 citing paper.

  • On the Robustness of Reward Models for Language Model Alignment cs.CL · 2025-05-12 · conditional · none · ref 2

    Adding a batch-wise sum-to-zero penalty to Bradley-Terry reward modeling makes reward models more robust to unseen prompts and responses, according to experiments across multiple model families and benchmarks.