Pith. sign in

REVIEW 3 cited by

Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.06000 v1 pith:IVASIGWS submitted 2024-12-08 cs.CL cs.LG

classification cs.CLcs.LG
keywords rlhfmodelsperformancedatamodelpolicyrewardcomputational
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study explores the scaling properties of Reinforcement Learning from Human Feedback (RLHF) in Large Language Models (LLMs). Although RLHF is considered an important step in post-training of LLMs, its scaling potential is still largely unknown. We systematically analyze key components in the RLHF framework--model size, data composition, and inference budget--and their impacts on performance. Our findings show that increasing data diversity and volume improves reward model performance, helping process-supervision models scale better. For policy training, more response samples per prompt boost performance initially but quickly plateau. And larger reward models offer modest gains in policy training. In addition, larger policy models benefit less from RLHF with a fixed reward model. Overall, RLHF scales less efficiently than pretraining, with diminishing returns from additional computational resources. Based on these observations, we propose strategies to optimize RLHF performance within computational limits.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tiny Reward Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    TinyRM shows that 400M-parameter bidirectional masked language models, tuned with FLAN-style prompting, DoRA, and layer freezing, outperform a 70B reward model on RewardBench reasoning and come close on safety.

  2. Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning

    cs.LG 2025-06 reject novelty 6.0 of 10

    From Open LLM Leaderboard data grouped by base model, the authors recover a three-factor ordering of LLM capabilities and claim instruction-following causally supports math reasoning.

  3. VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VARD fine-tunes diffusion models by backpropagating through a learned value function that assigns dense, differentiable reward estimates to every intermediate denoising step, with KL regularization keeping the model n...

Pith tools