REVIEW 2 major objections 5 minor 1 cited by
On the Robustness of Reward Models for Language Model Alignment
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that runaway hidden-state norm spread causes reward-model over-optimization and that a sum-to-zero batch regularizer cures it.
desk verdict BSR's empirical gains look real, but the paper's gradient derivation (Eq. 28) is wrong, so the norm-dispersion causal story doesn't hold. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Batch-wise sum-to-zero regularization (BSR): an additional squared term $\mathcal{L}_{\mathrm{BSR}} = \left(\frac{1}{2|B|}\sum_{i=1}^{|B|}\sum_{j\in\{w,l\}} r(x_i,y_{i,j})\right)^2$ added to the Bradley-Terry loss. Because its gradient is proportional to the reward itself, it pushes extreme positive and negative reward outliers back toward zero, damping the hidden-state norm dispersion that the paper identifies as the carrier of over-optimization; unlike logit normalization, it does not discard the magnitude information in reward prediction.
What would settle it
Train pooled reward models with and without BSR and evaluate them on freshly collected human-labeled preference pairs spanning several response styles; if the BSR model does not beat the plain Bradley-Terry model on those human labels, the claim that hidden-state norm dispersion is the main cause of over-optimization toward real preferences would be falsified.
Extended reading notes
Core claim
Starting from the factorization $r(x,y) = \|W_p\| \, \|h(x,y)\| \cos\psi$, the paper observes that the projection head norm stays near its initialization while the variance of $\|h(x,y)\|$ grows throughout training. It therefore claims that the BT objective's reward-margin maximization is effectively carried by inflating hidden-state norm differences, which is the same mechanism as over-confidence in classifiers. Adding BSR, a squared penalty on the batch sum of rewards, creates a gradient that pulls outlier rewards back toward zero proportionally to their magnitude; the paper shows empirically that this keeps hidden-state norms in a consistent range on unseen data. The claimed consequence is that BSR-trained reward models beat plain BT, hinge, logit-normalized, and margin-boosted baselines in all four over-optimization scenarios, and the robustness propagates to RLHF: the policy stays aligned to the gold reward model instead of stagnating, generation length drops by 40 percent, and the length-controlled AlpacaEval win rate rises by about 7 percent.
Load-bearing premise
The load-bearing assumption is that ArmoRM's preferences are a faithful stand-in for true human preferences; if ArmoRM carries the same verbosity or self-preference biases that the method targets, the measured robustness may be robustness to the wrong judge.
Editorial extensions
If this is right
- Adding BSR to the standard Bradley-Terry loss makes reward models more robust across unseen prompts, unseen response generators, and both simultaneously, on Llama-3 and Qwen2.5 backbones from 1B to 8B scale.
- Robustness transfers to RLHF: policies trained with RLOO on a BSR-regularized reward model keep improving on the gold preference model instead of stagnating, indicating less reward hacking.
- On RM-Bench, BSR's accuracy gain concentrates on hard preference pairs with subtle differences; at a suitable weight it exceeds the plain BT model by more than 5 percent on that harder subset.
- The regularized model is less verbose: RLOO with BSR produces 40 percent shorter generations than the SFT model while raising the length-controlled AlpacaEval 2.0 win rate by about 7 percent.
- The advantage of BSR grows with model size, suggesting that the regularization becomes more important as hidden dimensions and backbone capacity increase.
Reading between the lines
- The verbosity angle can be pushed further: since BSR keeps hidden-state norms stable and the paper ties norm outliers to reward gaming, a direct testable prediction is that BSR-trained reward models show a smaller reward-length correlation on held-out style distributions.
- The same norm-dispersion mechanism likely afflicts direct alignment objectives with implicit rewards, such as DPO-style losses, so a sum-to-zero-style regularizer could be adapted there; the paper does not test this.
- Because the controlled evaluation uses ArmoRM as the gold judge, the causal claim would be strengthened by repeating the four-scenario comparison with human preference labels; if the BSR advantage shrinks, part of the measured robustness is alignment to that judge's biases.
- A practical selection rule left implicit is to pick the regularization weight by hard-task accuracy, as the 8B setup does; that rule could transfer directly to other data scales and model families.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes batch-wise sum-to-zero regularization (BSR) as an add-on to the Bradley-Terry (BT) reward modeling loss. The authors argue that BT training induces excessive dispersion of hidden-state norms, which they claim is the main source of reward-model over-optimization. BSR is introduced to penalize extreme reward magnitudes, reduce this norm dispersion, and improve robustness to unseen prompts and response distributions. The paper evaluates BSR against several baselines in four generalization scenarios using ArmoRM as a gold preference model, assesses downstream RLHF training with RLOO, and reports experiments with an 8B model on RM-Bench and AlpacaEval 2.0, claiming consistent improvements including a 40% reduction in generation length and a 7% relative gain in win rate. Code, data, and models are released.
Significance. If the central claims hold, BSR is a simple, inexpensive regularizer with meaningful practical benefits for reward-model robustness and downstream RLHF alignment. The empirical study is broad: it covers multiple model families and sizes, defines a clean four-scenario evaluation of generalization, and includes external benchmarks (RM-Bench, AlpacaEval) in addition to the ArmoRM-based controlled setup. The release of code, data, and models strengthens reproducibility. However, the paper's core mechanistic account of BSR is undermined by an incorrect gradient derivation, and the causal claim that hidden-state norm dispersion is the main source of over-optimization is not established by the experiments. The empirical results remain valuable, but the theoretical framing requires substantial revision.
major comments (2)
- [Section 4.2, Eq. (28)] The gradient of L_BSR in Eq. (25) is miscomputed. With N = 2|B| rewards in the batch, L_BSR = ( (1/N) * sum_{k} r_k )^2, so for any example m, dL_BSR/dh_m = (2 * mean_r / N) * W_p = ( (1/|B|) * mean_r ) * W_p, which is identical for all examples in the batch. Eq. (28) instead states dL_BSR/dh(x_i,y_{i,j}) = (1/|B|) * r(x_i,y_{i,j}) * W_p, which corresponds to a per-example squared-reward penalty and is not the gradient of Eq. (25). Because the paper repeatedly justifies BSR as 'penalizing outliers with large magnitude' and attributes the reduced norm dispersion to gradient contributions proportional to individual rewards, this is not a minor typo: the described mechanism is not the mechanism of the implemented loss. The authors must either correct the derivation and reinterpret the regularizer as a batch-mean-centering term, or change the loss (e.g., to a per-example squared penalty) if the intended behavior is indeed outlier penalization.
- [Sections 4.1 and 3] The causal claim that excessive dispersion of hidden-state norms is the 'main source of over-optimization' is not supported by the evidence presented. Figure 2 shows a correlation between BT training and growing ||h(x,y_w) - h(x,y_l)||, and Figure 3 shows that BSR reduces norm dispersion, but no experiment directly manipulates norm dispersion while holding other factors fixed. Moreover, because the true BSR gradient is a common batch-level direction (see previous comment), the variance reduction in Figure 3 is an indirect side effect of mean-centering rather than a direct penalty on high-norm outliers. In addition, the four-scenario evaluation in Section 3 uses ArmoRM as the gold preference model r*; if ArmoRM carries its own verbosity or self-preference biases, then 'over-optimization' as measured here is alignment to ArmoRM, not necessarily to human preference. The authors should weaken the causal wording (e.g., to 'BSR empirically improves robustness') or add a controlled test that isolates norm dispersion, such as comparing BSR against an explicit per-example norm penalty or an L2 regularization on hidden states.
minor comments (5)
- [Section 5.3, Table 2] The table header contains 'BTBT-BSR' in the first column; this should read 'RMBT-BSR' for consistency with the rest of the paper.
- [Section 5.3, Table 1] The sentence 'RMBT-BSR experiences around 40% increase' appears to be a typo: the table shows that RMBT (the baseline) has the large train-to-eval effective-rank increase (+10.42), while RMBT-BSR is nearly stable (+0.14).
- [Section 2.2] The definition of over-optimization as an accuracy gap between in-domain and out-of-domain sets is non-standard; a brief sentence connecting this to the usual notion of reward over-optimization (reward increases while true reward degrades) would help the reader.
- [Section 5.3, Table 3] The average response length is reported without units; AlpacaEval typically reports token counts, and the normalization in Figure 5 is also not defined, so the plots are hard to interpret quantitatively.
- [Abstract and Section 5.3] The phrase 'with 8B RMs' could be misread as the policy being 8B; the RLHF experiments use a 1.5B policy (Qwen2.5-1.5B) with an 8B reward model, so the abstract would benefit from explicit clarification.
Circularity Check
No significant circularity: BSR gains are evaluated against external ArmoRM and RM-Bench/AlpacaEval, and self-citations are not load-bearing.
full rationale
The paper's central causal claim is that BT training inflates hidden-state norm dispersion, which drives over-optimization. That claim is supported by direct measurements (Figure 2) and by the correlation between hidden-state norm instability and out-of-domain degradation, not by the definition of the loss. BSR (Eq. 25) is a new regularizer whose robustness benefits are assessed against an external gold model, ArmoRM, and against external benchmarks (RM-Bench and AlpacaEval 2.0), so the central 'BSR improves robustness' result is not predetermined by its construction. The choice of ArmoRM as r* is a proxy assumption; even if ArmoRM carries its own biases, that is an evaluation-validity concern, not circular reasoning. Self-citations (ORPO, cross-lingual RM work, AlphaPO) appear only in related-work or example contexts and do not carry the argument; no uniqueness theorem or ansatz is imported from the authors' prior work. The skeptic's Eq. (28)-vs-Eq. (25) gradient observation is a mathematical correctness issue about whether BSR acts as a per-example outlier penalty, but it does not make any prediction equivalent to its inputs, so it does not constitute circularity.
Assumptions & free parameters
free parameters (1)
- lambda (BSR weight) =
10^-3 (also tested 10^-2, 10^-4)
assumptions (5)
- domain assumption ArmoRM's scores represent the true human preference distribution r*.
- domain assumption The reward is a linear projection of the last hidden state: r(x,y) = W_p^T h(x,y).
- domain assumption LLM hidden states have low effective rank at initialization, so initial hidden-state differences are small.
- domain assumption The four generalization scenarios (in-domain, prompt-disjoint, response-disjoint, mutual-disjoint) are the relevant axes of reward model distribution shift.
- domain assumption The Bradley-Terry model describes human preference in this setting.
Cite this review
Pith. "Pith review of On the Robustness of Reward Models for Language Model Alignment." pith.science (2026). https://pith.science/paper/PLKQWJ55
@misc{pith2026250507271,
author = {Pith},
title = {Pith review of: On the Robustness of Reward Models for Language Model Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/PLKQWJ55}},
note = {Machine review of arXiv:2505.07271}
}
read the original abstract
The Bradley-Terry (BT) model is widely practiced in reward modeling for reinforcement learning with human feedback (RLHF). Despite its effectiveness, reward models (RMs) trained with BT model loss are prone to over-optimization, losing generalizability to unseen input distributions. In this paper, we study the cause of over-optimization in RM training and its downstream effects on the RLHF procedure, accentuating the importance of distributional robustness of RMs in unseen data. First, we show that the excessive dispersion of hidden state norms is the main source of over-optimization. Then, we propose batch-wise sum-to-zero regularization (BSR) to enforce zero-centered reward sum per batch, constraining the rewards with extreme magnitudes. We assess the impact of BSR in improving robustness in RMs through four scenarios of over-optimization, where BSR consistently manifests better robustness. Subsequently, we compare the plain BT model and BSR on RLHF training and empirically show that robust RMs better align the policy to the gold preference model. Finally, we apply BSR to high-quality data and models, which surpasses state-of-the-art RMs in the 8B scale by adding more than 5% in complex preference prediction tasks. By conducting RLOO training with 8B RMs, AlpacaEval 2.0 reduces generation length by 40% while adding a 7% increase in win rate, further highlighting that robustness in RMs induces robustness in RLHF training. We release the code, data, and models: https://github.com/LinkedIn-XFACT/RM-Robustness.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
Preference-based LLM alignment under an unknown reward-preference link becomes a single-index model; three new algorithms converge to the optimal divergence-constrained policy without knowing the link.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s
Ahmadian, A., Cremer, C., Gall \'e , M., Fadaee, M., Kreutzer, J., Pietquin, O., \"U st \"u n, A., and Hooker, S. Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...
work page 2024
-
[3]
B., Lozhkov, A., Bakouch, E., Blázquez, G
Allal, L. B., Lozhkov, A., Bakouch, E., Blázquez, G. M., Penedo, G., Tunstall, L., Marafioti, A., Kydlíček, H., Lajarín, A. P., Srivastav, V., Lochner, J., Fahlgren, C., Nguyen, X.-S., Fourrier, C., Burtenshaw, B., Larcher, H., Zhao, H., Zakka, C., Morlon, M., Raffel, C., von Werra, L., and Wolf, T. Smollm2: When smol goes big -- data-centric training of ...
arXiv 2025
-
[4]
The falcon series of open language models, 2023
Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Étienne Goffinet, Hesslow, D., Launay, J., Malartic, Q., Mazzotta, D., Noune, B., Pannier, B., and Penedo, G. The falcon series of open language models, 2023. URL https://arxiv.org/abs/2311.16867
arXiv 2023
-
[5]
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022
arXiv 2022
-
[6]
G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pp.\ 2397--2430. PMLR, 2023
2023
-
[7]
Too much in common: Shifting of embeddings in transformer language models and its implications
Bi \'s , D., Podkorytov, M., and Liu, X. Too much in common: Shifting of embeddings in transformer language models and its implications. In Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (eds.), Proceedings of the 2021 Conference of the North American Chapter of the As...
2021
-
[8]
Bradley, R. A. and Terry, M. E. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39 0 (3/4): 0 324--345, 1952. ISSN 00063444. URL http://www.jstor.org/stable/2334029
arXiv 1952
Show all 78 references
-
[9]
K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T. T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P. J., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., ...
2023
-
[10]
H., Chen, S., Liu, Z., Jiang, F., and Wang, B
Chen, G. H., Chen, S., Liu, Z., Jiang, F., and Wang, B. Humans or LLM s as the judge? a study on judgement bias. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp.\ 8301--8327, Miam...
2024 doi
-
[11]
ODIN : Disentangled reward mitigates hacking in RLHF
Chen, L., Zhu, C., Chen, J., Soselia, D., Zhou, T., Goldstein, T., Huang, H., Shoeybi, M., and Catanzaro, B. ODIN : Disentangled reward mitigates hacking in RLHF . In Forty-first International Conference on Machine Learning, 2024 b . URL https://openreview.net/forum?id=zcIV8OQFVF
2024
-
[12]
E., Stoica, I., and Xing, E
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. Vicuna: An open-source chatbot impressing gpt-4 with 90\ URL https://lmsys.org/blog/2023-03-30-vicuna/
2023
-
[13]
F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017
2017
-
[14]
Support-vector networks
Cortes, C. Support-vector networks. Machine Learning, 1995
1995
-
[15]
Reward model ensembles help mitigate overoptimization
Coste, T., Anwar, U., Kirk, R., and Krueger, D. Reward model ensembles help mitigate overoptimization. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=dcjtMYkpXx
2024
-
[16]
Ultrafeedback: boosting language models with scaled ai feedback
Cui, G., Yuan, L., Ding, N., Yao, G., He, B., Zhu, W., Ni, Y., Xie, G., Xie, R., Lin, Y., Liu, Z., and Sun, M. Ultrafeedback: boosting language models with scaled ai feedback. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org, 2025
2025
-
[17]
Davidson, R. R. On extending the bradley-terry model to accommodate ties in paired comparison experiments. Journal of the American Statistical Association, 65 0 (329): 0 317--328, 1970. ISSN 01621459, 1537274X. URL http://www.jstor.org/stable/2283595
1970
-
[18]
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N., Chen, Y., Xu, B., Qin, Y., Hu, S., Liu, Z., Sun, M., and Zhou, B. Enhancing chat language models by scaling high-quality instructional conversations. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Lan...
2023 doi
-
[19]
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., Rodriguez, A., Gregerson, A., Spataru, A., Rozie...
2024 arXiv
-
[20]
Length-controlled alpacaeval: A simple debiasing of automatic evaluators
Dubois, Y., Liang, P., and Hashimoto, T. Length-controlled alpacaeval: A simple debiasing of automatic evaluators. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=CybBmzWBX0
2024
-
[21]
N., Jiao, J., Zhu, B., Gonzalez, J
Frick, E., Li, T., Chen, C., Chiang, W.-L., Angelopoulos, A. N., Jiao, J., Zhu, B., Gonzalez, J. E., and Stoica, I. How to evaluate reward models for rlhf. arXiv preprint arXiv:2410.14872, 2024
2024 arXiv
-
[22]
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J. Scaling laws for reward model overoptimization. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of ...
2023
-
[23]
and Irving, G
Gleave, A. and Irving, G. Uncertainty estimation for language reward models. arXiv preprint arXiv:2203.07472, 2022
2022 arXiv
-
[24]
Gupta, A., Tang, S., Song, Q., Zhu, S., Hong, J., Saha, A., Gupta, V., Lee, N., Kim, E., Zhu, J., Pillai, N., and Keerthi, S. S. Alpha PO - reward shape matters for LLM alignment. In Forty-second International Conference on Machine Learning, 2025. URL https://arxiv.org/abs/2501.03884
2025 arXiv
-
[25]
ORPO : Monolithic preference optimization without reference model
Hong, J., Lee, N., and Thorne, J. ORPO : Monolithic preference optimization without reference model. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp.\ 11170--11189, Miami, Florida...
2024
-
[26]
Cross-lingual transfer of reward models in multilingual alignment
Hong, J., Lee, N., Mart \'i nez-Casta \ n o, R., Rodr \'i guez, C., and Thorne, J. Cross-lingual transfer of reward models in multilingual alignment. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of ...
2025
-
[27]
Liger kernel: Efficient triton kernels for llm training, 2024
Hsu, P.-L., Dai, Y., Kothapalli, V., Song, Q., Tang, S., Zhu, S., Shimizu, S., Sahni, S., Ning, H., and Chen, Y. Liger kernel: Efficient triton kernels for llm training, 2024. URL https://arxiv.org/abs/2410.10989
2024 arXiv
-
[28]
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Hu, J., Wu, X., Zhu, Z., Xianyu, Wang, W., Zhang, D., and Cao, Y. Openrlhf: An easy-to-use, scalable and high-performance rlhf framework. arXiv preprint arXiv:2405.11143, 2024
2024 arXiv
-
[29]
The n+ implementation details of RLHF with PPO : A case study on TL ; DR summarization
Huang, S., Noukhovitch, M., Hosseini, A., Rasul, K., Wang, W., and Tunstall, L. The n+ implementation details of RLHF with PPO : A case study on TL ; DR summarization. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=kHO2ZTa8e3
2024
-
[30]
A generalized bradley-terry model: From group competition to individual skill
Huang, T.-k., Lin, C.-j., and Weng, R. A generalized bradley-terry model: From group competition to individual skill. In Saul, L., Weiss, Y., and Bottou, L. (eds.), Advances in Neural Information Processing Systems, volume 17. MIT Press, 2004. URL https://proceedings.neurips.c...
2004
-
[31]
Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mistral 7b, 2023 a
2023
-
[32]
Jiang, D., Ren, X., and Lin, B. Y. LLM -blender: Ensembling large language models with pairwise ranking and generative fusion. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volum...
2023 doi
-
[33]
Rank Correlation Methods
Kendall, M. Rank Correlation Methods. Griffin books on statistics. Hafner Publishing Company, 1962. URL https://books.google.co.kr/books?id=1whKAAAAMAAJ
1962
-
[34]
Evaluating robustness of reward models for mathematical reasoning, 2024
Kim, S., Kang, D., Kwon, T., Chae, H., Won, J., Lee, D., and Yeo, J. Evaluating robustness of reward models for mathematical reasoning, 2024. URL https://arxiv.org/abs/2410.01729
2024 arXiv
-
[35]
H., Gonzalez, J., Zhang, H., and Stoica, I
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, SOSP '23, pp.\ 611–626...
2023
-
[36]
Lambert, N., Morrison, J., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L. J. V., Liu, A., Dziri, N., Lyu, S., Gu, Y., Malik, S., Graf, V., Hwang, J. D., Yang, J., Bras, R. L., Tafjord, O., Wilhelm, C., Soldaini, L., Smith, N. A., Wang, Y., Dasigi, P., and Hajishi...
2024 arXiv
-
[37]
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L. J. V., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N. A., and Hajishirzi, H. R eward B ench: Evaluating reward models for language modeling. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Findi...
2025
-
[38]
Y., Zeng, L., Liu, J., Yan, R., He, J., Wang, C., Yan, S., Liu, Y., and Zhou, Y
Liu, C. Y., Zeng, L., Liu, J., Yan, R., He, J., Wang, C., Yan, S., Liu, Y., and Zhou, Y. Skywork- Reward : Bag of Tricks for Reward Modeling in LLMs , October 2024 a . URL http://arxiv.org/abs/2410.18451. arXiv:2410.18451 [cs]
2024 arXiv
-
[39]
Z., Liu, Y., Piot, B., Ittycheriah, A., Kumar, A., and Saleh, M
Liu, T., Xiong, W., Ren, J., Chen, L., Wu, J., Joshi, R., Gao, Y., Shen, J., Qin, Z., Yu, T., Sohn, D., Makarova, A., Liu, J. Z., Liu, Y., Piot, B., Ittycheriah, A., Kumar, A., and Saleh, M. RRM : Robust reward model training mitigates reward hacking. In The Thirteenth Interna...
2025
-
[40]
RM -bench: Benchmarking reward models of language models with subtlety and style
Liu, Y., Yao, Z., Min, R., Cao, Y., Hou, L., and Li, J. RM -bench: Benchmarking reward models of language models with subtlety and style. In The Thirteenth International Conference on Learning Representations, 2025 b . URL https://openreview.net/forum?id=QEHrmQPBdd
2025
-
[41]
Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer
Liu, Z., Lu, M., Zhang, S., Liu, B., Guo, H., Yang, Y., Blanchet, J., and Wang, Z. Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024 b . U...
2024
-
[42]
Sim PO : Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D. Sim PO : Simple preference optimization with a reference-free reward. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=3Tzcot1LKb
2024
-
[43]
K., Strouse, D., Sandholm, T., Salakhutdinov, R., Dragan, A., and McAleer, S
Moskovitz, T., Singh, A. K., Strouse, D., Sandholm, T., Salakhutdinov, R., Dragan, A., and McAleer, S. M. Confronting reward model overoptimization with constrained RLHF . In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/for...
2024
-
[44]
Faster, more efficient RLHF through off-policy asynchronous learning
Noukhovitch, M., Huang, S., Xhonneux, S., Hosseini, A., Agarwal, R., and Courville, A. Faster, more efficient RLHF through off-policy asynchronous learning. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=FhTAG591Ve
2025
-
[45]
W., Liu, J., Malik, S., Merrill, W., Miranda, L
OLMo, T., Walsh, P., Soldaini, L., Groeneveld, D., Lo, K., Arora, S., Bhagia, A., Gu, Y., Huang, S., Jordan, M., Lambert, N., Schwenk, D., Tafjord, O., Anderson, T., Atkinson, D., Brahman, F., Clark, C., Dasigi, P., Dziri, N., Guerquin, M., Ivison, H., Koh, P. W., Liu, J., Mal...
2025 arXiv
-
[46]
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapi...
2024 arXiv
-
[47]
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 0 27730--27744, 2022
2022
-
[48]
Y., Padmakumar, V., Sellam, T., Parikh, A., and He, H
Pang, R. Y., Padmakumar, V., Sellam, T., Parikh, A., and He, H. Reward gaming in conditional text generation. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)...
2023 doi
-
[49]
D., Ermon, S., and Finn, C
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=HPuSIXJaa9
2023
-
[50]
B., Finn, C., and Niekum, S
Rafailov, R., Chittepu, Y., Park, R., Sikchi, H., Hejna, J., Knox, W. B., Finn, C., and Niekum, S. Scaling laws for reward model overoptimization in direct alignment algorithms. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://...
2024
-
[51]
Zero: memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. Zero: memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC '20. IEEE Press, 2020. ISBN 9781728199986
2020
-
[52]
and Vetterli, M
Roy, O. and Vetterli, M. The effective rank: A measure of effective dimensionality. In 2007 15th European Signal Processing Conference, pp.\ 606--610, 2007
2007
-
[53]
Verbosity bias in preference labeling by large language models
Saito, K., Wachi, A., Wataoka, K., and Akimoto, Y. Verbosity bias in preference labeling by large language models. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following, 2023. URL https://openreview.net/forum?id=magEgFpK1y
2023
-
[54]
Skalse, J. M. V., Howe, N. H. R., Krasheninnikov, D., and Krueger, D. Defining and characterizing reward gaming. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=yb3HOXO3lX2
2022
-
[55]
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. Advances in neural information processing systems, 33: 0 3008--3021, 2020
2020
-
[56]
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023
2023
-
[57]
G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., Ferret, J., Liu, P., Tafti, P., Friesen, A., Casbon, M., Ramos, S., Kumar, R., Lan, C
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., Ferret, J., Liu, P., Tafti, P., Friesen, A., Casbon, M., Ramos, S., Kumar, R., Lan, C. L., Jerome, S., Tsitsulin, A., Vieillard, N., Stanczyk, P., ...
2024 arXiv
-
[58]
Team, M. N. Introducing mpt-30b: Raising the bar for open-source foundation models. Blog, 2023. www.mosaicml.com/blog/mpt-30b
2023
-
[59]
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N....
2023
-
[60]
Creating a coding assistant with starcoder
Tunstall, L., Lambert, N., Rajani, N., Beeching, E., Le Scao, T., von Werra, L., Han, S., Schmid, P., and Rush, A. Creating a coding assistant with starcoder. Hugging Face Blog, 2023. https://huggingface.co/blog/starchat
2023
-
[61]
E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., Werra, L
Tunstall, L., Beeching, E. E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., Werra, L. V., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A. M., and Wolf, T. Zephyr: Direct distillation of LM alignment. In First Conference on Language Modeling, 2024...
2024
-
[62]
Trl: Transformer reinforcement learning
von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., Huang, S., Rasul, K., and Gallouédec, Q. Trl: Transformer reinforcement learning. https://github.com/huggingface/trl, 2020
2020
-
[63]
Secrets of rlhf in large language models part ii: Reward modeling, 2024 a
Wang, B., Zheng, R., Chen, L., Liu, Y., Dou, S., Huang, C., Shen, W., Jin, S., Zhou, E., Shi, C., Gao, S., Xu, N., Zhou, Y., Fan, X., Xi, Z., Zhao, J., Wang, X., Ji, T., Yan, H., Shen, L., Chen, Z., Gui, T., Zhang, Q., Qiu, X., Huang, X., Wu, Z., and Jiang, Y.-G. Secrets of rl...
2024 arXiv
-
[64]
Interpretable preferences via multi-objective reward modeling and mixture-of-experts
Wang, H., Xiong, W., Xie, T., Zhao, H., and Zhang, T. Interpretable preferences via multi-objective reward modeling and mixture-of-experts. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2024, pp.\ 10582-...
2024 doi
-
[65]
Mitigating neural network overconfidence with logit normalization
Wei, H., Xie, R., Cheng, H., Feng, L., An, B., and Li, Y. Mitigating neural network overconfidence with logit normalization. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), Proceedings of the 39th International Conference on Machine Lea...
2022
-
[66]
Diff-erank: A novel rank-based metric for evaluating large language models
Wei, L., Tan, Z., Li, C., Wang, J., and Huang, W. Diff-erank: A novel rank-based metric for evaluating large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=nvn80cscVm
2024
-
[67]
Reuse your rewards: Reward model transfer for zero-shot cross-lingual alignment
Wu, Z., Balashankar, A., Kim, Y., Eisenstein, J., and Beirami, A. Reuse your rewards: Reward model transfer for zero-shot cross-lingual alignment. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language...
2024 doi
-
[68]
Wizard LM : Empowering large pre-trained language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., Lin, Q., and Jiang, D. Wizard LM : Empowering large pre-trained language models to follow complex instructions. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview....
2024
-
[69]
Qwen2 technical report, 2024 a
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., Yang, J., Xu, J., Zhou, J., Bai, J., He, J., Lin, J., Dang, K., Lu, K., Chen, K., Yang, K., Li, M....
2024 arXiv
-
[70]
Regularizing hidden states enables learning generalizable reward model for LLM s
Yang, R., Ding, R., Lin, Y., Zhang, H., and Zhang, T. Regularizing hidden states enables learning generalizable reward model for LLM s. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024 b . URL https://openreview.net/forum?id=jwh9MHEfmY
2024
-
[71]
Advancing LLM reasoning generalists with preference trees
Yuan, L., Cui, G., Wang, H., Ding, N., Wang, X., Deng, J., Shan, B., Chen, H., Xie, R., Lin, Y., Liu, Z., Zhou, B., Peng, H., Liu, Z., and Sun, M. Advancing LLM reasoning generalists with preference trees. In AI for Math Workshop @ ICML 2024, 2024. URL https://openreview.net/f...
2024
-
[72]
Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles
Zhai, Y., Zhang, H., Lei, Y., Yu, Y., Xu, K., Feng, D., Ding, B., and Wang, H. Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles. arXiv preprint arXiv:2401.00243, 2023
2023 arXiv
-
[73]
Mitigating reward overoptimization via lightweight uncertainty estimation
Zhang, X., Ton, J.-F., Shen, W., Wang, H., and Liu, Y. Mitigating reward overoptimization via lightweight uncertainty estimation. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024 a . URL https://openreview.net/forum?id=kYio3xH6eb
2024
-
[74]
From lists to emojis: How format bias affects model alignment, 2024 b
Zhang, X., Xiong, W., Chen, L., Zhou, T., Huang, H., and Zhang, T. From lists to emojis: How format bias affects model alignment, 2024 b . URL https://arxiv.org/abs/2409.11704
2024 arXiv
-
[75]
Pytorch fsdp: Experiences on scaling fully sharded data parallel
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., and Li, S. Pytorch fsdp: Experiences on scaling fully sharded data parallel....
2023
-
[76]
E., and Stoica, I
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging LLM -as-a-judge with MT -bench and chatbot arena. In Thirty-seventh Conference on Neural Information Processing Systems Da...
2023
-
[77]
RMB : Comprehensively benchmarking reward models in LLM alignment
Zhou, E., Zheng, G., Wang, B., Xi, Z., Dou, S., Bao, R., Shen, W., Xiong, L., Fan, J., Mou, Y., Zheng, R., Gui, T., Zhang, Q., and Huang, X. RMB : Comprehensively benchmarking reward models in LLM alignment. In The Thirteenth International Conference on Learning Representation...
2025
-
[78]
M., Stiennon, N., Wu, J., Brown, T
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences, 2020
2020
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.