REVIEW 3 major objections 5 minor 29 references
The Superalignment of Superhuman Intelligence with Large Language Models
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This position paper argues that superalignment can be achieved by an automatic loop in which an attacker exposes a model's weaknesses, a critic supplies scalable feedback, and the learner improves itself with minimal human oversight.
desk verdict A clear, honest position paper that defines superalignment as scalable learning from noisy labels and proposes an attacker-learner-critic loop; the framework is plausible but the conclusion slightly overstates how much existing self-improvement work verifies it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the attacker-learner-critic loop: an attacker writes adversarial queries that make the learner fail, a critic explains why the learner's responses are good or bad in natural language, and the learner updates itself from those textual critiques, with human experts occasionally intervening. The paper argues that textual critiques are especially promising because they do not require training a separate reward model, and it identifies weak-to-strong generalization, scalable oversight, and automated evaluation as the load-bearing capabilities that must hold for the loop to work.
What would settle it
A controlled experiment in which a strong learner is trained exclusively on critiques generated by a weaker critic on tasks too hard for humans would falsify the framework if the learner fails to improve or degrades over multiple rounds.
Extended reading notes
Core claim
The central claim is that superalignment reduces to a scalable learning loop rather than a static alignment procedure. The paper proposes that a learner model can be continuously improved by an attacker that discovers its weaknesses and a critic that generates textual critiques of the learner's responses; the learner trains on these critiques together with minimal human feedback, and the process repeats automatically. The paper further claims that weak-to-strong generalization, scalable oversight, and automated evaluation are the key research problems underlying this loop, and that existing self-improvement methods can be seen as early attempts toward superalignment that show positive results. This is a conceptual framework, not an experimental demonstration, and the paper leaves open the detailed implementation of each module.
Load-bearing premise
The framework assumes the critic can produce feedback that is genuinely faithful and useful for the learner, even though the critic is fallible and may be weaker than the learner; if the critiques are noisy or unlearnable, the loop cannot keep the model aligned.
Editorial extensions
If this is right
- If the loop works, superhuman models can be aligned through iterative self-improvement rather than through human annotation of every difficult task.
- Weak-to-strong generalization would let a weaker supervisor elicit and strengthen capabilities of a stronger model, reversing the usual teacher-student direction.
- Textual critiques could replace or augment reward models, reducing the need for explicit reward-model training in alignment.
- The loop's adversarial attack module would make evaluation dynamic and ongoing, catching weaknesses that static benchmarks miss.
- The existence of self-alignment, self-play, and self-refinement results suggests the pipeline is at least partially realizable today.
Reading between the lines
- Editorial inference: the framework implies that the critic's faithfulness, not the learner's capacity, is the main bottleneck, so improving critique quality should directly improve alignment outcomes.
- Editorial inference: the loop's long-run stability is not established; the model-collapse results the paper cites indicate that iterative training on generated data can degrade diversity and out-of-distribution performance unless fresh data or oversight enters.
- Editorial inference: a direct test would be to run the loop on olympiad-level mathematics or competitive programming, where human labels are unreliable, and measure whether the learner's performance increases across rounds.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that as LLMs approach or exceed human performance on complex tasks, conventional human-feedback alignment becomes unscalable, and proposes a learning-theoretic definition of superalignment: designing alignment algorithms that learn from noisy labels (pointwise or pairwise) when tasks are too complex for human annotation and the model is stronger than human experts. It identifies three key research problems—weak-to-strong generalization, scalable oversight, and evaluation—and then presents a conceptual framework consisting of an attacker, a learner, and a critic. The attacker generates adversarial queries to expose learner weaknesses, the learner improves from scalable critic feedback with minimal human intervention, and the critic produces textual critiques. The paper surveys related work on self-alignment, self-play, and self-refinement, and concludes that these existing works 'partially verif[y] the feasibility' of the framework.
Significance. As a research agenda, the paper is a useful synthesis: it gives a clear framing of superalignment as a noisy-label learning problem, identifies concrete open problems (e.g., critic faithfulness, learnable feedback forms, automatic weakness discovery), and honestly acknowledges negative results such as model collapse and self-preference. The proposed attacker-learner-critic loop is an organizing framework rather than an empirical claim, and for a position paper that is acceptable provided the claims are carefully calibrated. The paper's main weakness is that its conclusion overstates the degree to which existing work verifies feasibility; the cited evidence is mixed, and the central assumption about critic quality is left explicitly open. The paper contains no experiments or proofs, but this is not a defect for a position paper if the framing and open problems are the contribution.
major comments (3)
- [Section 4.4 and Section 5] The conclusion that existing self-alignment, self-play, and self-refinement works 'partially verif[y] the feasibility of the framework' is not supported by the evidence cited in the paper itself. Section 4.4 cites model collapse [60] and declines in output diversity and OOD generalization in iterative self-refinement loops [61–64], and reference [49] shows that LLM evaluators favor their own generations. These results indicate that the loop can amplify noise or reduce diversity under conditions the paper does not specify. The feasibility claim should be softened to 'these works provide partial evidence in restricted settings where automatic verification is available' or accompanied by explicit conditions under which the attacker-learner-critic loop is expected to improve rather than degrade.
- [Sections 4.2–4.3] The framework's central load-bearing assumption is that the critic can provide feedback that is simultaneously faithful, informative, and learnable by the learner. The paper itself lists this as an open problem: it asks 'how can we ensure and evaluate the faithfulness of a generated critic?' (Section 4.3) and 'how can the learner learn from such textual critics?' (Section 4.2). Because the attacker, learner, and critic can be versions of the same foundation model, the loop is informationally closed; if the critic is noisy or biased, the learner may amplify that bias rather than receive an independent alignment signal. The manuscript should either provide a minimal formal condition (e.g., a bound on critic error or a diversity-preservation mechanism) or explicitly reframe feasibility as an empirical hypothesis to be tested, rather than a claim partially verified by prior work.
- [Section 4.1] The attacker is trained using a reward signal derived from the critic, but the paper does not discuss how to distinguish weaknesses of the learner from artifacts or biases of the critic. If the critic systematically down-ranks certain legitimate responses, the attacker will learn to exploit those critic errors, and the loop will optimize against a flawed reward rather than against genuine alignment failures. The paper should discuss this potential circularity, for example by proposing the use of multiple critics, human spot-checking of adversarial queries, or validation of attacker-generated queries on independent evaluators.
minor comments (5)
- [Section 2.2] The word 'alignning' should be corrected to 'aligning' in the sentence describing PPO.
- [Section 3.1] In the sentence beginning 'however, it becomes much more complex in The setting of LLMs', 'The' should be lowercase.
- [Section 3.3] The phrase 'today's evaluation is heavily reliable on static evaluation' should be 'heavily relies on static evaluation'.
- [References] References [29] and [35] are the same CriticGPT paper (McAleese et al., 2024) and should be merged or cross-referenced to avoid duplication.
- [Section 5] The phrase 'that is worthy to study in near future' should read 'that is worthy of study in the near future'.
Circularity Check
No significant circularity: the paper's conceptual framework is not derived from its own outputs, and the cited prior systems are independent empirical evidence.
full rationale
This is a position paper that stipulates a definition of superalignment and presents a conceptual attacker-learner-critic framework with explicitly open research problems, rather than a derivation chain that reduces to its own inputs. No equation is reused as an output, no fitted parameter is renamed as a prediction, and no uniqueness theorem is imported from the authors' prior work. The conclusion's claim that existing self-alignment, self-play, and self-refinement works 'partially verif[y] the feasibility' (Section 5) is an inductive appeal to independently benchmarked systems, including STaR, SPIN, Self-Refine, and the authors' own SPAR, AutoDetect, and CritiqueLLM. Although several of these are authored by the present group, each has its own published results and benchmarks, so they are not constructed artifacts of the present paper's definition. Moreover, the paper candidly cites model collapse and diversity decline in self-consuming training loops (Section 4.4), which shows that the feasibility claim is not being forced by suppressing counterevidence. The open questions about critic faithfulness, learnability, and mixed human-AI feedback are acknowledged rather than assumed away. Overall, the central content is a research agenda with external empirical support, not a circular derivation.
Assumptions & free parameters
assumptions (3)
- domain assumption Superhuman AI systems will exist and will exceed human experts on many tasks, making human annotations noisy.
- domain assumption AI critic models can provide feedback that is scalable, faithful, and usable by a learner to improve.
- domain assumption Iterated self-refinement with model-generated data does not lead to model collapse or divergence.
Cite this review
Pith. "Pith review of The Superalignment of Superhuman Intelligence with Large Language Models." pith.science (2026). https://pith.science/paper/IXG6AV5J
@misc{pith2026241211145,
author = {Pith},
title = {Pith review of: The Superalignment of Superhuman Intelligence with Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/IXG6AV5J}},
note = {Machine review of arXiv:2412.11145}
}
read the original abstract
We have witnessed superhuman intelligence thanks to the fast development of large language models and multimodal language models. As the application of such superhuman models becomes more and more popular, a critical question arises here: how can we ensure superhuman models are still safe, reliable and aligned well to human values? In this position paper, we discuss the concept of superalignment from the learning perspective to answer this question by outlining the learning paradigm shift from large-scale pretraining, supervised fine-tuning, to alignment training. We define superalignment as designing effective and efficient alignment algorithms to learn from noisy-labeled data (point-wise samples or pair-wise preference data) in a scalable way when the task becomes very complex for human experts to annotate and the model is stronger than human experts. We highlight some key research problems in superalignment, namely, weak-to-strong generalization, scalable oversight, and evaluation. We then present a conceptual framework for superalignment, which consists of three modules: an attacker which generates adversary queries trying to expose the weaknesses of a learner model; a learner which will refine itself by learning from scalable feedbacks generated by a critic model along with minimal human experts; and a critic which generates critics or explanations for a given query-response pair, with a target of improving the learner by criticizing. We discuss some important research problems in each component of this framework and highlight some interesting research ideas that are closely related to our proposed framework, for instance, self-alignment, self-play, self-refinement, and more. Last, we highlight some future research directions for superalignment, including identification of new emergent risks and multi-dimensional alignment.
Reference graph
Works this paper leans on
-
[3]
Large language model alignment: A survey
9 Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. Large language model alignment: A survey. arXiv preprint arXiv:2309.15025 ,
-
[5]
Neural machine translation by jointly learning to align and translate
14 Yoshua Bengio Dzmitry Bahdanau, Kyunghyun Cho. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations, ICLR 2015 . 15 Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason E Weston. Self-rewarding language models. In Forty-first Int...
arXiv 2015
-
[6]
Kto: Model alignment as prospect theoretic optimization
18 Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. Kto: Model alignment as prospect theoretic optimization. arXiv preprint arXiv:2402.01306 ,
-
[7]
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
20 Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. In Forty-first International Conference on Machine Learning . 21 Samuel R Bowman, Jeeyoon Hyun, Ethan Perez, E...
-
[8]
Supervising strong learners by amplifying weak experts
22 Paul Christiano, Buck Shlegeris, and Dario Amodei. Supervising strong learners by amplifying weak experts. arXiv preprint arXiv:1810.08575,
-
[9]
24 Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg
Association for Computational Linguistics. 24 Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871 ,
-
[10]
Constitutional ai: Harmlessness from ai feedback
26 Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073,
-
[11]
Improving factuality and reasoning in language models through multiagent debate
27 Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. Improving factuality and reasoning in language models through multiagent debate. In Forty-first International Conference on Machine Learning . 28 Geoffrey Irving, Paul Christiano, and Dario Amodei. Ai safety via debate. arXiv preprint arXiv:1805.00899 ,
Show all 29 references
-
[13]
Training verifiers to solve math word problems
30 Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168,
-
[15]
35 Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike
Association for Computational Linguistics. 35 Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike. Llm critics help catch llm bugs. arXiv preprint arXiv:2407.00215 ,
-
[16]
Generative verifiers: Reward modeling as next-token prediction
36 Lunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi, Aviral Kumar, and Rishabh Agarwal. Generative verifiers: Reward modeling as next-token prediction. arXiv preprint arXiv:2408.15240 ,
-
[17]
AutoDetect: Towards a unified framework for automated weakness detection in large language models
37 Jiale Cheng, Yida Lu, Xiaotao Gu, Pei Ke, Xiao Liu, Yuxiao Dong, Hongning Wang, Jie Tang, and Minlie Huang. AutoDetect: Towards a unified framework for automated weakness detection in large language models. In Findings of the Association for Computational Linguistics: EMNLP...
2024
-
[18]
Language models learn to mislead humans via rlhf
39 Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R Bowman, He He, and Shi Feng. Language models learn to mislead humans via rlhf. arXiv preprint arXiv:2409.12822 ,
-
[19]
Universal and transferable adversarial attacks on aligned language models
41 Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043 ,
-
[20]
Unveiling the implicit toxicity in large language models
44 Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, and Minlie Huang. Unveiling the implicit toxicity in large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages 1322–1338, Singapore, December
2023
-
[21]
45 Zihuiwen Ye, Fraser Greenlee-Scott, Max Bartolo, Phil Blunsom, Jon Ander Campos, and Matthias Gall´ e
Association for Computational Linguistics. 45 Zihuiwen Ye, Fraser Greenlee-Scott, Max Bartolo, Phil Blunsom, Jon Ander Campos, and Matthias Gall´ e. Improving reward models with synthetic critiques. arXiv preprint arXiv:2405.20850 ,
-
[22]
Reinforcement learning for generative ai: A survey
46 Yuanjiang Cao, Quan Z Sheng, Julian McAuley, and Lina Yao. Reinforcement learning for generative ai: A survey. arXiv preprint arXiv:2308.14328,
-
[23]
Learning to refine with fine-grained natural language feedback
47 Manya Wadhwa, Xinyu Zhao, Junyi Jessy Li, and Greg Durrett. Learning to refine with fine-grained natural language feedback. In Findings of the Association for Computational Linguistics: EMNLP 2024 , pages 12281–12308, Miami, Florida, USA, November
2024
-
[24]
48 Yiqing Xie, Wenxuan Zhou, Pradyot Prakash, Di Jin, Yuning Mao, Quintin Fettes, Arya Talebzadeh, Sinong Wang, Han Fang, Carolyn Rose, et al
Association for Computational Linguistics. 48 Yiqing Xie, Wenxuan Zhou, Pradyot Prakash, Di Jin, Yuning Mao, Quintin Fettes, Arya Talebzadeh, Sinong Wang, Han Fang, Carolyn Rose, et al. Improving model factuality with fine-grained critique-based evaluator. arXiv preprint arXiv...
-
[25]
Bayesian calibration of win rate estimation with LLM evaluators
50 Yicheng Gao, Gonghan Xu, Zhe Wang, and Arman Cohan. Bayesian calibration of win rate estimation with LLM evaluators. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages...
2024
-
[26]
51 Jaehun Jung, Faeze Brahman, and Yejin Choi
Association for Computational Linguistics. 51 Jaehun Jung, Faeze Brahman, and Yejin Choi. Trust or escalate: Llm judges with provable guarantees for human agreement. arXiv preprint arXiv:2407.18370 ,
-
[27]
Generative judge for evaluating alignment
52 Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Pengfei Liu, et al. Generative judge for evaluating alignment. In The Twelfth International Conference on Learning Representations . 53 Lianghui Zhu, Xinggang Wang, and Xinlong Wang. Judgelm: Fine-tuned large language models...
-
[28]
Self-critiquing models for assisting human evaluators
54 William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. Self-critiquing models for assisting human evaluators. arXiv preprint arXiv:2206.05802 ,
-
[29]
Lm vs lm: Detecting factual errors via cross examination
59 Roi Cohen, May Hamri, Mor Geva, and Amir Globerson. Lm vs lm: Detecting factual errors via cross examination. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages 12621–12640,
2023
-
[30]
Panacea: Pareto alignment via preference adaptation for llms
66 Yifan Zhong, Chengdong Ma, Xiaoyuan Zhang, Ziran Yang, Haojun Chen, Qingfu Zhang, Siyuan Qi, and Yaodong Yang. Panacea: Pareto alignment via preference adaptation for llms. arXiv preprint arXiv:2402.02030 , 2024
2024 arXiv
-
[2021]
Evaluating large language models trained on code
32 Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374,
-
[2022]
A survey of large language models
2 Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223 ,
-
[2023]
Proximal policy optimization algorithms
10 John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 ,
-
[2024]
Model evaluation for extreme risks
7 Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, et al. Model evaluation for extreme risks. arXiv preprint arXiv:2305.15324,
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.