REVIEW 3 major objections 5 minor 1 cited by
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper proposes an AI-45° Law requiring capability and safety to advance at the same rate, and a Causal Ladder of Trustworthy AGI that organizes safety research into three layers, to guide a balanced road to AGI.
desk verdict Useful taxonomy, underdefined 45° law — worth peer review if the law is reframed as a normative heuristic. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The two load-bearing objects are the 45° line and the Causal Ladder. The 45° line is the geometric statement of the AI-45° Law: in a plane whose axes are AI capability and AI safety, ideal progress keeps the two in lockstep, so the trajectory stays on a line of slope one; the region below it represents safety lag and contains 'red line' existential risks, while the 'yellow line' marks early-warning thresholds. The Causal Ladder of Trustworthy AGI translates the Ladder of Causation's three rungs—association, intervention, counterfactual—into three layers of safety research: Approximate Alignment (fitting values from data, e.g., supervised fine-tuning and machine unlearning), Intervenable (verifiable and steerable inference, e.g., RLHF, mechanistic interpretability, scalable oversight), and Reflectable (self-reflection, world models, counterfactual interpretability). This correspondence is the mechanism that lets the paper classify a wide range of existing techniques into a single hierarchical structure.
What would settle it
Look for a concrete capability–safety coordinate system: if no consistent way exists to assign capability and safety scores to the same AI system such that the 45° line separates safe from unsafe trajectories, the geometric content of the law evaporates. A decisive test would be showing that safety is not a single scalar quantity—e.g., a system ranked both safer and less safe than another depending on which safety dimension is chosen—which would make the 45° slope undefined.
Extended reading notes
Core claim
The paper's central claim is that a trustworthy path to AGI requires capability and safety to be developed in parallel at comparable rates, and that the AI-45° Law makes this requirement explicit by placing ideal progress on a 45° line in a capability–safety coordinate system. The current trajectory, described as 'crippled AI', deviates far below that line; the space far below it is the region of existential 'red line' risks, while a 'yellow line' marks thresholds that would trigger stricter assurance before systems become dangerous. To make the law actionable, the paper introduces the Causal Ladder of Trustworthy AGI, a three-layer taxonomy—Approximate Alignment, Intervenable, and Reflectable—that maps existing safety techniques onto ascending levels of causal understanding, and a five-level Matrix of Trustworthy AGI that grades systems from perception trustworthiness up to collaboration trustworthiness, with dependence on the Reflectable Layer increasing at higher levels. The paper is a position piece: it is proposing a framework and a vocabulary for balanced AGI development, not reporting experimental results.
Load-bearing premise
The framework assumes AI capability and safety can be measured on two comparable axes, so that 'equal rates' along a 45° line is meaningful; the paper gives no units, metrics, or conversion for either dimension.
Editorial extensions
If this is right
- The 45° law gives a shared criterion for judging development roadmaps: any plan that lets capability outpace safety is, by definition, headed toward the red-line region.
- Red lines and yellow lines become concrete planning tools: a system below the yellow line needs only basic testing, while systems above it require substantially stronger assurance mechanisms.
- The Causal Ladder provides a common language for comparing safety techniques, so methods as different as machine unlearning, RLHF, and world models can be positioned relative to one another and to the trustworthiness level they support.
- The five-level matrix offers a staged target for AGI development, with each level (perception, reasoning, decision-making, autonomy, collaboration) building on the previous and leaning more heavily on the Reflectable Layer.
- The framework can guide governance: treating AI safety as a global public good and managing the full lifecycle become parts of the same balanced roadmap.
Reading between the lines
- The 45° framing invites a reformulation that could be tested: if defensible metrics for capability and safety existed, the law would predict that systems with equal capability but different safety scores would differ in accident and misuse rates; measuring that gradient would support or collapse the single-slope picture.
- The mapping to the Ladder of Causation suggests a research program: just as causal inference progressed from correlation to intervention to counterfactual reasoning, safety assurance could be graded by which causal rung its methods occupy, with reflectable methods treated as strictly harder than intervenable ones.
- The paper leaves open who sets the yellow-line thresholds, so a natural extension is a governance mechanism where threshold calibration is itself constrained by the red-line logic, preventing a race to lower standards.
- The framework's generality implies it could be checked on current models by attempting to place existing large language models in the five-level matrix and verifying whether the predicted dependence on reflection techniques matches observed safety failures.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes the AI-45° Law, which states that AI capability and safety should progress at the same rate, represented by a 45° line in a capability-safety coordinate system. It uses this law to define red lines for existential risks and yellow lines for early-warning thresholds. The paper also introduces the Causal Ladder of Trustworthy AGI, a three-layer taxonomy (Approximate Alignment, Intervenable, Reflectable) inspired by Pearl's Ladder of Causation, and defines five levels of trustworthy AGI (Perception, Reasoning, Decision-making, Autonomy, Collaboration). It closes with a set of governance measures. The manuscript is a position paper and explicitly defers empirical validation to future work (Section 6).
Significance. If the Causal Ladder taxonomy proves usable, it could serve as a useful organizing structure for AI safety research, grouping diverse existing methods into a coherent hierarchy and linking them to Pearl's causality levels. The paper is transparent about its limitations and avoids overclaiming empirical support. However, the central AI-45° Law as stated is not testable because capability and safety are not operationalized on a common scale; as a result, the claimed diagnostic (crippled AI) and prescriptive target (45° trajectory) are underdetermined. The taxonomic framework is the more defensible contribution, while the law currently functions as a normative heuristic rather than a scientific claim.
major comments (3)
- [§2.1, Figure 1] The AI-45° Law as stated is geometrically underdetermined because the two axes, capability and safety, have no defined units or metrics, and no argument is given that the notion of 'the same rate' is meaningful across incommensurable dimensions; under independent monotone rescaling of either axis, the 45° slope has no invariant meaning, so the red and yellow line regions in §§2.2–2.3 are not well-defined. The paper should either provide operational definitions of capability and safety progress (for example, specific benchmarks and safety evaluations with a defined conversion) or explicitly reframe the law as a normative visual heuristic rather than a descriptive scientific claim.
- [§4, §6] The paper defers all empirical validation to future work and offers no testable predictions or classification criteria; in particular, the placement of models in the Matrix of Trustworthy AGI (for example, OpenAI-o1 as a 'basic form of the Reflectable Layer') is asserted without a rubric, so the claimed practicality of the Causal Ladder cannot be assessed. The authors should specify measurable criteria for assigning methods and models to layers and levels, or explicitly state that the framework is a qualitative taxonomy rather than an empirically validated roadmap.
- [§3.1–§3.3] The correspondence between the three layers and Pearl's three levels (association, intervention, counterfactuals) is analogical, but the membership of techniques in layers appears ambiguous; for instance, RLHF is placed in the Intervenable Layer while supervised fine-tuning, which also shapes model behavior, is in the Approximate Alignment Layer, and the distinction is not formally defined. The paper should define the distinguishing criterion (for example, whether the method modifies the inference process or requires external intervention during deployment) to make the hierarchy reproducible.
minor comments (5)
- [§1.2] The phrase 'reactive approach' approach' contains a duplicated word and should read 'reactive approach'.
- [§3.3] The sentence 'it often falls to ensure safety and reliability' appears to contain a typo and should be 'it often fails to ensure safety and reliability'.
- [References] Reference [48] is missing its authors and full title, and reference [79] contains a broken DOI; both should be corrected before publication.
- [§4] The sentence 'At their core, they still focused on the Perception Trustworthiness level' is ungrammatical; 'focused' should be 'focus' or the sentence should be recast.
- [§5] The governance bullet list is not explicitly connected to the Causal Ladder or the 45° Law; adding a sentence linking each governance measure to the framework would strengthen the roadmap.
Circularity Check
No circularity: the AI-45° law is a normative stipulation rather than a derived prediction, and the Causal Ladder is an acknowledged analogy to Pearl's external Ladder of Causation.
full rationale
This is a position paper with no numerical fitting or derivation chain. The central claim, the AI-45° law, is introduced explicitly as a 'guiding principle' ('we introduce the AI-45° law as a guiding principle'), and the equal-rate 45° line is stipulated rather than derived from data; there is therefore no fitted input renamed as a prediction and no self-definitional equation. The Causal Ladder is transparently an organizational analogy to an external source ('inspired by Judea Pearl’s “Ladder of Causation”'), with each layer said to 'correspond' to a Pearl level; this is an imported taxonomy, not a renamed known result presented as a new empirical finding. The red and yellow lines are graphical regions defined relative to the 45° line ('can be more clearly illustrated as occupying the lower right region relative to the 45° line'), not empirical consequences. Many references are to the authors' own prior surveys and benchmarks, but none is load-bearing: no uniqueness theorem from the authors is invoked, no ansatz is smuggled in via citation, and the five trustworthiness levels are explicitly 'prospectively defined' rather than derived. The paper's own conclusion ('Future work will empirically validate the proposed framework') confirms that no hidden empirical fit is being reported. The lack of operational units for capability and safety is a testability and underdetermination concern, not a circularity concern.
Assumptions & free parameters
free parameters (1)
- 45 degree slope of the balance line =
1:1 (chosen, not fitted)
assumptions (4)
- domain assumption AI capability and safety can be measured and compared on a common coordinate system.
- domain assumption The three levels of Pearl's Ladder of Causation map meaningfully onto the three proposed trustworthiness layers.
- domain assumption Existential risks from advanced AI are real and the IDAIS red lines are a valid basis for limits.
- domain assumption Future AGI systems will be developed and can be guided by the proposed roadmap.
invented entities (3)
-
AI-45 degree law
-
Causal Ladder of Trustworthy AGI
-
Crippled AI
Cite this review
Pith. "Pith review of Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI." pith.science (2026). https://pith.science/paper/VWE7GHQ6
@misc{pith2026241214186,
author = {Pith},
title = {Pith review of: Towards AI-$45^\circ$ Law: A Roadmap to Trustworthy AGI},
year = {2026},
howpublished = {\url{https://pith.science/paper/VWE7GHQ6}},
note = {Machine review of arXiv:2412.14186}
}
abstract
Ensuring Artificial General Intelligence (AGI) reliably avoids harmful behaviors is a critical challenge, especially for systems with high autonomy or in safety-critical domains. Despite various safety assurance proposals and extreme risk warnings, comprehensive guidelines balancing AI safety and capability remain lacking. In this position paper, we propose the \textit{AI-\textbf{$45^{\circ}$} Law} as a guiding principle for a balanced roadmap toward trustworthy AGI, and introduce the \textit{Causal Ladder of Trustworthy AGI} as a practical framework. This framework provides a systematic taxonomy and hierarchical structure for current AI capability and safety research, inspired by Judea Pearl's ``Ladder of Causation''. The Causal Ladder comprises three core layers: the Approximate Alignment Layer, the Intervenable Layer, and the Reflectable Layer. These layers address the key challenges of safety and trustworthiness in AGI and contemporary AI systems. Building upon this framework, we define five levels of trustworthy AGI: perception, reasoning, decision-making, autonomy, and collaboration trustworthiness. These levels represent distinct yet progressive aspects of trustworthy AGI. Finally, we present a series of potential governance measures to support the development of trustworthy AGI.
Figures
Forward citations
Cited by 1 Pith paper
-
An Early Warning of Emerging Biosecurity Risks in Frontier LLMs
A bio-red-teaming model is reported to jailbreak 14 frontier LLMs into producing dangerous biosecurity outputs, but the claimed wet-lab physical verification was not actually carried out.
Reference graph
Works this paper leans on
-
[1]
https://www.anthropic.com/news/anthropics-responsible-scaling- policy, 2023
Responsible scaling policy. https://www.anthropic.com/news/anthropics-responsible-scaling- policy, 2023. Accessed: 2023-09-19
2023
-
[2]
https://idais.ai/dialogue/idais- beijing/, 2024
Consensus statement on red lines in artificial intelligence. https://idais.ai/dialogue/idais- beijing/, 2024. Accessed: 2024-03-10
2024
-
[3]
https://idais.ai/dialogue/idais-venice/, 2024
The global nature of ai risks makes it necessary to recognize ai safety as a global public good. https://idais.ai/dialogue/idais-venice/, 2024. Accessed: 2024-09-08
2024
-
[4]
https://assets.anthropic.com/m/24a47b00f10301cd/original/Anthropic- Responsible-Scaling-Policy-2024-10-15.pdf, 2024
Responsible scaling program updates. https://assets.anthropic.com/m/24a47b00f10301cd/original/Anthropic- Responsible-Scaling-Policy-2024-10-15.pdf, 2024. Accessed: 2024-10-15
2024
-
[5]
Current state of llm risks and ai guardrails
Suriya Ganesh Ayyamperumal and Limin Ge. Current state of llm risks and ai guardrails. arXiv preprint arXiv:2406.12934, 2024
arXiv 2024
-
[6]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023
arXiv 2023
-
[7]
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023
arXiv 2023
-
[8]
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022
arXiv 2022
Show all 102 references
-
[9]
Managing extreme ai risks amid rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, et al. Managing extreme ai risks amid rapid progress. Science, 384(6698):842–845, 2024
2024
-
[10]
Mechanistic interpretability for ai safety–a review
Leonard Bereska and Efstratios Gavves. Mechanistic interpretability for ai safety–a review. arXiv preprint arXiv:2404.14082, 2024
2024 arXiv
-
[11]
Diverse and effective red teaming with auto-generated rewards and multi-step reinforcement learning
Alex Beutel, Kai Xiao, Johannes Heidecke, and Lilian Weng. Diverse and effective red teaming with auto-generated rewards and multi-step reinforcement learning
-
[12]
Measuring progress on scalable oversight for large language models
Samuel R Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamil˙e Lukoši¯ut˙e, Amanda Askell, Andy Jones, Anna Chen, et al. Measuring progress on scalable oversight for large language models. arXiv preprint arXiv:2211.03540, 2022
2022 arXiv
-
[13]
Language models are few-shot learners
Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020
2005 arXiv
-
[14]
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Peter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar, et al. The malicious use of artificial intelligence: Forecasting, prevention, and mitigation. arXiv preprint arXiv:1802.07228, 2018
2018 arXiv
-
[15]
Weak-to-strong gener- alization: Eliciting strong capabilities with weak supervision.arXiv preprint arXiv:2312.09390, 2023
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschen- brenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al. Weak-to-strong gener- alization: Eliciting strong capabilities with weak supervision.arXiv preprint arXiv:2312.09390, 2023
2023 arXiv
-
[16]
Internlm2 technical report
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al. Internlm2 technical report. arXiv preprint arXiv:2403.17297, 2024. 9
2024 arXiv
-
[17]
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-V oss, Kather- ine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pa...
2021
-
[18]
Is power-seeking ai an existential risk? arXiv preprint arXiv:2206.13353, 2022
Joseph Carlsmith. Is power-seeking ai an existential risk? arXiv preprint arXiv:2206.13353, 2022
2022 arXiv
-
[19]
Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective
Meiqi Chen, Yixin Cao, Yan Zhang, and Chaochao Lu. Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective. Findings of the Association for Computational Linguistics: EMNLP, Miami, Florida, USA, November 12-16, 2024
2024
-
[20]
Cello: Causal evaluation of large vision- language models
Meiqi Chen, Bo Peng, Yan Zhang, and Chaochao Lu. Cello: Causal evaluation of large vision- language models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP, Miami, Florida, USA, November 12-16, 2024
2024
-
[21]
From imitation to introspection: Probing self-consciousness in language models
Sirui Chen, Shu Yu, Shengjie Zhao, and Chaochao Lu. From imitation to introspection: Probing self-consciousness in language models. arXiv preprint arXiv:2410.18819, 2024
2024 arXiv
-
[22]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF Conference on Computer Vis...
2024
-
[23]
Feder Cooper, Christopher A
A. Feder Cooper, Christopher A. Choquette-Choo, Miranda Bogen, Matthew Jagielski, Katja Filippova, Ken Ziyu Liu, Alexandra Chouldechova, Jamie Hayes, Yangsibo Huang, Niloofar Mireshghallah, Ilia Shumailov, Eleni Triantafillou, Peter Kairouz, Nicole Mitchell, Percy Liang, Danie...
2024
-
[24]
Towards guar- anteed safe ai: A framework for ensuring robust and reliable ai systems
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, et al. Towards guar- anteed safe ai: A framework for ensuring robust and reliable ai systems. arXiv preprint arXiv:2405.06624, 2024
2024 arXiv
-
[25]
Arti- ficial intelligence regulation: a framework for governance
Patricia Gomes Rêgo de Almeida, Carlos Denner dos Santos, and Josivania Silva Farias. Arti- ficial intelligence regulation: a framework for governance. Ethics and Information Technology, 23(3):505–525, 2021
2021
-
[26]
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. arXiv preprint arXiv:2002.06305, 2020
2002 arXiv
-
[27]
Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Xilin Wei, Songyang Zhang, Haodong Duan, Maosong Cao, et al. Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model. arXiv preprint arXiv:2401.16420, 2024
2024 arXiv
-
[28]
Attacks, defenses and evaluations for llm conversation safety: A survey
Zhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, and Yu Qiao. Attacks, defenses and evaluations for llm conversation safety: A survey. arXiv preprint arXiv:2402.09283, 2024
2024 arXiv
-
[29]
Clear: Character unlearning in textual and visual modalities
Alexey Dontsov, Dmitrii Korzh, Alexey Zhavoronkin, Boris Mikheev, Denis Bobkov, Aibek Alanov, Oleg Y Rogov, Ivan Oseledets, and Elena Tutubalina. Clear: Character unlearning in textual and visual modalities. arXiv preprint arXiv:2410.18057, 2024
-
[30]
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024
2024 arXiv
-
[31]
Who’s harry potter? approximate unlearning in llms
Ronen Eldan and Mark Russinovich. Who’s harry potter? approximate unlearning in llms. arXiv preprint arXiv:2310.02238, 2023. 10
2023 arXiv
-
[32]
Statement on ai risk, 2024
Center for AI Safety. Statement on ai risk, 2024. Accessed: 2023-05-30
2024
-
[33]
Counterintuitive behavior of social systems
Jay W Forrester. Counterintuitive behavior of social systems. Theory and decision, 2(2):109– 140, 1971
1971
-
[34]
Artificial intelligence, values, and alignment
Iason Gabriel. Artificial intelligence, values, and alignment. Minds and machines, 30(3):411– 437, 2020
2020
-
[35]
Mental models
Dedre Gentner and Albert L Stevens. Mental models. Psychology Press, 2014
2014
-
[36]
Mllmguard: A multi-dimensional safety evaluation suite for multimodal large language models, 2024
Tianle Gu, Zeyang Zhou, Kexin Huang, Dandan Liang, Yixu Wang, Haiquan Zhao, Yuanqi Yao, Xingge Qiao, Keqing Wang, Yujiu Yang, Yan Teng, Yu Qiao, and Yingchun Wang. Mllmguard: A multi-dimensional safety evaluation suite for multimodal large language models, 2024
2024
-
[37]
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber. Recurrent world models facilitate policy evolution. Advances in neural information processing systems, 31, 2018
2018
-
[38]
An overview of catastrophic ai risks
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside. An overview of catastrophic ai risks. arXiv preprint arXiv:2306.12001, 2023
2023 arXiv
-
[39]
Stabilizing translucencies: Governing ai transparency by standardization
Charlotte Högberg. Stabilizing translucencies: Governing ai transparency by standardization. Big Data & Society, 11(1):20539517241234298, 2024
2024
-
[40]
Curiosity-driven red-teaming for large language models
Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang, Aldo Pareja, James Glass, Akash Srivastava, and Pulkit Agrawal. Curiosity-driven red-teaming for large language models. arXiv preprint arXiv:2402.19464, 2024
2024 arXiv
-
[41]
Flames: Benchmarking value alignment of chinese large language models
Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, Xipeng Qiu, et al. Flames: Benchmarking value alignment of chinese large language models. arXiv preprint arXiv:2311.06899, 2023
2023 arXiv
-
[42]
From pixels to principles: A decade of progress and landscape in trustworthy computer vision
Kexin Huang, Yan Teng, Yang Chen, and Yingchun Wang. From pixels to principles: A decade of progress and landscape in trustworthy computer vision. Science and Engineering Ethics, 30(3):26, 2024
2024
-
[43]
Trustllm: Trustworthiness in large language models
Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, et al. Trustllm: Trustworthiness in large language models. arXiv preprint arXiv:2401.05561, 2024
2024 arXiv
-
[44]
When code isn’t law: rethinking regulation for artificial intelligence
Brian Judge, Mark Nitzberg, and Stuart Russell. When code isn’t law: rethinking regulation for artificial intelligence. Policy and Society, page puae020, 2024
2024
-
[45]
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020
2001 arXiv
-
[46]
Aligning large language models with representation editing: A control perspective
Lingkai Kong, Haorui Wang, Wenhao Mu, Yuanqi Du, Yuchen Zhuang, Yifei Zhou, Yue Song, Rongzhi Zhang, Kai Wang, and Chao Zhang. Aligning large language models with representation editing: A control perspective. arXiv preprint arXiv:2406.05954, 2024
2024 arXiv
-
[47]
Evaluations: autonomy and artificial intelligence: a threat or savior? Springer, 2017
William Frere Lawless and Donald A Sofge. Evaluations: autonomy and artificial intelligence: a threat or savior? Springer, 2017
2017
-
[48]
Learning to watermark llm-generated text via rein- forcement learning
VIA REINFORCEMENT LEARNING. Learning to watermark llm-generated text via rein- forcement learning
-
[49]
Deepfakes, phrenology, surveillance, and more! a taxonomy of ai privacy risks
Hao-Ping Lee, Yu-Ju Yang, Thomas Serban V on Davier, Jodi Forlizzi, and Sauvik Das. Deepfakes, phrenology, surveillance, and more! a taxonomy of ai privacy risks. InProceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–19
-
[50]
Trustworthy ai: From principles to practices
Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou. Trustworthy ai: From principles to practices. ACM Computing Surveys, 55(9):1–46, 2023. 11
2023
-
[51]
Inference- time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Inference- time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[52]
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. arXiv preprint arXiv:2402.05044, 2024
2024 arXiv
-
[53]
Controllable text generation for large language models: A survey
Xun Liang, Hanyu Wang, Yezhaohui Wang, Shichao Song, Jiawei Yang, Simin Niu, Jie Hu, Dan Liu, Shunyu Yao, Feiyu Xiong, et al. Controllable text generation for large language models: A survey. arXiv preprint arXiv:2408.12599, 2024
2024 arXiv
-
[54]
A survey of text watermarking in the era of large language models
Aiwei Liu, Leyi Pan, Yijian Lu, Jingjing Li, Xuming Hu, Xi Zhang, Lijie Wen, Irwin King, Hui Xiong, and Philip Yu. A survey of text watermarking in the era of large language models. ACM Computing Surveys, 57(2):1–36, 2024
2024
-
[55]
Don’t always say no to me: Benchmarking safety-related refusal in large vlm
Xin Liu, Zhichen Dong, Zhanhui Zhou, Yichen Zhu, Yunshi Lan, Jing Shao, Chao Yang, and Yu Qiao. Don’t always say no to me: Benchmarking safety-related refusal in large vlm. 2024
2024
-
[56]
Mm-safetybench: A benchmark for safety evaluation of multimodal large language models
Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. InEuropean Conference on Computer Vision, pages 386–403. Springer, 2025
2025
-
[57]
Query-relevant images jailbreak large multi-modal models
Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang, and Yu Qiao. Query-relevant images jailbreak large multi-modal models. arXiv preprint arXiv:2311.17600, 2023
2023 arXiv
-
[58]
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. Trustworthy llms: a survey and guideline for evaluating large language models’ alignment. arXiv preprint arXiv:2308.05374, 2023
2023 arXiv
-
[59]
Machine unlearning in generative ai: A survey
Zheyuan Liu, Guangyao Dou, Zhaoxuan Tan, Yijun Tian, and Meng Jiang. Machine unlearning in generative ai: A survey. arXiv preprint arXiv:2407.20516, 2024
2024 arXiv
-
[60]
Inference-time language model alignment via integrated value guidance
Zhixuan Liu, Zhanhui Zhou, Yuanfu Wang, Chao Yang, and Yu Qiao. Inference-time language model alignment via integrated value guidance. arXiv preprint arXiv:2409.17819, 2024
2024 arXiv
-
[61]
Causal interpretability for machine learning-problems, methods and evaluation
Raha Moraffah, Mansooreh Karami, Ruocheng Guo, Adrienne Raglin, and Huan Liu. Causal interpretability for machine learning-problems, methods and evaluation. ACM SIGKDD Explorations Newsletter, 22(1):18–33, 2020
2020
-
[62]
Large language models in cybersecurity: State-of-the-art
Farzad Nourmohammadzadeh Motlagh, Mehrdad Hajizadeh, Mehryar Majd, Pejman Najafi, Feng Cheng, and Christoph Meinel. Large language models in cybersecurity: State-of-the-art. arXiv preprint arXiv:2402.00891, 2024
2024 arXiv
-
[63]
Rule based rewards for language model safety
Tong Mu, Alec Helyar, Johannes Heidecke, Joshua Achiam, Andrea Vallone, Ian Kivlichan, Molly Lin, Alex Beutel, John Schulman, and Lilian Weng. Rule based rewards for language model safety. arXiv preprint arXiv:2411.01111, 2024
2024 arXiv
-
[64]
Accountability in artificial intelligence: what it is and how it works
Claudio Novelli, Mariarosaria Taddeo, and Luciano Floridi. Accountability in artificial intelligence: what it is and how it works. Ai & Society, 39(4):1871–1882, 2024
2024
-
[65]
Uniguard: Towards universal safety guardrails for jailbreak attacks on multimodal large language models
Sejoon Oh, Yiqiao Jin, Megha Sharma, Donghyun Kim, Eric Ma, Gaurav Verma, and Srijan Kumar. Uniguard: Towards universal safety guardrails for jailbreak attacks on multimodal large language models. arXiv preprint arXiv:2411.01703, 2024
2024 arXiv
-
[66]
GPT-4 technical report
OpenAI. GPT-4 technical report. March 2023
2023
-
[67]
Openai o1 system card, 2024
OpenAI. Openai o1 system card, 2024. Accessed: 2024-09-12
2024
-
[68]
Video generation models as world simulators, 2024
OpenAI. Video generation models as world simulators, 2024. Accessed: 2024-01-15. 12
2024
-
[69]
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:2773...
2022
-
[70]
A ‘biased’emerging governance regime for artificial intelligence? how ai ethics get skewed moving from principles to practices
Nicola Palladino. A ‘biased’emerging governance regime for artificial intelligence? how ai ethics get skewed moving from principles to practices. Telecommunications Policy, 47(5):102479, 2023
2023
-
[71]
Automatically correcting large language models: Surveying the landscape of diverse self-correction strategies
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang. Automatically correcting large language models: Surveying the landscape of diverse self-correction strategies. arXiv preprint arXiv:2308.03188, 2023
2023 arXiv
-
[72]
Automated red teaming with goat: the generative offensive agent tester
Maya Pavlova, Erik Brinkman, Krithika Iyer, Vitor Albiero, Joanna Bitton, Hailey Nguyen, Joe Li, Cristian Canton Ferrer, Ivan Evtimov, and Aaron Grattafiori. Automated red teaming with goat: the generative offensive agent tester. arXiv preprint arXiv:2410.01606, 2024
-
[73]
Causality
Judea Pearl. Causality. Cambridge university press, 2009
2009
-
[74]
The book of why: the new science of cause and effect
Judea Pearl and Dana Mackenzie. The book of why: the new science of cause and effect. Basic books, 2018
2018
-
[75]
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red teaming language models with language models. arXiv preprint arXiv:2202.03286, 2022
2022 arXiv
-
[76]
Dean: Deactivating the coupled neurons to mitigate fairness-privacy conflicts in large language models
Chen Qian, Dongrui Liu, Jie Zhang, Yong Liu, and Jing Shao. Dean: Deactivating the coupled neurons to mitigate fairness-privacy conflicts in large language models. arXiv preprint arXiv:2410.16672, 2024
2024 arXiv
-
[77]
Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models
Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu, and Jing Shao. Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models. arXiv preprint arXiv:2402.19465, 2024
2024 arXiv
-
[78]
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[79]
Natural language processing: transform- ing how machines understand human language (2023)
Abu Rayhan, Robert Kinzler, and Rajan Rayhan. Natural language processing: transform- ing how machines understand human language (2023). DOI: https://doi. org/10.13140/RG, 2(34900.99200)
2023
-
[80]
Identifying semantic induction heads to understand in-context learning
Jie Ren, Qipeng Guo, Hang Yan, Dongrui Liu, Quanshi Zhang, Xipeng Qiu, and Dahua Lin. Identifying semantic induction heads to understand in-context learning. arXiv preprint arXiv:2402.13055, 2024
2024 arXiv
-
[82]
Self-reflection in llm agents: Effects on problem-solving performance
Matthew Renze and Erhan Guven. Self-reflection in llm agents: Effects on problem-solving performance. arXiv preprint arXiv:2405.06682, 2024
2024 arXiv
-
[83]
Scaling laws for deep learning
Jonathan S Rosenfeld. Scaling laws for deep learning. arXiv preprint arXiv:2108.07686, 2021
2021 arXiv
-
[84]
Re- flexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Re- flexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[85]
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020
2020
-
[86]
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. Reinforcement learning: An introduction. MIT press, 2018. 13
2018
-
[87]
Cat-llm: Prompting large language models with text style definition for chinese article-style transfer
Zhen Tao, Dinghao Xi, Zhiyu Li, Liumin Tang, and Wei Xu. Cat-llm: Prompting large language models with text style definition for chinese article-style transfer. arXiv preprint arXiv:2401.05707, 2024
2024 arXiv
-
[88]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[89]
Value sensitive design and responsible innovation
Jeroen Van den Hoven. Value sensitive design and responsible innovation. Responsible innovation: Managing the responsible emergence of science and innovation in society, pages 75–83, 2013
2013
-
[90]
Interpretable counterfactual explanations guided by prototypes
Arnaud Van Looveren and Janis Klaise. Interpretable counterfactual explanations guided by prototypes. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pages 650–665. Springer, 2021
2021
-
[91]
Counterfactual explanations for machine learning: A review
Sahil Verma, John Dickerson, and Keegan Hines. Counterfactual explanations for machine learning: A review. arXiv preprint arXiv:2010.10596, 2:1, 2020
2010 arXiv
-
[92]
Decodingtrust: A comprehensive assess- ment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al. Decodingtrust: A comprehensive assess- ment of trustworthiness in gpt models. Advances in Neural Information Processing Systems, 36:31232...
2023
-
[93]
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024
2024 arXiv
-
[94]
Emu3: Next-token prediction is all you need
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al. Emu3: Next-token prediction is all you need. arXiv preprint arXiv:2409.18869, 2024
2024 arXiv
-
[95]
ai safety as global public goods
Yingchun Wang, Kai Jia, Jing Zhao, Ling Chen, Chunshen Qin, Yuan Yuan, Hongyu Fu, and Xingzhou Liang. “ai safety as global public goods” working report, 2024. Accessed: 2024-07-05
2024
-
[96]
Using the veil of ignorance to align ai systems with principles of justice
Laura Weidinger, Kevin R McKee, Richard Everett, Saffron Huang, Tina O Zhu, Martin J Chadwick, Christopher Summerfield, and Iason Gabriel. Using the veil of ignorance to align ai systems with principles of justice. Proceedings of the National Academy of Sciences , 120(18):e221...
2023
-
[97]
Efuf: Efficient fine-grained unlearning framework for mitigating hallucinations in multimodal large language models
Shangyu Xing, Fei Zhao, Zhen Wu, Tuo An, Weihao Chen, Chunhui Li, Jianbing Zhang, and Xinyu Dai. Efuf: Efficient fine-grained unlearning framework for mitigating hallucinations in multimodal large language models. arXiv preprint arXiv:2402.09801, 2024
2024 arXiv
-
[98]
Qwen2 technical report
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024
2024 arXiv
-
[99]
Huref: Human-readable fingerprint for large language models
Boyi Zeng, Lizheng Wang, Yuncong Hu, Yi Xu, Chenghu Zhou, Xinbing Wang, Yu Yu, and Zhouhan Lin. Huref: Human-readable fingerprint for large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2023
2023
-
[100]
The better angels of machine personality: How personality relates to llm safety
Jie Zhang, Dongrui Liu, Chen Qian, Ziyue Gan, Yong Liu, Yu Qiao, and Jing Shao. The better angels of machine personality: How personality relates to llm safety. arXiv preprint arXiv:2407.12344, 2024
2024 arXiv
-
[101]
Reef: Representation encoding fingerprints for large language models
Jie Zhang, Dongrui Liu, Chen Qian, Linfeng Zhang, Yong Liu, Yu Qiao, and Jing Shao. Reef: Representation encoding fingerprints for large language models. arXiv preprint arXiv:2410.14273, 2024. 14
2024 arXiv
-
[102]
Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization
Zhanhui Zhou, Jie Liu, Jing Shao, Xiangyu Yue, Chao Yang, Wanli Ouyang, and Yu Qiao. Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computation...
2024
-
[103]
Weak-to-strong search: Align large language models via searching over small language models
Zhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong, Chao Yang, and Yu Qiao. Weak-to-strong search: Align large language models via searching over small language models. arXiv preprint arXiv:2405.19262, 2024. 15
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.