REVIEW 4 major objections 5 minor 29 references
What AI evaluations for preventing catastrophic risks can and cannot do
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper argues that AI evaluations cannot establish upper bounds on model capabilities, so passing a safety evaluation is not evidence that a model lacks dangerous capabilities.
desk verdict Solid critique of evaluation-based safety cases, but the headline impossibility claim is asserted, not proven, and the paper's own best example cuts against it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the AI capability evaluation itself, treated as behavior-based measurement under the current paradigm. The load-bearing mechanism is under-elicitation: the gap between what a model can do and what an evaluator's specific prompts, scaffolding, and fine-tuning manage to elicit. From this gap the paper derives all three major limits, no upper bounds on capabilities, no reliable precursor-based forecasting, and no robust measurement of the propensities (objectives, drives, or instrumental goals) that would determine an autonomous system's behavior in novel situations. The paper also names the 'safety buffer' assumption, the idea that precursor capabilities appear early enough to trigger precautions, and shows that it depends on an unsupported difficulty gap between precursor and dangerous capabilities.
What would settle it
One concrete disconfirmation would be a demonstrated upper bound: for some dangerous capability, an evaluation result proven to be invariant across all plausible elicitation strategies, so that a failed test genuinely implies the model cannot perform the task. A second would be a validated method that measures a model's propensities and reliably predicts its behavior in novel situations; the paper claims no such approach exists.
Extended reading notes
Core claim
The paper's central claim is that, within the current paradigm of measuring AI by observable behavior, evaluations can establish lower bounds on capabilities and, in principle, assess certain misuse risks, but they cannot establish upper bounds on capabilities, cannot robustly forecast future capabilities, and cannot robustly assess risks from autonomous misaligned systems. The unifying reason is under-elicitation: when a model fails a task, that failure provides no strong evidence that the model lacks the capability, because alternative scaffolding, fine-tuning, post-training enhancements, or in-context learning might unlock it. Consequently, a safety case that relies on 'the model did not demonstrate dangerous behavior' is not a safety demonstration. The paper's conclusion is that evaluations should remain part of the governance toolkit but should not be the main basis for deciding that an AI system is safe.
Load-bearing premise
The conclusion depends on the claim that these limitations are fundamental and cannot be overcome within the current behavior-observation paradigm; if future evaluation methods could substantially reduce under-elicitation or measure internal propensities, the case for not relying on evaluations as the main safety mechanism would lose much of its force.
Editorial extensions
If this is right
- A model that passes a dangerous-capability evaluation has not been shown to lack that capability, so deployment decisions should not treat a pass as evidence of safety.
- Safety frameworks that rely on detecting precursor capabilities before dangerous ones appear are betting on an unverified assumption about the sequence in which capabilities emerge.
- Because evaluations can only give the latest point at which action must be taken, waiting for a failed evaluation before acting may be dangerous.
- For misalignment risk, post-training evaluation may be insufficient; a misaligned model could cause harm during training or internal use before any evaluation could catch it.
- The paper's own recommendations point to third-party audits, conservative red lines, defense-in-depth cybersecurity, misalignment monitoring, and research as complements to evaluations.
Reading between the lines
- If the paper is right, regulators should give negative evaluation results much less weight and should look for governance levers that do not depend on proving the absence of capabilities, such as process, security, and containment requirements.
- The under-elicitation evidence suggests a concrete measurement program: systematically varying scaffolding and post-training conditions across models to estimate how large the elicitation gap is; if the gap turned out to be small and predictable, some upper-bound reasoning might become practical.
- The paper's logic extends naturally to interpretability: evidence from internal representations could in principle fill part of the gap left by behavioral evaluation, since it does not depend on eliciting a behavior to detect a latent capability.
- A testable extension of the paper's position is that sudden jumps in measured benchmark performance from new elicitation methods will keep occurring, rather than diminishing over time.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that AI evaluations, despite being central to current safety-case approaches, are insufficient as a primary safeguard against catastrophic risks. It distinguishes what evaluations can do—establishing lower bounds on capabilities and, under favorable assumptions, assessing misuse risk—from what they cannot do: establish upper bounds on capabilities, reliably forecast future dangerous capabilities, and robustly assess misalignment and autonomy risks. The argument is illustrated with the CyberSecEval 2/Project Naptime example and supported by references to frontier-lab safety frameworks. The paper concludes that evaluations should not be the main determinant of policy decisions and offers preliminary recommendations: third-party audits, conservative red lines, cybersecurity defense-in-depth, misalignment monitoring, and research investment.
Significance. If the paper's conclusions were fully established, they would have substantial implications for AI governance, since several major labs currently use evaluation-based safety cases for frontier models. The paper makes a useful conceptual contribution by separating lower-bound evidence from upper-bound claims and by distinguishing human-misuse evaluation from autonomy and misalignment evaluation. It also gives a concrete, well-documented example of under-elicitation. However, the paper's strongest policy conclusion depends on a universal negative—that the limitations are fundamental and cannot be overcome within the current paradigm—and this is not established by the evidence presented. The paper is better as a critique of current evaluation practice than as a proof of inherent impossibility; that distinction should be reflected in its conclusions.
major comments (4)
- [Section 3 (opening) and Section 4] The central move from 'current evaluations do not establish upper bounds, reliably forecast, or robustly assess autonomous risk' to 'fundamental limitations that cannot be overcome within the current paradigm' is asserted, not demonstrated. A universal negative of this kind is load-bearing because it justifies the headline conclusion that evaluations should not be relied on as the main safety mechanism. The paper's own strongest example cuts against it: CyberSecEval 2 measured 5% exploitability, and Project Naptime raised the same model to 71–100% within the same behavioral-evaluation paradigm. That shows under-elicitation can be substantially reduced, not that it cannot be. Please either provide an explicit argument, or evidence, that no future evaluation method can overcome these limitations, or weaken the claim to 'are not currently overcome' and adjust the policy conclusions accordingly.
- [Section 3.3] The claim that 'We currently lack even theoretical approaches for measuring these propensities or determining if a system is truly aligned' is an unsupported enumeration of all possible future approaches, and it is load-bearing for the section's conclusion that misalignment risks cannot be robustly assessed. Likewise, the statement that honeypots 'never provide robust evidence of alignment' requires proof or a precise definition of 'robust' rather than an intuitive assertion. As written, the argument from absence of known methods is not sufficient for the strength of the conclusion.
- [Section 3.2] The claim that precursor-based forecasting 'rests on unjustifiable assumptions' is stronger than the evidence supports. The paper gives plausible failure modes (small gaps, simultaneous emergence, discontinuous progress) and notes that current frameworks do not justify the assumptions, but it does not show that no framework could ever justify them. Since the forecasting limitation is one of the three central 'cannot' claims, this should be reframed as 'currently unjustified' or supported by a more direct argument against the possibility of justification.
- [Section 5 vs Section 3] There is an internal tension between the structural claim that limitations 'cannot be overcome simply via more thorough evaluation' and the recommendations, which explicitly aim to improve evaluation practice: third-party audits 'can help identify potential issues and improve the overall evaluation process,' and further research is proposed to move away from the current paradigm. The paper should clarify the sense in which improvements are possible while the limitations remain fundamental, and should state what evidence would count against the fundamental-limitation claim. Without this, the conclusion that evaluations cannot be the main safety mechanism does not follow cleanly from the premises.
minor comments (5)
- [Section 2.3] The phrase 'serve ascoordination points' is missing a space and should read 'serve as coordination points.'
- [Section 3.1] The statement that 'There are no principled methods to tell whether capabilities are being optimally elicited' is given without citation or argument; please define 'principled' and discuss at least the known elicitation methods that were considered.
- [Section 3.1] The term 'upper bound' is used informally; consider defining it as a statistical statement about performance over a task distribution and clarifying what kind of evidence would count as an upper bound.
- [Section 2.2] The sentence 'It is extremely likely that current evaluations will fail to consider important threat vectors or under-elicit AI system capabilities' asserts near-certainty about all current evaluations; 'are likely to fail' would be more proportionate to the evidence presented.
- [Figure 2] The caption labels case (B) as 'Evaluations too infrequent failure,' but the text describes a smaller-than-expected gap between precursor and dangerous capabilities; the label and the description should be aligned.
Circularity Check
No significant circularity: the paper's claims are argued from external evidence and logical analysis, not derived from its own prior work or from fitted inputs.
full rationale
The paper is an argumentative policy analysis rather than a formal derivation, so most circularity patterns do not apply. The only self-citation is the authors' prior 'Declare and Justify' paper [1], referenced in the introduction to set context; it supplies no premise needed for the later conclusions, which rest on external examples (CyberSecEval 2, Project Naptime, SWE-bench, industry safety frameworks) and on logical points about behavioral observation. The claim that evaluations cannot establish upper bounds follows from the nature of finite behavioral tests: absence of observed behavior does not logically imply absence of capability. That is an analytic observation, not a result derived from an input that already contains the conclusion. The broader assertion that the limitations are fundamental and cannot be overcome within the current paradigm is an unsupported universal negative, and it is arguably in tension with the paper's own Project Naptime example, which shows within-paradigm improvement from 5% to 71-100% under-elicitation. That is a correctness or overclaim concern, not circularity: the conclusion is overbroad relative to the evidence, but it is not equivalent to its inputs by construction. No fitted parameters, no imported uniqueness theorems, and no ansatz-smuggling via self-citation appear. Hence score 0.
Assumptions & free parameters
assumptions (5)
- domain assumption The current paradigm measures AI capabilities based on observed behavior.
- domain assumption Catastrophic risks from advanced AI are a real target for governance and should be mitigated.
- domain assumption Frontier safety cases rely on the premises that absence of demonstrated capability implies absence of danger, and that precursor capabilities will be detected before dangerous ones.
- ad hoc to paper No principled method exists for optimal capability elicitation.
- ad hoc to paper No theoretical approaches exist for measuring propensities or alignment.
Cite this review
Pith. "Pith review of What AI evaluations for preventing catastrophic risks can and cannot do." pith.science (2026). https://pith.science/paper/C3N7F6T4
@misc{pith2026241208653,
author = {Pith},
title = {Pith review of: What AI evaluations for preventing catastrophic risks can and cannot do},
year = {2026},
howpublished = {\url{https://pith.science/paper/C3N7F6T4}},
note = {Machine review of arXiv:2412.08653}
}
read the original abstract
AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what these evaluations can and cannot tell us. Evaluations can establish lower bounds on AI capabilities and assess certain misuse risks given sufficient effort from evaluators. Unfortunately, evaluations face fundamental limitations that cannot be overcome within the current paradigm. These include an inability to establish upper bounds on capabilities, reliably forecast future model capabilities, or robustly assess risks from autonomous AI systems. This means that while evaluations are valuable tools, we should not rely on them as our main way of ensuring AI systems are safe. We conclude with recommendations for incremental improvements to frontier AI safety, while acknowledging these fundamental limitations remain unsolved.
Figures
Reference graph
Works this paper leans on
-
[1]
Declare and justify: Explicit assumptions in ai evaluations are necessary for effective regulation
Peter Barnett and Lisa Thiergart. Declare and justify: Explicit assumptions in ai evaluations are necessary for effective regulation. arXiv preprint arXiv:2411.12820, 2024
arXiv 2024
-
[2]
Safety Cases: How to Justify the Safety of Advanced AI Systems, March 2024
Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen. Safety Cases: How to Justify the Safety of Advanced AI Systems, March 2024. arXiv:2403.10462 [cs]
arXiv 2024
-
[3]
Safety case template for frontier ai: A cyber inability argument
Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett, Tomek Korbak, Jessica Wang, Benjamin Hilton, and Geoffrey Irving. Safety case template for frontier ai: A cyber inability argument. arXiv preprint arXiv:2411.08088, 2024
arXiv 2024
-
[4]
Anthropic’s Responsible Scaling Policy Version 1.0, 2023
Anthropic. Anthropic’s Responsible Scaling Policy Version 1.0, 2023
work page 2023
- [5]
- [6]
-
[7]
Evaluating Frontier Models for Dangerous Capabilities, April 2024
Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, 8 Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anc...
arXiv 2024
-
[8]
Llm agents can autonomously hack websites
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang. Llm agents can autonomously hack websites. arXiv preprint arXiv:2402.06664, 2024
arXiv 2024
Show all 29 references
-
[9]
LLM Agents can Autonomously Exploit One-day Vulnerabilities, April 2024
Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang. LLM Agents can Autonomously Exploit One-day Vulnerabilities, April 2024. arXiv:2404.08144 [cs]
2024 arXiv
-
[10]
Teams of LLM Agents can Exploit Zero-Day Vulnerabilities, June 2024
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang. Teams of LLM Agents can Exploit Zero-Day Vulnerabilities, June 2024. arXiv:2406.01637 [cs]
2024 arXiv
-
[11]
On the Conversa- tional Persuasiveness of Large Language Models: A Randomized Controlled Trial, March 2024
Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West. On the Conversa- tional Persuasiveness of Large Language Models: A Randomized Controlled Trial, March 2024. arXiv:2403.14380 [cs]
2024 arXiv
-
[12]
S. C. Matz, J. D. Teeny, S. S. Vaid, H. Peters, G. M. Harari, and M. Cerf. The potential of generative AI for personalized persuasion at scale. Scientific Reports, 14(1):4692, February 2024
2024
-
[13]
Black-Box Access is Insufficient for Rigorous AI Audits
Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin V on Hagen, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau...
2024 arXiv
-
[14]
Number 1
Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, and Henry Alexander Bradley.Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. Number 1. Rand Corporation, 2024
2024
-
[15]
Mistral CEO confirms ‘leak’ of new open source AI model nearing GPT-4 performance
Carl Franzen. Mistral CEO confirms ‘leak’ of new open source AI model nearing GPT-4 performance. https://venturebeat.com/ai/ mistral-ceo-confirms-leak-of-new-open-source-ai-model-nearing-gpt-4-performance/ ,
-
[16]
Coordinated pausing: An evaluation-based coordination scheme for frontier ai developers
Jide Alaga and Jonas Schuett. Coordinated pausing: An evaluation-based coordination scheme for frontier ai developers. arXiv preprint arXiv:2310.00374, 2023
2023 arXiv
-
[17]
Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models
Manish Bhatt, Sahana Chennabasappa, Yue Li, Cyrus Nikolaidis, Daniel Song, Shengye Wan, Faizan Ahmad, Cornelius Aschermann, Yaohui Chen, Dhaval Kapil, et al. Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models. arXiv preprint arXiv:2404.13161, 2024
2024 arXiv
-
[18]
Project Naptime: Evaluating Offensive Security Capabili- ties of Large Language Models
Sergei Glazunov and Mark Brand. Project Naptime: Evaluating Offensive Security Capabili- ties of Large Language Models. https://googleprojectzero.blogspot.com/2024/06/ project-naptime.html, 2024. [Accessed 22-11-2024]
2024
-
[19]
AI capabilities can be significantly improved without expensive retraining
Tom Davidson, Jean-Stanislas Denain, Pablo Villalobos, and Guillem Bas. AI capabilities can be significantly improved without expensive retraining. arXiv preprint arXiv:2312.07413, 2023
2023 arXiv
-
[20]
SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024
2024
-
[21]
SWE-bench leaderboard
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. SWE-bench leaderboard. https://www.swebench.com/, 2024. [Accessed 22-11-2024]
2024
-
[22]
GAIA Leaderboard - a Hugging Face Space by gaia-benchmark
Hugging Face. GAIA Leaderboard - a Hugging Face Space by gaia-benchmark. https: //huggingface.co/spaces/gaia-benchmark/leaderboard. [Accessed 25-11-2024]. 9
2024
-
[23]
HumanEval Benchmark (Code Generation)
Papers with Code. HumanEval Benchmark (Code Generation). https://paperswithcode. com/sota/code-generation-on-humaneval . [Accessed 25-11-2024]
2024
-
[24]
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, et al. A survey on in-context learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1107–1128, 2024
2024
-
[25]
Anthropic’s Responsible Scaling Policy, 2024
Anthropic. Anthropic’s Responsible Scaling Policy, 2024
2024
-
[26]
We need a Science of Evals
Marius Hobbhahn. We need a Science of Evals. https://www.apolloresearch.ai/blog/ we-need-a-science-of-evals . [Accessed 12-09-2024]
2024
-
[27]
Brown, and Francis Rhys Ward
Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F. Brown, and Francis Rhys Ward. AI Sandbagging: Language Models can Strategically Underperform on Evaluations, June 2024. arXiv:2406.07358 [cs]
2024 arXiv
-
[28]
Stress-testing capability elicitation with password-locked models
Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger. Stress-testing capability elicitation with password-locked models. arXiv preprint arXiv:2405.19550, 2024. 10
2024 arXiv
-
[2024]
[Accessed 22-11-2024]
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.