Pith. sign in

REVIEW 4 major objections 5 minor 29 references

What AI evaluations for preventing catastrophic risks can and cannot do

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper argues that AI evaluations cannot establish upper bounds on model capabilities, so passing a safety evaluation is not evidence that a model lacks dangerous capabilities.

desk verdict Solid critique of evaluation-based safety cases, but the headline impossibility claim is asserted, not proven, and the paper's own best example cuts against it. read the letter →

arxiv 2412.08653 v1 pith:C3N7F6T4 submitted 2024-11-26 cs.CY cs.AI

classification cs.CYcs.AI
keywords AIevaluationscatastrophicriskcapabilityelicitationunder-elicitationsafetycasesgovernanceprecursorcapabilitiesmisalignment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

AI capability evaluations are good at one thing: proving that a system can do at least what it did in the test. This paper argues that they cannot prove the opposite, that a system lacks a dangerous capability, because a failed test only shows that the evaluator's prompts and scaffolding did not elicit the capability. The same limit applies to forecasting: measuring 'precursor' abilities does not guarantee warning before dangerous abilities appear, since capabilities may emerge without predictable precursors. The authors conclude that evaluations are valuable for lower bounds, misuse assessment when evaluators have an edge, early warning, and coordination, but they cannot serve as the main mechanism for ensuring AI systems are safe.

What carries the argument

The central object is the AI capability evaluation itself, treated as behavior-based measurement under the current paradigm. The load-bearing mechanism is under-elicitation: the gap between what a model can do and what an evaluator's specific prompts, scaffolding, and fine-tuning manage to elicit. From this gap the paper derives all three major limits, no upper bounds on capabilities, no reliable precursor-based forecasting, and no robust measurement of the propensities (objectives, drives, or instrumental goals) that would determine an autonomous system's behavior in novel situations. The paper also names the 'safety buffer' assumption, the idea that precursor capabilities appear early enough to trigger precautions, and shows that it depends on an unsupported difficulty gap between precursor and dangerous capabilities.

What would settle it

One concrete disconfirmation would be a demonstrated upper bound: for some dangerous capability, an evaluation result proven to be invariant across all plausible elicitation strategies, so that a failed test genuinely implies the model cannot perform the task. A second would be a validated method that measures a model's propensities and reliably predicts its behavior in novel situations; the paper claims no such approach exists.

Watch

Extended reading notes

Core claim

The paper's central claim is that, within the current paradigm of measuring AI by observable behavior, evaluations can establish lower bounds on capabilities and, in principle, assess certain misuse risks, but they cannot establish upper bounds on capabilities, cannot robustly forecast future capabilities, and cannot robustly assess risks from autonomous misaligned systems. The unifying reason is under-elicitation: when a model fails a task, that failure provides no strong evidence that the model lacks the capability, because alternative scaffolding, fine-tuning, post-training enhancements, or in-context learning might unlock it. Consequently, a safety case that relies on 'the model did not demonstrate dangerous behavior' is not a safety demonstration. The paper's conclusion is that evaluations should remain part of the governance toolkit but should not be the main basis for deciding that an AI system is safe.

Load-bearing premise

The conclusion depends on the claim that these limitations are fundamental and cannot be overcome within the current behavior-observation paradigm; if future evaluation methods could substantially reduce under-elicitation or measure internal propensities, the case for not relying on evaluations as the main safety mechanism would lose much of its force.

Editorial extensions

If this is right

  • A model that passes a dangerous-capability evaluation has not been shown to lack that capability, so deployment decisions should not treat a pass as evidence of safety.
  • Safety frameworks that rely on detecting precursor capabilities before dangerous ones appear are betting on an unverified assumption about the sequence in which capabilities emerge.
  • Because evaluations can only give the latest point at which action must be taken, waiting for a failed evaluation before acting may be dangerous.
  • For misalignment risk, post-training evaluation may be insufficient; a misaligned model could cause harm during training or internal use before any evaluation could catch it.
  • The paper's own recommendations point to third-party audits, conservative red lines, defense-in-depth cybersecurity, misalignment monitoring, and research as complements to evaluations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper is right, regulators should give negative evaluation results much less weight and should look for governance levers that do not depend on proving the absence of capabilities, such as process, security, and containment requirements.
  • The under-elicitation evidence suggests a concrete measurement program: systematically varying scaffolding and post-training conditions across models to estimate how large the elicitation gap is; if the gap turned out to be small and predictable, some upper-bound reasoning might become practical.
  • The paper's logic extends naturally to interpretability: evidence from internal representations could in principle fill part of the gap left by behavioral evaluation, since it does not depend on eliciting a behavior to detect a latent capability.
  • A testable extension of the paper's position is that sudden jumps in measured benchmark performance from new elicitation methods will keep occurring, rather than diminishing over time.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper argues that AI evaluations, despite being central to current safety-case approaches, are insufficient as a primary safeguard against catastrophic risks. It distinguishes what evaluations can do—establishing lower bounds on capabilities and, under favorable assumptions, assessing misuse risk—from what they cannot do: establish upper bounds on capabilities, reliably forecast future dangerous capabilities, and robustly assess misalignment and autonomy risks. The argument is illustrated with the CyberSecEval 2/Project Naptime example and supported by references to frontier-lab safety frameworks. The paper concludes that evaluations should not be the main determinant of policy decisions and offers preliminary recommendations: third-party audits, conservative red lines, cybersecurity defense-in-depth, misalignment monitoring, and research investment.

Significance. If the paper's conclusions were fully established, they would have substantial implications for AI governance, since several major labs currently use evaluation-based safety cases for frontier models. The paper makes a useful conceptual contribution by separating lower-bound evidence from upper-bound claims and by distinguishing human-misuse evaluation from autonomy and misalignment evaluation. It also gives a concrete, well-documented example of under-elicitation. However, the paper's strongest policy conclusion depends on a universal negative—that the limitations are fundamental and cannot be overcome within the current paradigm—and this is not established by the evidence presented. The paper is better as a critique of current evaluation practice than as a proof of inherent impossibility; that distinction should be reflected in its conclusions.

major comments (4)
  1. [Section 3 (opening) and Section 4] The central move from 'current evaluations do not establish upper bounds, reliably forecast, or robustly assess autonomous risk' to 'fundamental limitations that cannot be overcome within the current paradigm' is asserted, not demonstrated. A universal negative of this kind is load-bearing because it justifies the headline conclusion that evaluations should not be relied on as the main safety mechanism. The paper's own strongest example cuts against it: CyberSecEval 2 measured 5% exploitability, and Project Naptime raised the same model to 71–100% within the same behavioral-evaluation paradigm. That shows under-elicitation can be substantially reduced, not that it cannot be. Please either provide an explicit argument, or evidence, that no future evaluation method can overcome these limitations, or weaken the claim to 'are not currently overcome' and adjust the policy conclusions accordingly.
  2. [Section 3.3] The claim that 'We currently lack even theoretical approaches for measuring these propensities or determining if a system is truly aligned' is an unsupported enumeration of all possible future approaches, and it is load-bearing for the section's conclusion that misalignment risks cannot be robustly assessed. Likewise, the statement that honeypots 'never provide robust evidence of alignment' requires proof or a precise definition of 'robust' rather than an intuitive assertion. As written, the argument from absence of known methods is not sufficient for the strength of the conclusion.
  3. [Section 3.2] The claim that precursor-based forecasting 'rests on unjustifiable assumptions' is stronger than the evidence supports. The paper gives plausible failure modes (small gaps, simultaneous emergence, discontinuous progress) and notes that current frameworks do not justify the assumptions, but it does not show that no framework could ever justify them. Since the forecasting limitation is one of the three central 'cannot' claims, this should be reframed as 'currently unjustified' or supported by a more direct argument against the possibility of justification.
  4. [Section 5 vs Section 3] There is an internal tension between the structural claim that limitations 'cannot be overcome simply via more thorough evaluation' and the recommendations, which explicitly aim to improve evaluation practice: third-party audits 'can help identify potential issues and improve the overall evaluation process,' and further research is proposed to move away from the current paradigm. The paper should clarify the sense in which improvements are possible while the limitations remain fundamental, and should state what evidence would count against the fundamental-limitation claim. Without this, the conclusion that evaluations cannot be the main safety mechanism does not follow cleanly from the premises.
minor comments (5)
  1. [Section 2.3] The phrase 'serve ascoordination points' is missing a space and should read 'serve as coordination points.'
  2. [Section 3.1] The statement that 'There are no principled methods to tell whether capabilities are being optimally elicited' is given without citation or argument; please define 'principled' and discuss at least the known elicitation methods that were considered.
  3. [Section 3.1] The term 'upper bound' is used informally; consider defining it as a statistical statement about performance over a task distribution and clarifying what kind of evidence would count as an upper bound.
  4. [Section 2.2] The sentence 'It is extremely likely that current evaluations will fail to consider important threat vectors or under-elicit AI system capabilities' asserts near-certainty about all current evaluations; 'are likely to fail' would be more proportionate to the evidence presented.
  5. [Figure 2] The caption labels case (B) as 'Evaluations too infrequent failure,' but the text describes a smaller-than-expected gap between precursor and dangerous capabilities; the label and the description should be aligned.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims are argued from external evidence and logical analysis, not derived from its own prior work or from fitted inputs.

full rationale

The paper is an argumentative policy analysis rather than a formal derivation, so most circularity patterns do not apply. The only self-citation is the authors' prior 'Declare and Justify' paper [1], referenced in the introduction to set context; it supplies no premise needed for the later conclusions, which rest on external examples (CyberSecEval 2, Project Naptime, SWE-bench, industry safety frameworks) and on logical points about behavioral observation. The claim that evaluations cannot establish upper bounds follows from the nature of finite behavioral tests: absence of observed behavior does not logically imply absence of capability. That is an analytic observation, not a result derived from an input that already contains the conclusion. The broader assertion that the limitations are fundamental and cannot be overcome within the current paradigm is an unsupported universal negative, and it is arguably in tension with the paper's own Project Naptime example, which shows within-paradigm improvement from 5% to 71-100% under-elicitation. That is a correctness or overclaim concern, not circularity: the conclusion is overbroad relative to the evidence, but it is not equivalent to its inputs by construction. No fitted parameters, no imported uniqueness theorems, and no ansatz-smuggling via self-citation appear. Hence score 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper rests on domain assumptions about the stakes and the current evaluation paradigm, plus two ad hoc assertions about the absence of elicitation and propensity-measurement methods. These two assertions are the weakest structural supports, as they are universal negatives presented without proof.

assumptions (5)
  • domain assumption The current paradigm measures AI capabilities based on observed behavior.
    The paper's limitation analysis is explicitly scoped to this paradigm in Section 3's introduction and the Conclusion.
  • domain assumption Catastrophic risks from advanced AI are a real target for governance and should be mitigated.
    The paper's motivation and recommendations presuppose this urgency; it is not argued for within the paper.
  • domain assumption Frontier safety cases rely on the premises that absence of demonstrated capability implies absence of danger, and that precursor capabilities will be detected before dangerous ones.
    Presented as the target of critique, referencing Anthropic RSP, OpenAI Preparedness, and Google DeepMind FSF in Section 1.
  • ad hoc to paper No principled method exists for optimal capability elicitation.
    Asserted in Section 3.1 without proof; it is load-bearing for the conclusion that upper bounds cannot be established.
  • ad hoc to paper No theoretical approaches exist for measuring propensities or alignment.
    Asserted in Section 3.3 without proof; it is load-bearing for the conclusion that misalignment risk cannot be robustly assessed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of What AI evaluations for preventing catastrophic risks can and cannot do." pith.science (2026). https://pith.science/paper/C3N7F6T4

@misc{pith2026241208653,
  author       = {Pith},
  title        = {Pith review of: What AI evaluations for preventing catastrophic risks can and cannot do},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/C3N7F6T4}},
  note         = {Machine review of arXiv:2412.08653}
}
read the original abstract

AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what these evaluations can and cannot tell us. Evaluations can establish lower bounds on AI capabilities and assess certain misuse risks given sufficient effort from evaluators. Unfortunately, evaluations face fundamental limitations that cannot be overcome within the current paradigm. These include an inability to establish upper bounds on capabilities, reliably forecast future model capabilities, or robustly assess risks from autonomous AI systems. This means that while evaluations are valuable tools, we should not rely on them as our main way of ensuring AI systems are safe. We conclude with recommendations for incremental improvements to frontier AI safety, while acknowledging these fundamental limitations remain unsolved.

Figures

Figures reproduced from arXiv: 2412.08653 by the authors.

Figure 1
Figure 1. Evaluations are performed such that the regions assumed to be safe are always overlapping. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Precursor-based capability forecasting and potential failure modes. Blue circles indicate [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

29 extracted references · 13 canonical work pages

  1. [1]

    Declare and justify: Explicit assumptions in ai evaluations are necessary for effective regulation

    Peter Barnett and Lisa Thiergart. Declare and justify: Explicit assumptions in ai evaluations are necessary for effective regulation. arXiv preprint arXiv:2411.12820, 2024

  2. [2]

    Safety Cases: How to Justify the Safety of Advanced AI Systems, March 2024

    Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen. Safety Cases: How to Justify the Safety of Advanced AI Systems, March 2024. arXiv:2403.10462 [cs]

  3. [3]

    Safety case template for frontier ai: A cyber inability argument

    Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett, Tomek Korbak, Jessica Wang, Benjamin Hilton, and Geoffrey Irving. Safety case template for frontier ai: A cyber inability argument. arXiv preprint arXiv:2411.08088, 2024

  4. [4]

    Anthropic’s Responsible Scaling Policy Version 1.0, 2023

    Anthropic. Anthropic’s Responsible Scaling Policy Version 1.0, 2023

  5. [5]

    Preparedness Framework (Beta), 2023

    OpenAI. Preparedness Framework (Beta), 2023

  6. [6]

    Frontier Safety Framework, 2024

    Google Deepmind. Frontier Safety Framework, 2024

  7. [7]

    Evaluating Frontier Models for Dangerous Capabilities, April 2024

    Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, 8 Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anc...

  8. [8]

    Llm agents can autonomously hack websites

    Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang. Llm agents can autonomously hack websites. arXiv preprint arXiv:2402.06664, 2024

Show all 29 references
  1. [9]

    LLM Agents can Autonomously Exploit One-day Vulnerabilities, April 2024

    Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang. LLM Agents can Autonomously Exploit One-day Vulnerabilities, April 2024. arXiv:2404.08144 [cs]

  2. [10]

    Teams of LLM Agents can Exploit Zero-Day Vulnerabilities, June 2024

    Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang. Teams of LLM Agents can Exploit Zero-Day Vulnerabilities, June 2024. arXiv:2406.01637 [cs]

  3. [11]

    On the Conversa- tional Persuasiveness of Large Language Models: A Randomized Controlled Trial, March 2024

    Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West. On the Conversa- tional Persuasiveness of Large Language Models: A Randomized Controlled Trial, March 2024. arXiv:2403.14380 [cs]

  4. [12]

    S. C. Matz, J. D. Teeny, S. S. Vaid, H. Peters, G. M. Harari, and M. Cerf. The potential of generative AI for personalized persuasion at scale. Scientific Reports, 14(1):4692, February 2024

  5. [13]

    Black-Box Access is Insufficient for Rigorous AI Audits

    Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin V on Hagen, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau...

  6. [14]

    Number 1

    Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, and Henry Alexander Bradley.Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. Number 1. Rand Corporation, 2024

  7. [15]

    Mistral CEO confirms ‘leak’ of new open source AI model nearing GPT-4 performance

    Carl Franzen. Mistral CEO confirms ‘leak’ of new open source AI model nearing GPT-4 performance. https://venturebeat.com/ai/ mistral-ceo-confirms-leak-of-new-open-source-ai-model-nearing-gpt-4-performance/ ,

  8. [16]

    Coordinated pausing: An evaluation-based coordination scheme for frontier ai developers

    Jide Alaga and Jonas Schuett. Coordinated pausing: An evaluation-based coordination scheme for frontier ai developers. arXiv preprint arXiv:2310.00374, 2023

  9. [17]

    Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models

    Manish Bhatt, Sahana Chennabasappa, Yue Li, Cyrus Nikolaidis, Daniel Song, Shengye Wan, Faizan Ahmad, Cornelius Aschermann, Yaohui Chen, Dhaval Kapil, et al. Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models. arXiv preprint arXiv:2404.13161, 2024

  10. [18]

    Project Naptime: Evaluating Offensive Security Capabili- ties of Large Language Models

    Sergei Glazunov and Mark Brand. Project Naptime: Evaluating Offensive Security Capabili- ties of Large Language Models. https://googleprojectzero.blogspot.com/2024/06/ project-naptime.html, 2024. [Accessed 22-11-2024]

  11. [19]

    AI capabilities can be significantly improved without expensive retraining

    Tom Davidson, Jean-Stanislas Denain, Pablo Villalobos, and Guillem Bas. AI capabilities can be significantly improved without expensive retraining. arXiv preprint arXiv:2312.07413, 2023

  12. [20]

    SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024

    Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024

  13. [21]

    SWE-bench leaderboard

    Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. SWE-bench leaderboard. https://www.swebench.com/, 2024. [Accessed 22-11-2024]

  14. [22]

    GAIA Leaderboard - a Hugging Face Space by gaia-benchmark

    Hugging Face. GAIA Leaderboard - a Hugging Face Space by gaia-benchmark. https: //huggingface.co/spaces/gaia-benchmark/leaderboard. [Accessed 25-11-2024]. 9

  15. [23]

    HumanEval Benchmark (Code Generation)

    Papers with Code. HumanEval Benchmark (Code Generation). https://paperswithcode. com/sota/code-generation-on-humaneval . [Accessed 25-11-2024]

  16. [24]

    A survey on in-context learning

    Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, et al. A survey on in-context learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1107–1128, 2024

  17. [25]

    Anthropic’s Responsible Scaling Policy, 2024

    Anthropic. Anthropic’s Responsible Scaling Policy, 2024

  18. [26]

    We need a Science of Evals

    Marius Hobbhahn. We need a Science of Evals. https://www.apolloresearch.ai/blog/ we-need-a-science-of-evals . [Accessed 12-09-2024]

  19. [27]

    Brown, and Francis Rhys Ward

    Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F. Brown, and Francis Rhys Ward. AI Sandbagging: Language Models can Strategically Underperform on Evaluations, June 2024. arXiv:2406.07358 [cs]

  20. [28]

    Stress-testing capability elicitation with password-locked models

    Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger. Stress-testing capability elicitation with password-locked models. arXiv preprint arXiv:2405.19550, 2024. 10

  21. [2024]

    [Accessed 22-11-2024]

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.