REVIEW 4 major objections 5 minor 2 cited by
Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper argues that AI regulation should require developers to declare and justify the assumptions behind evaluation-based safety cases, and to halt development when those assumptions cannot be justified.
desk verdict A coherent policy argument for forcing explicit assumptions in AI evaluation safety cases; the halt trigger is under-specified, but the core proposal is sound and deserves serious peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is an explicit assumptions inventory for evaluation-based safety cases. For existing models it contains comprehensive threat modeling, proxy task validity, and adequate capability elicitation; for future models it adds comprehensive coverage of future threat vectors, validity and necessity of precursor capabilities, adequate elicitation of precursors, a sufficient compute gap between precursors and dangerous capabilities, comprehensive tracking of capability inputs, and accurate capability forecasts. The regulatory mechanism is the declare-and-justify requirement: the developer publishes the inventory and its justifications, third-party experts review them, and a judgment of inadequate justification triggers the same halt as a demonstrated unacceptable danger.
What would settle it
A concrete test would be an audit in which independent expert panels review the same developer safety case and are asked to judge whether the declared assumptions are adequately justified; if their verdicts diverge systematically, the proposed stop trigger cannot be applied consistently and the regime's central mechanism fails.
Extended reading notes
Core claim
The paper's central claim is that evaluation-based safety cases are valid only if a specified set of assumptions hold, and that this conditionality should be made legally explicit. Concretely: a finding that a model did not perform a dangerous task does not establish that the model lacks the dangerous capability unless the threat model was comprehensive, the proxy task was a necessary prerequisite, and elicitation was adequate; likewise, forecasts that future models will be safe require justifiable precursor definitions, a sufficient compute gap, and accurate forecasting. Because several of these assumptions cannot currently be justified with high confidence, the paper concludes that regulators should treat the absence of adequate justification as a stop condition equivalent to a failed evaluation, and should require developers to declare and justify assumptions as part of their safety case.
Load-bearing premise
The load-bearing premise is that third-party experts and regulators can reliably judge when a developer's justification of an assumption is adequate; the paper proposes such review but supplies no criteria or method for making that judgment.
Editorial extensions
If this is right
- An evaluation report that finds no dangerous capability would no longer count as evidence of safety unless the developer also justifies the threat model, proxy validity, and elicitation assumptions.
- Regulators would have a defined stop trigger: development halts not only on red-line capability results but also when assumptions are missing or judged inadequately justified.
- Developer safety frameworks would need to grow from capability thresholds to full assumption inventories with third-party review.
- Forecasting-based safety cases would need to demonstrate precursor validity, a sufficient compute gap, and an accurate forecasting track record before they justify continued training.
- Because many assumptions cannot currently be justified, evaluation-based claims that a system is safe would become provisional rather than established.
Reading between the lines
- An extension the paper leaves implicit is that the same declare-and-justify requirement could be applied symmetrically to the evaluators and auditing bodies themselves, forcing them to disclose their own elicitation and threat-modeling assumptions.
- The proposal implies a testable governance experiment: jurisdictions that adopt assumption disclosure could be compared with those that do not, tracking whether reported confidence matches actual incident rates.
- If adequacy judgments become a bottleneck, the practical effect may be to shift regulatory weight from measuring capabilities to judging the quality of arguments, which would require a shared standard for what counts as adequate justification.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that AI evaluations used in safety cases rest on a set of assumptions—comprehensive threat modeling, proxy task validity, adequate capability elicitation, necessity of precursor capabilities, sufficient compute gap, and accurate forecasting—and that many of these are currently difficult or impossible to justify. It proposes that AI regulation based on evaluations should require developers to explicitly state and justify these assumptions, subject the justifications to review by third-party experts, and halt development if the justifications are judged inadequate or if evaluations reveal unacceptable danger. The paper distinguishes assumptions for evaluating existing models from those for forecasting future models, and it separates misuse risk from autonomous misalignment risk.
Significance. This is a well-structured and clearly argued position paper that makes a useful conceptual contribution to AI governance. Its main strength is a systematic taxonomy of assumptions in evaluation-based safety cases, with a nuanced treatment of which assumptions might be justifiable for misuse risk versus autonomous risk. The paper is appropriately hedged and cites relevant developer policies and prior critiques. However, the central regulatory proposal is not operationalized: the adequacy standard for justifications is left undefined, and the paper does not address how the proposed mechanism would avoid perverse incentives or function in practice. Because these gaps bear directly on the paper's main policy recommendation, they require attention before the manuscript is suitable for publication.
major comments (4)
- [Section 4] The regulatory trigger depends on the concept of 'inadequate justifications' and the standard 'hold with very high probability,' but the paper provides no criteria, confidence thresholds, or decision procedure for making these judgments. Section 4 states that justifications should be 'assessed by third-party experts' and 'judged to be inadequate,' yet it does not specify what evidence counts as adequate, how to resolve disagreements among experts, or how to calibrate 'very high probability.' Without these details, the mechanism is unenforceable and could be gamed by strategic declarations. The authors should either provide an operational definition of adequacy or explicitly defer to a defined regulatory process with accountability mechanisms.
- [Section 3.1.1 and Section 4] The proposal encounters a circularity for the comprehensive threat modeling assumption: judging whether the assumption is adequately justified requires knowing whether all threat vectors have been considered, which is exactly the knowledge the evaluation was supposed to supply. The paper does not explain how third-party experts could verify completeness of threat coverage without circularity. The authors should discuss concrete methods (e.g., independent red-teaming, formal threat-model decomposition, or historical calibration against past adversarial discoveries) that could break this circularity, or explicitly acknowledge that this assumption cannot be verified and adjust the proposal accordingly.
- [Sections 3.1.1, 3.2.5, and 4] The paper's own conclusions are in tension with its proposed halt rule. Section 3.1.1 states that comprehensive threat modeling 'cannot be robustly justified' for autonomous risk, and Section 3.2.5 states that the compute-gap assumption 'cannot be robustly justified.' Under Section 4's rule that development 'should not continue if the assumptions are not judged to hold with very high probability,' this would mandate an immediate halt to frontier development. That may be the authors' intended conclusion, but the paper does not discuss this implication, nor does it consider whether a graded response (e.g., deployment with precautions, enhanced monitoring, or conditional approval) would be more appropriate when some assumptions are unjustified but evaluations show no current danger. The paper should clarify the intended consequences and specify a fallback if the assumptions are only partially justified.
- [Section 4] The proposal does not address regulatory feasibility or perverse incentives. A 'declare and justify' regime will only work if regulators can verify that the list of assumptions is complete and that justifications are not elaborate but empty theater. The paper does not discuss how to audit for completeness, how to protect genuinely sensitive threat models while ensuring public scrutiny, or what happens if developers simply declare that they have 'high confidence' without sufficient evidence. These implementation challenges are central to the paper's claim that the approach is 'a practical path towards more effective governance,' so the authors should at least outline a verification and audit framework or acknowledge the need for further policy research.
minor comments (5)
- [Abstract] There are typographical errors in the abstract: 'import ant' should be 'important' and 'assumption s' should be 'assumptions.'
- [Section 2] In the fourth item of the evaluation workflow, 'such a s' should be 'such as.'
- [References] References [17], [18], [22], [27], and [34] have URLs inserted directly in the text; they should be moved to the bibliography or formatted consistently with the other references.
- [Section 4] The phrase 'high-risk AI systems' is used without definition; the paper should clarify whether this refers to a particular regulatory risk tier, the systems exceeding the 'red lines' threshold, or something else.
- [General] The paper would be strengthened by a worked example of a developer's justification and a third-party expert review, illustrating how the adequacy standard would be applied in practice.
Circularity Check
No circularity: the paper is a policy argument with no fitted parameters, predictions, or load-bearing self-citations.
full rationale
The paper does not present a derivation chain, a fitted model, or a quantitative prediction. Its central claim is a regulatory proposal: if regulation is based on AI evaluations, developers should state and justify key assumptions, and development should halt when those justifications are inadequate. This is an argument supported by references to external developer policies (Anthropic RSP, OpenAI Preparedness Framework, Google DeepMind Frontier Safety Framework) and prior critiques of AI evaluations. No parameter is fitted to data and then renamed as a prediction. No cited uniqueness theorem or prior result by the same authors is invoked to force a conclusion. The paper's own observations that assumptions such as comprehensive threat modeling and sufficient compute gap 'cannot be robustly justified' are offered as limitations of current evaluation practice, not as conclusions derived from those assumptions. The potential epistemic circularity noted in the skeptic reading—judging whether threat modeling is comprehensive requires knowing all threat vectors, which is the very knowledge evaluations are supposed to supply—is explicitly acknowledged in Section 3.1.1 as a difficulty of justification, not disguised as a result. The proposal's undefined adequacy standard is a policy gap, not a circular derivation. Therefore the paper is self-contained as an argument and receives a score of 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Catastrophic risks from AI warrant new regulation.
- domain assumption AI evaluations are and will remain a central pillar of regulation.
- domain assumption Third-party experts can reliably assess the adequacy of justifications.
- domain assumption Explicit disclosure of assumptions improves regulatory decisions rather than merely creating compliance paperwork.
Cite this review
Pith. "Pith review of Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation." pith.science (2026). https://pith.science/paper/IJQ5HV5P
@misc{pith2026241112820,
author = {Pith},
title = {Pith review of: Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation},
year = {2026},
howpublished = {\url{https://pith.science/paper/IJQ5HV5P}},
note = {Machine review of arXiv:2411.12820}
}
read the original abstract
As AI systems advance, AI evaluations are becoming an important pillar of regulations for ensuring safety. We argue that such regulation should require developers to explicitly identify and justify key underlying assumptions about evaluations as part of their case for safety. We identify core assumptions in AI evaluations (both for evaluating existing models and forecasting future models), such as comprehensive threat modeling, proxy task validity, and adequate capability elicitation. Many of these assumptions cannot currently be well justified. If regulation is to be based on evaluations, it should require that AI development be halted if evaluations demonstrate unacceptable danger or if these assumptions are inadequately justified. Our presented approach aims to enhance transparency in AI development, offering a practical path towards more effective governance of advanced AI systems.
Forward citations
Cited by 2 Pith papers
-
A Conceptual Framework for AI Capability Evaluations
A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.
-
What AI evaluations for preventing catastrophic risks can and cannot do
AI evaluations can establish lower bounds on capabilities but cannot establish upper bounds, forecast future capabilities robustly, or assess misalignment risk, so they should not be the primary basis for AI safety decisions.
Reference graph
Works this paper leans on
-
[1]
Frontier AI Regulation: Managing Emerging Risks to Pu blic Safety, November 2023
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jad e Leung, Cullen O’Keefe, Jess Whit- tlestone, Shahar Avin, Miles Brundage, Justin Bullock, Dun can Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewi s Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Y onadav Shavit, Divya Siddarth, Robert Trager, and Kevin Wol...
arXiv 2023
-
[2]
Managing extreme AI risks am id rapid progress
Y oshua Bengio, Geoffrey Hinton, Andrew Y ao, Dawn Song, Pieter Abbeel, Trevor Darrell, Y u- val Noah Harari, Y a-Qin Zhang, Lan Xue, Shai Shalev-Shwartz , Gillian Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atılım Güne¸ s Baydin, Sheila M cIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Ru ssell, Daniel Kahneman, ...
work page 2024
-
[3]
International Scientific Report on the Safety of Advanced AI
Y ohsua Bengio, Daniel Privitera, Tamay Besiroglu, Rish i Bommasani, Stephen Casper, Y ejin Choi, Danielle Goldfarb, Hoda Heidari, Leila Khalatbari, S hayne Longpre, et al. International Scientific Report on the Safety of Advanced AI . Department for Science, Innovation and Tech- nology, 2024
work page 2024
-
[4]
An o verview of catastrophic AI risks
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside. An o verview of catastrophic AI risks. arXiv preprint arXiv:2306.12001 , 2023
arXiv 2023
-
[5]
OpenAI chief seeks new Microsoft funds to build ‘superintelligence.’
M Murgia. OpenAI chief seeks new Microsoft funds to build ‘superintelligence.’. The Finan- cial Times, 2023
work page 2023
-
[6]
Anthropic’s Responsible Scaling Policy, 20 23
Anthropic. Anthropic’s Responsible Scaling Policy, 20 23
-
[7]
Preparedness Framework (Beta), 2023
OpenAI. Preparedness Framework (Beta), 2023
2023
-
[8]
Frontier Safety Framework, 2024
Google Deepmind. Frontier Safety Framework, 2024
2024
Show all 37 references
-
[9]
Model evaluation for extreme risks, Septem- ber 2023
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary P huong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljun g, Noam Kolt, Lewis Ho, Di- vya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gab riel, Vijay Bolina, Jack Clark, Y oshua Ben...
2023 arXiv
-
[10]
Evaluating Frontier Models for Dangerous Capabilities, April 2024
Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cog an, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Y annis Assael, Sar ah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, L ewis Ho, Sharon Lin, Sebastian Farquhar...
2024
-
[11]
LLM Agents can Autonomously Exploit One-day Vulnerabilities, April 2024
Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang. LLM Agents can Autonomously Exploit One-day Vulnerabilities, April 2024. arXiv:2404. 08144 [cs]
2024
-
[12]
Teams of LLM Agents can Exploit Zero-Day Vulnerabilities, June 2024
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and D aniel Kang. Teams of LLM Agents can Exploit Zero-Day Vulnerabilities, June 2024. arXiv:24 06.01637 [cs]
2024
-
[13]
Soice, Rafael Rocha, Kimberlee Cordova, Micha el Specter, and Kevin M
Emily H. Soice, Rafael Rocha, Kimberlee Cordova, Micha el Specter, and Kevin M. Es- velt. Can large language models democratize access to dual- use biotechnology?, June 2023. arXiv:2306.03809 [cs]
2023 arXiv
-
[14]
Prioritizing High-Consequence Biolo gical Capabilities in Evaluations of Artificial Intelligence Models, May 2024
Jaspreet Pannu, Doni Bloomfield, Alex Zhu, Robert MacKn ight, Gabe Gomes, Anita Cicero, and Thomas Inglesby. Prioritizing High-Consequence Biolo gical Capabilities in Evaluations of Artificial Intelligence Models, May 2024
2024
-
[15]
On the Conver- sational Persuasiveness of Large Language Models: A Random ized Controlled Trial, March
Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallo tti, and Robert West. On the Conver- sational Persuasiveness of Large Language Models: A Random ized Controlled Trial, March
-
[16]
S. C. Matz, J. D. Teeny, S. S. V aid, H. Peters, G. M. Harari , and M. Cerf. The potential of generative AI for personalized persuasion at scale. Scientific Reports , 14(1):4692, February 2024
2024
-
[17]
AI Safety Institute Signs Agreements Regarding AI Safety Re- search, Testing and Evaluation With Anthropic and OpenAI — n ist.gov
U.S. AI Safety Institute Signs Agreements Regarding AI Safety Re- search, Testing and Evaluation With Anthropic and OpenAI — n ist.gov. https://www.nist.gov/news-events/news/2024/08/us-ai -safety-institute-signs-agreements-regardi [Accessed 12-09-2024]
2024
-
[18]
Rishi Sunak promised to make AI safe
Vincent Manancourt, Gian V olpicelli, and Mohar Chatte rjee. Rishi Sunak promised to make AI safe. Big Tech’s not playing ball. https://www.politico.eu/article/rishi-sunak-ai-test ing-tech-ai-safety-institute/ ,
-
[19]
Reasons to doubt the impact of AI risk ev aluations
Gabriel Mukobi. Reasons to doubt the impact of AI risk ev aluations. arXiv preprint arXiv:2408.02565, 2024
2024 arXiv
-
[21]
Evaluation in artificial intel ligence: from task-oriented to ability- oriented measurement
José Hernández-Orallo. Evaluation in artificial intel ligence: from task-oriented to ability- oriented measurement. Artificial Intelligence Review , 48(3):397–447, October 2017
2017
-
[22]
Evaluating AI Evaluation: Perils and Pros pects, July 2024
John Burden. Evaluating AI Evaluation: Perils and Pros pects, July 2024. arXiv:2407.09221 [cs]
2024 arXiv
-
[23]
Beyond static AI evaluations: advancing human interaction evaluations for llm harms and risks
Lujain Ibrahim, Saffron Huang, Lama Ahmad, and Markus A nderljung. Beyond static AI evaluations: advancing human interaction evaluations for llm harms and risks. arXiv preprint arXiv:2405.10632, 2024
2024 arXiv
-
[24]
We need a Science of Evals
Marius Hobbhahn. We need a Science of Evals. https://www.apolloresearch.ai/blog/we-need-a-scienc e-of-evals. [Accessed 12-09-2024]
2024
-
[25]
Secur- ing AI Model W eights: Preventing Theft and Misuse of Frontie r Models
Sella Nevo, Dan Lahav, Ajay Karpur, Y ogev Bar-On, and He nry Alexander Bradley. Secur- ing AI Model W eights: Preventing Theft and Misuse of Frontie r Models . Number 1. Rand Corporation, 2024. 6
2024
-
[26]
AI Safety Institute approach to evaluations, 2024
UK Department for Science, Innovation and Technology. AI Safety Institute approach to evaluations, 2024
2024
-
[27]
OpenAI Red Teaming Network, 2023
OpenAI. OpenAI Red Teaming Network, 2023
2023
-
[28]
Compute trends across three eras of machine learni ng
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu , Marius Hobbhahn, and Pablo Vil- lalobos. Compute trends across three eras of machine learni ng. In 2022 International Joint Conference on Neural Networks (IJCNN) , pages 1–8. IEEE, 2022
2022
-
[29]
AI capabilities can be significantly improved without expensive retraining
Tom Davidson, Jean-Stanislas Denain, Pablo Villalobo s, and Guillem Bas. AI capabilities can be significantly improved without expensive retraining . arXiv preprint arXiv:2312.07413 , 2023
2023 arXiv
-
[30]
Guidelines for capability elicitation, 2024
METR. Guidelines for capability elicitation, 2024
2024
-
[31]
Brown, and Francis Rhys Ward
Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samue l F. Brown, and Francis Rhys Ward. AI Sandbagging: Language Models can Strategically Underperf orm on Evaluations, June 2024. arXiv:2406.07358 [cs]
2024 arXiv
-
[32]
Black- Box Access is Insufficient for Rigorous AI Audits
Stephen Casper, Carson Ezell, Charlotte Siegmann, Noa m Kolt, Taylor Lynn Curtis, Ben- jamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin V on Hagen, Silas Alberti, Alan C han, Qinyi Sun, Michael Gerovitch, David...
2024 arXiv
-
[33]
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barre t Zoph, Sebastian Borgeaud, Dani Y ogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682 , 2022
2022 arXiv
-
[34]
Predictability and surprise in large generative models
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda As kell, Y untao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, et al. Predictability and surprise in large generative models. In Proceedings of the 2022 ACM Conference on Fairness, Account ability, an...
2022
-
[35]
Sb 1047: Safe and secure innovation for fr ontier artificial intelligence models act., 2024
Scott Wiener. Sb 1047: Safe and secure innovation for fr ontier artificial intelligence models act., 2024. 7
2024
-
[36]
Evaluating Predictions of Model Behaviour
Alan Chan. Evaluating Predictions of Model Behaviour. https://www.governance.ai/post/evaluating-predictio ns-of-model-behaviour,
-
[37]
[Accessed 12-09-2024]
2024
-
[2024]
arXiv:2403.14380 [cs]
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.