{"id":"c944602c-5b1b-42ab-af2b-58a582b87e99","arxiv_id":"2412.08653","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI evaluations can establish lower bounds on capabilities but cannot establish upper bounds, forecast future capabilities robustly, or assess misalignment risk, so they should not be the primary basis for AI safety decisions.","lead":"This paper argues that AI capability evaluations, while useful for showing what AI systems can do, cannot prove what they cannot do, cannot reliably predict future capabilities, and cannot robustly assess risks from autonomous systems. It concludes that governments and labs should not treat evaluations as the main safety mechanism for frontier AI.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The policy conclusion rests on an unproven impossibility claim: 'fundamental limitations that cannot be overcome within the current paradigm' is asserted, not demonstrated.","rationale":"The reader's weakest-assumption analysis identifies exactly the same load-bearing concern: the paper asserts, rather than proves, that the limitations are fundamental and cannot be overcome within the current paradigm. My stress test agrees with that assessment. The paper's internal logic supports the weaker, conditional claim that current evaluations provide only lower bounds and are unreliable for upper bounds, forecasting, and autonomous-risk assessment; that weaker claim is well-argued with concrete examples. But the policy conclusion that evaluations 'cannot reasonably form the exclusive basis' for safety decisions depends on the stronger impossibility claim. Since that claim is not established, the verdict should remain conditional: accept the critique of current evaluation practice, but require the authors either to supply a rigorous impossibility argument or to weaken the conclusion to 'current evaluations do not suffice and we lack evidence that they will suffice in the near term.' No change to the reader's conditional verdict is needed, because the reader already flags this as the decisive issue.","tokens_in":7656,"tokens_out":2756,"duration_ms":28352,"concrete_test":"Make the impossibility claim precise and test it formally. Fix a formal definition of the 'current paradigm' as black-box behavioral evaluations with a bounded elicitation budget, define a model class and a capability predicate, and attempt to prove a theorem that no finite evaluation in this class can certify an upper bound on that capability with any reliability guarantee. If a counterexample exists (e.g., exhaustive search over a small closed task space produces a guaranteed upper bound), the 'cannot be overcome' claim is false as stated; if a proof can be constructed, the concern is resolved. In parallel, attempt a systematic literature search for evaluation methods that combine behavioral tests with interpretability or mechanistic checks to see whether any proposed method already targets the gap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most load-bearing step is the move from 'current evaluations do not establish upper bounds, reliably forecast, or robustly assess autonomous risk' to 'these are fundamental limitations that cannot be overcome within the current paradigm' (Section 3, repeated in the Conclusion). The latter universal negative is what justifies the headline policy conclusion that evaluations should not be relied on as the main safety mechanism. But the paper provides no proof of impossibility, and its own evidence cuts the other way: the CyberSecEval 2 / Project Naptime example (Section 3.1) shows that a new scaffolding method raised measured exploit rates from 5% to 71-100% on the same model. That is a dramatic improvement in under-elicitation within the same 'behavioral evaluation' paradigm, not a demonstration that such improvements are impossible. Similarly, the paper's recommendations for third-party audits, red lines, and research implicitly assume evaluation practice can improve. The valid core is that currently deployed evaluations fall short; the stronger claim that they fundamentally cannot improve is load-bearing, unsupported, and not entailed by the examples. If future methods substantially reduce under-elicitation or give partial propensity measurements, the conclusion that evaluations cannot be a primary safety mechanism loses most of its force.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that AI evaluations, despite being central to current safety-case approaches, are insufficient as a primary safeguard against catastrophic risks. It distinguishes what evaluations can do—establishing lower bounds on capabilities and, under favorable assumptions, assessing misuse risk—from what they cannot do: establish upper bounds on capabilities, reliably forecast future dangerous capabilities, and robustly assess misalignment and autonomy risks. The argument is illustrated with the CyberSecEval 2/Project Naptime example and supported by references to frontier-lab safety frameworks. The paper concludes that evaluations should not be the main determinant of policy decisions and offers preliminary recommendations: third-party audits, conservative red lines, cybersecurity defense-in-depth, misalignment monitoring, and research investment.","tokens_in":94,"tokens_out":5748,"duration_ms":76439,"significance":"If the paper's conclusions were fully established, they would have substantial implications for AI governance, since several major labs currently use evaluation-based safety cases for frontier models. The paper makes a useful conceptual contribution by separating lower-bound evidence from upper-bound claims and by distinguishing human-misuse evaluation from autonomy and misalignment evaluation. It also gives a concrete, well-documented example of under-elicitation. However, the paper's strongest policy conclusion depends on a universal negative—that the limitations are fundamental and cannot be overcome within the current paradigm—and this is not established by the evidence presented. The paper is better as a critique of current evaluation practice than as a proof of inherent impossibility; that distinction should be reflected in its conclusions.","major_comments":[{"comment":"The central move from 'current evaluations do not establish upper bounds, reliably forecast, or robustly assess autonomous risk' to 'fundamental limitations that cannot be overcome within the current paradigm' is asserted, not demonstrated. A universal negative of this kind is load-bearing because it justifies the headline conclusion that evaluations should not be relied on as the main safety mechanism. The paper's own strongest example cuts against it: CyberSecEval 2 measured 5% exploitability, and Project Naptime raised the same model to 71–100% within the same behavioral-evaluation paradigm. That shows under-elicitation can be substantially reduced, not that it cannot be. Please either provide an explicit argument, or evidence, that no future evaluation method can overcome these limitations, or weaken the claim to 'are not currently overcome' and adjust the policy conclusions accordingly.","section":"Section 3 (opening) and Section 4"},{"comment":"The claim that 'We currently lack even theoretical approaches for measuring these propensities or determining if a system is truly aligned' is an unsupported enumeration of all possible future approaches, and it is load-bearing for the section's conclusion that misalignment risks cannot be robustly assessed. Likewise, the statement that honeypots 'never provide robust evidence of alignment' requires proof or a precise definition of 'robust' rather than an intuitive assertion. As written, the argument from absence of known methods is not sufficient for the strength of the conclusion.","section":"Section 3.3"},{"comment":"The claim that precursor-based forecasting 'rests on unjustifiable assumptions' is stronger than the evidence supports. The paper gives plausible failure modes (small gaps, simultaneous emergence, discontinuous progress) and notes that current frameworks do not justify the assumptions, but it does not show that no framework could ever justify them. Since the forecasting limitation is one of the three central 'cannot' claims, this should be reframed as 'currently unjustified' or supported by a more direct argument against the possibility of justification.","section":"Section 3.2"},{"comment":"There is an internal tension between the structural claim that limitations 'cannot be overcome simply via more thorough evaluation' and the recommendations, which explicitly aim to improve evaluation practice: third-party audits 'can help identify potential issues and improve the overall evaluation process,' and further research is proposed to move away from the current paradigm. The paper should clarify the sense in which improvements are possible while the limitations remain fundamental, and should state what evidence would count against the fundamental-limitation claim. Without this, the conclusion that evaluations cannot be the main safety mechanism does not follow cleanly from the premises.","section":"Section 5 vs Section 3"}],"minor_comments":[{"comment":"The phrase 'serve ascoordination points' is missing a space and should read 'serve as coordination points.'","section":"Section 2.3"},{"comment":"The statement that 'There are no principled methods to tell whether capabilities are being optimally elicited' is given without citation or argument; please define 'principled' and discuss at least the known elicitation methods that were considered.","section":"Section 3.1"},{"comment":"The term 'upper bound' is used informally; consider defining it as a statistical statement about performance over a task distribution and clarifying what kind of evidence would count as an upper bound.","section":"Section 3.1"},{"comment":"The sentence 'It is extremely likely that current evaluations will fail to consider important threat vectors or under-elicit AI system capabilities' asserts near-certainty about all current evaluations; 'are likely to fail' would be more proportionate to the evidence presented.","section":"Section 2.2"},{"comment":"The caption labels case (B) as 'Evaluations too infrequent failure,' but the text describes a smaller-than-expected gap between precursor and dangerous capabilities; the label and the description should be aligned.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a policy-oriented computer-society venue, and it explicitly builds on the authors' prior 'Declare and Justify' paper with appropriate citation. The main issue is overreach in the impossibility claims; I would not reject because the taxonomy and concrete examples are useful, but the conclusions need to be calibrated to the evidence. No concerns about undisclosed conflicts are visible from the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"One thing to know: this is a useful, clearly written position paper arguing that AI evaluations cannot carry the weight of safety cases. The CyberSecEval 2 / Project Naptime example and the lower-bounds/upper-bounds distinction are worth citing. But the headline 'cannot' overstates what the argument actually supports.\n\nThe genuinely new content is the organizing matrix—timing (current vs future models) by risk type (misuse vs misalignment)—and the concrete failure-mode diagrams for precursor-based forecasting (Figure 2). The synthesis of known limitations (under-elicitation, sandbagging, unknown unknowns) into a critique of Anthropic RSP, OpenAI Preparedness, and DeepMind FSF is cogent and well grounded in cited examples. The paper is honest about uncertainty and its preliminary recommendations (third-party audits, conservative red lines, defense in depth, monitoring, research) are sensible and incremental.\n\nThe soft spot is the load-bearing claim, repeated in the abstract and conclusion, that these limitations are 'fundamental' and 'cannot be overcome within the current paradigm.' That is asserted, not demonstrated. The paper says 'There are no principled methods' and 'We currently lack even theoretical approaches,' but absence of evidence is not evidence of impossibility. Worse, the paper's own best example cuts the other way: Project Naptime increased measured exploit rates from 5% to 71–100% on the same model with better scaffolding—a dramatic reduction in under-elicitation within the same behavioral-evaluation paradigm. That does not refute the impossibility of upper bounds, but it does undercut the universal claim that improvements can't happen. The authors' recommendations also implicitly assume evaluation practice can improve; audits and red lines only make sense if evaluations get better. If the paper had claimed 'current evaluations are insufficient and we don't yet know whether they can ever be sufficient,' the conclusion would follow. As written, the policy recommendation—don't rely on evaluations as the main safety mechanism—rests on a universal negative the examples don't entail.\n\nWho should read it: anyone working on AI governance, safety cases, or evaluation methodology. It's a credible piece that deserves serious engagement, but a referee should push the authors to qualify or split the impossibility claims. The core critique is solid; the overclaim is fixable.\n\nRecommendation: send to peer review. It should not be desk-rejected. With revision this could be a useful reference point for the field.","headline":"Solid critique of evaluation-based safety cases, but the headline impossibility claim is asserted, not proven, and the paper's own best example cuts against it.","tokens_in":8381,"tokens_out":2861,"would_cite":true,"duration_ms":24922,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that AI evaluations cannot establish upper bounds on model capabilities, so passing a safety evaluation is not evidence that a model lacks dangerous capabilities.","keywords":["AI evaluations","catastrophic risk","capability elicitation","under-elicitation","safety cases","AI governance","precursor capabilities","misalignment"],"falsifier":"One concrete disconfirmation would be a demonstrated upper bound: for some dangerous capability, an evaluation result proven to be invariant across all plausible elicitation strategies, so that a failed test genuinely implies the model cannot perform the task. A second would be a validated method that measures a model's propensities and reliably predicts its behavior in novel situations; the paper claims no such approach exists.","tokens_in":7458,"feed_emoji":"⚠️","tokens_out":7186,"duration_ms":57988,"temperature":0.7,"pith_summary":"AI capability evaluations are good at one thing: proving that a system can do at least what it did in the test. This paper argues that they cannot prove the opposite, that a system lacks a dangerous capability, because a failed test only shows that the evaluator's prompts and scaffolding did not elicit the capability. The same limit applies to forecasting: measuring 'precursor' abilities does not guarantee warning before dangerous abilities appear, since capabilities may emerge without predictable precursors. The authors conclude that evaluations are valuable for lower bounds, misuse assessment when evaluators have an edge, early warning, and coordination, but they cannot serve as the main mechanism for ensuring AI systems are safe.","feed_headline":"A model that passes safety checks may still harbor dangerous skills","feed_subtitle":"The paper argues evals only set lower bounds on capabilities, so safety cases cannot rest on them.","key_machinery":"The central object is the AI capability evaluation itself, treated as behavior-based measurement under the current paradigm. The load-bearing mechanism is under-elicitation: the gap between what a model can do and what an evaluator's specific prompts, scaffolding, and fine-tuning manage to elicit. From this gap the paper derives all three major limits, no upper bounds on capabilities, no reliable precursor-based forecasting, and no robust measurement of the propensities (objectives, drives, or instrumental goals) that would determine an autonomous system's behavior in novel situations. The paper also names the 'safety buffer' assumption, the idea that precursor capabilities appear early enough to trigger precautions, and shows that it depends on an unsupported difficulty gap between precursor and dangerous capabilities.","core_discovery":"The paper's central claim is that, within the current paradigm of measuring AI by observable behavior, evaluations can establish lower bounds on capabilities and, in principle, assess certain misuse risks, but they cannot establish upper bounds on capabilities, cannot robustly forecast future capabilities, and cannot robustly assess risks from autonomous misaligned systems. The unifying reason is under-elicitation: when a model fails a task, that failure provides no strong evidence that the model lacks the capability, because alternative scaffolding, fine-tuning, post-training enhancements, or in-context learning might unlock it. Consequently, a safety case that relies on 'the model did not demonstrate dangerous behavior' is not a safety demonstration. The paper's conclusion is that evaluations should remain part of the governance toolkit but should not be the main basis for deciding that an AI system is safe.","pith_inferences":["If the paper is right, regulators should give negative evaluation results much less weight and should look for governance levers that do not depend on proving the absence of capabilities, such as process, security, and containment requirements.","The under-elicitation evidence suggests a concrete measurement program: systematically varying scaffolding and post-training conditions across models to estimate how large the elicitation gap is; if the gap turned out to be small and predictable, some upper-bound reasoning might become practical.","The paper's logic extends naturally to interpretability: evidence from internal representations could in principle fill part of the gap left by behavioral evaluation, since it does not depend on eliciting a behavior to detect a latent capability.","A testable extension of the paper's position is that sudden jumps in measured benchmark performance from new elicitation methods will keep occurring, rather than diminishing over time."],"forward_implications":["A model that passes a dangerous-capability evaluation has not been shown to lack that capability, so deployment decisions should not treat a pass as evidence of safety.","Safety frameworks that rely on detecting precursor capabilities before dangerous ones appear are betting on an unverified assumption about the sequence in which capabilities emerge.","Because evaluations can only give the latest point at which action must be taken, waiting for a failed evaluation before acting may be dangerous.","For misalignment risk, post-training evaluation may be insufficient; a misaligned model could cause harm during training or internal use before any evaluation could catch it.","The paper's own recommendations point to third-party audits, conservative red lines, defense-in-depth cybersecurity, misalignment monitoring, and research as complements to evaluations."],"supporting_citations":[{"why":"Supplies the initial benchmark result showing a low observed rate of vulnerable-code exploitation.","marker":"[17]"},{"why":"Introduces the alternative scaffolding that raised the same model's observed exploitation rate to near-total, the paper's key evidence of under-elicitation.","marker":"[18]"},{"why":"Shows post-training enhancements can improve capabilities without retraining, supporting the claim that failed evaluations are not stable.","marker":"[19]"},{"why":"Provides the in-context learning mechanism by which partially present component skills might combine, supporting the upper-bound argument.","marker":"[24]"},{"why":"A frontier safety framework whose responsible-scaling logic depends on the 'safety buffer' assumption the paper critiques.","marker":"[4]"},{"why":"A preparedness framework that implicitly relies on precursor capabilities and evaluation triggers.","marker":"[5]"},{"why":"A frontier safety framework that explicitly discusses precursor capabilities and safety buffers, serving as the target of the forecasting critique.","marker":"[6]"},{"why":"A later version of a frontier safety policy that continues to depend on evaluation-based safety coverage.","marker":"[25]"},{"why":"Documents sandbagging, the behavior that makes negative evaluations unreliable for misalignment assessment.","marker":"[27]"},{"why":"Tests capability elicitation with password-locked models, showing limits of fine-tuning-based elicitation for novel capabilities.","marker":"[28]"}],"fun_headline_variants":["AI evals can't prove safety, only show known dangers","Why passing AI safety tests doesn't mean the model is safe","Evals set lower bounds, so safety cases are shaky","Under-elicitation: the flaw that makes AI evals unreliable","Don't bet on evals to catch catastrophic AI risks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion depends on the claim that these limitations are fundamental and cannot be overcome within the current behavior-observation paradigm; if future evaluation methods could substantially reduce under-elicitation or measure internal propensities, the case for not relying on evaluations as the main safety mechanism would lose much of its force.","fun_headline_variants_meta":{"raw":{"variants":["AI evals can't prove safety, only show known dangers","Why passing AI safety tests doesn't mean the model is safe","Evals set lower bounds, so safety cases are shaky","Under-elicitation: the flaw that makes AI evals unreliable","Don't bet on evals to catch catastrophic AI risks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000427,"raw_usage":{"total_tokens":2121,"prompt_tokens":819,"completion_tokens":1302,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":1217}},"tokens_in":435,"tokens_out":1302,"duration_ms":9240,"temperature":1.0,"reasoning_tokens":1217,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:52:22.046434+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One concrete disconfirmation would be a demonstrated upper bound: for some dangerous capability, an evaluation result proven to be invariant across all plausible elicitation strategies, so that a failed test genuinely implies the model cannot perform the task. A second would be a validated method that measures a model's propensities and reliably predicts its behavior in novel situations; the paper claims no such approach exists.","supporting_citations":[{"cited_title":"Project Naptime: Evaluating Offensive Security Capabili- ties of Large Language Models","cited_arxiv_id":null,"evidence_quote":"Introduces the alternative scaffolding that raised the same model's observed exploitation rate to near-total, the paper's key evidence of under-elicitation."},{"cited_title":"Anthropic’s Responsible Scaling Policy Version 1.0, 2023","cited_arxiv_id":null,"evidence_quote":"A frontier safety framework whose responsible-scaling logic depends on the 'safety buffer' assumption the paper critiques."},{"cited_title":"Preparedness Framework (Beta), 2023","cited_arxiv_id":null,"evidence_quote":"A preparedness framework that implicitly relies on precursor capabilities and evaluation triggers."},{"cited_title":"Frontier Safety Framework, 2024","cited_arxiv_id":null,"evidence_quote":"A frontier safety framework that explicitly discusses precursor capabilities and safety buffers, serving as the target of the forecasting critique."},{"cited_title":"Anthropic’s Responsible Scaling Policy, 2024","cited_arxiv_id":null,"evidence_quote":"A later version of a frontier safety policy that continues to depend on evaluation-based safety coverage."}],"review_version":1}