{"id":"338325f4-662b-4895-93ee-63aca0f8ed6b","arxiv_id":"2502.05791","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Applying Assurance 2.0 to a cyber-misuse safety case, the authors show that high top-level confidence requires extremely high confidence in every component, and propose an LLM-based Delphi for eliciting those component probabilities.","lead":"This paper applies the Assurance 2.0 safety case methodology to a frontier AI cyber-misuse argument and finds that reaching 95% confidence in a small safety case would require about 99.3% confidence in every component. It also introduces an LLM-run Delphi process for estimating probabilities of safety case doubts, and a scoring formula for prioritizing which doubts to resolve first.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 99.3% per-component requirement in §7.1 assumes mutual independence; §7.3 concedes the fragment's sub-claims share budget, team, and information, so positive correlation could substantially lower the needed confidence, making the headline challenge an artifact of the independence assumption.","rationale":"The reader's weakest_assumption correctly identifies the independence assumption in the propagation formulas as load-bearing. My stress-test focuses on that same assumption because the numerical headline of §7.1 (99.3%) is the paper's most cited quantitative finding and is a direct algebraic consequence of Eq. (7) and Eq. (9). The paper's own §7.3 provides the admission that the fragment's three sub-claims are not independent, which means the condition under which the headline is derived is not met for the very case study used. Positive dependence among sub-claims increases the probability of their conjunction, so the required marginal confidence can be materially lower than the product-method value. The sum-of-doubts method is even more conservative, as it is a union-bound lower bound; Bloomfield and Rushby, as cited in the paper, caution that absolute values carry little significance. Therefore the headline 'requires near-perfect confidence' is an upper-bound-style challenge that depends on an assumption known to be false, not a robust property of conjunctive safety cases. The paper does disclose the limitation, which is why the verdict remains CONDITIONAL rather than REJECT. A concrete Bayesian-network or copula calculation with plausible correlations would settle whether the effect is large enough to change the practical message. I agree with the reader that this is the weakest assumption; no other concern, such as LLM calibration, is as central to the 99.3% claim, since that arithmetic uses illustrative probabilities rather than the Delphi outputs.","tokens_in":40026,"tokens_out":7378,"duration_ms":70702,"concrete_test":"Recompute the required per-component confidence p for the C2.2.1 fragment under a positively dependent structure. Build a Bayesian network in which C2.2.1.1, C2.2.1.2, and C2.2.1.3 share a common parent node representing organizational reliability (budget, team, and information quality), with conditional probabilities tuned so the pairwise correlation between leaf claims is approximately 0.4. Set each assigned probability (leaf posteriors, provenance side-claims, and W2.2.1) to p, propagate via the product method, and solve for the p that gives P(C2.2.1) = 0.95. If the required p drops below 0.985, the 99.3% headline is not robust to plausible dependencies and should be rephrased as conditional on independence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numerical result of §7.1, that achieving 95% confidence in the C2.2.1 fragment requires ~99.3% confidence in each of the seven assigned probabilities, is derived from Eq. (7) (product method) and Eq. (9) (sum of doubts method) under the explicit assumption that the side-claim and all sub-claims are mutually independent. The paper itself states in §7.3 that the three sub-claims C2.2.1.1, C2.2.1.2, and C2.2.1.3 are not fully independent because they share budget, team, and information. Positive dependence raises the probability of the conjunction relative to the product of the marginals. In the limiting case of perfect positive correlation, 95% confidence in the top-level claim is achievable with only ~95% confidence in each component, not 99.3%. Even moderate correlation, plausible for a shared safety team, will move the required p substantially. The executive summary and §7.1 present the 99.3% figure without this caveat front and center, so the challenge result can be read as a general property of conjunctive safety cases rather than a conditional result under an independence assumption that is already known to fail for the worked example. The paper does disclose the limitation later, but the sensitivity of the headline number to this assumption is not quantified, making the independence assumption the most load-bearing element of the paper's central quantitative claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the Assurance 2.0 methodology to confidence assessment for a frontier AI cyber-misuse inability safety case. It introduces a purely LLM-implemented Delphi method for eliciting leaf-node probabilities, validates this pipeline against 100 resolved Metaculus forecasting questions, and then uses Assurance 2.0's product and sum-of-doubts propagation methods on a seven-component fragment of the case. The paper's central quantitative finding is that achieving 95% confidence in the fragment's top-level claim requires approximately 99.3% confidence in each assigned component probability under both propagation methods. It also proposes a defeater-prioritisation scoring scheme and recommendations for communicating confidence to executives. The authors conclude that absolute probabilistic confidence is very difficult to achieve for frontier AI safety cases, while arguing that the process of producing such valuations can still improve safety cases.","tokens_in":40285,"tokens_out":7438,"duration_ms":71877,"significance":"If the central result is taken at face value, the paper provides a concrete, policy-relevant warning: even a small conjunctive safety case with only seven components demands near-perfect confidence in every leaf and side claim before a 95% top-level confidence can be reported. This is a useful counterweight to casual uses of quantitative safety-case confidence. The paper's strengths include its explicit statement of the propagation assumptions, its worked examples, the availability of code for the Delphi pipeline, and its unusually candid discussion of limitations, including the known violation of the independence assumption in the worked fragment. The two weak points are load-bearing: the 99.3% headline figure is conditional on mutual independence, and the calibration evidence for the LLM Delphi method rests on a confounded, post-hoc benchmark. The paper is a valuable applied contribution, but these two issues need to be addressed before the quantitative claims can be regarded as fully supported.","major_comments":[{"comment":"The headline claim that 95% confidence in the C2.2.1 fragment requires ~99.3% confidence in each of the seven components is derived under the explicit assumption that the side-claim and sub-claims are mutually independent. Section 7.3 concedes that C2.2.1.1, C2.2.1.2, and C2.2.1.3 share budget, team, and information, so they are positively correlated. With positive correlation, the probability of the conjunction is higher than the product of the marginal probabilities, so the per-component confidence needed to reach 95% top-level confidence can be substantially lower than 99.3%. Because this number is presented in the Executive Summary and Section 10.1 as the paper's main challenge result, the authors should quantify the sensitivity to plausible correlations (e.g., a simple common-factor model) or clearly reframe the 99.3% figure as an independence-conditional upper bound rather than a general property of conjunctive safety cases.","section":"Executive Summary and Section 7.1 (Eqs. 7, 9); Section 7.3"},{"comment":"The claim that the LLM-based Delphi pipeline is 'better calibrated' than Metaculus is not supported by the evidence as presented. The 100 questions are a post-hoc selected set of resolved Metaculus questions; Metaculus's forecasts incorporate information up to question close, whereas the LLM is restricted to an October 2023 knowledge cutoff, creating a systematic information asymmetry. No confidence interval or statistical test is reported for the calibration difference, and the selection of questions is not pre-registered. The paper acknowledges some of these limitations, but the stated conclusion still overstates the strength of the evidence. The authors should either temper the calibration claim to 'preliminary and not directly comparable' or provide a benchmark with matched information sets and a pre-specified question set.","section":"Section 4.2 (Figure 5E)"},{"comment":"The statement that 'for both methods, the confidence for assigned probabilities must be around 99.3%' conflates a sufficient condition with a necessary condition in the sum-of-doubts method. Equation (9) is a lower bound on the confidence of the top-level claim, so setting that lower bound to 0.95 gives a sufficient condition on the reported bound, not a necessary condition on the actual probability of the top-level claim. The product method, by contrast, gives an exact value under the independence assumption. The paper should clarify this logical asymmetry, since a reader may otherwise conclude that the sum-of-doubts method imposes a hard requirement that it does not actually establish.","section":"Section 7.1 (Eq. 9 and following)"}],"minor_comments":[{"comment":"In the paragraph beginning 'Where safety engineering standards exist', the phrase 'the developers’ of frontier AI systems' contains a typo and should read 'the developers of frontier AI systems'.","section":"Section 2.1"},{"comment":"The text cites p = 0.053 from a Mann-Whitney U test as a 'strong tendency', but this is not significant at the conventional 0.05 level. Please report an effect size and confidence interval, and avoid language that implies a robust finding.","section":"Section 4.2 (Figure 5C)"},{"comment":"The sentence 'we simply round such results up to 0' is imprecise: rounding -0.5 up to 0 is not standard rounding. The text should say that the lower-bound result is truncated or set to 0 when it is negative.","section":"Section 7.1 (Figure 7)"},{"comment":"Figures 9 and 10 are referenced in the text but appear only as placeholders in the manuscript; please ensure the actual visualisations are included in the final version.","section":"Section 9 (Figures 9 and 10)"},{"comment":"The choice of pseudo-counts (=10) for the credible intervals is not justified and is not varied in a sensitivity analysis. A short robustness check would strengthen the presentation of the uncertainty estimates.","section":"Section 4.2 (Figure 5D)"},{"comment":"The text says 'we show 3 examples of what these questions look like', but the figure caption lists five example questions. Please align the text and the figure.","section":"Section 4.2 (Figure 5F caption)"},{"comment":"The prioritisation formula depends on weighting factors W_Probability, W_Impact, and W_Effort, but the paper does not suggest how these weights should be set beyond noting that they vary by team. Although Section 8.6 acknowledges the limitation, a brief discussion of calibration or sensitivity analysis would make the proposed method more actionable.","section":"Section 8.4 (Eq. 10)"}],"recommendation":"major_revision","confidential_remarks":"This is a genuinely useful applied paper with an unusually honest treatment of its own limitations. The two issues that block acceptance are the unquantified sensitivity of the headline 99.3% figure to the independence assumption and the confounded calibration benchmark for the LLM Delphi method. Both are fixable within the manuscript's scope: the first with a simple correlation sensitivity analysis or a more careful framing, and the second with tempered claims or a matched-information benchmark. I would not reject the paper, but I would not accept it in its current form; the quantitative claims need to be made conditional in a way that is visible in the Executive Summary, not only in the later limitations sections."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one for the method, not the headline number. The genuinely new piece is the LLM-implemented Delphi pipeline: 50 GPT-4o-mini instances, five rounds, consensus stopping, inverse-variance weighting, code on GitHub. Benchmarking on 100 resolved Metaculus questions is a reasonable first check, and using it to elicit probabilities for the cyber-misuse defeaters is a sensible extension. The paper is also unusually candid about limitations: it flags the timing asymmetry in the Metaculus comparison, the selection bias, the dependence between sub-claims in §7.3, and the fact that Bloomfield and Rushby themselves warn against absolute valuations. That honesty is real.\n\nThe propagation arithmetic is correct under the stated assumptions. Seven conjunctive components at equal confidence p give top confidence 7p−6 by sum of doubts and p^7 by product; hitting 0.95 needs p≈0.993. That is a legitimate illustration of how steep conjunctive safety cases get.\n\nNow the soft spots, in proportion. The calibration claim is stronger than the evidence: 100 self-selected questions, Metaculus mean taken at the last open timestamp while the LLM is cutoff in October 2023, no confidence bounds on the comparison, and general forecasting is not the same as forecasting novel cyberattack defeaters. The paper acknowledges most of this, but the executive summary still says the pipeline is better calibrated, which overstates it. Second, the stress-test concern about independence is partly right but not fatal. The paper does disclose the independence assumption and even notes in §7.3 that the C2.2.1 sub-claims share budget, team, and information. What it does not do is quantify how much positive correlation lowers the required per-node confidence. The language \"must be around 99.3%\" is too strong: for sum of doubts it is a conservative sufficient bound, not a necessary condition, and for product positive dependence can relax the requirement. So the headline number is real but conditional. Third, \"Contextual Doubt\" is speculative and flagged as preliminary; fine as a suggestion, not a result.\n\nWho gets value: assurance practitioners, AI governance researchers, and people working on safety-case templates. It deserves a serious referee. I would send it out with a request for sensitivity analysis on the correlation assumption and a toned-down calibration claim, but desk rejection would be wrong.","headline":"A useful, honest application of Assurance 2.0 to a frontier AI cyber safety case: the LLM Delphi pipeline is the real contribution, and the 99.3% propagation result is conditional on assumptions the paper mostly discloses.","tokens_in":40883,"tokens_out":2967,"would_cite":true,"duration_ms":32325,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"To hold 95 percent confidence in a seven-part frontier AI cyber-safety case, each part must run about 99.3 percent confident.","keywords":["AI safety cases","frontier AI","confidence assessment","Assurance 2.0","cyber misuse","LLM-based Delphi method","defeaters","probabilistic reasoning"],"falsifier":"Recompute the required per-component confidence for the same C2.2.1 fragment after replacing the independence assumption with realistic positive correlations among the three sub-claims, or after allowing the top claim to be supported by any one of several diverse sub-arguments; if a 95% top-level confidence is then reached with per-component confidences well below about 99.3%, the challenge result does not generalise beyond its stated assumptions.","tokens_in":39760,"feed_emoji":"🛡️","tokens_out":10808,"duration_ms":93212,"temperature":0.7,"pith_summary":"The paper asks whether a frontier AI developer can attach a credible number to the claim that deploying a system poses no unacceptable cyber risk. It applies the Assurance 2.0 methodology to a seven-component fragment of a cyber-misuse 'inability' argument, propagating leaf-level probabilities upward by two rules: the product method and the sum-of-doubts method. The central quantitative finding is that reaching 95% confidence in the fragment's top claim requires roughly 99.3% confidence in every one of the seven assigned probabilities, under both propagation rules. The authors conclude that credible absolute probabilistic confidence in frontier AI safety cases is very hard to achieve, though the process of assigning and propagating confidence can sharpen the argument itself. They also propose an LLM-run Delphi process for generating reproducible leaf-node confidence values and a scoring rule for prioritising which doubts, or defeaters, to investigate first.","feed_headline":"95% certainty in an AI safety case needs 99.3% per component","feed_subtitle":"Even a 7-part cyber-safety argument requires near-certainty in every element, so frontier AI assurance faces a hard bar.","key_machinery":"The load-bearing machinery is a pair of arithmetic propagation rules applied to a conjunctive argument tree. In the product method, the confidence of a claim is the product of the confidences of its side-claim and all its sub-claims; in the sum-of-doubts method, doubt in a claim is bounded above by the sum of doubts in its supports, so confidence becomes the sum of component confidences minus the number of components, rounded up to zero if negative. A second piece of machinery is an LLM-based Delphi pipeline: 50 instances of a large language model act as expert forecasters, iterating over up to five rounds until their estimates converge below a standard-deviation threshold, with final probabilities weighted by each expert's consistency across rounds. This pipeline is used to estimate how likely each safety-case defeater is to be sustained, producing reproducible input probabilities for the propagation rules.","core_discovery":"Working from a cyber-misuse inability safety case, the paper isolates the C2.2.1 fragment in which a 7-day maximum time to detect a novel AI-enabled attack and take the system offline is decomposed into three sub-claims plus a side-claim, each supported by evidence. It assigns illustrative posterior probabilities $P(C|E)$ and side-claim strengths, then propagates them through the argument using Assurance 2.0's product and sum-of-doubts rules. Both rules give the same threshold for 95% top-level confidence: if all assigned probabilities are equal, each must be about $0.993$ ($99.3\\%$). Because the fragment is only a small part of a full safety case, the required per-element confidence for the whole case would be even higher, and the paper argues this makes reliable absolute probabilistic valuation of frontier AI safety cases very difficult. The paper therefore presents probabilistic valuation as most useful for comparative 'what if' analysis rather than as an absolute go/no-go number.","pith_inferences":["The 99.3% figure depends on an independence assumption that the paper itself flags as false for its own fragment; modelling the shared budget, team, and information among the three C2.2.1 sub-claims is the natural next computation and could move the threshold in either direction.","A regulator who simply demands 95% top-level probability on a large conjunctive safety case may be demanding near-impossible per-leaf numbers; a more workable requirement would target the argument's structure, such as diverse independent sub-arguments or a probabilistic top-level claim.","The LLM Delphi results on generic forecasting questions suggest a reproducible route to leaf probabilities, but the paper does not test that pipeline on actual safety-case claims about future AI behaviour; a hybrid human-LLM panel would be the direct next test."],"forward_implications":["A full frontier AI safety case, being much larger than the seven-component fragment, would require even higher per-element confidences to reach any fixed top-level threshold under these propagation rules.","The sum-of-doubts rule can return zero confidence on realistic fragments, so an absolute numerical verdict from these methods alone should not be treated as a deploy/no-deploy test.","The LLM Delphi method provides a reproducible, auditable procedure for generating leaf-node probabilities, which is useful for third-party evaluation even though the probabilities remain estimates.","Prioritising defeaters by their probability of being sustained, their impact on top-level confidence, and the estimated effort to resolve them can shorten the path to finding decisive flaws in a safety case.","Using probabilistic claims inside the argument itself, or structuring the case with diverse redundant sub-arguments, may lower the required per-element confidence, but the paper leaves that as future work."],"supporting_citations":[{"why":"Provides the Assurance 2.0 framework and the product and sum-of-doubts propagation equations used to obtain the 99.3% threshold.","marker":"[2]"},{"why":"Supplies the cyber-misuse inability safety case template, including the C2.2.1 fragment and the defeaters that are analysed.","marker":"[3]"},{"why":"Introduces the Delphi consensus method that the paper adapts to LLM experts for leaf-node probability elicitation.","marker":"[17]"},{"why":"Gives methodological guidance on Delphi panel design used to structure the iterative LLM forecasting pipeline.","marker":"[18]"},{"why":"Establishes safety cases for frontier AI as the object whose confidence the paper assesses.","marker":"[1]"},{"why":"Documents validation concerns with quantitative confidence methods, motivating the paper's cautious stance on absolute probabilistic values.","marker":"[6]"}],"fun_headline_variants":["95% AI safety case requires 99.3% per component","Why absolute AI safety confidence is hard to quantify","Safety case math: per-part 99.3% for 95% overall","Probabilistic AI safety: comparative use, not absolute","Frontier AI safety: near-certainty needed in every claim"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The numerical threshold assumes each claim is supported only by one fixed conjunction of independent sub-claims and a side-claim, so that confidence multiplies or doubts add; if the sub-claims overlap, share causes, or offer alternative support paths, the 99.3% figure changes.","fun_headline_variants_meta":{"raw":{"variants":["95% AI safety case requires 99.3% per component","Why absolute AI safety confidence is hard to quantify","Safety case math: per-part 99.3% for 95% overall","Probabilistic AI safety: comparative use, not absolute","Frontier AI safety: near-certainty needed in every claim"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000428,"raw_usage":{"total_tokens":2215,"prompt_tokens":994,"completion_tokens":1221,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":1133}},"tokens_in":610,"tokens_out":1221,"duration_ms":10766,"temperature":1.0,"reasoning_tokens":1133,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T17:58:16.774605+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the required per-component confidence for the same C2.2.1 fragment after replacing the independence assumption with realistic positive correlations among the three sub-claims, or after allowing the top claim to be supported by any one of several diverse sub-arguments; if a 95% top-level confidence is then reached with per-component confidences well below about 99.3%, the challenge result does not generalise beyond its stated assumptions.","supporting_citations":[{"cited_title":"Assessing Confidence with Assurance 2.0","cited_arxiv_id":"2205.04522","evidence_quote":"Provides the Assurance 2.0 framework and the product and sum-of-doubts propagation equations used to obtain the 99.3% threshold."},{"cited_title":"Dalkey and O","cited_arxiv_id":null,"evidence_quote":"Introduces the Delphi consensus method that the paper adapts to LLM experts for leaf-node probability elicitation."},{"cited_title":"Khodyakov, S","cited_arxiv_id":null,"evidence_quote":"Gives methodological guidance on Delphi panel design used to structure the iterative LLM forecasting pipeline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents validation concerns with quantitative confidence methods, motivating the paper's cautious stance on absolute probabilistic values."}],"review_version":1}