{"id":"f809edc7-b4f1-4550-b1e5-b3fbab4a263b","arxiv_id":"2607.11999","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A coordinated eight-component insurance infrastructure could make affirmative AI-agent coverage with billion-dollar limits achievable by 2030.","lead":"This report argues that insurers can underwrite AI agents at billion-dollar scales by 2030, but only if the industry builds shared infrastructure for data, standards, pricing, and claims. It lays out an eight-component 'AI insurance stack' and special instruments for catastrophic frontier-AI risks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pricing premise is unvalidated: performance-evaluation and red-team scores are assumed to behave like actuarial data (§II.6.A, §II.5.B), but no evidence links them to realized losses.","rationale":"Reader identified same concern. My stress-test confirms. The report's central claim has multiple supporting assumptions (coordination, capacity, data pooling), but the most load-bearing is that forward-looking evals can substitute for actuarial data. This is explicitly asserted in §II.6.A and §II.5.B. It is not derived from data or independent validation. The report's own appendix and footnote qualify the available data, and the AIUC-1 standard conflict is disclosed but does not cure the absence of calibration. A conditional verdict is appropriate: the blueprint is actionable but the pricing premise must be demonstrated before the 'achievable by 2030' assertion can be accepted. No change to verdict from my pass.","tokens_in":43044,"tokens_out":3128,"duration_ms":35549,"concrete_test":"Retrospective calibration test: take a cohort of enterprise AI-agent deployments with pre-policy performance-evaluation scores (hallucination/refusal rates, red-team severity) and at least 12 months of subsequent claims/loss data (frequency, severity, loss ratio). Rank by score; compute realized loss ratio and mean severity per decile; report correlation (e.g., Spearman) and a simple credibility-weighted pricing table. If top-decile vs bottom-decile loss ratios do not separate, or if the correlation is not significantly positive, the §II.6.A formula is not supported. For robustness, re-score a random subset with independent hidden probes to test whether public scores are inflated by overfitting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('billions of coverage by 2030') requires insurers to earn premiums adequate for losses. The report's pricing engine supplies this via expected loss = severity × probability × usage, where probability is 'filled' by performance evaluations described as 'structurally similar to an actuarial table' (§II.6.A) and red teaming is held to bound tail severity (§II.5.B). This substitution is load-bearing but unvalidated. Test scores measure behavior under evaluation distributions; they predict real-world claim frequency only if test conditions are representative and scores are not gamed. The report offers no retrospective or prospective evidence of such calibration. Its own Appendix 1 cautions that the incident/usage index 'is too noisy and too short in temporal coverage to support strong conclusions' and that construct validity is fragile; public incident data rarely includes dollar losses (n.36). Worse, evals used for pricing are prone to Goodhart: once scores become rating inputs, policyholders and vendors can optimize for the tests rather than for loss reduction. Red teaming reveals possible failure modes, not their probability. Thus if eval scores do not track loss ratios, the formula has no empirical basis, pricing is guesswork, and the 'achievable by 2030' conclusion is unsupported. This is not a claim of fraud; it is an identified gap in the argument's load-bearing empirical premise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The report argues that affirmative insurance coverage for AI agents, with limits reaching billions, is achievable by 2030 if the insurance industry coordinates to build an eight-component 'AI insurance stack.' The stack spans incident data collection, CAT modeling, standards, contract design, risk selection, pricing, monitoring, and claims management. The paper draws on historical analogies (UL, the Closed Claims Project, nuclear insurance, cyber insurance) and proposes technical underwriting tools such as performance evaluations, red teaming, and telemetry-based monitoring. It also sketches a separate 'AI CAT' layer for societal-scale risks, including mutuals, CAT bonds, and government backstops. The central claim is explicitly conditional on insurers being able to price and select risk without mature actuarial data, using evaluation scores and red-team results as substitutes. The report is transparent about some limitations, particularly in Appendix 1 and footnote 4, but those limitations bear directly on the load-bearing pricing premise.","tokens_in":43360,"tokens_out":2778,"duration_ms":30376,"significance":"This is a timely and unusually concrete policy/industry blueprint, with detailed recommendations assigned to specific actors (carriers, reinsurers, modelers, regulators). Its value lies in framing AI agent risk as an insurability problem and proposing a coordinated infrastructure response, rather than in new empirical results. The historical precedents are relevant and well chosen, and the report is honest about several of its own evidential weaknesses — the fragile incident/usage index, the absence of dollar-loss data in public incident reports, and the linear approximation of an unavailable IMF model. If the proposed stack were built and its pricing premise validated, the report could serve as a useful planning document for insurers and policymakers. However, the central 'achievable by 2030' claim rests on an unvalidated substitution of evaluation scores for actuarial data, and the GDP-loss estimate is based on an unverifiable approximation. These need to be either substantially supported or carefully de-emphasized before the report can carry its conclusions.","major_comments":[{"comment":"The pricing formula (expected loss = severity × probability × usage) relies on performance evaluations and red teaming to estimate probability and severity, with §II.6.A asserting these outputs are 'structurally similar to an actuarial table' and §II.5.B proposing them as substitutes for loss data. This is load-bearing for the 'billions by 2030' conclusion, yet the report provides no evidence that evaluation scores or red-team findings predict claim frequency or severity. Appendix 1 acknowledges the incident/usage index is 'too noisy and too short in temporal coverage to support strong conclusions' and that its construct validity is 'fragile,' and n.36 admits public incident data 'rarely if ever include exact dollar costs.' The report also does not address Goodhart dynamics: once scores become rating inputs, they will be optimized. A concrete validation strategy, pilot-study results, or","section":"§II.6.A and §II.5.B"},{"comment":"The $200 billion US GDP figure is presented in the Key Takeaways and Extended Summary as an estimate of what depends on insurers, but it is derived from a linear approximation of an IMF model to which the authors do not have access, under an explicit quasi-linearity assumption. The footnote discloses this, but the headline presentation does not. Because this number is used as a policy stake, it should be labeled as an illustrative sensitivity calculation, not an estimate, and should not appear without its caveats in the executive summary.","section":"Footnote 4 / Extended Summary"},{"comment":"Several authors are employees of the Artificial Intelligence Underwriting Company (AIUC), which maintains AIUC-1, one of the four standards the report discusses and implicitly recommends for underwriting use. The same company produced the GDP estimate. Footnote 29 discloses the AIUC-1 conflict and states other authors take no position on its relative merits, which is good practice, but the report would benefit from a more prominent competing-interests statement and from making clear which specific claims and recommendations are authored by AIUC-affiliated authors. This is not a claim of misconduct, but a policy report recommending a standard with a direct financial stake should make the conflict visible to a reader who only reads the executive summary.","section":"§I.2, §II.3.D, author list"},{"comment":"The historical analogies (UL, Closed Claims Project, INPO) are used as implicit support for the claim that a coordinated stack will reduce losses and enable profitable underwriting. The paper does not address a key disanalogy: those precedents involved physical systems with relatively stable failure modes, whereas AI agents are continuously updated, general-purpose, and susceptible to adversarial optimization of the very evaluation metrics proposed as pricing inputs. At minimum, the report should explicitly discuss why Goodharting of standardized evaluations is not expected to replicate the UL experience, and what safeguards (e.g., hidden test sets, surprise audits) would be deployed.","section":"§I.5 and §II.3.E"}],"minor_comments":[{"comment":"The appendix is commendably honest about biases, but the interpretation section could more explicitly say that the 80% decline in incident-to-usage ratio is not a reliable trend for underwriting. Consider adding a one-line summary box for practitioners.","section":"Appendix 1"},{"comment":"The footnote beginning 'Reviewers note that...' appears to be an internal review note that was accidentally left in the manuscript. Please remove or convert to a normal authorial remark.","section":"§II.1.B, footnote 14"},{"comment":"The paragraph beginning 'This stifles institutional learning...' appears to be cut off mid-sentence ('there is n...'). Complete the sentence and ensure no accidental truncations elsewhere.","section":"§II.8.D"},{"comment":"The claim that 'we expect no more than fifty people are studying them full-time, worldwide' is an unsourced numerical claim. Either provide a citation or remove the specific number.","section":"§II.2.A"},{"comment":"The comparison table would be stronger with a column or footnote explicitly listing the disclosure/conflict status of each standard's relationship to the authors. The current footnote is in the text, not the table.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"This is a policy/blueprint paper, not an empirical contribution. The main risk is overclaiming: the headline 'achievable by 2030' and the $200bn GDP figure are both more conditional than the executive summary suggests. The authors are clearly aware of the fragility of their empirical inputs (Appendix 1, footnote 4, n.36). With a careful reframing of those two items as forward-looking scenarios rather than demonstrated estimates, and with explicit discussion of the evaluation-to-loss calibration problem, I believe the paper could be publishable as a serious policy proposal. The AIUC conflict is disclosed, but the prominence of AIUC-derived products and estimates in the paper warrants a stronger editorial note. I am not recommending rejection because the core proposal is constructive, detailed, and largely framed as a call to action rather than a proof."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead this if you care about how insurance could shape AI risk governance. It's not an academic paper; it's an industry blueprint with a policy argument, but it's a serious one. What's new is the synthesis: eight components—incident data, CAT modeling, standards, contract design, risk selection, pricing, monitoring, claims—organized into a stack specifically for agentic AI, plus an AI CAT menu of mutuals, CAT bonds, liability regimes, and government backstops. Historical precedents like Underwriters Laboratories and the Closed Claims Project are used well. The authors are also refreshingly candid: they flag that the incident-to-usage index is too noisy and short for strong conclusions, that the GDP estimate is a linear approximation of an IMF model they don't have, and they disclose that AIUC employees helped write it while recommending the AIUC-1 standard. That transparency matters.\n\nThe soft spot is real and load-bearing. The pricing engine relies on performance-evaluation scores and red teaming as 'quasi-actuarial data' (§II.6.A), but there's no evidence these signals predict actual loss ratios. Test conditions aren't necessarily representative; once scores become rating inputs, Goodharting is likely. Red teaming bounds severity only if the discovered failures approximate real-world tail events. The report acknowledges the data are fragile but then treats the eval-based pricing as the foundation of the 2030 target. It's also not clear where the loss-severity priors come from for novel agent failure modes. This isn't fraud—it's an honest gap, and the authors say so implicitly—but it's the difference between a plausible case and a demonstrated one.\n\nThe conflict of interest is disclosed but worth weighing: three authors, including the lead, work for the company that maintains AIUC-1, and the report recommends that standard. That doesn't invalidate the analysis, but independent validation and public release of Appendix 1's data and code would raise my confidence substantially.\n\nBottom line: this is a genuinely useful read for insurers, policymakers, and AI-governance researchers. The stack is coherent and the recommendations are concrete. The empirical support for the headline claim is thin, but the paper is honest about many of its limits. I'd send it to peer review, flag the pricing premise for scrutiny, and ask for the supporting data. It's the kind of big-picture, actionable synthesis that gets more useful when people push on the weak points.\n\nRecommendation: engage with it, but read the appendix caveats first.","headline":"A serious, well-structured blueprint for AI-agent insurance with a load-bearing but unvalidated pricing premise; deserves a real referee.","tokens_in":43969,"tokens_out":2672,"would_cite":true,"duration_ms":29487,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Affirmative insurance for AI agents with billion-dollar limits is achievable by 2030 if insurers build shared infrastructure.","keywords":["AI insurance","AI agents","silent coverage","affirmative coverage","accumulation risk","catastrophe modeling","red teaming","insurability"],"falsifier":"Collect a cohort of several hundred insured AI-agent deployments, record pre-deployment performance-evaluation scores and red-team findings, then track claims frequency and severity for two years. If scores show no correlation with claims after controlling for usage and sector—or if scores are easily gamable—the stack's pricing engine is falsified and the 2030 achievability claim collapses.","tokens_in":42951,"feed_emoji":"🛡️","tokens_out":4447,"duration_ms":45120,"temperature":0.7,"pith_summary":"The report argues that the emerging AI agent economy—projected to move trillions of dollars by 2030—is currently insured through unpriced 'silent coverage' inside cyber, professional, and general liability policies, leaving insurers exposed and businesses unprotected. It claims insurability is trending the wrong way: agent capabilities are outpacing reliability, incident severity is climbing, and dependence on a few foundation-model providers creates correlated-loss risk. Against that backdrop, the report's central claim is that affirmative, explicitly priced AI coverage with enterprise limits in the billions is achievable by 2030, but only if the industry coordinates on shared infrastructure. The blueprint is an eight-component 'AI insurance stack'—incident data pooling, catastrophe modeling, standards, contract design, risk selection, pricing, monitoring, and claims management. A sympathetic reader would care because, if true, insurance becomes a lever for safe adoption of a transformative technology rather than a brake on it.","feed_headline":"Billion-dollar AI agent coverage is achievable by 2030","feed_subtitle":"A coordinated eight-part insurance stack—shared data, standards, stress tests—can turn unpriced AI risk into insurable products.","key_machinery":"The central object is the 'AI insurance stack,' eight interlocking components spanning the whole policy lifecycle, with the claims-to-underwriting feedback loop as the connective tissue. The identity that makes the argument run is the pricing formula expected loss = severity × probability × usage, where severity is bounded by red teaming, probability is estimated from performance-evaluation scores used as 'quasi-actuarial data,' and usage is measured through telemetry partnerships. Standards and certifications act as the underwriting shortcut that ties the components together, much as fire-rating schedules let property insurers price communities at a glance.","core_discovery":"The paper's central discovery is a design: a complete insurance infrastructure that lets carriers underwrite AI agents before actuarial data matures. The load-bearing move is to treat performance evaluations and red teaming as 'quasi-actuarial data'—structured, forward-looking signals about failure rates and tail severity that can be substituted for historical loss tables. On this basis, the paper argues that a coordinated eight-component stack can turn today's silent, unpriced AI exposure into affirmative coverage: explicit definitions and exclusions, system-specific risk selection, usage-based pricing, ongoing monitoring, and a claims-forensics feedback loop. The paper also argues that soc","pith_inferences":["Testable extension: if evaluation scores and red-team results do not predict subsequent claims, the pricing engine fails quietly; a cohort study linking pre-deployment test scores to claims over several hundred policies would settle this within a few years.","The report's declining incident-per-usage claim implies premiums should track AI usage volume more than revenue or headcount; usage-telemetry pilots can test whether usage is indeed the better exposure basis.","The cold-start logic implies consolidation around shared standards and databases will need to happen before limits can reach the billions, and that antitrust constraints on joint policy language may be the quiet bottleneck.","Government disclosure mandates, which the report treats as a complement to industry action, may end up being the binding condition; voluntary pooling alone has repeatedly failed in cyber."],"forward_implications":["Silent AI exposure inside legacy cyber, D&O, E&O, and general liability policies would be replaced by affirmative endorsements or standalone policies, giving portfolio managers and regulators visibility into concentration.","Premiums would differentiate on the basis of system specifications, safeguards, and evaluation scores, so heavy users stop subsidizing light users and responsible deployers pay less.","Accumulation risk—especially dependence on a handful of upstream model providers—would be carved out and managed with exclusions, sub-limits, and catastrophe models rather than blanket denials.","Standards and third-party audits would become de facto insurability signals, creating a market race-to-the-top in agent safety similar to crashworthiness ratings in auto insurance.","Societal-scale frontier AI risk would move to alternative risk-transfer vehicles—mutuals, CAT bonds, and government backstops—because private markets alone cannot carry that tail."],"fun_headline_variants":["Blueprint turns AI agents into insurable risk by 2030","AI insurance stack: from unpriced risk to billion-dollar coverage","Insurers can underwrite AI agents by 2030 with this blueprint","AI agents become insurable with an eight-part stack"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Insurers can obtain reliable, non-gamable signals—from performance evaluations, red teaming, and telemetry—that predict expected AI-agent losses well enough to price and select risk, despite the absence of mature actuarial data.","fun_headline_variants_meta":{"raw":{"variants":["Blueprint turns AI agents into insurable risk by 2030","AI insurance stack: from unpriced risk to billion-dollar coverage","Insurers can underwrite AI agents by 2030 with this blueprint","AI agents become insurable with an eight-part stack"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00075,"raw_usage":{"total_tokens":3208,"prompt_tokens":808,"completion_tokens":2400,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":2329}},"tokens_in":552,"tokens_out":2400,"duration_ms":14609,"temperature":1.0,"reasoning_tokens":2329,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:42:36.270856+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a cohort of several hundred insured AI-agent deployments, record pre-deployment performance-evaluation scores and red-team findings, then track claims frequency and severity for two years. If scores show no correlation with claims after controlling for usage and sector—or if scores are easily gamable—the stack's pricing engine is falsified and the 2030 achievability claim collapses.","supporting_citations":[],"review_version":2}