Pith. sign in

REVIEW 2 major objections 6 minor 49 references

The paper claims that fragmented AI safety thresholds can be harmonized into a small set of auditable floors, with a non-zero TLO completion rule for cyber misuse and a 5x/3-month rate break for automated AI R&D.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 21:16 UTC pith:32HA2Z65

load-bearing objection Concrete, honest floors for cyber and AI R&D; the harm-model numbers behind them are shakier than the floors themselves. the 2 major comments →

arxiv 2607.16112 v1 pith:32HA2Z65 submitted 2026-07-17 cs.AI

Harmonizing AI Safety Thresholds

classification cs.AI
keywords AI safety thresholdscapability thresholdscyber riskbioriskautomated AI R&Dexpected harmbenchmark evaluationharmonization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to show that the safety thresholds published by frontier AI companies, which differ in scope and specificity, can be translated into shared, third-party-checkable minimum floors. For cyber misuse, it proposes a floor that triggers whenever a model completes at least one full multi-step attack chain on a standardized evaluation range at a fixed compute budget, a condition that two April 2026 frontier models would have met. For automated AI R&D, it proposes a floor based on observable benchmark progress: a model trips the threshold if it drives capability growth at least five times the baseline trend for at least three months. For biorisk, the paper argues that no defensible quantitative floor can be stated yet, but this is itself a useful diagnosis because it pinpoints the missing piece: per-stage uplift estimates that are unsaturated, community-validated, and full-chain. The method matters because it lets regulators and competitors audit whether thresholds have been crossed without trusting internal tests.

Core claim

The central claim is that the expected-harm framework for misuse risks and the rate-of-progress framework for automated R&D can turn heterogeneous threshold language into common quantitative floors. Applying the framework yields two operational floors: a cyber floor defined by non-zero full-chain completion on the TLO evaluation at a pinned 100M-token budget, with the Clopper-Pearson lower confidence bound above zero across pre-registered trials; and an R&D floor defined by a measured improvement rate at least five times the baseline trend, sustained for at least three months. The same analysis shows that no analogous bio floor can be responsibly fixed until an unsaturated, independently val

What carries the argument

The expected-harm equation E[H_j] = N_j × P_success,j × h_j for misuse risks, combined with the release-condition exposure scalar s_r that multiplies total AI-enabled harm, and, for automated R&D, the rate ratio r_n/r_0 comparing a model's measured improvement rate to the baseline trend with a three-month sustain condition. These turn threshold language into quantities that can be computed from public benchmark scores and evaluation runs.

Load-bearing premise

The load-bearing premise is that TLO full-chain completion rates measure the probability of real-world cyber-attack success; if the evaluation range is not representative of operational intrusion paths, or if the paper's uncited release-condition scalars are wrong, the quantitative harm figures and release verdicts do not follow.

What would settle it

A replicated study showing no correlation between TLO completion rates and real-world intrusion outcomes, or an independent audit of a trusted-access regime that measures effective exposure above the implied ~1.5% of public-API level for a Mythos-class model, would disprove the central quantitative claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If adopted, the cyber floor would have triggered enhanced safeguards for both April 2026 frontier models concurrently, closing the gap where one company's framework fired and the other's did not.
  • The R&D floor is checkable by any third party with access to benchmark scores and release dates, making threshold enforcement feasible without proprietary evaluations.
  • Biorisk harmonization shifts from arguing about threshold wording to commissioning the missing per-stage uplift elicitation, which is a concrete, fundable task.
  • Companies can keep stricter internal thresholds while sharing the common floor, reducing the race-to-the-bottom pressure.
  • The binary non-zero TLO completion rule sidesteps disputes over arbitrary continuous probability cutoffs while still marking a real capability regime transition.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • One testable extension is to apply the 5x/three-month rule retrospectively to historical frontier-model releases to see which prior models would have tripped it; this would calibrate the floor's strictness against the public record.
  • The same rate-of-progress logic could be adapted to domains other than AI R&D, such as autonomous scientific discovery, whenever a composite capability index exists.
  • The cyber floor's non-zero rung may quickly become uninformative as models routinely clear several completions; the paper's two higher rungs, autonomous discovery and end-to-end hardened intrusion, will then do the real work, and those are harder to audit.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper proposes a framework for converting frontier AI companies' heterogeneous capability thresholds into comparable quantitative 'floors' across three risk domains. For cyber, it combines an expected-harm model (N×P×H) with TLO cyber-range evaluations, proposing a minimum floor of non-zero full-chain TLO completion at a pinned 100M-token budget, and claims that public API access to current models generates additional expected harm above the USD 1B annual-harm reference even after neutralizing low-tier commodity pathways. For biorisk, it applies the same methodology but concludes that no defensible quantitative floor can yet be set, identifying missing full-chain validated KRIs and per-stage uplift elicitation as the binding gap. For automated AI R&D, it proposes a floor that triggers safeguards when a model causes AI progress at least five times the baseline trend for at least three months, which it argues is consistent with the existing threshold language of Anthropic, Google DeepMind, and OpenAI under stated assumptions.

Significance. If the framework and its calibrations hold, the paper makes a useful contribution to threshold harmonization: it gives a concrete translation procedure from heterogeneous policy language to auditable quantitative triggers; the cyber TLO tripwire and the AI R&D rate-of-progress floor are falsifiable and partially third-party checkable; and the biorisk analysis is a clear negative diagnosis that tells the field what evidence is missing. The paper is unusually explicit about its limitations, consistently labels author-calibrated priors as such, and does not overclaim in the biorisk section. The main value lies in the translation methodology and the proposed floors, not in the point estimates of dollar harm. The significance is real but tempered by calibration uncertainty in the cyber harm model and by the conditional nature of the AI R&D threshold equivalence.

major comments (2)
  1. [§3.7, Tables XIV/XV/XVII, §3.11 and Appendix B.5] The 'Public-API binary' conclusion—that public API access exceeds the USD 1B annual-harm reference across the whole anchor range—is not robust to the P_success uncertainty that the paper itself identifies. The headline calculation uses AT3 P_AI = 0.60 from TLO 6/10, while the only human-validated kill-chain decomposition (Table XVII, from Barrett et al. OC3 SME Ransomware) gives 0.084, a sevenfold gap. Since E[H] scales linearly with P_success, switching to 0.084 lowers the AT3–AT5 contribution from about $18.25B to roughly $2.5B, and a real-world P_success near 0.01 would put public API below the USD 1B reference. Table XV varies only the aggregate baseline H0, not the capability proxy, and no Monte Carlo propagation is provided. At minimum, the paper needs a sensitivity analysis over AT3 P_AI and should soften the claim that the Public-API binary is robust, or justify why TLO 6/10 is a
  2. [§5.1 and Appendix A, Eq. (14)] The claim that the 5×/3-month floor is the minimum common denominator of the three companies' automated AI R&D thresholds rests on two stated assumptions: that 'several months' in OpenAI's threshold means exactly three months, and that r2024 = r2018-2024. These assumptions are not derived from evidence; 'several' in ordinary usage can denote four to six months, and OpenAI's own example ('sped up to just 4 weeks') suggests a shorter interval. If m>3, OpenAI's threshold implies more additional progress than Anthropic's, so the proposed floor is not the unique common denominator but one candidate harmonization. The paper should either provide a justification for the 3-month interpretation or present the floor as a possible harmonized value, rather than as the definite minimum extracted from the companies' language.
minor comments (6)
  1. [§5.3] The illustrative calculation for r0 uses a two-point slope between Claude 3 Opus and Claude 3.7 Sonnet, whereas §5.2 says a trend line should be fitted over time. A two-point estimate is highly sensitive to endpoint choice; the paper should either fit a proper regression or explicitly label the calculation as an illustration and provide sensitivity to the baseline window.
  2. [§3.7] The text states that an observed 2/10 carries an exact interval of 'roughly 0.03 to 0.56'. The Clopper–Pearson 95% interval for 2/10 is approximately 0.025 to 0.556, so the reported lower endpoint is slightly off.
  3. [Table VI] The entry '0/10≈0.00' uses an approximation symbol where the value is exactly zero; minor inconsistency.
  4. [§4.2.6] The proposed biorisk floor uses 'Anthropic's 2× signal' as an example threshold, but no justification is given for why 2× is the appropriate harmonized value across companies, nor how it relates to the 5×/3-month cyber or R&D floors.
  5. [References] Several references are incomplete or inconsistently formatted, e.g., 'Zhang et al. Bountybench, 2025a' lacks a full author list, and 'TitanCA' has no individual authors listed. Please standardize.
  6. [§3.9 and §5.3] The model referred to as 'Mythos Preview' in the cyber section and 'Claude Mythos Preview' in the R&D section should be named consistently, and its manufacturer (Anthropic) should be identified at first use in each section.

Circularity Check

1 steps flagged

Automated AI R&D floor is a restatement of OpenAI's threshold with 'several months' set to 3; cyber/bio domains are not circular.

specific steps
  1. renaming known result [Section 5.1–5.2 and Appendix A (Eqs. 14–16)]
    "We examine the thresholds defined by the main frontier AI companies and extract the minimum common denominator, which they have implicitly accepted. We then derive a quantitative threshold that captures this minimum common denominator... Under the assumption that 'several months' in OpenAI's formulation means three months, the Anthropic and OpenAI automated AI R&D thresholds are equivalent... Equating the two shows that the rate proposed by OpenAI, maintained for 3 months, yields the same additional progress as the rate proposed by Anthropic, maintained for 1 year."

    The proposed floor 'at least five times faster than trend during at least three months' is OpenAI's threshold ('1/5th the wall-clock time... sustainably for several months') with 'several months' stipulated to be 3 and the rate expressed relative to trend. Appendix A does not derive 5 and 3 months from progress data; it assumes m=3 and r2024=r2018-2024, then obtains equality. Thus the harmonized floor is the input policy language restated in quantitative coordinates; the claim that it 'covers' all three companies follows by construction, not from independent evidence. This is a transparent translation/operationalization rather than an independent derivation.

full rationale

The paper is largely self-contained and transparent. The cyber floor is independently grounded in the empirical zero-to-nonzero TLO transition, and the paper explicitly says its justification does not depend on the dollar calculation; the biorisk section is a diagnostic that identifies missing evidence rather than a derived floor. No load-bearing self-citation is present: Murray et al. and Barrett et al. are not authored by this paper's authors. The one partial circularity is the automated AI R&D threshold: the 5x/3-month formulation is extracted from OpenAI's threshold language under a stipulated reading of 'several months' and an assumed equality of 2024 and 2018-2024 trends, and Appendix A's equivalence is algebraically forced by those assumptions. Because the paper openly frames this as extracting a minimum common denominator and because the other two domains retain independent content, the score is moderate rather than severe.

Axiom & Free-Parameter Ledger

6 free parameters · 8 axioms · 0 invented entities

The paper introduces no new physical or mechanistic entities. Its analytical constructs—AT1–AT5 attacker tiers, exposure scalars, and the kill-chain phase matrix—are re-labeled or adapted from cited prior work and are already parameterized in the free-parameter ledger. The main burden is therefore carried by author-calibrated priors and structural modeling choices, not by invented entities.

free parameters (6)
  • Cyber baseline pathway allocation (N0, P0, h per AT1–AT5) = Back-fit to sum to USD 500B; Table XIII
    Ten rows are constructed to sum to the global baseline anchor, so the per-pathway entries are not independent empirical estimates.
  • AI-enabled uplift multipliers (q and P_AI per pathway) = Table XIV, e.g. q=0.10, P_AI=0.60 for AT3
    Author-calibrated priors for how much AI adds attack-equivalent volume and success probability; not sourced to incident data or elicitation.
  • Release-condition exposure scalars s_r,j = Table XVI, e.g. Trusted API 0.001–0.090 depending on tier
    Explicitly described as uncited author priors; they determine the Trusted-API verdict and the required safeguard performance figures.
  • Kill-chain phase probabilities p_k,j = Table XVII, AT3 product ≈ 0.084
    Mostly author-calibrated, with AT3 values borrowed from Barrett et al. OC3 SME Ransomware; the column product differs from the TLO-based P_success by ~7×, flagged as open.
  • AI R&D baseline trend slope r0 = 16.36 ECI points/year (Claude 3 Opus → Claude 3.7 Sonnet)
    The trend slope is chosen from a specific company's model series and time window; alternative choices (2018–2024 vs 2024, industry vs company) would change the floor determination.
  • Threshold constants for AI R&D: 5× and 3 months = 5× trend for at least 3 months
    These constants are taken from OpenAI's threshold language and the assumption that 'several months' means three months; they are not independently derived from risk or evidence.
axioms (8)
  • domain assumption Expected harm is the correct primitive for misuse thresholds.
    Adopted from Koessler et al. (2024) without independent justification; drives the entire cyber and bio framework.
  • domain assumption Linear release-condition scaling: E[HAI,r] = s_r × E[HAI].
    Section 2.1, Eq. (3). Treats exposure as a simple multiplicative scalar per release condition, with no interaction or nonlinearity.
  • domain assumption TLO full-chain completion rate is a valid proxy for real-world P_success.
    Section 3.6 uses 6/10 TLO completions as P_AI for AT3 pathways. The paper acknowledges it is a proxy, not a direct operational estimate.
  • domain assumption Kill chains are AND-gates at chokepoints, with OR-gates at K2/K5.
    Sections 3.3.1 and 4.2.2. The AND/OR structure is a modeling choice; the paper notes the OR-gate assumption for bio acquisition and delivery is contested.
  • domain assumption ECI benchmark scores quantify AI progress and can be fit with a trend.
    Section 5.2. The R&D floor depends on ECI as a measure of progress and on the chosen trend window; benchmark and harness drift are acknowledged limitations.
  • domain assumption Defensive AI uplift is set to zero in the cyber calculation.
    Section 3.11 states the model estimates gross offense and sets defensive AI to zero; the paper acknowledges this is not sign-neutral.
  • domain assumption For biorisk, the modern baseline of successful sophisticated bioattacks is effectively zero and cannot be empirically validated.
    Section 4.2.3. This frames the whole bio analysis as counterfactual and makes validation impossible; the paper is explicit about this but relies on it.
  • ad hoc to paper 'Several months' in OpenAI's threshold means three months, and r2024 = r2018-2024.
    Appendix A needs these equalities to show Anthropic and OpenAI thresholds are equivalent. The paper flags the first as an assumption but does not defend it with evidence.

pith-pipeline@v1.3.0-alltime-deepseek · 29144 in / 12768 out tokens · 110025 ms · 2026-08-01T21:16:50.621440+00:00 · methodology

0 comments
read the original abstract

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.

Figures

Figures reproduced from arXiv: 2607.16112 by Luis F. Lafuerza, Markov Grey, Matthew Ball, Wilber Sean Anterola.

Figure 1
Figure 1. Figure 1: Illustration of the automated AI R&D threshold. [PITH_FULL_IMAGE:figures/full_fig_p018_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

49 extracted references · 10 linked inside Pith

  1. [1]

    2024 , publisher =

    Alexanian, Tessa and Langenkamp, Max , title =. 2024 , publisher =

  2. [2]

    2026 , eprint=

    Estimating the Social Cost of Corporate Data Breaches , author=. 2026 , eprint=

  3. [3]

    2025 , month = may, type =

    Claude Opus 4 System Card , institution =. 2025 , month = may, type =

  4. [4]

    Responsible Scaling Policy, Version 3.0 , year =

  5. [5]

    2026 , month = apr, type =

    Claude Mythos Preview System Card , institution =. 2026 , month = apr, type =

  6. [6]

    Responsible Scaling Policy, Version 3.1 , year =

  7. [7]

    Responsible Scaling Policy, Version 3.2 , year =

  8. [8]

    and Murray, M

    Barrett, S. and Murray, M. and Quarks, O. and Smith, M. and Krys, J. and Campos, S. and Tlaie Boria, A. and Touzet, C. and Hayrapet, S. and Heiding, F. and Nevo, O. and Swanda, A. and Aguirre, J. and Gershovich, A. B. and Clay, E. and Fetterman, R. and Fritz, M. and Juarez, M. and Mavroudis, V. and Papadatos, H. , title =. 2025 , howpublished =. 2512.0886...

  9. [9]

    2025 , institution =

  10. [10]

    and others , title =

    Dauparas, J. and others , title =. Science , volume =

  11. [11]

    2025 , url =

    Internet Crime Report 2024 , institution =. 2025 , url =

  12. [12]

    and Feldman, T

    Feldman, J. and Feldman, T. , title =. 2025 , howpublished =. 2509.02610 , archivePrefix=

  13. [13]

    and Payne, W

    Folkerts, L. and Payne, W. and Inman, S. and Giavridis, P. and Skinner, J. and Deverett, S. and Aung, J. and Zorer, E. and Schmatz, M. and Ghanem, M. and Wilkinson, J. and Steer, A. and Hong, V. and Wang, J. , title =. 2026 , howpublished =. 2603.11214 , archivePrefix=

  14. [14]

    Frontier Safety Framework, Version 3.0 , year =

  15. [15]

    2025 , month = oct, howpublished =

    Gopal, Anjali and Guest, Oliver and Besiroglu, Tamay and. 2025 , month = oct, howpublished =

  16. [16]

    Biothreat Creation Process Framework , year =

  17. [17]

    2025 , howpublished =

    Hattoh and others , title =. 2025 , howpublished =. 2505.17154 , archivePrefix=

  18. [18]

    Anson Ho and Jean-Stanislas Denain and David Atanasov and Samuel Albanie and Rohin Shah , year=. A. 2512.00193 , archivePrefix=

  19. [19]

    2025 , url =

    Cost of a Data Breach Report 2025 , institution =. 2025 , url =

  20. [20]

    and others , title =

    Inagaki, T. and others , title =. 2023 , howpublished =. 2304.10267 , archivePrefix=

  21. [21]

    2025 , howpublished =

    Khlaaf, Heidy and West, Sarah Myers , title =. 2025 , howpublished =. 2504.15088 , archivePrefix=

  22. [22]

    and Schuett, J

    Koessler, L. and Schuett, J. and Anderljung, M. , title =. 2024 , howpublished =. 2406.14713 , archivePrefix=

  23. [23]

    and Halstead, J

    Lukosiute, K. and Halstead, J. and Righetti, L. , title =. 2026 , howpublished =. 2603.20570 , archivePrefix=

  24. [24]

    2025 , url =

    Common Elements of Frontier. 2025 , url =

  25. [25]

    and Lucas, Caleb and Guest, Ella , title =

    Mouton, Christopher A. and Lucas, Caleb and Guest, Ella , title =. 2024 , institution =

  26. [26]

    2025 , howpublished =

    Zhang, Zaixi and Chakraborty, Souradip and Bedi, Amrit Singh and others , title =. 2025 , howpublished =. 2510.15975 , archivePrefix=

  27. [27]

    2025 , journal =

    Qu, Yuanhao and Huang, Kaixuan and Yin, Ming and others , title =. 2025 , journal =. doi:10.1038/s41551-025-01463-z , url =

  28. [28]

    2025 , howpublished =

    Cong, Le and others , title =. 2025 , howpublished =. 2510.14861 , archivePrefix=

  29. [29]

    and others , title =

    Swanson, Kyle and Wu, Wesley and Bulaong, Nash L. and others , title =. 2025 , journal =. doi:10.1038/s41586-025-09442-9 , url =

  30. [30]

    2026 , howpublished =

  31. [31]

    Greg , title =

    Brent, Roger and McKelvey, Jr., T. Greg , title =. 2025 , howpublished =. 2506.13798 , archivePrefix=

  32. [32]

    2025 , institution =

    Murray, Malcolm and Barrett, Steve and Papadatos, Henry and Quarks, Otter and Smith, Matt and Tlaie Boria, Alejandro and Touzet, Chlo\'e and Campos, Sim\'eon , title =. 2025 , institution =. 2512.08844 , archivePrefix=

  33. [33]

    Biodefense in the Age of Synthetic Biology , publisher =

  34. [34]

    Preparedness Framework, Version 2 , year =

  35. [35]

    2025 , month = dec, type =

  36. [36]

    2026 , month = apr, type =

  37. [37]

    , title =

    Righetti, L. , title =. 2025 , institution =

  38. [38]

    2025 , institution =

    Rodriguez, Mikel and Popa, Raluca Ada and Flynn, Four and Liang, Lihao and Dafoe, Allan and Wang, Anna , title =. 2025 , institution =. 2503.11917 , archivePrefix=

  39. [39]

    , title =

    Sandbrink, J. , title =. 2023 , howpublished =. 2306.13952 , archivePrefix=

  40. [40]

    Virology Capabilities Test (VCT) , year =

  41. [41]

    Departmental Guidance on Valuation of a Statistical Life in Economic Analysis , year =

  42. [42]

    2026 , url =

    Our Evaluation of Claude Mythos Preview's Cyber Capabilities , institution =. 2026 , url =

  43. [43]

    2026 , url =

    How Fast Is Autonomous. 2026 , url =

  44. [44]

    2025 , month = oct, howpublished =

    Ziosi and others , title =. 2025 , month = oct, howpublished =

  45. [45]

    and others , title =

    Zhang, Andy K. and others , title =. 2024 , howpublished =. 2408.08926 , archivePrefix=

  46. [46]

    Zhang and others , title =

  47. [47]

    2025 , howpublished =

    Wang, Zhun and Song, Dawn and others , title =. 2025 , howpublished =. 2506.02548 , archivePrefix=

  48. [48]

    2026 , howpublished =

    Zhang and others , title =. 2026 , howpublished =. 2604.17860 , archivePrefix=

  49. [49]

    2026 , howpublished =

    Chauvin, Anson and Barry, Josh and Denain, Jean-Stanislas and Ho, Anson , title =. 2026 , howpublished =