REVIEW 2 major objections 6 minor 49 references
The paper claims that fragmented AI safety thresholds can be harmonized into a small set of auditable floors, with a non-zero TLO completion rule for cyber misuse and a 5x/3-month rate break for automated AI R&D.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 21:16 UTC pith:32HA2Z65
load-bearing objection Concrete, honest floors for cyber and AI R&D; the harm-model numbers behind them are shakier than the floors themselves. the 2 major comments →
Harmonizing AI Safety Thresholds
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the expected-harm framework for misuse risks and the rate-of-progress framework for automated R&D can turn heterogeneous threshold language into common quantitative floors. Applying the framework yields two operational floors: a cyber floor defined by non-zero full-chain completion on the TLO evaluation at a pinned 100M-token budget, with the Clopper-Pearson lower confidence bound above zero across pre-registered trials; and an R&D floor defined by a measured improvement rate at least five times the baseline trend, sustained for at least three months. The same analysis shows that no analogous bio floor can be responsibly fixed until an unsaturated, independently val
What carries the argument
The expected-harm equation E[H_j] = N_j × P_success,j × h_j for misuse risks, combined with the release-condition exposure scalar s_r that multiplies total AI-enabled harm, and, for automated R&D, the rate ratio r_n/r_0 comparing a model's measured improvement rate to the baseline trend with a three-month sustain condition. These turn threshold language into quantities that can be computed from public benchmark scores and evaluation runs.
Load-bearing premise
The load-bearing premise is that TLO full-chain completion rates measure the probability of real-world cyber-attack success; if the evaluation range is not representative of operational intrusion paths, or if the paper's uncited release-condition scalars are wrong, the quantitative harm figures and release verdicts do not follow.
What would settle it
A replicated study showing no correlation between TLO completion rates and real-world intrusion outcomes, or an independent audit of a trusted-access regime that measures effective exposure above the implied ~1.5% of public-API level for a Mythos-class model, would disprove the central quantitative claim.
If this is right
- If adopted, the cyber floor would have triggered enhanced safeguards for both April 2026 frontier models concurrently, closing the gap where one company's framework fired and the other's did not.
- The R&D floor is checkable by any third party with access to benchmark scores and release dates, making threshold enforcement feasible without proprietary evaluations.
- Biorisk harmonization shifts from arguing about threshold wording to commissioning the missing per-stage uplift elicitation, which is a concrete, fundable task.
- Companies can keep stricter internal thresholds while sharing the common floor, reducing the race-to-the-bottom pressure.
- The binary non-zero TLO completion rule sidesteps disputes over arbitrary continuous probability cutoffs while still marking a real capability regime transition.
Where Pith is reading between the lines
- One testable extension is to apply the 5x/three-month rule retrospectively to historical frontier-model releases to see which prior models would have tripped it; this would calibrate the floor's strictness against the public record.
- The same rate-of-progress logic could be adapted to domains other than AI R&D, such as autonomous scientific discovery, whenever a composite capability index exists.
- The cyber floor's non-zero rung may quickly become uninformative as models routinely clear several completions; the paper's two higher rungs, autonomous discovery and end-to-end hardened intrusion, will then do the real work, and those are harder to audit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a framework for converting frontier AI companies' heterogeneous capability thresholds into comparable quantitative 'floors' across three risk domains. For cyber, it combines an expected-harm model (N×P×H) with TLO cyber-range evaluations, proposing a minimum floor of non-zero full-chain TLO completion at a pinned 100M-token budget, and claims that public API access to current models generates additional expected harm above the USD 1B annual-harm reference even after neutralizing low-tier commodity pathways. For biorisk, it applies the same methodology but concludes that no defensible quantitative floor can yet be set, identifying missing full-chain validated KRIs and per-stage uplift elicitation as the binding gap. For automated AI R&D, it proposes a floor that triggers safeguards when a model causes AI progress at least five times the baseline trend for at least three months, which it argues is consistent with the existing threshold language of Anthropic, Google DeepMind, and OpenAI under stated assumptions.
Significance. If the framework and its calibrations hold, the paper makes a useful contribution to threshold harmonization: it gives a concrete translation procedure from heterogeneous policy language to auditable quantitative triggers; the cyber TLO tripwire and the AI R&D rate-of-progress floor are falsifiable and partially third-party checkable; and the biorisk analysis is a clear negative diagnosis that tells the field what evidence is missing. The paper is unusually explicit about its limitations, consistently labels author-calibrated priors as such, and does not overclaim in the biorisk section. The main value lies in the translation methodology and the proposed floors, not in the point estimates of dollar harm. The significance is real but tempered by calibration uncertainty in the cyber harm model and by the conditional nature of the AI R&D threshold equivalence.
major comments (2)
- [§3.7, Tables XIV/XV/XVII, §3.11 and Appendix B.5] The 'Public-API binary' conclusion—that public API access exceeds the USD 1B annual-harm reference across the whole anchor range—is not robust to the P_success uncertainty that the paper itself identifies. The headline calculation uses AT3 P_AI = 0.60 from TLO 6/10, while the only human-validated kill-chain decomposition (Table XVII, from Barrett et al. OC3 SME Ransomware) gives 0.084, a sevenfold gap. Since E[H] scales linearly with P_success, switching to 0.084 lowers the AT3–AT5 contribution from about $18.25B to roughly $2.5B, and a real-world P_success near 0.01 would put public API below the USD 1B reference. Table XV varies only the aggregate baseline H0, not the capability proxy, and no Monte Carlo propagation is provided. At minimum, the paper needs a sensitivity analysis over AT3 P_AI and should soften the claim that the Public-API binary is robust, or justify why TLO 6/10 is a
- [§5.1 and Appendix A, Eq. (14)] The claim that the 5×/3-month floor is the minimum common denominator of the three companies' automated AI R&D thresholds rests on two stated assumptions: that 'several months' in OpenAI's threshold means exactly three months, and that r2024 = r2018-2024. These assumptions are not derived from evidence; 'several' in ordinary usage can denote four to six months, and OpenAI's own example ('sped up to just 4 weeks') suggests a shorter interval. If m>3, OpenAI's threshold implies more additional progress than Anthropic's, so the proposed floor is not the unique common denominator but one candidate harmonization. The paper should either provide a justification for the 3-month interpretation or present the floor as a possible harmonized value, rather than as the definite minimum extracted from the companies' language.
minor comments (6)
- [§5.3] The illustrative calculation for r0 uses a two-point slope between Claude 3 Opus and Claude 3.7 Sonnet, whereas §5.2 says a trend line should be fitted over time. A two-point estimate is highly sensitive to endpoint choice; the paper should either fit a proper regression or explicitly label the calculation as an illustration and provide sensitivity to the baseline window.
- [§3.7] The text states that an observed 2/10 carries an exact interval of 'roughly 0.03 to 0.56'. The Clopper–Pearson 95% interval for 2/10 is approximately 0.025 to 0.556, so the reported lower endpoint is slightly off.
- [Table VI] The entry '0/10≈0.00' uses an approximation symbol where the value is exactly zero; minor inconsistency.
- [§4.2.6] The proposed biorisk floor uses 'Anthropic's 2× signal' as an example threshold, but no justification is given for why 2× is the appropriate harmonized value across companies, nor how it relates to the 5×/3-month cyber or R&D floors.
- [References] Several references are incomplete or inconsistently formatted, e.g., 'Zhang et al. Bountybench, 2025a' lacks a full author list, and 'TitanCA' has no individual authors listed. Please standardize.
- [§3.9 and §5.3] The model referred to as 'Mythos Preview' in the cyber section and 'Claude Mythos Preview' in the R&D section should be named consistently, and its manufacturer (Anthropic) should be identified at first use in each section.
Circularity Check
Automated AI R&D floor is a restatement of OpenAI's threshold with 'several months' set to 3; cyber/bio domains are not circular.
specific steps
-
renaming known result
[Section 5.1–5.2 and Appendix A (Eqs. 14–16)]
"We examine the thresholds defined by the main frontier AI companies and extract the minimum common denominator, which they have implicitly accepted. We then derive a quantitative threshold that captures this minimum common denominator... Under the assumption that 'several months' in OpenAI's formulation means three months, the Anthropic and OpenAI automated AI R&D thresholds are equivalent... Equating the two shows that the rate proposed by OpenAI, maintained for 3 months, yields the same additional progress as the rate proposed by Anthropic, maintained for 1 year."
The proposed floor 'at least five times faster than trend during at least three months' is OpenAI's threshold ('1/5th the wall-clock time... sustainably for several months') with 'several months' stipulated to be 3 and the rate expressed relative to trend. Appendix A does not derive 5 and 3 months from progress data; it assumes m=3 and r2024=r2018-2024, then obtains equality. Thus the harmonized floor is the input policy language restated in quantitative coordinates; the claim that it 'covers' all three companies follows by construction, not from independent evidence. This is a transparent translation/operationalization rather than an independent derivation.
full rationale
The paper is largely self-contained and transparent. The cyber floor is independently grounded in the empirical zero-to-nonzero TLO transition, and the paper explicitly says its justification does not depend on the dollar calculation; the biorisk section is a diagnostic that identifies missing evidence rather than a derived floor. No load-bearing self-citation is present: Murray et al. and Barrett et al. are not authored by this paper's authors. The one partial circularity is the automated AI R&D threshold: the 5x/3-month formulation is extracted from OpenAI's threshold language under a stipulated reading of 'several months' and an assumed equality of 2024 and 2018-2024 trends, and Appendix A's equivalence is algebraically forced by those assumptions. Because the paper openly frames this as extracting a minimum common denominator and because the other two domains retain independent content, the score is moderate rather than severe.
Axiom & Free-Parameter Ledger
free parameters (6)
- Cyber baseline pathway allocation (N0, P0, h per AT1–AT5) =
Back-fit to sum to USD 500B; Table XIII
- AI-enabled uplift multipliers (q and P_AI per pathway) =
Table XIV, e.g. q=0.10, P_AI=0.60 for AT3
- Release-condition exposure scalars s_r,j =
Table XVI, e.g. Trusted API 0.001–0.090 depending on tier
- Kill-chain phase probabilities p_k,j =
Table XVII, AT3 product ≈ 0.084
- AI R&D baseline trend slope r0 =
16.36 ECI points/year (Claude 3 Opus → Claude 3.7 Sonnet)
- Threshold constants for AI R&D: 5× and 3 months =
5× trend for at least 3 months
axioms (8)
- domain assumption Expected harm is the correct primitive for misuse thresholds.
- domain assumption Linear release-condition scaling: E[HAI,r] = s_r × E[HAI].
- domain assumption TLO full-chain completion rate is a valid proxy for real-world P_success.
- domain assumption Kill chains are AND-gates at chokepoints, with OR-gates at K2/K5.
- domain assumption ECI benchmark scores quantify AI progress and can be fit with a trend.
- domain assumption Defensive AI uplift is set to zero in the cyber calculation.
- domain assumption For biorisk, the modern baseline of successful sophisticated bioattacks is effectively zero and cannot be empirically validated.
- ad hoc to paper 'Several months' in OpenAI's threshold means three months, and r2024 = r2018-2024.
read the original abstract
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.
Figures
Reference graph
Works this paper leans on
-
[1]
2024 , publisher =
Alexanian, Tessa and Langenkamp, Max , title =. 2024 , publisher =
2024
-
[2]
2026 , eprint=
Estimating the Social Cost of Corporate Data Breaches , author=. 2026 , eprint=
2026
-
[3]
2025 , month = may, type =
Claude Opus 4 System Card , institution =. 2025 , month = may, type =
2025
-
[4]
Responsible Scaling Policy, Version 3.0 , year =
-
[5]
2026 , month = apr, type =
Claude Mythos Preview System Card , institution =. 2026 , month = apr, type =
2026
-
[6]
Responsible Scaling Policy, Version 3.1 , year =
-
[7]
Responsible Scaling Policy, Version 3.2 , year =
-
[8]
Barrett, S. and Murray, M. and Quarks, O. and Smith, M. and Krys, J. and Campos, S. and Tlaie Boria, A. and Touzet, C. and Hayrapet, S. and Heiding, F. and Nevo, O. and Swanda, A. and Aguirre, J. and Gershovich, A. B. and Clay, E. and Fetterman, R. and Fritz, M. and Juarez, M. and Mavroudis, V. and Papadatos, H. , title =. 2025 , howpublished =. 2512.0886...
arXiv 2025
-
[9]
2025 , institution =
2025
-
[10]
and others , title =
Dauparas, J. and others , title =. Science , volume =
-
[11]
2025 , url =
Internet Crime Report 2024 , institution =. 2025 , url =
2024
-
[12]
Feldman, J. and Feldman, T. , title =. 2025 , howpublished =. 2509.02610 , archivePrefix=
Pith/arXiv arXiv 2025
-
[13]
Folkerts, L. and Payne, W. and Inman, S. and Giavridis, P. and Skinner, J. and Deverett, S. and Aung, J. and Zorer, E. and Schmatz, M. and Ghanem, M. and Wilkinson, J. and Steer, A. and Hong, V. and Wang, J. , title =. 2026 , howpublished =. 2603.11214 , archivePrefix=
arXiv 2026
-
[14]
Frontier Safety Framework, Version 3.0 , year =
-
[15]
2025 , month = oct, howpublished =
Gopal, Anjali and Guest, Oliver and Besiroglu, Tamay and. 2025 , month = oct, howpublished =
2025
-
[16]
Biothreat Creation Process Framework , year =
-
[17]
Hattoh and others , title =. 2025 , howpublished =. 2505.17154 , archivePrefix=
Pith/arXiv arXiv 2025
-
[18]
Anson Ho and Jean-Stanislas Denain and David Atanasov and Samuel Albanie and Rohin Shah , year=. A. 2512.00193 , archivePrefix=
-
[19]
2025 , url =
Cost of a Data Breach Report 2025 , institution =. 2025 , url =
2025
-
[20]
Inagaki, T. and others , title =. 2023 , howpublished =. 2304.10267 , archivePrefix=
Pith/arXiv arXiv 2023
-
[21]
Khlaaf, Heidy and West, Sarah Myers , title =. 2025 , howpublished =. 2504.15088 , archivePrefix=
Pith/arXiv arXiv 2025
-
[22]
Koessler, L. and Schuett, J. and Anderljung, M. , title =. 2024 , howpublished =. 2406.14713 , archivePrefix=
Pith/arXiv arXiv 2024
-
[23]
Lukosiute, K. and Halstead, J. and Righetti, L. , title =. 2026 , howpublished =. 2603.20570 , archivePrefix=
arXiv 2026
-
[24]
2025 , url =
Common Elements of Frontier. 2025 , url =
2025
-
[25]
and Lucas, Caleb and Guest, Ella , title =
Mouton, Christopher A. and Lucas, Caleb and Guest, Ella , title =. 2024 , institution =
2024
-
[26]
Zhang, Zaixi and Chakraborty, Souradip and Bedi, Amrit Singh and others , title =. 2025 , howpublished =. 2510.15975 , archivePrefix=
arXiv 2025
-
[27]
Qu, Yuanhao and Huang, Kaixuan and Yin, Ming and others , title =. 2025 , journal =. doi:10.1038/s41551-025-01463-z , url =
-
[28]
Cong, Le and others , title =. 2025 , howpublished =. 2510.14861 , archivePrefix=
arXiv 2025
-
[29]
Swanson, Kyle and Wu, Wesley and Bulaong, Nash L. and others , title =. 2025 , journal =. doi:10.1038/s41586-025-09442-9 , url =
-
[30]
2026 , howpublished =
2026
-
[31]
Brent, Roger and McKelvey, Jr., T. Greg , title =. 2025 , howpublished =. 2506.13798 , archivePrefix=
Pith/arXiv arXiv 2025
-
[32]
Murray, Malcolm and Barrett, Steve and Papadatos, Henry and Quarks, Otter and Smith, Matt and Tlaie Boria, Alejandro and Touzet, Chlo\'e and Campos, Sim\'eon , title =. 2025 , institution =. 2512.08844 , archivePrefix=
arXiv 2025
-
[33]
Biodefense in the Age of Synthetic Biology , publisher =
-
[34]
Preparedness Framework, Version 2 , year =
-
[35]
2025 , month = dec, type =
2025
-
[36]
2026 , month = apr, type =
2026
-
[37]
, title =
Righetti, L. , title =. 2025 , institution =
2025
-
[38]
Rodriguez, Mikel and Popa, Raluca Ada and Flynn, Four and Liang, Lihao and Dafoe, Allan and Wang, Anna , title =. 2025 , institution =. 2503.11917 , archivePrefix=
Pith/arXiv arXiv 2025
-
[39]
Sandbrink, J. , title =. 2023 , howpublished =. 2306.13952 , archivePrefix=
Pith/arXiv arXiv 2023
-
[40]
Virology Capabilities Test (VCT) , year =
-
[41]
Departmental Guidance on Valuation of a Statistical Life in Economic Analysis , year =
-
[42]
2026 , url =
Our Evaluation of Claude Mythos Preview's Cyber Capabilities , institution =. 2026 , url =
2026
-
[43]
2026 , url =
How Fast Is Autonomous. 2026 , url =
2026
-
[44]
2025 , month = oct, howpublished =
Ziosi and others , title =. 2025 , month = oct, howpublished =
2025
-
[45]
Zhang, Andy K. and others , title =. 2024 , howpublished =. 2408.08926 , archivePrefix=
Pith/arXiv arXiv 2024
-
[46]
Zhang and others , title =
-
[47]
Wang, Zhun and Song, Dawn and others , title =. 2025 , howpublished =. 2506.02548 , archivePrefix=
arXiv 2025
-
[48]
Zhang and others , title =. 2026 , howpublished =. 2604.17860 , archivePrefix=
Pith/arXiv arXiv 2026
-
[49]
2026 , howpublished =
Chauvin, Anson and Barry, Josh and Denain, Jean-Stanislas and Ho, Anson , title =. 2026 , howpublished =
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.