Pith. sign in

REVIEW 2 major objections 5 minor 42 references

Frontier AI needs Basel-style buffers and a sector-wide early-warning system, not only single-model safety reviews.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 01:44 UTC pith:6M4CPMEU

load-bearing objection Solid, carefully scoped design paper that packages known pieces into a named two-layer macro-prudential system; the buffer metrics are constructive placeholders, not demonstrated anchors, but the author flags the disanalogies and the work still deserves referees. the 2 major comments →

arxiv 2607.03542 v1 pith:6M4CPMEU submitted 2026-07-03 cs.CY

Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI

classification cs.CY
keywords macro-prudential AI governancefrontier AIearly warning systemsafety buffersECARBasel III analoguesSIAIcorrelated risk
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that frontier-AI governance has the same structural gaps banking faced before 2008: discovering a risk does not guarantee action, and reviewing one model does not manage correlated build-up across labs. It proposes MEWRS, a two-layer macro-prudential system aimed first at labs' internal research and production systems rather than public products. Layer A routes structured reports on dual-use capabilities, autonomy indicators, and security compromises through a government clearinghouse to domain-specific defender groups with pre-committed playbooks. Layer B turns three buffer metrics—Effective Compute-at-Risk, Cumulative Red-Team Hours, and Alignment Robustness Score—into operational controls so that faster capability scaling automatically tightens safeguards, the way risk-weighted assets drive capital ratios. The framework maps six Basel III mechanisms onto AI analogues, names seven failure modes with mitigations, and sketches red-team exercises for validation. A sympathetic reader cares because the design targets the cascade risk no single lab can see alone.

Core claim

A two-layer macro-prudential early warning and response system for developer-internal frontier AI—Layer A's finder-coordinator-defender reporting pipeline plus Layer B's ECAR/CRTH/ARS buffer calibration, with SIAI tiering and counter-cyclical controls keyed to a Capability Surge Index—can detect correlated risk build-ups across the sector and create pre-committed off-ramps before a cascade unfolds.

What carries the argument

MEWRS: Layer A (finder-coordinator-defender routing of structured reports) coupled to Layer B (three quantitative safety buffers—ECAR = compute × autonomy × reach, CRTH weighted red-team hours, ARS robustness score—that auto-calibrate operational controls).

Load-bearing premise

The three buffer metrics can be measured well enough to guide real policy without being gamed into uselessness, even though AI capability is not a conserved balance-sheet quantity and there is no market price signal or lender-of-last-resort backstop.

What would settle it

A multi-lab red-team/blue-team exercise that compares MEWRS-guided response against status-quo response on the same scenarios (insider exfiltration, supply-chain compromise, cross-lab capability surge) and measures whether time-to-triage, mitigation uptake, and dwell time improve enough for the buffers to drive real control changes.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes MEWRS, a two-layer macro-prudential early-warning and response system for developer-internal frontier AI (as distinct from externally released products). Layer A adapts a finder-coordinator-defender pipeline: structured reports on dual-use capabilities, autonomy/loss-of-control indicators, security compromises, and alignment-integrity failures are routed through a jurisdiction-specific government clearinghouse (illustrated with BIS/CAISI) to domain-specific defender working groups with pre-committed playbooks. Layer B defines three buffer metrics—ECAR = C(m)·A(m)·R(m) (Eq. 1), CRTH as independence- and access-weighted red-team hours (Eq. 2), and ARS as empirical robustness under adversarial/distribution-shift conditions—and maps them, together with SIAI tiering and a Capability Surge Index, onto six Basel III mechanisms (Table 1). The paper supplies a reporting schema outline, seven failure modes with mitigations (including gaming via standardised factor tables, dual disclosure, and third-party audit), an Appendix B worked example, and a red-team/blue-team validation plan. The central claim is that standardised aggregation of these buffers plus cross-lab correlation signals can detect sector-level risk build-up and create pre-committed off-ramps before cascades.

Significance. If the design is workable, the paper fills a genuine gap: most frontier-AI governance work is micro-prudential (single-model evaluations, RSPs, release decisions), while correlated multi-lab dynamics remain under-specified. The contribution is a concrete, implementable architecture rather than a pure taxonomy. Strengths include deliberate scoping to internal deployments, an explicit functional (not literal) Basel mapping with existing AI-governance analogues flagged in Table 1, seven failure modes with concrete mitigations (especially gaming and coordinator abuse), and candid disanalogies in §9 (no conserved quantity, no market-discipline channel, compressed cycles, no lender-of-last-resort). The voluntary-to-mandatory scaffolding argument and the exercise-based validation plan make the proposal falsifiable in principle. These features make the manuscript a useful reference point for subsequent pilot design even if the metrics require substantial calibration work.

major comments (2)
  1. §8 Assumptions (ii) and §9 Limits of the Basel analogy: the load-bearing claim that Layer B buffers will 'automatically trigger stronger safeguards' and create pre-committed off-ramps (§1 Contribution; Abstract; §10) rests on ECAR/CRTH/ARS being operationalisable well enough to guide policy without being gamed out of usefulness. The paper itself states that capability is not a conserved balance-sheet quantity, there is no price-signal/market-discipline channel, cycles run in weeks-to-quarters, and there is no lender-of-last-resort. Eq. (1) multiplies C by relative [0,1] factors A and R assigned with confidence intervals (Appendix B); the worked example keys buffers to the conservative upper end of those intervals using hand-chosen floors (5e25 weighted FLOP, 1000 CRTH hours, ARS 0.75). The only anchors offered are the standardised factor tables, dual disclosure, and third-party audit rig
  2. §4.1–4.3 and Appendix B: free parameters (CRTH α/ι weights, illustrative ECAR/CRTH/ARS floors, A(m)/R(m) point estimates and intervals) are acknowledged as placeholders, yet the operational-control mapping in the worked example treats them as if they already produce determinate tier placements. The paper needs either (a) an explicit statement that all numerical thresholds are purely illustrative and that no claim of calibrated policy guidance is made until pilot exercises produce them, or (b) a minimal calibration protocol (data sources, decision rules for standardised factor tables, success criteria for the red-team/blue-team programme in §8) that would make the buffer-to-control translation reproducible. As written, a reader cannot distinguish a design sketch from a ready-to-pilot specification.
minor comments (5)
  1. Figure 1 (Appendix A) is described in the text but the caption and flow labels would benefit from explicit numbering of the feedback loops (report → buffer recalibration; telemetry → new Layer A reports) so that the two-layer coupling is immediately visible.
  2. Table 1, 'Living wills' row: the note that Anthropic's RSP v3.0 replaced pause/halt commitments with Frontier Safety Roadmaps is useful; a one-sentence clarification of what an AI 'resolution plan' would contain beyond existing RSP language would strengthen the mapping.
  3. §3 Jurisdictional structure: the BIS/CAISI assignment is correctly labelled an institutional hypothesis, but a short footnote on equivalent EU AI Office / UK AISI intake pathways would make the multilateral Basel-style claim more concrete for non-US readers.
  4. Eq. (2) access weights (α = 1.0 / 0.7 / 0.4 / 0.2) are stated as illustrative placeholders; flagging them as such in the equation environment itself (not only in the surrounding prose) would reduce the risk of later citation as calibrated constants.
  5. Glossary (Appendix C) is helpful; adding CSI and SIAI expansions in the main text at first use (they appear before the glossary) would improve readability.

Circularity Check

0 steps flagged

No circularity: MEWRS is a constructive policy proposal whose buffer metrics are defined for future calibration, not derived predictions that reduce to fitted inputs.

full rationale

The paper proposes a two-layer governance architecture (Layer A reporting pipeline; Layer B ECAR/CRTH/ARS buffers) by functional analogy to Basel III. Equation (1) defines ECAR(m)=C(m)·A(m)·R(m) with A and R as relative [0,1] factors; CRTH and ARS are likewise constructive definitions. The Appendix B.1 worked example uses hand-chosen illustrative floors and confidence intervals, explicitly labelled as such, and does not present any fitted parameter as a discovered prediction. The six Basel mappings (Table 1) are structural translations of function, not tautologies. Self-positioning against FAS (2024), RSPs, Reuel et al., and Basel is comparative literature placement, not a load-bearing uniqueness theorem or self-citation chain that forces the result. There are no empirical fits, no self-definitional loops of the form 'X derives Y where X is defined via Y', and no ansatz smuggled in via prior author work. The paper is therefore self-contained as a design proposal; circularity score is zero.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 6 invented entities

The central claim is a governance-design claim, so its load rests on domain assumptions about AI systemic risk and institutional transferability, plus invented quantitative entities (the three buffers and related indices) whose operational content is not yet independently measured. Free parameters are the illustrative weights and floors used to show how buffers would bind. No machine-checked or empirical calibration anchors these quantities yet; the paper treats them as infrastructure to be falsified in pilots.

free parameters (5)
  • CRTH access weights α (full internals=1.0, fine-tuning=0.7, elevated API=0.4, black-box=0.2)
    Explicitly labeled illustrative placeholders in §4.2; buffer requirements in the worked example depend on these hand-chosen weights.
  • CRTH independence weights ι (e.g. internal 0.3, contracted 0.7, national institute 1.0)
    Chosen in Appendix B.1 example; heavily discount internal hours and determine whether CRTH clears the floor.
  • Illustrative top-tier ECAR threshold (5×10^25 weighted FLOP)
    Appendix B.1 tier boundary is illustrative; Model X’s tier placement depends on this hand-set cutoff.
  • Illustrative CRTH floor (1000 weighted hours) and ARS floor (0.75) for top-tier deployments
    Appendix B.1; jointly determine elevated buffer band and operational controls in the example.
  • Autonomy A(m) and reach R(m) point estimates and intervals for a deployment
    Assigned by evaluator judgment relative to sector frontier; conservative upper ends drive ECAR buffers under the opacity-is-costly rule (Appendix B).
axioms (5)
  • domain assumption Frontier-AI sector risk is structurally analogous enough to pre-2008 banking systemic risk that Basel-style macro-prudential mechanisms transfer in function.
    Load-bearing framing of Abstract and §1; §9 lists disanalogies but still relies on functional transfer for the central design claim.
  • ad hoc to paper ECAR(m)=C(m)·A(m)·R(m) with A,R relative to the sector frontier is a useful operational proxy for systemic blast radius.
    Equation (1) §4.1; multiplicative form and relative scaling are design choices, not derived from prior measurement theory.
  • domain assumption Major governments or international bodies are receptive to financial-style measures for AI, and major labs will comply with adopted rules.
    Stated as required assumptions (i) and (iii) in §8; non-compliance/rogue actors deferred to future work.
  • domain assumption A government clearinghouse can observe cross-lab interlinkage channels (shared architectures, data, compute, evals, tools, personnel) that no single lab sees.
    §6 correlation-signal argument; paper admits the measurement infrastructure does not yet exist and must be built by the reporting pipeline itself.
  • domain assumption BIS as legal intake and NIST CAISI as technical evaluator are a feasible U.S. institutional hypothesis for a voluntary pilot.
    §3 Jurisdictional structure; paper notes neither currently holds mandatory AI incident-reporting authority and policy has moved deregulatory.
invented entities (6)
  • MEWRS (Macro-Prudential Early Warning and Response System) no independent evidence
    purpose: Name and package the two-layer coordinated-response plus safety-buffer architecture for internal frontier AI.
    Core invented institutional system; independent evidence only via future pilots/exercises, not existing deployments.
  • Effective Compute-at-Risk (ECAR) no independent evidence
    purpose: Proxy systemic blast radius of a misaligned model as risk-weighted-assets analogue.
    New metric defined in §4.1; no external calibrated standard or published measurement series yet.
  • Cumulative Red-Team Hours (CRTH) no independent evidence
    purpose: Independence- and access-weighted testing effort as residual-uncertainty / volatility-haircut analogue.
    New metric §4.2; weights uncalibrated; not an established industry KPI.
  • Alignment Robustness Score (ARS) no independent evidence
    purpose: Aggregate consistency of safety-relevant outcomes under adversarial and distribution-shift conditions.
    New score §4.3; standing evaluation suite not specified beyond description; no public benchmark series.
  • Systemically Important AI Institution (SIAI) designation no independent evidence
    purpose: G-SIB-style tier imposing stricter buffers on labs with systemically consequential footprint.
    New designation §4.4; related to but distinct from EU GPAI-with-systemic-risk; no existing SIAI regime.
  • Capability Surge Index (CSI) no independent evidence
    purpose: Public reference rate of sector-aggregate frontier ECAR change to guide counter-cyclical buffer discretion.
    Introduced in §5 as credit-to-GDP-gap analogue; not computed on real data in the paper.

pith-pipeline@v1.1.0-grok45 · 20888 in / 4427 out tokens · 44315 ms · 2026-07-12T01:44:26.073051+00:00 · methodology

0 comments
read the original abstract

Frontier-AI governance today faces a problem structurally analogous to the one banking regulation faced pre-2008, and which post-2008 reforms (Basel III, Dodd-Frank) have since addressed. Two gaps recur: discovering a risk is not tantamount to acting on it, and individual-model review is unlike managing correlated build-up across the sector. Drawing on the Basel III framework and the U.S. financial-stability architecture, I propose a macro-prudential early warning and response system ("MEWRS") for internal frontier AI. These are systems deployed for labs' own internal research, testing, and production workflows, as distinct from externally released products. Layer A adapts the finder-coordinator-defender early-warning model to route structured reports on dual-use capabilities, autonomy indicators, and security compromises through a government clearinghouse to domain-specific defender working groups. Layer B calibrates operational controls via three quantitative buffer metrics, namely Effective Compute-at-Risk (ECAR), Cumulative Red-Team Hours (CRTH), and an Alignment Robustness Score (ARS), so that faster capability scaling automatically triggers stronger safeguards, analogously to how risk-weighted assets drive capital ratios under Basel III. I outline the reporting schema, map six Basel III mechanisms onto AI-governance analogues, identify seven failure modes with concrete mitigations, and sketch an exercise-based validation plan. MEWRS is designed to detect correlated risk build-ups across the frontier-AI sector and create pre-committed off-ramps before a cascade unfolds.

Figures

Figures reproduced from arXiv: 2607.03542 by Pranav Mehta.

Figure 1
Figure 1. Figure 1: MEWRS two-layer architecture. Layer A routes structured reports through a government clearinghouse to domain-specific defender working groups. Layer B translates three headline metrics into operational buffer requirements, with telemetry feeding back into Layer A. B. ECAR Calibration Notes The ECAR definition in (1) is deliberately coarse. In practice, each factor is assigned a point estimate together with… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

42 extracted references · 10 linked inside Pith

  1. [1]

    Secure, governable chips: Using on-chip mechanisms to manage national security risks from AI & advanced computing

    Aarne, O., Fist, T., and Withers, C. Secure, governable chips: Using on-chip mechanisms to manage national security risks from AI & advanced computing. Technical report, Center for a New American Security, January 8, 2024. https://www.cnas.org/publications/reports/secure-governable-chips

  2. [2]

    and Restrepo, P

    Acemoglu, D. and Restrepo, P. The race between man and machine: Implications of technology for growth, factor shares, and employment. American Economic Review, 108(6):1488--1542, 2018

  3. [3]

    Frontier AI regulation: Managing emerging risks to public safety

    Anderljung, M., Barnhart, J., Korinek, A., Leung, J., O'Keefe, C., Whittlestone, J., et al. Frontier AI regulation: Managing emerging risks to public safety. arXiv preprint arXiv:2307.03718, 2023

  4. [4]

    Responsible Scaling Policy, Version 1.0

    Anthropic. Responsible Scaling Policy, Version 1.0. Technical report, Anthropic, September 19, 2023. https://www-cdn.anthropic.com/1adf000c8f675958c2ee23805d91aaade1cd4613/responsible-scaling-policy.pdf

  5. [5]

    Responsible Scaling Policy, Version 3.3

    Anthropic. Responsible Scaling Policy, Version 3.3. Technical report, Anthropic, effective May 26, 2026. Current version in the RSP v3.x series; Version 3.0 was effective February 24, 2026. https://www.anthropic.com/responsible-scaling-policy

  6. [6]

    Basel III: Finalising post-crisis reforms

    Basel Committee on Banking Supervision. Basel III: Finalising post-crisis reforms. Technical report, Bank for International Settlements, 2017. https://www.bis.org/bcbs/publ/d424.pdf

  7. [7]

    About BIS

    Bureau of Industry and Security. About BIS. U.S. Department of Commerce, 2026. https://www.bis.gov/about-bis, accessed 2026

  8. [8]

    The foundation model transparency index

    Bommasani, R., Klyman, K., Longpre, S., Kapoor, S., Maslej, N., Xiong, B., et al. The foundation model transparency index. arXiv preprint arXiv:2310.12941, 2023

  9. [9]

    The Brussels Effect: How the European Union Rules the World

    Bradford, A. The Brussels Effect: How the European Union Rules the World. Oxford University Press, 2020

  10. [10]

    Toward trustworthy AI development: Mechanisms for supporting verifiable claims

    Brundage, M., Avin, S., Wang, J., Belfield, H., Krueger, G., Hadfield, G., et al. Toward trustworthy AI development: Mechanisms for supporting verifiable claims. arXiv preprint arXiv:2004.07213, 2020

  11. [11]

    and Trager, R

    Bucknall, B. and Trager, R. Structured access for third-party research on frontier AI models: Investigating researchers' model access requirements. Technical report, Centre for the Governance of AI, 2023

  12. [12]

    and Moss, D

    Carpenter, D. and Moss, D. A. (eds.) Preventing Regulatory Capture: Special Interest Influence and How to Limit It. Cambridge University Press, 2014

  13. [13]

    L., Bucknall, B., et al

    Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., et al. Black-box access is insufficient for rigorous AI audits. In ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2024

  14. [14]

    What failure looks like

    Christiano, P. What failure looks like. AI Alignment Forum, 2019. https://www.alignmentforum.org/posts/HBxe6wdjxK239zajf/what-failure-looks-like

  15. [15]

    Sector Risk Management Agencies

    Cybersecurity and Infrastructure Security Agency. Sector Risk Management Agencies. U.S. Department of Homeland Security, 2026. https://www.cisa.gov/topics/critical-infrastructure-security-and-resilience/critical-infrastructure-sectors/sector-risk-management-agencies, accessed 2026

  16. [16]

    What multipolar failure looks like, and robust agent-agnostic processes (RAAPs)

    Critch, A. What multipolar failure looks like, and robust agent-agnostic processes (RAAPs). LessWrong, 2021. https://www.lesswrong.com/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic

  17. [17]

    Regulatory capture: A review

    Dal B\'o, E. Regulatory capture: A review. Oxford Review of Economic Policy, 22(2):203--225, 2006

  18. [18]

    Executive Order 14110: Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence

    Executive Office of the President. Executive Order 14110: Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence. Federal Register, 88(210):75191--75226, October 30, 2023. Rescinded January 20, 2025 by Executive Order 14148. https://www.federalregister.gov/documents/2023/11/01/2023-24283/

  19. [19]

    Executive Order 14179: Removing Barriers to American Leadership in Artificial Intelligence

    Executive Office of the President. Executive Order 14179: Removing Barriers to American Leadership in Artificial Intelligence. Federal Register, 90(20):8741, January 23, 2025. https://www.federalregister.gov/documents/2025/01/31/2025-02172/

  20. [20]

    The General-Purpose AI Code of Practice

    European Commission. The General-Purpose AI Code of Practice. Published July 10, 2025; assessed as adequate by the Commission and the AI Board. Details compliance with EU AI Act Articles 53 and 55. https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai

  21. [21]

    Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act)

    European Parliament and Council. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, 2024. General-purpose AI model obligations under Chapter V, including Articles 53 and 55, became applicable August 2, 2025; Article 50 transparency obligations apply from ...

  22. [22]

    and Newman, A

    Farrell, H. and Newman, A. L. Weaponized interdependence: How global economic networks shape state coercion. International Security, 44(1):42--79, 2019

  23. [23]

    Comprehensive Capital Analysis and Review 2020 Summary Instructions

    Board of Governors of the Federal Reserve System. Comprehensive Capital Analysis and Review 2020 Summary Instructions. Technical report, Federal Reserve, March 2020. https://www.federalreserve.gov/publications/comprehensive-capital-analysis-and-review-summary-instructions-2020.htm

  24. [24]

    Creating an early warning system for AI-powered threats to national security and public safety

    Federation of American Scientists. Creating an early warning system for AI-powered threats to national security and public safety. Policy report, Federation of American Scientists, June 26, 2024. https://fas.org/publication/an-early-warning-system-for-ai/

  25. [25]

    Policy measures to address systemically important financial institutions

    Financial Stability Board. Policy measures to address systemically important financial institutions. Technical report, FSB, November 2011

  26. [26]

    Guidance on nonbank financial company determinations

    Financial Stability Oversight Council. Guidance on nonbank financial company determinations. Federal Register, 88:80110--80131, November 17, 2023. https://www.federalregister.gov/documents/2023/11/17/2023-25053/guidance-on-nonbank-financial-company-determinations

  27. [27]

    Authority to require supervision and regulation of certain nonbank financial companies: Proposed interpretive guidance

    Financial Stability Oversight Council. Authority to require supervision and regulation of certain nonbank financial companies: Proposed interpretive guidance. Federal Register, 91:15551--15566, March 30, 2026. https://www.federalregister.gov/documents/2026/03/30/2026-06114/authority-to-require-supervision-and-regulation-of-certain-nonbank-financial-companies

  28. [28]

    Hadfield, G. K. and Clark, J. Regulatory markets: The future of AI governance. arXiv preprint arXiv:2304.04914, 2023

  29. [29]

    A., and Zilberman, N

    Heim, L., Fist, T., Egan, J., Huang, S., Zekany, S., Trager, R., Osborne, M. A., and Zilberman, N. Governing through the cloud: The intermediary role of compute providers in AI regulation. arXiv preprint arXiv:2403.08501, 2024

  30. [30]

    Sleeper agents: Training deceptive LLMs that persist through safety training

    Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., et al. Sleeper agents: Training deceptive LLMs that persist through safety training. arXiv preprint arXiv:2401.05566, 2024

  31. [31]

    Frontier models are capable of in-context scheming

    Meinke, A., Schoen, B., Scheurer, J., Balesni, M., Shah, R., and Hobbhahn, M. Frontier models are capable of in-context scheming. arXiv preprint arXiv:2412.04984, 2024

  32. [32]

    Center for AI Standards and Innovation (CAISI), 2026

    National Institute of Standards and Technology. Center for AI Standards and Innovation (CAISI), 2026. https://www.nist.gov/caisi, accessed 2026

  33. [33]

    AI Risk Management Framework (AI RMF 1.0)

    National Institute of Standards and Technology. AI Risk Management Framework (AI RMF 1.0). NIST AI 100-1, January 2023

  34. [34]

    Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile

    National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1, July 2024. https://doi.org/10.6028/NIST.AI.600-1

  35. [35]

    Preparedness Framework, Version 2

    OpenAI. Preparedness Framework, Version 2. Technical report, OpenAI, April 15, 2025. Supersedes the December 2023 Beta. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf

  36. [36]

    Open problems in technical AI governance

    Reuel, A., Bucknall, B., Casper, S., Fist, T., Soder, L., Aarne, O., et al. Open problems in technical AI governance. arXiv preprint arXiv:2407.14981, 2024

  37. [37]

    What does it take to catch a Chinchilla? Verifying rules on large-scale neural network training via compute monitoring

    Shavit, Y. What does it take to catch a Chinchilla? Verifying rules on large-scale neural network training via compute monitoring. arXiv preprint arXiv:2303.11341, 2023

  38. [38]

    Model evaluation for extreme risks

    Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., et al. Model evaluation for extreme risks. arXiv preprint arXiv:2305.15324, 2023

  39. [39]

    Stigler, G. J. The theory of economic regulation. Bell Journal of Economics and Management Science, 2(1):3--21, 1971

  40. [40]

    F., and Ward, F

    van der Weij, T., Hofst\"atter, F., Jaffe, O., Brown, S. F., and Ward, F. R. AI sandbagging: Language models can strategically underperform on evaluations. arXiv preprint arXiv:2406.07358, 2024

  41. [41]

    Poisoning language models during instruction tuning

    Wan, A., Wallace, E., Shen, S., and Klein, D. Poisoning language models during instruction tuning. In International Conference on Machine Learning (ICML), 2023

  42. [42]

    Winning the Race: America's AI Action Plan

    Executive Office of the President. Winning the Race: America's AI Action Plan. The White House, July 2025. https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf