Pith. sign in

REVIEW 2 major objections 4 minor 39 references

Restricting access to a dual-use AI model is precautionary only if it delays harmful actors more than defenders; the paper derives a unique adversary-substitution threshold above which broad release beats controlled access.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 04:01 UTC pith:QRMTHSA2

load-bearing objection A real formal contribution on actor-specific release timing whose headline reversal rests on an exponential assumption the paper discloses but does not test. the 2 major comments →

arxiv 2607.22957 v1 pith:QRMTHSA2 submitted 2026-07-24 cs.CY cs.GT

Who Does Withholding Delay? A Game-Theoretic Model of Open-Weight AI Release Under Asymmetric Proliferation

classification cs.CY cs.GT
keywords open-weight AI releaseaccess inversionasymmetric proliferationsubstitute acquisitiondefender-first windowrelease reviewgame-theoretic modelAI safety policy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish when withholding an open-weight AI model actually buys safety. Its central condition is actor-specific: restriction is precautionary only if it delays harmful actors more than defenders. It proves that, in a linear benchmark with independent exponential substitute-acquisition times, there is a unique adversary-substitution rate above which broad release overtakes controlled access, and it characterizes when a defender-first window and removable safeguards have value. The point of the model is to force release debates to estimate concrete quantities—who gets a substitute when, what capability release adds to whom, and how fast defenders can deploy protection—rather than argue about openness in the abstract.

Core claim

The central discovery is a set of closed-form conditions for when controlled access, defender-first sequencing, safeguarded open weights, and minimally restricted open weights should be preferred. Under the assumption that restricted-access substitute times are independent exponentials with hazards λ_S and λ_D, the discounted access exposure of each population is λ_i/[ρ(λ_i+ρ)], so restriction creates a positive adversary access advantage exactly when λ_S > λ_D. Over a finite horizon H, immediate release adds capability q_i e^{-λ_i H} to population i, meaning release empowers the slower-substituting group most when usefulness is equal. In the linear benchmark, the difference between broad re

What carries the argument

The engine is the pair of independent exponential substitute-acquisition times for sophisticated adversaries and defenders (T_S∼Exp(λ_S), T_D∼Exp(λ_D)) together with the discounted-access function F(λ)=λ/(λ+ρ). The exponential form turns each policy's welfare into closed-form occupancy probabilities; the ratio F(λ_S)/F(λ_D) controls access inversion, e^{-λ_i H} controls finite-horizon empowerment, and the same F enters the unique threshold λ*_S through θ = −ρΨ(0)/(α q_S). This machinery converts actor-by-actor substitution speed into a policy ranking.

Load-bearing premise

The load-bearing assumption is that each actor's wait for an adequate substitute, under restriction, follows a simple exponential clock and that those clocks tick independently for adversaries and defenders.

What would settle it

A longitudinal dataset of real model releases recording, for each actor class, the first effective-access date under both restricted and open policies: if the empirical share acquiring by horizon H deviates materially from 1−e^{−λ_i H}, or if λ_S and λ_D are positively correlated during global events, the linear benchmark's unique threshold λ*_S = ρθ/(1−θ) would not describe the world.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • When adversaries obtain substitutes faster than defenders (λ_S > λ_D), withholding gives adversaries a discounted access advantage, so delay-based justifications for control fail.
  • Under equal usefulness, immediate release adds more finite-horizon capability to defenders than to sophisticated adversaries; with unequal usefulness, the ratio condition q_D/q_S > e^{-(λ_S−λ_D)H} governs.
  • If endpoint conditions hold, there is a unique adversary-substitution threshold: below λ*_S control is preferred, above it broad release is preferred, with λ*_S = ρθ/(1−θ).
  • A defender-first window is valuable only when selected defenders deploy protection before adversaries substitute or the scheduled public release; its success probability μ/(μ+λ_S)(1−e^{−(μ+λ_S)τ}) falls as adversary substitution rises.
  • Removable safeguards are worth keeping when the deterred opportunistic misuse δ m_O exceeds the friction cost β f d_O plus lost benefits and irreversibility differences; otherwise minimally restricted release wins.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the exponential-hazard assumption fails—say, actual substitute times follow a Weibull or common-shock process—the unique threshold may become a range or disappear; the same actor-delay accounting would still apply, but the closed forms would need rederivation.
  • Applied to compute-export controls, the model suggests effectiveness should be measured by how much a control moves the effective substitute-access time of target actors, not by shipment volumes or license denials alone.
  • A natural test: prospectively record four release milestones (announcement, hosted availability, weight availability, actor-specific deployment) across many models and compare realized first-effective-access times to the exponential benchmark; systematic deviation would refute the quantitative threshold.
  • The model implies release decisions should be revisited whenever a foreign substitute release or an inference-cost drop changes any λ_i; the paper gestures at this but leaves the re-review trigger unspecified.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper develops a game-theoretic model of open-weight AI release in which a laboratory chooses among controlled access, a defender-first window, safeguarded open weights, and minimally restricted open weights, while sophisticated adversaries, opportunistic adversaries, and distributed defenders differ in their ability to acquire substitutes. The main analytic results are: access inversion (Prop. 1), asymmetric empowerment (Prop. 2), a unique adversary-substitution threshold above which broad release beats control in a linear benchmark (Prop. 3), a defensive network externality condition (Prop. 4), and a credibility condition for defender-first windows (Prop. 5). A deterministic numerical implementation solves the full four-policy comparison under a convex harm function and reports policy shares over nested parameter boxes. The paper applies the framework to recent release and incident cases and concludes that release reviews should estimate actor-specific substitution times, marginal capability gains, deployment rates, defensive reach, newly enabled misuse, and nonrecallable losses.

Significance. If the results hold, the paper makes a useful conceptual contribution: it formalizes the intuitive but often-neglected point that withholding is precautionary only when it delays harmful actors more than defenders, and it derives concrete conditions under which restriction can backfire. The paper is unusually transparent about its assumptions and limitations, explicitly labeling its welfare parameters as illustrative and providing a reproducible deterministic sensitivity design. The closed-form propositions are correct under the stated exponential benchmark, and the numerical state-occupancy formulas check out. The main weakness is that every headline analytic result depends on the exponential independent acquisition assumption in eq. (4), and the paper does not provide robustness for non-exponential or correlated acquisition processes. Because the operational recommendation directs practitioners to estimate substitution times, this distributional dependence is not merely technical; it affects whether the stated policy-ranking conditions are sufficient or even meaningful in real settings.

major comments (2)
  1. [§5.2 and Propositions 1–3, 5] The central analytic results rest on the assumption in eq. (4) that T_S ~ Exp(λ_S), T_D ~ Exp(λ_D), and T_S ⊥ T_D. The paper discloses this in §5.2 and §11, but it does not provide robustness for non-exponential or correlated substitution processes. This is load-bearing rather than cosmetic: for general T_i, A_i = L_i(ρ)/ρ where L_i is the Laplace transform, so the sign of access inversion is governed by L_S(ρ) − L_D(ρ), not by E[T_i] or λ_i. A mean-preserving spread can reverse Prop. 1's conclusion, the finite-horizon comparison in Prop. 2 depends on survival functions rather than means, and Prop. 3's monotonicity and threshold uniqueness use the exponential likelihood ratio. The practitioner rule in §1 tells reviewers to estimate 'actor-specific substitution times,' but means are insufficient. Please add a robustness analysis with non-exponential distributions (e.g., Gamma, Weibull, or
  2. [§6.3, eqs. (24)–(26)] The unique threshold λ*_S = ρθ/(1−θ) is derived from the exponential functional form F(λ)=λ/(λ+ρ). The paper notes that the threshold exists only when 0<θ<1, but θ itself is defined through Ψ(0), which is computed using the exponential F. If the exponential assumption is relaxed, Ψ(λ_S) need not be strictly increasing and the threshold need not be unique; under non-proportional hazards, multiple crossings are possible. The paper should state the general condition in terms of the Laplace transforms L_S(ρ) and L_D(ρ), and, if possible, identify a class of distributions (e.g., monotone likelihood ratio) under which the threshold property survives. Without such a statement, the proposition's practical relevance for release reviews is unclear.
minor comments (4)
  1. [§3.5.1, eq. (1)] The state-capability counterexample is presented as a welfare difference but is not integrated into the formal model of §5. Consider marking it explicitly as a heuristic example or deriving it from the same welfare function with stated assumptions.
  2. [§7, Fig. 6] The nested-box sensitivity analysis varies parameter ranges but not distributional assumptions. The caption and text are clear that these are deterministic parameter designs, but a reader might over interpret the narrow/reference/wide shares as robustness to model form. A sentence noting that the boxes test parameter bounds, not the exponential assumption, would help.
  3. [Fig. 1 caption] The phrase 'full weights (open circle: promised)' is ambiguous. Clarify that the open circle indicates a future scheduled weight release that had not occurred as of the cutoff.
  4. [§4] The term 'asymmetric proliferation' is used in the title and introduction but is not formally defined until the drone discussion. Define it explicitly at first use, perhaps in §4, to avoid ambiguity with 'asymmetric empowerment'.

Circularity Check

0 steps flagged

No circularity: the analytic propositions are parameter-free derivations from stated exponential assumptions; the numerical results are explicitly illustrative and deterministic.

full rationale

The paper's central claims are derived, not fitted. Proposition 1 (eq. 15) follows from eq. (4) by the Laplace transform of an exponential; Proposition 2 (eq. 18) is the survival function under the same assumption; Proposition 3's unique threshold is derived algebraically from the linear benchmark (eqs. 21-27); and Proposition 5's window probability is a direct exponential integral (eq. 29). The proofs are shown in-line and do not invoke any prior result by the same author. There are no self-citations at all: the closest antecedent, Landolt et al. [11], is cited as related work and is not load-bearing. The numerical section explicitly states that 'Only the ranking and the comparative statics carry meaning; the absolute numbers are normalized' and that the parameter values in Table 3 are 'illustrative assumptions.' The parameter-box scans are deterministic low-discrepancy designs, not fits to data, and the paper repeatedly warns that λ_S and λ_D remain unmeasured (Section 8: 'The evidence leaves λ_S and λ_D open'). The fact that the headline policy conclusion reflects the welfare definition is a modeling choice, not a circular derivation: the model defines welfare as expected discounted harm minus benefits and then evaluates policies under that definition. The exponential independence assumption is disclosed as a maintained assumption (Section 5.2), and the skeptical concern that non-exponential or correlated substitution times could change the results is a robustness limitation, not circularity. No step reduces to its own inputs by construction, so the appropriate score is 0.

Axiom & Free-Parameter Ledger

14 free parameters · 8 axioms · 0 invented entities

The analytic results are derived from a small set of distributional and commitment assumptions; the computational results depend on the illustrative Table 3 calibration and the parameter-box bounds. No parameters are fitted to data, which keeps circularity low but leaves the operational quantities unmeasured.

free parameters (14)
  • λ_D = 0.70
    Illustrative baseline defender substitute-acquisition rate (Table 3); not estimated from data.
  • λ_S = 0.95
    Illustrative baseline sophisticated-adversary acquisition rate (Table 3); not estimated from data.
  • μ = 2.50
    Selected-defender deployment rate for the pre-release window (Table 3).
  • q_S = 1.10
    Direct capability uplift for sophisticated adversaries (Table 3).
  • q_D = 0.95
    Direct capability uplift for defenders (Table 3).
  • η = 0.55
    Defensive network productivity (Table 3).
  • n_C, n_P, n_open = 0.18, 0.48, 1.00
    Defensive reach by policy (Table 3).
  • m_O = 0.62
    Opportunistic-misuse flow cost (Table 3).
  • δ = 0.62
    Safeguard deterrence share (Table 3).
  • f = 0.13
    Legitimate-use friction from safeguards (Table 3).
  • b_C, b_P, b_G, b_O = 0.08, 0.17, 0.34, 0.40
    Benefit flows by policy (Table 3).
  • I_C, I_G, I_O = 0, 0.38, 0.52
    Nonrecallable one-time costs by policy (Table 3).
  • ρ = 0.35
    Discount rate (Table 3).
  • γ = 1.60
    Harm curvature in the convex damage function (Table 3, eq. 32).
axioms (8)
  • domain assumption Substitute-acquisition times are independent exponentials (eq. 4).
    Used in Propositions 1–3 and 5; acknowledged as a tractability assumption in §5.2 and §11.
  • domain assumption No common shocks across actors or channels (competing-risks extension in eq. 5).
    Independence is maintained in §5.1–5.2 and listed as a limitation in §11.
  • domain assumption The laboratory can commit to the announced tier and timing.
    Maintained assumption in §5.2; no limited-commitment or time-inconsistency analysis.
  • domain assumption Access is binary and an acquired substitute is adequate for the capability under study.
    Maintained assumption in §5.2; partial or task-specific substitutes are deferred to future work.
  • domain assumption Welfare is expected discounted flow plus a one-time irreversibility term (eq. 12).
    The model ranks policies under stated tail-cost assumptions rather than deriving catastrophic social value from first principles (§5.2).
  • domain assumption Follower acquisition payoffs are separable across actors (eq. 7–9).
    Stage-two payoff structure in §5.1; a coupled contest is listed as a limitation in §11.
  • standard math Standard calculus and Laplace transform of exponential distributions.
    Required for Propositions 1–3 and 5 proofs.
  • domain assumption Convex damage function h(Δ) = (κ_h/γ) log(1 + e^{γΔ}) (eq. 32).
    Introduced in §7 for the computational model; not derived from microfoundations.

pith-pipeline@v1.3.0-alltime-deepseek · 19933 in / 16629 out tokens · 136955 ms · 2026-08-01T04:01:18.637163+00:00 · methodology

0 comments
read the original abstract

Restricting access to a dual-use AI model is precautionary only if it delays harmful actors more than defenders. That condition varies across actors: a state agency or organized criminal group may obtain a substitute through theft, distillation, intermediated access, independent development, or a foreign release, while a small utility or open-source maintainer may have no comparable route. We model a laboratory choosing among controlled access, a defender-first window, safeguarded open weights, and minimally restricted open weights. Access inversion occurs when restriction gives an access advantage to adversaries that obtain effective substitutes faster than defenders. Asymmetric empowerment occurs when immediate release adds the most capability to populations least likely to possess a substitute. The policy ranking also depends on relative usefulness, opportunistic misuse, offense-defense conversion, defensive spillovers, safeguard friction, and nonrecallable losses. A linear benchmark yields a unique adversary-substitution threshold above which broad release overtakes control when the endpoint conditions hold. A defender-first window has value when selected defenders deploy protection before adversaries catch up, and removable safeguards remain useful when they deter enough opportunistic misuse. A nonlinear implementation gives each release tier a nonempty policy region. Three nested 2,048-point deterministic designs assess sensitivity to parameter bounds, and a separate grid examines actor-specific deployment delays after release. Release, cyber-evaluation, and incident-response cases identify the quantities a release review should estimate: actor-specific substitution times, marginal capability gains, deployment rates, defensive reach, newly enabled misuse, and nonrecallable losses.

Figures

Figures reproduced from arXiv: 2607.22957 by Daniel Commey.

Figure 1
Figure 1. Figure 1: Observed release paths for five prominent model releases with authoritative primary-source dates [4, 12, 16, 18, 24]. Same-day weight publication did not remove deployment differences across these cases; requirements ranged from small distillations to server-scale systems. K3’s scheduled weight date was after the July 24 cutoff and had not yet occurred. The cases span staged, same-day, multi-size, and anno… view at source ↗
Figure 2
Figure 2. Figure 2: Observed access asymmetry in the July 2026 incident [8, 19]. OpenAI reported that its evaluation models operated with reduced cyber refusals. Hugging Face reported that unnamed commercial API models blocked analysis of live attack artifacts, after which it used self-hosted GLM 5.2. The two accounts involve different access conditions, and the rejected API providers remain unnamed. The analysis was complete… view at source ↗
Figure 3
Figure 3. Figure 3: Model-selected policy under the illustrative calibration as adversary substitution and opportunistic misuse vary. All four policies occupy nonempty regions. Restriction is preferred when misuse is high and adversary substitution is slow; broader access emerges as substitution accelerates or misuse falls. The black diamond is the baseline and the dashed line marks equal adversary and defender substitution r… view at source ↗
Figure 4
Figure 4. Figure 4: Two policy-boundary checks at baseline misuse. Panel A varies defensive network productivity; stronger returns to distributed participation expand safeguarded open release, consistent with Proposition 4. Panel B varies direct offense–defense conversion; offense￾favoring conversion preserves sequencing over a larger region. Letters denote the four policies listed in the legend. Faster adversary substitution… view at source ↗
Figure 5
Figure 5. Figure 5: Deterministic parameter-box scan over 2,048 low-discrepancy points. The design jointly varies 𝜆𝐷 ∈ [0.35, 1.20], 𝜆𝑆/𝜆𝐷 ∈ [0.30, 3.00], 𝑚𝑂 ∈ [0.10, 1.40], 𝜂 ∈ [0.10, 1.20], 𝑞𝑆/𝑞𝐷 ∈ [0.60, 1.80], 𝜇 ∈ [1, 4], 𝛿 ∈ [0.30, 0.85], 𝑓 ∈ [0.04, 0.25], 𝐼𝐺 ∈ [0.15, 0.65], 𝐼𝐶 ∈ [0, 0.65], and the pre-release delay cost in [0.03, 0.20]. The minimal-release irreversibility 𝐼𝑂 stays at its baseline 0.52, so the design inc… view at source ↗
Figure 6
Figure 6. Figure 6: Two scope checks beyond the reference design. Panel A repeats the same 2,048-point low-discrepancy design over nested narrow, reference, and wide parameter boxes. The shares change materially with the bounds, while each policy remains welfare-maximizing somewhere across the three scopes. Panel B introduces deterministic artifact-to-effective-use delays after broad release at the baseline calibration. The d… view at source ↗
Figure 7
Figure 7. Figure 7: Descriptive evidence from UK AISI’s July 2026 cyber evaluations [29]. Panel A reproduces the reported release-date lag to a comparably performing closed model. Panel B reproduces estimated costs for a 100-million-token cyber-range run at advertised first-party prices. These observations sit upstream of actor-specific acquisition and realized attack cost. created a confidentiality boundary. Estimating refus… view at source ↗
Figure 8
Figure 8. Figure 8: Welfare difference relative to controlled access as adversary substitution changes at three levels of opportunistic misuse. All panels share the same vertical scale; the horizontal zero line is controlled access. 20 [PITH_FULL_IMAGE:figures/full_fig_p020_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 4 canonical work pages · 3 internal anchors

  1. [1]

    Testimony before the U.S

    Sam Altman. Testimony before the U.S. Senate Committee on the Judiciary. Written testimony, June 2023. URL https: //openai.com/global-affairs/testimony-of-sam-altman-before-the-us-senate/. Accessed 2026-07-24

  2. [2]

    The case for targeted regulation

    Anthropic. The case for targeted regulation. Policy essay, October 2024. URL https://www.anthropic.com/news/ the-case-for-targeted-regulation. Accessed 2026-07-24

  3. [3]

    Interim measures for the management of generative artificial intelligence services

    Cyberspace Administration of China and six other agencies. Interim measures for the management of generative artificial intelligence services. Order No. 15; official Chinese text, July 2023. URL https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm. Effective 2023-08-15; accessed 2026-07-24

  4. [4]

    DeepSeek-R1 release

    DeepSeek. DeepSeek-R1 release. Technical release announcement, January 2025. URL https://api-docs.deepseek.com/news/ news250120. Accessed 2026-07-24

  5. [5]

    DeepSeek-R1: IncentivizingreasoningcapabilityinLLMsviareinforcementlearning.arXivpreprintarXiv:2501.12948,

    DeepSeek-AI. DeepSeek-R1: IncentivizingreasoningcapabilityinLLMsviareinforcementlearning.arXivpreprintarXiv:2501.12948,

  6. [6]

    Red skies ahead: Russia planning for its drone-driven army of tomorrow.Mil- itary Review, 106(1):1–10, 2026

    Ian DuPont, Ben Vranian, and Bryan Powers. Red skies ahead: Russia planning for its drone-driven army of tomorrow.Mil- itary Review, 106(1):1–10, 2026. URL https://www.armyupress.army.mil/Journals/Military-Review/English-Edition-Archives/ January-February-2026/Red-Skies-Ahead/

  7. [7]

    The Open-Weight Paradox: Why Restricting Access to AI Models May Undermine the Safety It Seeks to Protect

    Vinicius Santana Gomes. The open-weight paradox: Why restricting access to AI models may undermine the safety it seeks to protect. arXiv preprint arXiv:2604.17413, 2026. doi: 10.48550/arXiv.2604.17413. URL https://arxiv.org/abs/2604.17413

  8. [8]

    Security incident disclosure—july 2026

    Hugging Face. Security incident disclosure—july 2026. Security incident report, July 2026. URL https://huggingface.co/blog/ security-incident-july-2026. Accessed 2026-07-24

  9. [9]

    Onthesocietalimpactofopenfoundationmodels.arXivpreprintarXiv:2403.07918,

    Sayash Kapoor, Rishi Bommasani, Kevin Klyman, Shayne Longpre, Ashwin Ramaswami, Peter Cihon, Aspen Hopkins, Kevin Bankston,StellaBiderman,MirandaBogen,etal. Onthesocietalimpactofopenfoundationmodels.arXivpreprintarXiv:2403.07918,

  10. [10]

    Mapping the miltech war: Eight lessons from ukraine’s battlefield

    Bohdan Kostiuk, Daryna-Maryna Patiuk, Anastasiya Shapochkina, and Élie Tenenbaum. Mapping the miltech war: Eight lessons from ukraine’s battlefield. Focus stratégique 132, French Institute of International Relations (Ifri), February 2026. URL https://www.ifri.org/en/studies/mapping-miltech-war-eight-lessons-ukraines-battlefield

  11. [11]

    The Oracle's Gambit: A Game-Theoretic Framework for Responsible AI Release

    Christoph R. Landolt, Tobias Lorenz, Marta Kwiatkowska, and Mario Fritz. The oracle’s gambit: A game-theoretic framework for responsible ai release.arXiv preprint arXiv:2607.05442, 2026. doi: 10.48550/arXiv.2607.05442. URL https://arxiv.org/abs/2607. 05442

  12. [12]

    Introducing Llama 3.1: Our most capable models to date

    Meta AI. Introducing Llama 3.1: Our most capable models to date. Technical release announcement, July 2024. URL https: //ai.meta.com/blog/meta-llama-3-1/. Accessed 2026-07-24

  13. [13]

    Global AI governance action plan

    Ministry of Foreign Affairs of the People’s Republic of China. Global AI governance action plan. Official policy statement, July 2025. URL https://www.mfa.gov.cn/eng/zy/gb/202507/t20250729_11679232.html. Accessed 2026-07-24

  14. [14]

    Chair’s statement of the 2026 world artificial intelligence conference and high-level meeting on global AI governance

    Ministry of Foreign Affairs of the People’s Republic of China. Chair’s statement of the 2026 world artificial intelligence conference and high-level meeting on global AI governance. Official conference statement, July 2026. URL https://www.mfa.gov.cn/eng/xw/ zyxw/202607/t20260717_11984715.html. Conference held 2026-07-17 to 2026-07-20; accessed 2026-07-24

  15. [15]

    Why Open Source? A Game-Theoretic Analysis of the AI Race

    Andjela Mladenovic, Aaron Courville, and Gauthier Gidel. Why open source? a game-theoretic analysis of the ai race.arXiv preprint arXiv:2604.16227, 2026. doi: 10.48550/arXiv.2604.16227. URL https://arxiv.org/abs/2604.16227

  16. [16]

    Kimi K3: Open frontier intelligence

    Moonshot AI. Kimi K3: Open frontier intelligence. Technical release announcement, July 2026. URL https://www.kimi.com/blog/ kimi-k3. Accessed 2026-07-24; full weight release announced for 2026-07-27

  17. [17]

    Managing misuse risk for dual-use foundation models

    National Institute of Standards and Technology. Managing misuse risk for dual-use foundation models. Technical Report NIST AI 800-1 2pd (Second Public Draft), U.S. Department of Commerce, January 2025. URL https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI. 800-1.ipd2.pdf

  18. [18]

    Introducing gpt-oss

    OpenAI. Introducing gpt-oss. Technical release, August 2025. URL https://openai.com/index/introducing-gpt-oss/. Accessed 2026-07-24

  19. [19]

    OpenAI and Hugging Face partner to address security incident during model evaluation

    OpenAI. OpenAI and Hugging Face partner to address security incident during model evaluation. Preliminary security incident disclosure, July 2026. URL https://openai.com/index/hugging-face-model-evaluation-security-incident/. Accessed 2026-07-24

  20. [20]

    Garrett M. Searle. Tactical reconnaissance strike in ukraine: A mandate for the U.S. Army.Infantry, pages 38–45, Spring 2025. URL https://www.lineofdeparture.army.mil/Journals/Infantry/Infantry-Archive/Spring-2025/ Tactical-Reconnaissance-Strike-in-Ukraine/. 18 Who Does Withholding Delay? Daniel Commey

  21. [21]

    Wei, Christoph Winter, Mackenzie Arnold, Seán Ó hÉigeartaigh, Anton Korinek, et al

    Elizabeth Seger, Noemi Dreksler, Richard Moulange, Emily Dardaman, Jonas Schuett, K. Wei, Christoph Winter, Mackenzie Arnold, Seán Ó hÉigeartaigh, Anton Korinek, et al. Open-sourcing highly capable foundation models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives.arXiv preprint arXiv:2311.09227, 2023. doi: 10...

  22. [22]

    Slater, Michael Purcell, and Andrew M

    Matthew R. Slater, Michael Purcell, and Andrew M. Del Gaudio, editors.Considering Russia: Emergence of a Near Peer Competitor. Marine Corps University Press, Quantico, VA, 2017. URL https://www.govinfo.gov/content/pkg/GOVPUB-D214-PURL-gpo83374/ pdf/GOVPUB-D214-PURL-gpo83374.pdf

  23. [23]

    The gradient of generative ai release: Methods and considerations

    Irene Solaiman. The gradient of generative ai release: Methods and considerations. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 111–122. Association for Computing Machinery, 2023. doi: 10.1145/3593013. 3593981. URL https://doi.org/10.1145/3593013.3593981

  24. [24]

    Release strategies and the social impacts of language models.arXiv preprint arXiv:1908.09203,

    Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. Release strategies and the social impacts of language models.arXiv preprint arXiv:1908.09203,

  25. [25]

    Winning the race: America’s AI action plan

    The White House. Winning the race: America’s AI action plan. Technical report, Executive Office of the President, July 2025. URL https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf. Accessed 2026-07-24

  26. [26]

    White house launches gold eagle initiative for unprecedented cybersecurity vulner- ability coordination

    The White House. White house launches gold eagle initiative for unprecedented cybersecurity vulner- ability coordination. Official announcement, July 2026. URL https://www.whitehouse.gov/releases/2026/07/ white-house-launches-gold-eagle-initiative-for-unprecedented-cybersecurity-vulnerability-coordination/. Accessed 2026-07-24

  27. [27]

    National security presidential memorandum/NSPM-11: Artificial intelligence in the national se- curity enterprise

    The White House. National security presidential memorandum/NSPM-11: Artificial intelligence in the national se- curity enterprise. Presidential memorandum, June 2026. URL https://www.whitehouse.gov/presidential-actions/2026/06/ national-security-presidential-memorandum-nspm-11/. Accessed 2026-07-24

  28. [28]

    Managing risks from increasingly capable open-weight ai systems

    UK AI Security Institute. Managing risks from increasingly capable open-weight ai systems. Technical blog, August 2025. URL https://www.aisi.gov.uk/blog/managing-risks-from-increasingly-capable-open-weight-ai-systems. Accessed 2026-07-24

  29. [29]

    How far behind the frontier are leading open weight models on cyber? Technical blog, July 2026

    UK AI Security Institute. How far behind the frontier are leading open weight models on cyber? Technical blog, July 2026. URL https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber. Accessed 2026-07-24

  30. [30]

    Departmentofcommerceannouncesrescissionofbiden-eraartificial intelligence diffusion rule, strengthens chip-related export controls

    U.S.DepartmentofCommerce,BureauofIndustryandSecurity. Departmentofcommerceannouncesrescissionofbiden-eraartificial intelligence diffusion rule, strengthens chip-related export controls. Press release, May 2025. URL https://www.bis.gov/press-release/ department-commerce-announces-rescission-biden-era-artificial-intelligence-diffusion-rule-strengthens. Acce...

  31. [31]

    Department of Commerce, Bureau of Industry and Security

    U.S. Department of Commerce, Bureau of Industry and Security. Guidance regarding enforcement of license requirements for advanced computing items for entities headquartered in country group D:5 and Macau. Official export-control guidance, May 2026. URL https://media.bis.gov/media/documents/bis-guidance-may-31-2026.pdf. Accessed 2026-07-24

  32. [32]

    Department of Commerce, Bureau of Industry and Security

    U.S. Department of Commerce, Bureau of Industry and Security. Revision to license review policy for advanced computing commodities. Final rule, 91 FR 1684, January 2026. URL https://www.federalregister.gov/documents/2026/01/15/2026-00789/ revision-to-license-review-policy-for-advanced-computing-commodities. Effective 2026-01-15; accessed 2026-07-24

  33. [33]

    Estimating worst-case frontier risks of open-weight llms

    Eric Wallace, Olivia Watkins, Miles Wang, Kai Chen, and Chris Koch. Estimating worst-case frontier risks of open-weight llms. arXiv preprint arXiv:2508.03153, 2025. doi: 10.48550/arXiv.2508.03153. URL https://arxiv.org/abs/2508.03153

  34. [34]

    The model openness framework: Promoting completeness and openness for reproducibility, transparency, and usability in artificial intelligence.arXiv preprint arXiv:2403.13784, 2024

    Matt White, Ibrahim Haddad, Cailean Osborne, Xiao-Yang Yanglet Liu, Ahmed Abdelmonsef, Sachin Varghese, and Arnaud Le Hors. The model openness framework: Promoting completeness and openness for reproducibility, transparency, and usability in artificial intelligence.arXiv preprint arXiv:2403.13784, 2024. doi: 10.48550/arXiv.2403.13784. URL https://arxiv.or...

  35. [35]

    The economics of AI foundation models: Openness, competition, and governance.arXiv preprint arXiv:2510.15200, 2025

    Fasheng Xu, Xiaoyu Wang, Wei Chen, and Karen Xie. The economics of AI foundation models: Openness, competition, and governance.arXiv preprint arXiv:2510.15200, 2025. doi: 10.48550/arXiv.2510.15200. URL https://arxiv.org/abs/2510.15200

  36. [36]

    Open source AI is the path forward

    Mark Zuckerberg. Open source AI is the path forward. Meta policy essay, July 2024. URL https://about.fb.com/news/2024/07/ open-source-ai-is-the-path-forward/. Accessed 2026-07-24. 19 Who Does Withholding Delay? Daniel Commey A Additional welfare slices Defender pre-release Safeguarded open weights Minimally restricted weights Misuse = 0.20 0.5 1 2 3 -2.5 ...

  37. [2019]

    URL https://arxiv.org/abs/1908.09203

    doi: 10.48550/arXiv.1908.09203. URL https://arxiv.org/abs/1908.09203

  38. [2024]

    URL https://arxiv.org/abs/2403.07918

    doi: 10.48550/arXiv.2403.07918. URL https://arxiv.org/abs/2403.07918

  39. [2025]

    URL https://arxiv.org/abs/2501.12948

    doi: 10.48550/arXiv.2501.12948. URL https://arxiv.org/abs/2501.12948