{"id":"7fe2ece8-414f-4b88-9a7e-4c2955ab9722","arxiv_id":"2607.22957","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":14,"one_line_summary":"For dual-use AI, withholding helps only if it delays harmful actors more than defenders; the paper derives a substitution-rate threshold that decides when open release beats control.","lead":"Open-weight AI releases are usually debated as open-versus-closed, but this paper argues the real question is who gets a substitute model faster if you withhold yours. It derives a formal threshold above which broad release beats control, and gives labs and regulators an actor-by-actor accounting framework instead of a slogan.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Exponential independent acquisition times drive the headline reversal; without non-exponential or correlated robustness, the unique threshold and access-inversion sign are not established for real actors.","rationale":"The reader's weakest-assumption diagnosis is exactly right and load-bearing: eq. (4) is not a peripheral technical convenience but the source of the closed-form formulas that make the headline claims crisp. The paper's own limitations sections (§5.2, §11) concede that common shocks and non-exponential forms are excluded, yet no sensitivity analysis tests these exclusions. Since the conditional verdict was already based on this gap and the paper openly labels the calibration illustrative, my stress test does not move the verdict; it sharpens the reason for conditionality by showing that even the sign of access inversion can reverse under equal-mean non-exponential distributions. The proposed check would settle whether the missing robustness is a mathematical edge case or a practical failure. I agree with the reader's framing and see no additional hidden flaw; the deterministic reproducible code and nested parameter designs are real supporting evidence, but they vary parameters within the exponential model and therefore cannot address the distributional concern.","tokens_in":20263,"tokens_out":4700,"duration_ms":51157,"concrete_test":"Recompute Propositions 1–3 and the Fig. 3/5 policy sweeps with T_S and T_D drawn from Weibull distributions matched to the baseline means (shape 0.5 and 2.0), plus a shared Poisson shock specification to violate independence. Check two quantities: (i) whether sign(A_S − A_D) matches the mean-ordering in each cell; (ii) whether Ψ(λ_S) remains monotone and the threshold unique. If either fails, the decision rule must be revised to require full distributional estimates (survival functions), not just mean substitution times.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Every headline analytic result — access inversion (Prop. 1), asymmetric empowerment (Prop. 2), the unique proliferation-reversal threshold (Prop. 3), and the pre-release window probability (Prop. 5) — rests on eq. (4): T_S ~ Exp(λ_S), T_D ~ Exp(λ_D), T_S ⊥ T_D. The paper discloses this in §5.2 and §11, but provides no non-exponential or correlated-shock robustness. That makes the central operational claim ('estimate actor-specific substitution times') incomplete: means are not sufficient.\n\nConcretely, A_i = L_i(ρ)/ρ, where L_i is the Laplace transform of T_i, so the sign of access inversion is governed by L_S(ρ) - L_D(ρ), not by E[T_i]. A mean-preserving spread increases A_i because e^{-ρt} is convex, so a defender population with a slow mean but heterogeneous acquisition can have larger discounted access than a faster, more homogeneous adversary — reversing Prop. 1's conclusion. Similarly, ΔC_i(H) = q_i S_i(H), so the inequality ΔC_D(H) > ΔC_S(H) can hold at some horizons and fail at others unless both survival functions are exponential. Proposition 3's uniqueness uses the exponential likelihood ratio; under non-proportional hazards, Ψ(λ_S) need not be monotone and the threshold may be non-unique. Thus the paper's policy ranking is conditional on a distributional assumption that is untested and unlikely to hold uniformly across theft, distillation, foreign release, and deployment channels.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a game-theoretic model of open-weight AI release in which a laboratory chooses among controlled access, a defender-first window, safeguarded open weights, and minimally restricted open weights, while sophisticated adversaries, opportunistic adversaries, and distributed defenders differ in their ability to acquire substitutes. The main analytic results are: access inversion (Prop. 1), asymmetric empowerment (Prop. 2), a unique adversary-substitution threshold above which broad release beats control in a linear benchmark (Prop. 3), a defensive network externality condition (Prop. 4), and a credibility condition for defender-first windows (Prop. 5). A deterministic numerical implementation solves the full four-policy comparison under a convex harm function and reports policy shares over nested parameter boxes. The paper applies the framework to recent release and incident cases and concludes that release reviews should estimate actor-specific substitution times, marginal capability gains, deployment rates, defensive reach, newly enabled misuse, and nonrecallable losses.","tokens_in":20678,"tokens_out":6315,"duration_ms":66296,"significance":"If the results hold, the paper makes a useful conceptual contribution: it formalizes the intuitive but often-neglected point that withholding is precautionary only when it delays harmful actors more than defenders, and it derives concrete conditions under which restriction can backfire. The paper is unusually transparent about its assumptions and limitations, explicitly labeling its welfare parameters as illustrative and providing a reproducible deterministic sensitivity design. The closed-form propositions are correct under the stated exponential benchmark, and the numerical state-occupancy formulas check out. The main weakness is that every headline analytic result depends on the exponential independent acquisition assumption in eq. (4), and the paper does not provide robustness for non-exponential or correlated acquisition processes. Because the operational recommendation directs practitioners to estimate substitution times, this distributional dependence is not merely technical; it affects whether the stated policy-ranking conditions are sufficient or even meaningful in real settings.","major_comments":[{"comment":"The central analytic results rest on the assumption in eq. (4) that T_S ~ Exp(λ_S), T_D ~ Exp(λ_D), and T_S ⊥ T_D. The paper discloses this in §5.2 and §11, but it does not provide robustness for non-exponential or correlated substitution processes. This is load-bearing rather than cosmetic: for general T_i, A_i = L_i(ρ)/ρ where L_i is the Laplace transform, so the sign of access inversion is governed by L_S(ρ) − L_D(ρ), not by E[T_i] or λ_i. A mean-preserving spread can reverse Prop. 1's conclusion, the finite-horizon comparison in Prop. 2 depends on survival functions rather than means, and Prop. 3's monotonicity and threshold uniqueness use the exponential likelihood ratio. The practitioner rule in §1 tells reviewers to estimate 'actor-specific substitution times,' but means are insufficient. Please add a robustness analysis with non-exponential distributions (e.g., Gamma, Weibull, or","section":"§5.2 and Propositions 1–3, 5"},{"comment":"The unique threshold λ*_S = ρθ/(1−θ) is derived from the exponential functional form F(λ)=λ/(λ+ρ). The paper notes that the threshold exists only when 0<θ<1, but θ itself is defined through Ψ(0), which is computed using the exponential F. If the exponential assumption is relaxed, Ψ(λ_S) need not be strictly increasing and the threshold need not be unique; under non-proportional hazards, multiple crossings are possible. The paper should state the general condition in terms of the Laplace transforms L_S(ρ) and L_D(ρ), and, if possible, identify a class of distributions (e.g., monotone likelihood ratio) under which the threshold property survives. Without such a statement, the proposition's practical relevance for release reviews is unclear.","section":"§6.3, eqs. (24)–(26)"}],"minor_comments":[{"comment":"The state-capability counterexample is presented as a welfare difference but is not integrated into the formal model of §5. Consider marking it explicitly as a heuristic example or deriving it from the same welfare function with stated assumptions.","section":"§3.5.1, eq. (1)"},{"comment":"The nested-box sensitivity analysis varies parameter ranges but not distributional assumptions. The caption and text are clear that these are deterministic parameter designs, but a reader might over interpret the narrow/reference/wide shares as robustness to model form. A sentence noting that the boxes test parameter bounds, not the exponential assumption, would help.","section":"§7, Fig. 6"},{"comment":"The phrase 'full weights (open circle: promised)' is ambiguous. Clarify that the open circle indicates a future scheduled weight release that had not occurred as of the cutoff.","section":"Fig. 1 caption"},{"comment":"The term 'asymmetric proliferation' is used in the title and introduction but is not formally defined until the drone discussion. Define it explicitly at first use, perhaps in §4, to avoid ambiguity with 'asymmetric empowerment'.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well-organized and unusually honest, and the analytic results are correct under the stated exponential benchmark. However, the distributional robustness issue is central to the paper's operational claim, and the current text offers only a limitation paragraph rather than a fix or a clear boundary on the claim's scope. I would support acceptance after the authors either provide non-exponential/correlated robustness or explicitly rescope the operational conclusions to the exponential case and adjust the practitioner rule accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honest take: this is a genuine contribution—the first model I know that makes actor-specific substitution timing the central variable in release decisions, and it proves real results rather than hand-waving. Propositions 1–5 are correct under the stated assumptions: access inversion, asymmetric empowerment, the unique proliferation-reversal threshold, and the window bound all follow cleanly. The numerical implementation is deterministic, nested, and the code and data are public. That buys a lot of trust.\n\nThe paper is also honest about its own gaps. Sections 5.2 and 11 admit that the exponential independent acquisition assumption is doing the work, and that the empirical quantities needed to apply the decision rule are unmeasured. Those are disclosed limitations, not hidden ones. But disclosure is not robustness, and the stress-test note is right: every headline result goes through the Laplace transform of the acquisition time, so means are not sufficient. A mean-preserving spread in defender acquisition time increases discounted access because e^{-ρt} is convex, which can flip access inversion. Proposition 2's ordering can hold at some horizons and fail at others. Proposition 3's uniqueness uses the exponential likelihood ratio; under non-proportional hazards the threshold need not be unique. This is a serious gap, not a nitpick, because the paper's operational claim is \"estimate actor-specific substitution times.\" That advice only works if the distribution family is roughly right.\n\nThe policy cases and empirical observations are too thin to validate anything—the paper says so itself, and five purposive releases plus one evaluator's cyber scores do not pin down λ_S or λ_D. So the paper is a framework, not an empirical result. That is okay if it is read that way. One oddity: the acknowledgment credits GPT-5.6 Sol as a conversation partner. It does not change the assessment, but a referee might want clarification.\n\nWho should read it: governance researchers and release reviewers who want a precise vocabulary for the delay question. It will not settle any current decision, but it sharpens what evidence would settle one.\n\nMy recommendation: send it out. It needs a referee who checks the math and pushes for non-exponential robustness, but the theory is meaningful and the honesty is refreshing. Conditional acceptance with major revision would be right.","headline":"A real formal contribution on actor-specific release timing whose headline reversal rests on an exponential assumption the paper discloses but does not test.","tokens_in":21110,"tokens_out":1950,"would_cite":true,"duration_ms":22051,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Restricting access to a dual-use AI model is precautionary only if it delays harmful actors more than defenders; the paper derives a unique adversary-substitution threshold above which broad release beats controlled access.","keywords":["open-weight AI release","access inversion","asymmetric proliferation","substitute acquisition","defender-first window","release review","game-theoretic model","AI safety policy"],"falsifier":"A longitudinal dataset of real model releases recording, for each actor class, the first effective-access date under both restricted and open policies: if the empirical share acquiring by horizon H deviates materially from 1−e^{−λ_i H}, or if λ_S and λ_D are positively correlated during global events, the linear benchmark's unique threshold λ*_S = ρθ/(1−θ) would not describe the world.","tokens_in":20168,"feed_emoji":"⏳","tokens_out":6260,"duration_ms":61481,"temperature":0.7,"pith_summary":"The paper tries to establish when withholding an open-weight AI model actually buys safety. Its central condition is actor-specific: restriction is precautionary only if it delays harmful actors more than defenders. It proves that, in a linear benchmark with independent exponential substitute-acquisition times, there is a unique adversary-substitution rate above which broad release overtakes controlled access, and it characterizes when a defender-first window and removable safeguards have value. The point of the model is to force release debates to estimate concrete quantities—who gets a substitute when, what capability release adds to whom, and how fast defenders can deploy protection—rather than argue about openness in the abstract.","feed_headline":"One threshold decides when AI withholding backfires","feed_subtitle":"A game-theoretic model finds the adversary-substitution rate beyond which open release beats controlled access.","key_machinery":"The engine is the pair of independent exponential substitute-acquisition times for sophisticated adversaries and defenders (T_S∼Exp(λ_S), T_D∼Exp(λ_D)) together with the discounted-access function F(λ)=λ/(λ+ρ). The exponential form turns each policy's welfare into closed-form occupancy probabilities; the ratio F(λ_S)/F(λ_D) controls access inversion, e^{-λ_i H} controls finite-horizon empowerment, and the same F enters the unique threshold λ*_S through θ = −ρΨ(0)/(α q_S). This machinery converts actor-by-actor substitution speed into a policy ranking.","core_discovery":"The central discovery is a set of closed-form conditions for when controlled access, defender-first sequencing, safeguarded open weights, and minimally restricted open weights should be preferred. Under the assumption that restricted-access substitute times are independent exponentials with hazards λ_S and λ_D, the discounted access exposure of each population is λ_i/[ρ(λ_i+ρ)], so restriction creates a positive adversary access advantage exactly when λ_S > λ_D. Over a finite horizon H, immediate release adds capability q_i e^{-λ_i H} to population i, meaning release empowers the slower-substituting group most when usefulness is equal. In the linear benchmark, the difference between broad re","pith_inferences":["If the exponential-hazard assumption fails—say, actual substitute times follow a Weibull or common-shock process—the unique threshold may become a range or disappear; the same actor-delay accounting would still apply, but the closed forms would need rederivation.","Applied to compute-export controls, the model suggests effectiveness should be measured by how much a control moves the effective substitute-access time of target actors, not by shipment volumes or license denials alone.","A natural test: prospectively record four release milestones (announcement, hosted availability, weight availability, actor-specific deployment) across many models and compare realized first-effective-access times to the exponential benchmark; systematic deviation would refute the quantitative threshold.","The model implies release decisions should be revisited whenever a foreign substitute release or an inference-cost drop changes any λ_i; the paper gestures at this but leaves the re-review trigger unspecified."],"forward_implications":["When adversaries obtain substitutes faster than defenders (λ_S > λ_D), withholding gives adversaries a discounted access advantage, so delay-based justifications for control fail.","Under equal usefulness, immediate release adds more finite-horizon capability to defenders than to sophisticated adversaries; with unequal usefulness, the ratio condition q_D/q_S > e^{-(λ_S−λ_D)H} governs.","If endpoint conditions hold, there is a unique adversary-substitution threshold: below λ*_S control is preferred, above it broad release is preferred, with λ*_S = ρθ/(1−θ).","A defender-first window is valuable only when selected defenders deploy protection before adversaries substitute or the scheduled public release; its success probability μ/(μ+λ_S)(1−e^{−(μ+λ_S)τ}) falls as adversary substitution rises.","Removable safeguards are worth keeping when the deterred opportunistic misuse δ m_O exceeds the friction cost β f d_O plus lost benefits and irreversibility differences; otherwise minimally restricted release wins."],"fun_headline_variants":["One number decides if AI withholding backfires","The AI release math: fast substitutes flip the choice","When open weights beat controlled AI: a threshold","Why AI secrecy can actually speed up bad actors","Game theory: when to open-source your AI model"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is that each actor's wait for an adequate substitute, under restriction, follows a simple exponential clock and that those clocks tick independently for adversaries and defenders.","fun_headline_variants_meta":{"raw":{"variants":["One number decides if AI withholding backfires","The AI release math: fast substitutes flip the choice","When open weights beat controlled AI: a threshold","Why AI secrecy can actually speed up bad actors","Game theory: when to open-source your AI model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000156,"raw_usage":{"total_tokens":1081,"prompt_tokens":795,"completion_tokens":286,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":228}},"tokens_in":539,"tokens_out":286,"duration_ms":3710,"temperature":1.0,"reasoning_tokens":228,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:01:18.637163+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A longitudinal dataset of real model releases recording, for each actor class, the first effective-access date under both restricted and open policies: if the empirical share acquiring by horizon H deviates materially from 1−e^{−λ_i H}, or if λ_S and λ_D are positively correlated during global events, the linear benchmark's unique threshold λ*_S = ρθ/(1−θ) would not describe the world.","supporting_citations":[],"review_version":1}