{"id":"9e695d67-134b-4b51-8c78-a64a8dd6e264","arxiv_id":"2506.02035","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A policy framework for shaping access to AI cyber tools to favor defenders, with three access approaches and implementation guidance.","lead":"This report proposes 'differential access', a governance strategy for AI cyber capabilities that gives defenders prioritized access while restricting attackers. It lays out three approaches, a six-step selection process, and four example schemes for frontier AI developers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The proposed asymmetry depends on a persistent frontier-vs-open-source capability gap that is assumed but never demonstrated; if open-weight models are within one capability level on high-risk cyber tasks, differential access cannot deliver the claimed pro-defender tilt.","rationale":"The reader identifies enforceability as the weakest assumption: adversaries may obtain equivalent capabilities through open-source models, theft, or independent development. My concern is closely related but more specific: even if all access controls work perfectly, the scheme only helps if the gated frontier capability is sufficiently ahead of what adversaries can legally or illegally obtain elsewhere. The report itself supplies reasons to doubt a large persistent gap, including the near-frontier performance of open-weight models, the low cost of downstream scaffolding, and the observation in Section 1.3 that open-source models may already match or exceed frontier capabilities. This is not an internal logical inconsistency; the authors acknowledge the risk. It is an unquantified empirical premise that is load-bearing for the entire framework. I do not think this warrants changing the reader's CONDITIONAL verdict, because the paper is explicitly a framework or proposal rather than a demonstrated empirical result, and the stated limitations already call for further threat modeling. The conditionality should be understood as including a requirement to demonstrate, with current evaluations, that a meaningful capability gap exists on high-risk tasks. My proposed check would settle this directly: compare frontier and open-weight cyber capability levels under realistic scaffolding. If the gap is absent, the central 'asymmetry by design' claim loses its foundation; if the gap is present and durable, the proposal gains a concrete, falsifiable basis. This is why I agree only partially with the reader: the reader's enforceability concern is real, but the capability-gap premise is even more fundamental, since even perfect enforcement cannot produce asymmetry without a gap.","tokens_in":56286,"tokens_out":4007,"duration_ms":46401,"concrete_test":"Compile UK AISI-style cyber capability evaluations for the best proprietary frontier models and the best open-weight models (e.g., DeepSeek, Llama, Qwen) on a shared task suite covering high-risk capabilities such as vulnerability discovery, exploit development, and operational security. Measure each system both raw and with standard scaffolding and fine-tuning. If the best open-weight system is within one UK AISI capability level of the gated frontier system on the tasks used to justify Manage or Deny by Default, then the claimed asymmetric advantage is unsupported; the framework would need to demonstrate a durable gap or identify capabilities that cannot be scaffolded from open weights.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that differential access can tilt the offense-defense balance by giving defenders advanced AI cyber capabilities while denying them to attackers. For this to hold, the gated frontier systems must be materially more capable than what adversaries can obtain elsewhere. Section 1.1 concedes that adversaries can tap open-source models just behind the frontier, steal advanced models, or develop their own capabilities, and Section 1.3 admits that when open-source models match or exceed a foundation model's capabilities, developers may prefer Promote Access over restriction. That makes the argument conditional on a persistent capability gap, but the report offers no evidence for it. Worse, Section 2.1 emphasizes that downstream fine-tuning, scaffolding, and tool integration can significantly increase capability at a fraction of the cost, and footnote 55 reports that simple scaffolding lifts a model's CTF score from 72% to 95%. This suggests the effective capability an adversary can assemble from open weights plus tooling may be far closer to the frontier than raw model evaluations imply. If the gap is small or transient, Manage Access and Deny by Default merely delay attacker access while imposing administrative burdens on defenders, eroding the claimed asymmetry. Section 1.5 adds a related internal risk: defenders granted access become high-value targets, but no analysis quantifies credential theft, insider compromise, or leakage from those defenders. The framework is coherent as a proposal, but its central benefit rests on an unquantified empirical premise about capability gaps and adversarial adaptation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This policy report proposes \"differential access\" as a strategy for frontier AI developers and policymakers to shape access to AI-enabled cyber (AIxCyber) capabilities in order to give cyber defenders an asymmetric advantage over attackers. It defines three approaches along a restriction continuum—Promote Access, Manage Access, and Deny by Default—and offers a six-step selection process based on model capability level, defender maturity and criticality, strategic considerations, and technical infrastructure. It then presents four illustrative schemes (CNI innovators accelerator, dual-use authorization for security researchers, rapid response force of keystone defenders, and high-capability adversarial testing as a service) and discusses future research needs and limitations.","tokens_in":56537,"tokens_out":3599,"duration_ms":35594,"significance":"If the central claim holds, the framework could provide a practical vocabulary and decision structure for an important policy problem: how to distribute increasingly capable AI cyber tools without handing them to attackers. The report's strengths are its clarity, its systematic synthesis of existing taxonomies (UK AISI cyber capability levels, Pattern Labs offensive capabilities), the explicit naming of defender archetypes, and a transparent acknowledgment of unresolved issues. It also usefully stresses that access policy should not be driven solely by misuse prevention but should actively prioritize defenders. However, the significance is conditional: the promised pro-defender asymmetry is asserted rather than demonstrated, and the report itself concedes that the necessary threat modeling and technical-control evidence are not yet in place. The paper is best read as a research agenda or a framework proposal, not as a validated policy model.","major_comments":[{"comment":"The central claim that differential access can tilt the offense-defense balance assumes a persistent capability gap between gated frontier systems and what adversaries can obtain elsewhere, but the report provides no evidence for this gap and several of its own statements undercut it. Section 1.1 concedes that \"access restrictions cannot fully prevent adversaries from gaining access to similar capabilities via tapping open-source models just behind the frontier,\" and Section 1.3 states that when open-source models \"already match or exceed\" a foundation model's capabilities, developers may prefer Promote Access. More importantly, Section 2.1 emphasizes that \"downstream fine-tuning, scaffolding, tool integration, and other enhancements can significantly increase the capability level of a foundation model for a fraction of the cost,\" and footnote 55 reports that simple scaffolding raises an intercode-CTF score from 72% to 95%. If the effective capability an adversary can assemble from open weights plus tooling is close to the frontier, then Manage Access and Deny by Default delay attackers only marginally while burdening defenders, eroding the claimed asymmetry. The authors should either supply empirical evidence of a persistent, material capability gap on high-risk cyber tasks or explicitly reframe the paper's contribution as a conditional framework whose benefit depends on such a gap.","section":"§1.1, §1.3, §2.1, footnote 55"},{"comment":"The process for selecting among the three approaches is underdetermined because the report never provides a worked threat model, and the authors themselves identify this as a limitation. Section 7.1 states that \"without rigorous, domain-specific models, it is difficult to justify or calibrate levels of differential access,\" and Section 7.2 says that \"identifying specific AI-enabled threats is beyond the scope of this report.\" But Step 1 of the selection process (Section 1.2) requires assessing risk of misuse to choose an approach, and the four schemes in Section 6 are each premised on a concrete threat scenario (CNI attacks, exploit proliferation, accelerated patch cycles, offensive overhang). Without at least one detailed worked example that quantifies threat likelihood, impact, and the marginal effect of access restrictions, the framework cannot demonstrate when Manage Access, rather than Promote Access or Deny by Default, is the rational choice. The report should include a concrete threat-modeling case study or explicitly present the framework as a heuristic requiring additional threat analysis before operational use.","section":"§1.2, §7.1, §7.2"},{"comment":"The report identifies but does not analyze the risk that defenders granted differential access become high-value targets whose compromise would leak the very capabilities the scheme aims to restrict. Section 1.5 notes that \"the defenders who are granted differential access to these advanced capabilities will be valuable targets for malicious actors,\" and Section 3.5 lists credential theft and unauthorized repurposing among the risks for Force Multipliers. However, the paper provides no quantitative or probabilistic assessment of model theft, insider compromise, or leakage from vetted defenders. If leakage rates are non-negligible, differential access could function as a deliberate distribution channel for attackers rather than a barrier, reversing the claimed asymmetry. The authors should analyze this failure mode—for example, by estimating the additional exposure created by granting access to many defenders and comparing it with the reduction in attacker access achieved by restricting public release.","section":"§1.5, §3.5"},{"comment":"The Manage Access and Deny by Default approaches depend on technical controls whose reliability at scale is not established. Section 5.2 proposes input/output classifiers, circuit breakers, unlearning, and distillation as means of restricting capabilities, but immediately cautions that \"more research is needed to demonstrate if these methods can successfully limit a model's capabilities while leaving it sufficiently useful for narrow tasks.\" If these controls fail against jailbreaks or are easily circumvented, the distinction between Manage Access and full public release disappears, and the entire framework loses its operational meaning. The paper should either provide evidence from existing classifier or unlearning evaluations that these controls can be made reliable, or explicitly state that the framework's viability is contingent on future technical advances and is not yet ready for deployment.","section":"§5.2, Cyber-Tool Provision"}],"minor_comments":[{"comment":"The ordering of the three approaches is inconsistent: the TOC lists Promote Access (1.3), Manage Access (1.4), and Deny by Default (1.5), but the introductory text in Section 1.2 lists them as \"Promote Access, Deny by Default, and Manage Access.\" Please align the order.","section":"§1.2"},{"comment":"In the sample tiered access table, the Tier 1 entry says users can \"fully automate clearly malicious cyber operations\"; for defenders this is presumably intended to mean offensive operations in authorized contexts, but the phrase is confusing and should be reworded.","section":"§1.4, Table 1"},{"comment":"The examples of Keystone Defenders include specific company names and market-share figures, but the relevance of some statistics (e.g., Apple's device count) to access decisions is not made clear. A sentence linking each statistic to the defender's criticality or maturity would improve readability.","section":"§3.3"},{"comment":"Scheme A says the primary bottlenecks are around the product development and adoption lifecycle \"rather than novel technical methods to control/promote access,\" but the preceding sections place heavy weight on technical infrastructure; a brief reconciliation would help.","section":"§6.1"},{"comment":"The section on government policy mentions DARPA's AIxCC but does not cite it; adding a reference would be helpful for readers who want to follow up.","section":"§7.1"}],"recommendation":"major_revision","confidential_remarks":"This is a policy white paper rather than a technical research contribution. The editors may wish to assess whether the journal's scope accommodates such framework proposals. The main barrier to acceptance is the gap between the strong claim in the Executive Summary and the report's own admission (Sections 7.1 and 7.2) that the threat modeling needed to calibrate the framework does not yet exist. A revision that either supplies the missing evidence or explicitly downgrades the claim to a conditional research agenda would be more defensible. The report is well written and honest about its limitations, which counts in its favor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a well-structured policy proposal, not an empirical study, and on those terms it mostly works. The genuine contribution is a clean synthesis: three access approaches (Promote, Manage, Deny) arrayed by capability level, three defender archetypes (Keystone, Low-Maturity Critical, Force Multiplier), and a concrete process for matching them. The four illustrative schemes are practical and varied, and the technical infrastructure section gives developers something to argue with. The paper is also honest: it explicitly flag the need for better threat modeling (Section 7.2) and concedes that adversaries can obtain similar capabilities through open-source models, theft, or their own development (Section 1.1).\n\nThe soft spot is the one the stress-test names, and it is load-bearing. The entire asymmetry argument depends on a persistent gap between gated frontier capabilities and what an adversary can assemble from open weights plus tooling. That gap is asserted, never measured. Worse, the paper itself provides evidence that the gap may be narrow: footnote 55 reports that simple scaffolding lifts a model's CTF score from 72% to 95%. If equivalent capability is cheap to assemble, then Manage and Deny just delay attacker access while imposing administrative costs on defenders, and the intended tilt toward defense evaporates. The paper also acknowledges, but does not analyze, that defenders with access become high-value targets for credential theft and insider compromise. These are not fatal flaws in a framework document, but they mean the central claim is conditional in a way the executive summary does not convey.\n\nThe citation pattern is fine, drawing on UK AISI levels, Pattern Labs, and structured access literature without forcing conclusions. No fitted parameters or circular claims. For what it is, it is a serious piece of thinking.\n\nI would take this paper seriously in peer review. It deserves referee time, but the reviewers should push for a revision that either provides evidence on the size and persistence of the capability gap or scopes the claims accordingly. The framework is useful even if the asymmetry remains unproven; just don't let the headline overclaim.\n\nWho benefits: AI governance researchers, frontier lab policy teams, and cybersecurity practitioners thinking about access regimes. I'd bring it to a reading group; it generates good discussion even if we don't all buy the central premise.","headline":"A usable policy framework for differential access to AI cyber capabilities, but its central pro-defender asymmetry rests on an unquantified capability gap that the paper itself does not close.","tokens_in":57052,"tokens_out":1599,"would_cite":true,"duration_ms":18612,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Targeted access to AI cyber tools can tilt the offense-defense balance toward defenders.","keywords":["differential access","AI-enabled cyber","cyber defense","frontier AI governance","access control","capability assessment","defender prioritization","dual-use capabilities"],"falsifier":"Compare two groups of comparable organizations facing the same AI-enabled attack surface: one receives vetted early access to a restricted AIxCyber defensive tool, the other does not. If the treated group shows no measurable improvement in vulnerability discovery or response time, or if attackers obtain equivalent capabilities within the same window via open models, theft, or self-development, the paper's core claim fails.","tokens_in":56093,"feed_emoji":"🛡️","tokens_out":5181,"duration_ms":51453,"temperature":0.7,"pith_summary":"The paper proposes that deliberate, differentiated access to AI-enabled cyber capabilities can give defenders an asymmetric advantage over attackers. It introduces three approaches—Promote Access, Manage Access, and Deny by Default—that form a continuum from open release to tightly controlled access, and it argues that even the most restrictive option should still advantage defenders. The report supplies a six-step process for choosing an approach based on a model's cyber capability level, the defender's maturity and criticality, and strategic and technical implementation considerations. It also gives four concrete schemes. If the framework is right, frontier AI developers and policymakers have a way to avoid both unrestricted release and blanket denial as AI cyber capabilities grow.","feed_headline":"AI cyber tools should go to defenders first, report argues","feed_subtitle":"Capability level and defender maturity, not blanket bans, would decide who gets access.","key_machinery":"The machinery is a two-dimensional classification: model cyber capability levels (from technical non-expert through nation-state skill) paired with defender levels defined by maturity and criticality. On that grid sit the three access approaches—Promote Access, Manage Access, and Deny by Default—and the selection process that moves from capability assessment to goals, defender selection, approach choice, and then strategic and technical controls. The framework does the work of translating a vague worry about AI-enabled cyber attacks into a set of concrete access decisions and implementation levers.","core_discovery":"The paper's central discovery is a decision framework, not a new result about model behavior: access to AIxCyber capabilities should be shaped by capability level, defender role, and goals. The core claim is that current safeguards are insufficient because they ask whether a capability can be misused, not whether the right defenders can adopt it safely. The paper argues that three approaches—Promote Access, Manage Access, and Deny by Default—form a continuum, with increasing restrictiveness tied to higher capability levels and misuse risk, and that defender access remains a priority even at the most restrictive end. Four illustrative schemes show how the framework applies at different capability and defender levels, from accelerators for critical-infrastructure innovators to tightly controlled red-teaming as a service.","pith_inferences":["If differential access works for cyber, the same 'gate by actor maturity and criticality' pattern could transfer to other dual-use AI domains, with different capability taxonomies and defender archetypes.","A testable prediction follows: vetted defenders given early access should patch vulnerability classes faster and detect more intrusions than comparable defenders without access; program records could verify this.","Attackers will likely respond by targeting the defenders who hold privileged access, so any access scheme must budget for protecting those defenders, not only for internal misuse monitoring.","Because open-source models just behind the frontier erode exclusive access, the durable value of differential access may lie in the service layer—running adversary emulation and vulnerability discovery for defenders—rather than in controlling model weights."],"forward_implications":["Frontier AI developers can use the six-step process to decide, for a given model or derivative product, which access approach matches its measured cyber capability level.","Deny by Default does not have to mean no defensive benefit; developers can offer capability-as-a-service or reduced-capability tools to vetted defenders.","Manage Access tiering can give Keystone Defenders and vetted Force Multipliers earlier access to dual-use capabilities while keeping them from less mature actors.","Technical infrastructure—identity checks, capability classifiers, monitoring, and privacy-preserving logging—determines whether an approach is actually implementable, not just whether it sounds good.","Government incentives and safe-harbor policies can accelerate adoption among critical but under-resourced defenders such as critical-infrastructure operators."],"supporting_citations":[{"why":"Shows existing safeguards focus on preventing misuse rather than empowering defenders, motivating the need for differential access.","marker":"Footnote 2"},{"why":"Supplies the capability-level scoring system used as a proxy for cyber risk in the framework.","marker":"Footnote 4"},{"why":"Demonstrates that scaffolding and tool use can make foundation models execute complex attack sequences, supporting the need to consider downstream capability increases.","marker":"Footnote 5"},{"why":"Provides the offensive cyber capability taxonomy that the paper adapts into paired offensive and defensive capability areas.","marker":"Footnote 6"},{"why":"Discusses security levels and protection of model weights, informing the feasibility constraints on restrictive access approaches.","marker":"Footnote 23"},{"why":"Shows large performance gains from basic scaffolding in cyber tasks, motivating the defensive acceleration lever in technical infrastructure.","marker":"Footnote 55"},{"why":"Indicates the large community of security researchers available for a managed-access dual-use authorization scheme.","marker":"Footnote 63"}],"fun_headline_variants":["Differential AI access aims to tilt cyber balance to defenders","Promote, manage, or deny: AI access that favors defenders","Three-tier access plan puts cyber defenders first","AI cyber tools: access model favors defenders over attackers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The scheme assumes access restrictions can actually be maintained—that attackers cannot cheaply obtain the same or equivalent AI cyber capabilities from open-source models, theft, or their own development; if they can, withholding tools from defenders only makes defense weaker.","fun_headline_variants_meta":{"raw":{"variants":["Differential AI access aims to tilt cyber balance to defenders","Promote, manage, or deny: AI access that favors defenders","Three-tier access plan puts cyber defenders first","AI cyber tools: access model favors defenders over attackers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000488,"raw_usage":{"total_tokens":2347,"prompt_tokens":835,"completion_tokens":1512,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":451,"completion_tokens_details":{"reasoning_tokens":1447}},"tokens_in":451,"tokens_out":1512,"duration_ms":11752,"temperature":1.0,"reasoning_tokens":1447,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:07:28.495513+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare two groups of comparable organizations facing the same AI-enabled attack surface: one receives vetted early access to a restricted AIxCyber defensive tool, the other does not. If the treated group shows no measurable improvement in vulnerability discovery or response time, or if attackers obtain equivalent capabilities within the same window via open models, theft, or self-development, the paper's core claim fails.","supporting_citations":[],"review_version":1}