{"id":"66c505b0-c3d4-4935-8334-9b1904291a82","arxiv_id":"2604.23065","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A four-category disclosure framework for internal frontier AI deployments, covering capabilities, usage, safety mitigations, and governance.","lead":"The paper proposes a disclosure framework for frontier AI companies deploying advanced models internally, covering capabilities, usage, safety mitigations, and governance. A smart generalist would read it to understand what transparency standards AI labs may face under emerging regulation.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Reader's concern is reasonable but cannot be sharpened without full text; the load-bearing question is whether risk-mitigation is demonstrated or merely asserted per category.","rationale":"The reader correctly identified the load-bearing concern: the paper's usefulness hinges on whether disclosure is demonstrated to be net-beneficial and risk-mitigable, not merely asserted to be so. Since only the abstract is available, no more specific concern can be identified. The reader's UNVERDICTED verdict with LOW confidence is appropriate—there is insufficient information to assess novelty, soundness, or the depth of the risk analysis. My only refinement is to suggest that the concrete verification step focus on the 'capabilities' category specifically, as it is where the disclosure-risk tension is most acute and where the framework is most likely to either succeed or fail. If the full text provides substantive worked examples with specific mitigation mechanisms, the paper's contribution is real; if it only lists risks and asserts manageability, the framework is a taxonomy rather than actionable guidance. Either way, the abstract-only review cannot resolve this, so the verdict should remain UNVERDICTED.","tokens_in":1479,"tokens_out":1074,"duration_ms":32386,"concrete_test":"Obtain the full text and examine the 'capabilities' category section specifically. Check whether the paper provides at least one concrete worked example of a disclosure item where it (a) identifies a specific risk, (b) proposes a specific mitigation mechanism, and (c) explains why that mechanism plausibly reduces the risk without nullifying the disclosure's value. If no such worked example exists for any category, the framework's actionability claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that its four-category framework (capabilities, usage, safety mitigations, governance) provides actionable disclosure guidance that fills a gap between broad principles and concrete practice. The most load-bearing assumption is that, for each category, the paper substantively analyzes how disclosure-related risks (e.g., revealing proprietary capability information, enabling misuse) can be mitigated—rather than merely listing risks and asserting they are manageable. The abstract states the authors 'consider how disclosure-related risks can be mitigated,' but the verb 'consider' is doing significant work. If the full text's risk analysis is cursory or non-quantitative for any category, then the framework's recommendation for that category is not actionable: a developer reading it would know what to disclose but not whether the disclosure is safe to make. This is especially critical for the 'capabilities' category, where the tension between useful disclosure and competitive/misuse risk is sharpest. Without the full text, I cannot confirm whether the analysis rises above assertion, so I cannot identify a concern more specific than the reader's. The reader's weakest_assumption correctly targets this gap.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This manuscript proposes a four-category disclosure framework (capabilities, usage, safety mitigations, governance) for frontier AI developers' internal deployments of highly capable models. The authors position the framework as filling a gap between broad transparency principles and concrete disclosure guidance, arguing that internal deployments—particularly for AI R&D automation—currently face limited external oversight. For each category, the paper states it analyses benefits, limitations, and risk mitigations. The framework is intended to inform both public documents (e.g., system cards) and private regulatory reports.","tokens_in":1647,"tokens_out":925,"duration_ms":40805,"significance":"The topic is timely and policy-relevant. Internal deployments of frontier models for AI R&D automation are an emerging governance gap, and structured disclosure guidance could serve multiple stakeholders. However, I must be transparent: only the abstract was available for review. The full text was not provided, which substantially limits my ability to assess whether the framework's categories are well-justified, whether the risk analysis is substantive per category, or whether counterarguments are addressed. The strengths claimed in the abstract—structured per-category analysis of both benefits and risks, dual public/private applicability—are appropriate design goals for a governance framework, but I cannot confirm they are delivered.","major_comments":[{"comment":"The single most load-bearing claim is that, for each of the four categories, the paper substantively analyses how disclosure-related risks can be mitigated rather than merely listing risks and asserting they are manageable. The abstract's verb 'consider' is doing significant work here. Without the full text, I cannot verify whether the risk-mitigation analysis is demonstrated with concrete mechanisms, examples, or evidence—or whether it remains at the level of assertion. If, for any category (especially 'capabilities,' where the tension between useful disclosure and competitive/misuse risk is sharpest), the mitigation analysis is cursory, then the framework's recommendation for that category is not actionable: a developer would know what to disclose but not whether the disclosure is safe to make. This is the central correctness-risk concern and must be verified in full text before a firm","section":null},{"comment":"The abstract states that the framework 'could be used by developers to inform both public transparency documents... and private periodic reports required under emerging frontier AI regulation.' Whether this dual-use claim is supported depends on whether the paper provides concrete mapping from each category to specific disclosure contexts (public vs. private), addressing the different risk tolerances and audiences involved. A framework that does not distinguish what is appropriate for public vs. private reporting would be substantially weaker. Full-text verification needed.","section":null},{"comment":"The four-category structure (capabilities, usage, safety mitigations, governance) is presented as the framework's organising contribution, but the abstract provides no justification for why these four categories are necessary and sufficient. Whether the paper grounds this taxonomy in prior literature, empirical examples, or stakeholder analysis is unknown from the abstract alone. If the categories are asserted without justification, the framework's novelty claim is weakened.","section":null}],"minor_comments":[{"comment":"Abstract: 'It is essential, therefore, that developers provide evidence that internally deployed models are safe' — the strength of this normative claim ('essential') should be supported in the full text with argumentation, not assumed.","section":null},{"comment":"Abstract: the phrase 'recent work has highlighted the risks' should be backed by specific citations in the full text; the abstract gives no indication of the prior literature being engaged with.","section":null},{"comment":"The abstract does not specify whether the framework is intended primarily for a policy audience, a technical audience, or both. Clarifying the intended readership would help frame the analysis.","section":null}],"recommendation":"uncertain","confidential_remarks":"I was provided only with the abstract and reader/skeptic notes, not the full manuscript text. My assessment is necessarily limited to what can be inferred from the abstract. The three major comments identify the load-bearing points that must be verified in the full text before a substantive recommendation can be made. If the full text delivers substantive per-category risk-mitigation analysis with concrete mechanisms, grounds the four-category taxonomy, and maps categories to public vs. private disclosure contexts, the paper could warrant minor revision. If the risk analysis is cursory or the taxonomy unjustified, major revision would be appropriate. I recommend obtaining a full-text review before a final editorial decision."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for engaging with the abstract and for identifying the three most important questions the full paper must answer. We address each below. Our overarching response is that the concerns raised are legitimate and important, and we believe the full manuscript addresses them, but we acknowledge that we cannot fully resolve the referee's uncertainty without the referee reading the complete text.","responses":[{"response":"This is the most important concern and we agree it is load-bearing. The full manuscript does not merely list risks and assert they are manageable. For each of the four categories, we identify specific disclosure-related risks and then discuss concrete mitigation strategies. For the capabilities category specifically—which the referee correctly identifies as the sharpest tension—we address the competitive and misuse risks by distinguishing between information appropriate for private regulatory reporting (where confidentiality protections exist) and public disclosure (where they do not), and by proposing specific scoping mechanisms such as describing capability evaluation methodologies and results at a level of granularity sufficient for oversight without revealing reproducible details. We also discuss differential disclosure thresholds depending on the capability level. That said, we accept the referee's implicit challenge: if any reviewer finds the mitigation analysis for a given category insufficiently concrete after reading the full text, we will strengthen it. We commit to revisiting the capabilities section to ensure the mitigation mechanisms are as specific as possible, with worked examples where feasible.","revision_made":"partial","referee_comment":"Risk-mitigation analysis per category: whether the paper provides concrete mechanisms, examples, or evidence for mitigating disclosure-related risks, or merely asserts manageability. Special concern about the 'capabilities' category where disclosure vs. competitive/misuse tension is sharpest."},{"response":"The referee is right that a framework which does not distinguish public from private disclosure contexts would be substantially weaker. The full manuscript does provide this mapping. For each category, we discuss what information is appropriate for public transparency documents (e.g., system cards) versus private regulatory reports, and we explicitly address how the different risk tolerances and audiences of these two contexts affect disclosure recommendations. The distinction is woven throughout the per-category analysis rather than confined to a single section, because the appropriate disclosure level depends on the interaction between the category of information and the reporting context. We will ensure this mapping is as explicit as possible—potentially adding a summary table cross-referencing categories against disclosure contexts—to make the dual-use claim easier to verify.","revision_made":"partial","referee_comment":"Dual-use claim: whether the paper provides concrete mapping from each category to public vs. private disclosure contexts, addressing different risk tolerances and audiences."},{"response":"We agree that a taxonomy presented without justification would weaken the paper's contribution. The four categories are not asserted arbitrarily. The full manuscript grounds them in a review of existing transparency frameworks and disclosure practices (including system cards, frontier model regulations, and voluntary commitments), identifying what these frameworks already cover and where gaps exist specifically for internal deployments. The categories emerged from analyzing what information stakeholders—regulators, researchers, civil society—would need to assess the safety of internal deployments, combined with examination of what developers have historically disclosed (or failed to disclose) in analogous contexts. We do not claim the taxonomy is formally complete or exhaustive; we position it as a practically motivated organizing structure. If the referee, upon reading the full text, finds the justification insufficient, we will strengthen the literature grounding and add explicit discussion of alternative category structures we considered and rejected.","revision_made":"partial","referee_comment":"Justification for the four-category taxonomy: whether the categories are grounded in prior literature, empirical examples, or stakeholder analysis, or merely asserted."}],"tokens_in":1176,"tokens_out":1417,"duration_ms":38130,"standing_objections":["The referee's three major comments all explicitly state that full-text verification is needed before a firm assessment can be made. We cannot resolve this objection through a rebuttal alone; the referee must read the complete manuscript. We note that the review appears to have been conducted on the abstract only, and we respectfully request that the full text be reviewed before a final recommendation is issued."]},"desk_editor":{"model":"glm-5.2","letter":"Bottom line: this is a policy framework paper proposing four disclosure categories (capabilities, usage, safety mitigations, governance) for frontier AI developers' internal deployments. The gap it targets is real — there's plenty of broad transparency rhetoric but little concrete guidance on what to actually report — and a structured framework is a legitimate contribution if the analysis holds up in the full text. I can only assess the abstract, so my confidence is low across the board. The reader's scores and framing are fair given that constraint. The novelty claim (4.0) is about right: four standard governance categories organized into a disclosure framework is useful but not surprising. The significance (5.0) is reasonable — it could inform regulatory reporting standards, but that depends entirely on adoption, which is speculative. The weakest assumption the reader identifies is the right one to worry about: the abstract says the authors 'consider how disclosure-related risks can be mitigated,' and 'consider' is doing a lot of work. If the full text merely lists risks and asserts they're manageable without substantive analysis per category, then the framework tells developers what to disclose but not whether it's safe to do so. The capabilities category is where this tension is sharpest — useful capability disclosure directly conflicts with competitive and misuse concerns. Whether the paper handles this well is the make-or-break question, and I can't answer it from the abstract. The stress-test concern is well-posed but also can't be sharpened further without the full text. I don't see red flags in the abstract — it's coherent, scoped reasonably, and doesn't overclaim. But I also don't see evidence that the risk analysis rises above assertion. This is a paper that deserves a serious referee who can read the full text and assess whether the per-category analysis is substantive. The policy audience is clear: regulators, frontier AI governance teams, and standards bodies. If the risk treatment is rigorous, it's a solid contribution to an underdeveloped area. If it's cursory, it's a framework sketch dressed up as guidance. Worth sending out for full review.","headline":"Policy framework on internal deployment disclosure — useful gap-filling, but abstract-only review limits assessment of whether risk analysis is substantive or asserted","tokens_in":2098,"tokens_out":505,"would_cite":false,"duration_ms":38814,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Four-category framework for disclosing internal AI deployments","keywords":[],"falsifier":"If disclosure in one or more categories (e.g., capabilities) systematically creates more risk than benefit for developers, the framework's recommendations for that category would not be actionable without further qualification.","tokens_in":1592,"feed_emoji":"🔍","tokens_out":637,"duration_ms":31221,"temperature":0.7,"pith_summary":"The paper argues that frontier AI developers deploying highly capable models internally — for instance, to automate AI research and development — face a transparency gap: there is broad agreement that some disclosure is needed, but little concrete guidance on what exactly to disclose. The authors propose a four-category disclosure framework covering capabilities, usage, safety mitigations, and governance. For each category, they weigh the benefits of disclosure against its limitations and discuss how disclosure-related risks (such as exposing proprietary information or enabling misuse) might be mitigated. The framework is intended to be practical: developers could use it to structure both public-facing system cards and private reports submitted to regulators under emerging frontier AI oversight regimes.","feed_headline":"What should AI labs disclose about models they use internally?","feed_subtitle":"A four-category framework — capabilities, usage, safety, governance — aims to turn broad transparency principles into concrete disclosure.","key_machinery":"The paper's central object is the four-category disclosure framework itself: a partition of the disclosure space into capabilities (what the model can do), usage (how and where it is being used internally), safety mitigations (what precautions are in place), and governance (who oversees the deployment and how). The argumentative mechanism is a category-by-category cost-benefit analysis that treats each domain as requiring its own disclosure calculus rather than a uniform rule.","core_discovery":"The central contribution is a structured, four-part taxonomy — capabilities, usage, safety mitigations, and governance — that specifies what information frontier AI developers should disclose about models they deploy internally. The authors position this as bridging the gap between high-level transparency principles and actionable disclosure checklists, arguing that each category has distinct benefits and risks that must be weighed separately rather than treating disclosure as a single binary choice.","pith_inferences":[],"forward_implications":["Regulators drafting frontier AI reporting requirements could adopt or adapt the four-category structure as a minimum reporting standard for internal deployments.","AI developers could use the framework as a template for voluntary transparency disclosures, potentially reducing pressure for more prescriptive mandates.","If the framework becomes widely adopted, it could create comparability across developers' disclosures, enabling external analysts to benchmark internal deployment safety practices.","The category-by-category risk-benefit approach could be extended to other deployment contexts, such as models deployed to enterprise customers or used in critical infrastructure."],"fun_headline_variants":["A 4-part framework for AI labs to disclose internal model use","How AI labs should disclose internal deployments of frontier models","Making AI transparency concrete: what to disclose about internal deployments","Capabilities, usage, safety, governance: disclosing internal AI use","What AI developers must reveal about internal model deployments"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The framework assumes that meaningful disclosure about internal deployments is feasible and net-beneficial — that the transparency gains outweigh risks like revealing proprietary capability information or enabling misuse — but the paper may not fully demonstrate that this balance holds for every category.","fun_headline_variants_meta":{"raw":{"variants":["A 4-part framework for AI labs to disclose internal model use","How AI labs should disclose internal deployments of frontier models","Making AI transparency concrete: what to disclose about internal deployments","Capabilities, usage, safety, governance: disclosing internal AI use","What AI developers must reveal about internal model deployments"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1058,"prompt_tokens":395,"completion_tokens":663,"prompt_tokens_details":null},"tokens_in":395,"tokens_out":663,"duration_ms":17775,"temperature":1.0,"reasoning_tokens":618,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-04T16:01:04.379335+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If disclosure in one or more categories (e.g., capabilities) systematically creates more risk than benefit for developers, the framework's recommendations for that category would not be actionable without further qualification.","supporting_citations":[],"review_version":2}