REVIEW 2 major objections 7 minor 1 cited by
Position Paper: Model Access should be a Key Concern in AI Governance
T0 review · 2 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper argues that model access decisions—who gets to use, inspect, or modify an AI model—are a governance problem serious enough to warrant a dedicated research field.
desk verdict A well-scoped, honest position paper that names a useful field; the empirical agenda has construct-validity issues that the authors partly acknowledge, but the paper deserves serious engagement. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The operative framework is a three-part decomposition of access: model aspects (the components and information a developer controls, such as code, weights, training data, and outputs), access styles (the permissions and restrictions attached to an aspect, such as sampling, inspecting, fine-tuning, and modifying), and access groups (the parties who receive access, from internal staff to governments to the general public). This taxonomy replaces the binary 'open vs closed' framing with a structured question—who gets which style of access to which aspect, under what conditions—and it underlies all four recommendations. The recommendations convert the taxonomy into governance practice: extend safety evaluations across access styles, ask frontier companies to adopt 'Responsible Access Policies', build government evaluation and coordination capacity, and pursue international consensus on empirically-driven access decisions.
What would settle it
A controlled comparison in which evaluation organizations run the recommended access-style evaluations on several frontier models and find either that risks and benefits do not differ across access styles, or that decision-makers' release choices remain unchanged when presented with the evidence, would undercut the paper's central claim.
Extended reading notes
Core claim
The paper's central claim is that the downstream consequences of an AI model depend heavily on the manner and audience of access, and that these decisions are currently being made without adequate evidence or conceptual clarity. It argues that miscalibrated access can either magnify misuse and accident risks—for example by making models easier to jailbreak, allowing unretractable worldwide spread, and reducing oversight—or create serious opportunity costs, such as underused capabilities, unequal distribution of benefits from AI, and delays in safety research. To structure this problem, the paper introduces a taxonomy of model aspects, access styles, and access groups, and it sets out six open research problems around defining access elements, evaluating risks and benefits, navigating trade-offs, achieving impact, and future-proofing decisions. The paper does not claim to solve these problems; it claims that a coordinated, empirically-driven field can make them tractable.
Load-bearing premise
The paper's load-bearing premise is that more targeted empirical research will materially improve access governance decisions, rather than being ignored by decision-makers or failing to produce usable signals, and this premise is asserted rather than demonstrated.
Editorial extensions
If this is right
- If evaluations are run across access styles, decision-makers can estimate the marginal risk or benefit of each access grant before committing to irreversible releases such as open-weight distribution.
- Frontier AI companies can make access governance explicit by adding Responsible Access Policies to existing safety frameworks, with criteria for granting, restricting, rolling back, and withholding access.
- Governments and AI safety institutes can build empirical capacity to evaluate access risks, reducing reliance on voluntary corporate self-governance.
- International coordination can prevent a 'race to the bottom' in which one jurisdiction's permissive release creates global, persistent risk.
- A shared vocabulary of aspects, styles, and groups could make future legislation and corporate policy more precise than current model-category-based regulation.
Reading between the lines
- Editorial inference: the paper's 'irreversibility' emphasis implies that governance effort should be prioritized by reversibility—weight releases are more urgent to regulate than API sampling policies, because a bad decision can be retracted in one case and not the other.
- Editorial inference: the taxonomy could generalize beyond AI to other domains with digital access gradations, such as release of biological sequences, cyber exploit disclosure, or data sharing, where the same who-gets-what-in-which-form question arises.
- Editorial inference: a testable extension of the field's premise would be to build benchmark suites that score access regimes on evaluated risk and benefit dimensions, then track whether release decisions shift in response to those scores.
- Editorial inference: if the field matures, one would expect to see measurable divergence between jurisdictions that adopt evidence-based access review and those that do not, in terms of both safety incidents and realized benefits.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that current governance of AI model access is underdeveloped and proposes a new research field, Model Access Governance, to generate empirical evidence for access decisions. It defines model access in terms of model aspects, access styles, and access groups; explains the risks of miscalibrated access (jailbreak and augmentation, irreversible spread, loss of oversight, and opportunity costs); reviews the limits of existing research (lack of data, inadequate concepts, narrow scope); and makes four sets of time-bound recommendations for evaluation organizations, frontier AI companies, governments, and international bodies. Two appendices list open problems and review related work. The central claim is that targeted empirical research on access styles can materially improve governance decisions.
Significance. This is a well-structured and useful agenda-setting paper. Its value lies in making a clear taxonomy of model access, grounding recommendations in existing actor capacities (such as AISIs and Responsible Scaling Policies), and proposing concrete time horizons for action. It is transparent about the lack of empirical evidence and the provisional nature of its taxonomy, and its recommendations are actionable and falsifiable in principle. The paper does not overclaim that its proposed field will succeed; it argues for a shift in research priorities, which is an appropriate epistemic posture for a position paper. The main weakness, developed below, is that the proposed empirical program presupposes that 'access style' is a well-defined intervention despite the paper's own caveats about leaky access categories; this issue is load-bearing for the paper's central promise of evidence-based access governance.
major comments (2)
- [§3.1 (Short-term) with §2.2 and Appendix B] The proposed 6-month evaluations in §3.1 treat 'access style' as a well-defined, assignable independent variable. The paper itself concedes in §2.2 that access styles are 'not necessarily distinct' and that users 'may be able to extract information relating to one component of a model by accessing another'; Appendix B then cites model-stealing results (refs 47–52) showing that API access can approximate weight access under some conditions. Under these conditions, two deployments offering the same nominal access style can have very different effective access, and two different nominal styles can converge, so the marginal uplift of 'access style X' is not identified. Since the paper's central promise is empirically driven access governance, the paper should make construct validity and the measurement of effective access a first-phase research priority, or at minimum frame the planned evaluations as exploratory pilots whose leakage and equivalence assumptions are explicitly reported. This is not fatal to the paper's thesis, but it needs to be addressed before the recommendations can be adopted as stated.
- [§2.3] Proposition (iii) — that further targeted research will help decision-makers govern access more effectively — is the load-bearing premise of the paper, but it is asserted rather than systematically argued. The three supporting reasons (uncertainty is high, institutions are in place, AI research has affected policy) are not sufficient: high uncertainty makes evaluations informative only if they target decision-relevant quantities; existing institutions provide capacity but not a mechanism for evidence uptake; and the examples of compute governance and structured access influencing policy are analogical rather than demonstrations of evidence-driven access governance. The paper should engage directly with skeptical scenarios, such as evaluations that are not robust, not predictive of real-world harm, or ignored under commercial or political pressure, and should specify what kind of evidence would count as success or what institutional mechanisms would ensure results are used. A position paper need not prove the premise, but it should do more than state a belief.
minor comments (7)
- [§2.2] In the paragraph beginning 'We think that model access governance is unlikely to go well by default', the sentence 'First, companies developing may lack the incentives' appears to be missing an object; it should read 'companies developing AI models may lack the incentives'.
- [§2.1] In the first bullet list, the sentence 'it might be possible to use unsecured legacy models could be used to jailbreak them' contains a duplicated predicate; it should be rephrased, for example, as 'unsecured legacy models could be used to jailbreak them'.
- [Appendix A] The appendix subsections are labeled '4.1–4.6' even though the main text has only Sections 1–3; this numbering is confusing and should be changed to A.1–A.6 or similar.
- [§3.1–3.4] The time-horizon labels are inconsistent: §3.1 uses '18+ months', while §3.2 and §3.3 use '18 months +' and §3.4 uses '36 months +'; these should be standardized.
- [References] Carlini et al., 'Stealing part of a production language model', appears as both reference [31] and reference [52]; the duplicate should be removed or consolidated.
- [§3.1] The phrase 'in an reversible style' contains a grammar error; it should read 'in a reversible style'.
- [Acknowledgements] The name 'Guarav Sett' appears to be a misspelling of 'Gaurav Sett'.
Circularity Check
No circularity: the paper makes policy recommendations and identifies research gaps, with no derivation that reduces to its own inputs.
full rationale
This is a position paper rather than a derivation. Its central claim is that model access governance is under-studied and that targeted empirical research could improve access decisions. That claim is an argument about priorities, not a result derived from fitted parameters, definitions, or self-cited theorems. The paper's taxonomy of access styles is explicitly provisional, and its citations to the authors' own prior work (e.g., [17] for access styles and [37] for Responsible Access Policies) are transparent and non-load-bearing: the recommendations do not depend on those cited papers being true, only on their proposed concepts being useful. The paper even acknowledges a key limitation in Section 2.2 that access styles may not be distinct because information can be extracted across components, which cuts against, rather than circularly supports, the clean evaluative program proposed later. That tension is a construct-validity concern, not a circularity. No step in the argument equates a prediction with a fitted input, imports uniqueness from a self-citation, or renames a known result as a derivation. The paper is self-contained as a policy argument, so the appropriate score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Misgoverning model access can have serious negative consequences.
- domain assumption Further targeted research can generate evidence that materially improves access governance.
- domain assumption The taxonomy of access styles (sampling, inspecting, fine-tuning, modifying) is a useful analytic starting point.
Cite this review
Pith. "Pith review of Position Paper: Model Access should be a Key Concern in AI Governance." pith.science (2026). https://pith.science/paper/ZOG6UQFN
@misc{pith2026241200836,
author = {Pith},
title = {Pith review of: Position Paper: Model Access should be a Key Concern in AI Governance},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZOG6UQFN}},
note = {Machine review of arXiv:2412.00836}
}
read the original abstract
The downstream use cases, benefits, and risks of AI systems depend significantly on the access afforded to the system, and to whom. However, the downstream implications of different access styles are not well understood, making it difficult for decision-makers to govern model access responsibly. Consequently, we spotlight Model Access Governance, an emerging field focused on helping organisations and governments make responsible, evidence-based access decisions. We outline the motivation for developing this field by highlighting the risks of misgoverning model access, the limitations of existing research on the topic, and the opportunity for impact. We then make four sets of recommendations, aimed at helping AI evaluation organisations, frontier AI companies, governments and international bodies build consensus around empirically-driven access governance.
Figures
Forward citations
Cited by 1 Pith paper
-
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
With copyable pre-release evidence, any dual-use release rule that keeps legitimate utility q must leave worst-case attacker assistance at least Γ(q)>0, so useful capability, reliable safety, and open access cannot coexist.
Reference graph
Works this paper leans on
-
[1]
The Bletch- ley Declaration by Countries Attending the AI Safety Summit, 1-2 November 2023
Department for Science, Innovation and Technology, Foreign, Commonwealth & Development Office, and Prime Minister’s Office, 10 Downing Street. The Bletch- ley Declaration by Countries Attending the AI Safety Summit, 1-2 November 2023 . URL: https://www.gov.uk/government/publications/ai-safety-summit-2023-the-bletchley- declaration/the-bletchleydeclaration...
work page 2023
-
[2]
Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence
The White House. Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence . https://www.whitehouse.gov/briefing-room/presidential- actions/2023/10/30/executive-order-on-the-safe-secure-and-trustworthy-development-and- use-of-artificial-intelligence/. Oct. 2023
work page 2023
-
[3]
Regulation 2024/1689 of the European Parliament and of the Council
Council of European Union. Regulation 2024/1689 of the European Parliament and of the Council . http : / / eur - lex . europa . eu / legal - content / EN / TXT / ?qid = 1416170084502&uri=CELEX:32014R0269. June 2024
work page 2024
-
[4]
Frontier AI Regulation: Managing Emerging Risks to Public Safety
Markus Anderljung et al. Frontier AI Regulation: Managing Emerging Risks to Public Safety
-
[5]
Structured access: an emerging paradigm for safe AI deployment
Toby Shevlane. “Structured access: an emerging paradigm for safe AI deployment”. In: arXiv preprint arXiv:2201.05159 (2022)
arXiv 2022
-
[6]
Near to Mid-term Risks and Opportunities of Open Source Generative AI
Francisco Eiras et al. “Near to Mid-term Risks and Opportunities of Open Source Generative AI”. In: arXiv preprint arXiv:2404.17047 (2024)
arXiv 2024
-
[7]
Elizabeth Seger et al. Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives . 2023. arXiv: 2311.09227 [cs.CY]. URL: https://arxiv.org/abs/2311.09227
arXiv 2023
-
[8]
On the Societal Impact of Open Foundation Models
Sayash Kapoor et al. On the Societal Impact of Open Foundation Models . 2024. arXiv: 2403.07918 [cs.CY]. URL: https://arxiv.org/abs/2403.07918
arXiv 2024
Show all 63 references
-
[9]
Open-Source
David Evan Harris. How to Regulate Unsecured “Open-Source” AI: No Exemptions. Tech. rep. Tech Policy Press, December 4 2023
2023
-
[10]
Open Source AI Is the Path Forward | Meta
Mark Zuckerberg. Open Source AI Is the Path Forward | Meta. https://about.fb.com/ news/2024/07/open- source- ai- is- the- path- forward/ . [Accessed 27-07-2024]. 2024
2024
-
[11]
Joint Statement on AI Safety and Openness — open.mozilla.org
Mozilla. Joint Statement on AI Safety and Openness — open.mozilla.org . https://open. mozilla.org/letter/. [Accessed 27-07-2024]. October 31, 2023
2024
-
[13]
Responsible Scaling Policies (https://metr.org/blog/2023-09-26-rsp/)
METR. Responsible Scaling Policies (https://metr.org/blog/2023-09-26-rsp/) . Tech. rep. METR, 26 September 2023
2023
-
[14]
Beyond Open vs
Jon Bateman et al. Beyond Open vs. Closed: Emerging Consensus and Key Questions for Foundation AI Model Governance. https://carnegieendowment.org/research/2024/ 07 / beyond - open - vs - closed - emerging - consensus - and - key - questions - for - foundation-ai-model-governan...
2024
-
[15]
AI governance: a research agenda
Allan Dafoe. “AI governance: a research agenda”. In: Governance of AI Program, Future of Humanity Institute, University of Oxford: Oxford, UK 1442 (2018), p. 1443
2018
-
[16]
Towards a Framework for Openness in Foundation Models: Pro- ceedings from the Columbia Convening on Openness in Artificial Intelligence
Adrien Basdevant et al. Towards a Framework for Openness in Foundation Models: Pro- ceedings from the Columbia Convening on Openness in Artificial Intelligence. 2024. arXiv: 2405.15802 [cs.SE]. URL: https://arxiv.org/abs/2405.15802
2024 arXiv
-
[17]
Bucknall and Robert F
Benjamin S. Bucknall and Robert F. Trager. Structured access for third-party research on frontier AI models: Investigating researchers’ model access requirements. Tech. rep. Oxford Martin School of Governance, Oct. 2023. 7
2023
-
[18]
Ethical and social risks of harm from language models
Laura Weidinger et al. “Ethical and social risks of harm from language models”. In: arXiv preprint arXiv:2112.04359 (2021)
2021 arXiv
-
[19]
Advanced AI evaluations at AISI: May update | AISI Work — aisi.gov.uk
Technical Staff. Advanced AI evaluations at AISI: May update | AISI Work — aisi.gov.uk . https://www.aisi.gov.uk/work/advanced- ai- evaluations- may- update . [Ac- cessed 19-08-2024]. 20 May 2024
2024
-
[20]
OpenAI o1 system card
OpenAI. OpenAI o1 system card . https://openai.com/index/openai- o1- system- card/. [Accessed 14-09-2024]. 2024
2024
-
[21]
About Cygnet — grayswan.ai
Gray Swan. About Cygnet — grayswan.ai. https://docs.grayswan.ai/introduction. [Accessed 14-09-2024]. 2024
2024
-
[22]
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! 2023
Xiangyu Qi et al. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! 2023. arXiv: 2310.03693 [cs.CL]. URL: https://arxiv.org/ abs/2310.03693
2023 arXiv
-
[23]
Teams of LLM Agents can Exploit Zero-Day Vulnerabilities
Richard Fang et al. Teams of LLM Agents can Exploit Zero-Day Vulnerabilities. 2024. arXiv: 2406.01637 [cs.MA]. URL: https://arxiv.org/abs/2406.01637
2024 arXiv
-
[24]
The Simple Macroeconomics of AI
Daron Acemoglu. The Simple Macroeconomics of AI. Tech. rep. National Bureau of Economic Research, 2024
2024
-
[25]
Typhoon: Thai Large Language Models
Kunat Pipatanakul et al. Typhoon: Thai Large Language Models. 2023. arXiv: 2312.13951 [cs.CL]. URL: https://arxiv.org/abs/2312.13951
2023 arXiv
-
[26]
Systematic Inequalities in Language Technology Performance across the World’s Languages
Damian Blasi, Antonios Anastasopoulos, and Graham Neubig. “Systematic Inequalities in Language Technology Performance across the World’s Languages”. In: arXiv preprint arXiv:2110.06733 (2021)
2021 arXiv
-
[27]
Black-Box Access is Insufficient for Rigorous AI Audits
Stephen Casper et al. Black-Box Access is Insufficient for Rigorous AI Audits . 2024. DOI: 10.1145/3630106.3659037 . arXiv: 2401.14446 [cs.CY]. URL: https://arxiv.org/ abs/2401.14446
2024
-
[28]
Open Sourcing the AI Revolution: Framing the debate on open source, artificial intelligence and regulation — demos.co.uk
James Ball and Carl Miller. Open Sourcing the AI Revolution: Framing the debate on open source, artificial intelligence and regulation — demos.co.uk . https : / / demos . co . uk / research/open- sourcing- the- ai- revolution- framing- the- debate- on- open- source - artificia...
2024
-
[29]
The Open Source AI Definition – draft v
Open Source Initiative. The Open Source AI Definition – draft v. 0.0.3. [Accessed 01-08-2024]. Oct. 2023
2024
-
[30]
The Gradient of Generative AI Release: Methods and Considerations
Irene Solaiman. The Gradient of Generative AI Release: Methods and Considerations. 2023. arXiv: 2302.04844 [cs.CY]. URL: https://arxiv.org/abs/2302.04844
2023 arXiv
-
[32]
AI firms mustn’t govern themselves, say ex-members of OpenAI’s board — economist.com
Helen Toner and Tasha McCauley. AI firms mustn’t govern themselves, say ex-members of OpenAI’s board — economist.com . https : / / www . economist . com / by - invitation / 2024 / 05 / 26 / ai - firms - mustnt - govern - themselves - say - ex - members - of - openais-board. [A...
2024
-
[33]
All eyes on Sacramento: SB 1047 and the AI Safety Debate
Scott Kohler. All eyes on Sacramento: SB 1047 and the AI Safety Debate . https : / / carnegieendowment . org / posts / 2024 / 09 / california - sb1047 - ai - safety - regulation?lang=en. [Accessed 14-09-2024]. 11 September 2024
2024
-
[34]
Computing Power and the Governance of Artificial Intelligence
Girish Sastry et al. “Computing Power and the Governance of Artificial Intelligence”. In:arXiv preprint arXiv:2402.08797 (2024)
2024 arXiv
-
[35]
Evaluating AI Evaluation: Perils and Prospects
John Burden. Evaluating AI Evaluation: Perils and Prospects . 2024. arXiv: 2407.09221 [cs.AI]. URL: https://arxiv.org/abs/2407.09221
2024 arXiv
-
[36]
Reasons to Doubt the Impact of AI Risk Evaluations
Gabriel Mukobi. Reasons to Doubt the Impact of AI Risk Evaluations . 2024. arXiv: 2408. 02565 [cs.CY]. URL: https://arxiv.org/abs/2408.02565
2024 arXiv
-
[37]
AI Safety Frameworks Should Include Procedures for Model Access Decisions
Edward Kembery and Tom Reed. AI Safety Frameworks Should Include Procedures for Model Access Decisions. 2024. arXiv: 2411.10547 [cs.CY] . URL: https://arxiv.org/abs/ 2411.10547
2024 arXiv
-
[38]
Envisioning a Global Regime Complex to Govern Artificial Intelligence — carnegieendowment.org
Emma Klein and Stewart Patrick. Envisioning a Global Regime Complex to Govern Artificial Intelligence — carnegieendowment.org. https://carnegieendowment.org/research/ 2024 / 03 / envisioning - a - global - regime - complex - to - govern - artificial - intelligence?lang=en. [Ac...
2024
-
[39]
Towards an international regulatory framework for AI safety: lessons from the IAEA’s nuclear safety regulations
Seokki Cha. “Towards an international regulatory framework for AI safety: lessons from the IAEA’s nuclear safety regulations”. In:Humanities and Social Sciences Communications 11.1 (2024), pp. 1–13
2024
-
[40]
Verification methods for international AI agreements
Akash R Wasil et al. “Verification methods for international AI agreements”. In:arXiv preprint arXiv:2408.16074 (2024)
2024 arXiv
-
[41]
Beyond Privacy Trade-offs with Structured Transparency
Andrew Trask et al. Beyond Privacy Trade-offs with Structured Transparency. 2024. arXiv: 2012.08347 [cs.CR]. URL: https://arxiv.org/abs/2012.08347
2024 arXiv
-
[42]
A safe harbor for ai evaluation and red teaming
Shayne Longpre et al. “A safe harbor for ai evaluation and red teaming”. In: arXiv preprint arXiv:2403.04893 (2024)
2024 arXiv
-
[43]
Auditing large language models: a three-layered approach
Jakob Mökander et al. “Auditing large language models: a three-layered approach”. In: AI and Ethics (2023), pp. 1–31
2023
-
[44]
Conformity assessments and post-market monitoring: a guide to the role of auditing in the proposed European AI regulation
Jakob Mökander et al. “Conformity assessments and post-market monitoring: a guide to the role of auditing in the proposed European AI regulation”. In: Minds and Machines 32.2 (2022), pp. 241–268
2022
-
[45]
Ethics-based auditing to develop trustworthy AI
Jakob Mökander and Luciano Floridi. “Ethics-based auditing to develop trustworthy AI”. In: Minds and Machines 31.2 (2021), pp. 323–327
2021
-
[46]
Operationalising AI governance through ethics-based auditing: an industry case study
Jakob Mökander and Luciano Floridi. “Operationalising AI governance through ethics-based auditing: an industry case study”. In: AI and Ethics 3.2 (2023), pp. 451–468
2023
-
[47]
Stealing Machine Learning Models via Prediction APIs
Florian Tramèr et al. Stealing Machine Learning Models via Prediction APIs. 2016. arXiv: 1609.02943 [cs.CR]. URL: https://arxiv.org/abs/1609.02943
2016 arXiv
-
[48]
High Accuracy and High Fidelity Extraction of Neural Networks
Matthew Jagielski et al. “High Accuracy and High Fidelity Extraction of Neural Networks”. In: 29th USENIX Security Symposium (USENIX Security 20). USENIX Association, Aug. 2020, pp. 1345–1362. ISBN : 978-1-939133-17-5. URL: https://www.usenix.org/conference/ usenixsecurity20/p...
2020
-
[49]
Model Reconstruction from Model Explanations
Smitha Milli et al. Model Reconstruction from Model Explanations. 2018. arXiv: 1807.05185 [stat.ML]. URL: https://arxiv.org/abs/1807.05185
2018 arXiv
-
[50]
Leaky DNN: Stealing Deep-Learning Model Secret with GPU Context- Switching Side-Channel
Junyin Wei et al. “Leaky DNN: Stealing Deep-Learning Model Secret with GPU Context- Switching Side-Channel”. In: 2020 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN) (2020), pp. 125–137. URL: https : / / api . semanticscholar.org/CorpusI...
2020
-
[51]
Grey-box Extraction of Natural Language Models
Santiago Zanella-Beguelin et al. “Grey-box Extraction of Natural Language Models”. In: Proceedings of the 38th International Conference on Machine Learning . Ed. by Marina Meila and Tong Zhang. V ol. 139. Proceedings of Machine Learning Research. PMLR, July 2021, pp. 12278–122...
2021
-
[52]
Stealing Part of a Production Language Model
Nicholas Carlini et al. Stealing Part of a Production Language Model. 2024. arXiv: 2403. 06634 [cs.CR]. URL: https://arxiv.org/abs/2403.06634
2024 arXiv
-
[53]
RAND Report: Securing AI Model Weights
Sella Nevo et al. “RAND Report: Securing AI Model Weights”. In: (2024)
2024
-
[54]
Vulnerability Detection in Open Source Software: An Introduction
Stuart Millar. Vulnerability Detection in Open Source Software: An Introduction. 2022. arXiv: 2203.16428 [cs.CR]. URL: https://arxiv.org/abs/2203.16428
2022 arXiv
-
[55]
Considerations for Governing Open Foundation Models
Rishi Bommasani et al. Considerations for Governing Open Foundation Models. Tech. rep. Stanford HAI, 2024
2024
-
[56]
NTIA AI Open Model Weights Request for Comments
Misc. NTIA AI Open Model Weights Request for Comments. https://www.regulations. gov/document/NTIA-2023-0009-0001/comment . [Accessed 17-09-2024]. 26 February 2024
2023
-
[57]
Introducing the Frontier Safety Framework — deepmind.google
Anca Dragan, Helen King, and Allan Dafoe. Introducing the Frontier Safety Framework — deepmind.google. https://deepmind.google/discover/blog/introducing- the- frontier-safety-framework/. [Accessed 17-09-2024]. 17 May 2024
2024
-
[58]
OpenAI Preparedness Framework
OpenAI. OpenAI Preparedness Framework . https://cdn.openai.com/openai-preparedness- framework-beta.pdf. [Accessed 17-09-2024]. 2023
2024
-
[59]
A Grading Rubric for AI Safety Frame- works
Jide Alaga, Jonas Schuett, and Markus Anderljung. A Grading Rubric for AI Safety Frame- works. 2024. arXiv: 2409.08751 [cs.CY]. URL: https://arxiv.org/abs/2409.08751
2024 arXiv
-
[60]
Responsible Scaling Policy Updates — anthropic.com
Anthropic. Responsible Scaling Policy Updates — anthropic.com. https://www.anthropic. com/rsp-updates. [Accessed 21-10-2024]. 2024
2024
-
[61]
Third-party testing as a key ingredient of AI policy — anthropic.com
Anthropic. Third-party testing as a key ingredient of AI policy — anthropic.com . https : //www.anthropic.com/news/third-party-testing . [Accessed 22-10-2024]. 9
2024
-
[62]
Meta withholds advanced AI model from EU amid regulatory uncertainty — nationaltechnology.co.uk
Jonathan Easton. Meta withholds advanced AI model from EU amid regulatory uncertainty — nationaltechnology.co.uk. https://nationaltechnology.co.uk/Meta_Witholds_ Advanced_AI_Model_From_EU.php. [Accessed 17-10-2024]
2024
-
[63]
NTIA Supports Open Models to Promote AI Innovation | National Telecommunications and Information Administration — ntia.gov
Office of Public Affairs NTIA. NTIA Supports Open Models to Promote AI Innovation | National Telecommunications and Information Administration — ntia.gov . https://www. ntia . gov / press - release / 2024 / ntia - supports - open - models - promote - ai - innovation. [Accessed...
2024
-
[64]
sufficiently detailed summary about the content used for training
NTIA. Dual-Use Foundation Models with Widely Available Model Weights. https://www. ntia.gov/sites/default/files/publications/ntia- ai- open- model- report. pdf. [Accessed 17-09-2024]. 30 July 2024. 10 A Some Open Problems in Model Access Governance This section gestures to six...
2024
-
[2023]
URL: https://arxiv.org/abs/2307.03718
arXiv: 2307.03718 [cs.CY]. URL: https://arxiv.org/abs/2307.03718
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.