REVIEW 3 major objections 4 minor 37 references
BioTIER argues that the choice is not between safe and useful models: a small, precisely defined tier of catastrophic-risk biology should be gated, while the rest stays open, and it supplies a 542-prompt benchmark to measure both.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 01:57 UTC pith:C3VQKJXI
load-bearing objection A solid, well-executed refusal benchmark whose numbers you should treat as provisional until the prompts and labels are released. the 3 major comments →
BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
BioTIER's central discovery is a measurement, not a mechanism: when the same 542 prompts are put to 52 current models, refusal of high-risk biological content spans a 90-point spread, while permission of benign content is mostly compressed near the ceiling. Models that refuse the most dangerous prompts also over-refuse close-to-boundary benign prompts, but only modestly — most still answer over 85% of the benign set. One anonymized high-risk theme is a near-universal blind spot across the panel, and four frontier models re-run weeks later shifted by 12 to 28 points on refusal while barely moving on benign answering. The paper reads these results as evidence that a graduated, tiered access po
What carries the argument
The load-bearing object is BioTIER's three-tier risk taxonomy — Catastrophe Avoidance (CA), Biomedical DURC (BD) and Related Biology (RB) — realized as 542 manually written prompts with metadata linking each prompt to themes, agent types, select-agent tags, and presumed actor skill. The taxonomy is what makes differentiated access policies thinkable: a model can refuse CA universally, permit BD to verified researchers, and answer RB freely. The benchmark is dual-component: BioTIER-refuse (CA+BD) scores whether a model refuses, while BioTIER-permit (RB) scores whether it answers, so under- and over-refusal are measured on the same scale.
Load-bearing premise
All 542 prompts carry expert risk labels, and every compliance number inherits those labels; if the CA/BD/RB boundary is mis-calibrated — which the paper concedes is possible — the 90-point spread may measure label noise as much as genuine policy differences.
What would settle it
Have an independent panel of experts re-label all 542 prompts from scratch; if the resulting compliance ranking of the 52 models shifts materially, the instrument's discriminative claim fails. A second decisive check: take the prompts one model answers, verify whether the answers are actually accurate and actionable — the paper measures refusal, not the quality of what passes — and see whether the highest- and lowest-compliance models still separate.
If this is right
- A model can be both safe and useful: the paper shows a model that raised high-risk refusal by 28 percentage points between runs kept its benign-answer rate nearly unchanged.
- A single model's strict refusal policy provides little ecosystem-level protection: querying just six models yields answers to roughly 97.5% of the dangerous prompts, supporting the case for a shared minimum safety baseline.
- Topical blind spots can be localized: one anonymized category accounted for about 42–45% of all failures among the strongest refusers and was answered by nearly every other model, so targeted mitigations are possible.
- Refusal behavior drifts over time: four frontier models re-evaluated weeks later shifted by 12 to 28 percentage points on refusal while barely changing on benign answering, making longitudinal tracking necessary.
Where Pith is reading between the lines
- If the tier taxonomy became an industry baseline, the ensemble-elicitation number — six models answer ~97.5% of dangerous prompts — implies that enforcement must be coordinated; any developer that ignores the baseline re-creates the model-shopping loophole.
- A natural extension the paper flags but does not resolve is quality-graded scoring: the same binary refusal score groups a model that answers a CA prompt with accurate, actionable detail together with one that answers with vague talk, so real-world risk could differ even for identical compliance numbers.
- The near-universal blind-spot theme suggests the failure is not per-model policy but something shared, possibly the way dangerous concepts appear in training corpora; auditing what models saw for that theme would test whether retraining is needed, not just refusal tuning.
- The tiered-spectrum logic should transfer to chemical, radiological, and other catastrophic-risk domains, where a tiny high-risk fraction is embedded in a large legitimate literature; a parallel benchmark would let the same differentiated-access logic be tested there.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. BioTIER introduces a three-tier risk taxonomy (CA, BD, RB) and a 542-prompt benchmark to measure whether LLMs refuse catastrophic-risk and dual-use biological queries while permitting benign and close-to-boundary content. The authors evaluate 52 models from 10 developers, reporting a BioTIER-refuse compliance spread from 5.6% (DeepSeek-V3.1) to 96.4% (Claude Sonnet 4.6), a compressed BioTIER-permit distribution with Anthropic models over-refusing the most, a topical blind spot in anonymized theme CA-1, substantial within-model temporal drift on re-runs of four frontier models, and an ensemble-elicitation analysis of cross-model refusal complementarity.
Significance. If the CA/BD/RB labels are accepted as valid, BioTIER is a practically useful instrument: it jointly measures under- and over-refusal, localizes specific topical gaps, and provides longitudinal tracking. The evaluation design has real strengths: 52 models, 10 epochs per prompt, a three-member model-grader jury with majority voting, explicit re-runs of four frontier models, and a candid limitations section. The ensemble-elicitation analysis is a distinctive contribution that quantifies the 'model shopping' risk. However, the benchmark's central claim—that the 5.6–96.4% spread is a discriminative measurement of biosafety-relevant refusal—depends entirely on the expert assignment of 542 prompts to CA, BD, or RB. That assignment is not independently verifiable as submitted: the dataset is not publicly released, no inter-rater reliability is reported, and the paper concedes that the risk-threshold boundary involves subjective judgment. These issues do not invalidate the work, but they are load-bearing for the benchmark's validity and for the policy conclusions drawn from the model ranking.
major comments (3)
- [§3.1.2, §5.4] The central empirical claim—the 5.6–96.4% BioTIER-refuse spread and the per-model ranking—inherits every compliance score from the expert assignment of 542 prompts to CA, BD, or RB. Section 3.1.2 reports triple-consensus approval by 15 SMEs but provides no inter-rater reliability statistic, no blind re-labeling of a held-out subset, and no quantitative agreement on the CA/BD/RB boundary. Section 5.4 concedes that 'expert classification involves subjective judgment around risk thresholds.' If that boundary is miscalibrated, the ranking in Figures 4–6 and the CA-1 blind-spot analysis reflect labeler priors as much as model behavior. This is a correctness risk, not merely a reproducibility limitation. Please report per-theme inter-rater agreement (e.g., Fleiss' κ) on a held-out set, provide the label-boundary guide used by SMEs, and make the prompt-label mapping available to reviewers under
- [Abstract / Data Availability] The abstract and introduction state 'We release BioTIER,' but the Data Availability section says the data are 'available upon reasonable request to organizations with a track record in AI safety research,' with no DOI, repository, or release version. The benchmark artifact—the 542 prompts and their labels—is therefore not actually released, and no independent party can reproduce the reported compliance rates, check for contamination, or evaluate the taxonomy. For a benchmark paper, the instrument should be accessible at least to reviewers and preferably through a gated but well-defined release mechanism (e.g., a DOI with controlled access, a detailed access policy, and a versioned release). Without this, the paper's central claims are not independently verifiable.
- [§3.3] Compliance is measured by a hierarchical procedure in which most responses are graded by a majority vote of three LLM graders. The paper reports that the jury reached consensus in ~96% of cases and that this is 'similar to human vs. model scoring comparisons,' but no human-grader agreement data are presented. Since every compliance percentage in §4.2 and §4.3 depends on this grader, any systematic grader bias toward or against refusing responses would shift the entire ranking. Please report a quantitative human-model agreement study on a random sample (e.g., Cohen's or Fleiss' kappa per set and per theme), and specify how disagreement among the three graders was handled in the remaining ~4% of cases.
minor comments (4)
- [References / §3.3] The Inspect AI URL is given as 'https://github.com/UKGovernmentBEIS/inspect ai' with a space; it should be 'inspect_ai' or a proper URL.
- [§4.2, Tables S5–S6] The CA-1 theme is anonymized, and Tables S5 and S6 use anonymized labels without a key. Since the dataset is not released, the reader cannot assess what content this 'stark blind spot' actually covers or compare it with the theme descriptions in Table S1. Even a high-level qualitative description (without dangerous detail) would help.
- [§3.2.3] The qualitative characteristics (detail, creativity, skill, practicality) are discretized into bins with boundaries at <2, 2–3.5, >3.5. The choice of these boundaries is not justified; since these bins are used in the PCA interpretation (§4.5), a sensitivity check or a continuous-score version would strengthen the claim.
- [§4.2] The paper reports large temporal drift for three of four re-run models, and Section 5.4 notes that models were evaluated on different dates. The title claim of a 'highly discriminative' ranking would be more robust if the paper explicitly quantified how much of the cross-model spread could be explained by evaluation date; the re-run data suggest this is non-negligible.
Circularity Check
No significant circularity: BioTIER is a measurement benchmark whose compliance scores are observed against expert-defined labels, not derived from fitted inputs or self-citations.
full rationale
BioTIER is a measurement instrument rather than a derivation. The paper defines CA/BD/RB sets by expert judgment (§3.1.1), writes 542 prompts, and then measures model refusal/answer rates against those labels (§3.3, §4.2). The headline compliance spread (5.6%–96.4%) is an observed distribution, not the output of any fitted model or equation whose inputs are the labels. The CA-1 'blind spot' is likewise an empirical finding: the theme is defined in the taxonomy, and model answer rates on that theme are measured independently; the finding does not feed back into label assignment. The only self-citation of note is [21] (VCT, overlapping authors) for the 10-epoch evaluation protocol; this is a minor methodological reference, not a load-bearing premise, and the protocol is externally reproducible. Section 5.4's concession that 'expert classification involves subjective judgment around risk thresholds' flags a possible construct-validity problem with the labels, but a miscalibrated boundary would make the benchmark measure the wrong construct—it would not make the measurement circular. No fitted input is renamed as a prediction; no uniqueness theorem or ansatz is imported from the authors' prior work. Hence no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (2)
- Compliance majority threshold =
>=7/10 epochs
- Qualitative characteristic bin boundaries =
low:<2, medium:2-3.5, high:>3.5
axioms (4)
- domain assumption Expert CA/BD/RB labels are correct
- domain assumption LLM grader jury approximates human judgment
- domain assumption Refusal of CA/BD content reduces misuse risk
- domain assumption Ensemble model-shopping threat model
read the original abstract
As large language models become increasingly capable, concerns about their potential to assist with biological misuse continue to grow. Prioritization of safety differs across the model ecosystem, with some models freely providing high-risk information that could be misused, and others refusing benign scientific content, potentially hindering legitimate research. Both failures stem from a lack of targeted mitigation to distinguish the most dangerous information from broader scientific content. To address this, we introduce BioTIER (Biological Targeted Information for Exclusion and Refusal), a benchmark designed to enable more targeted biological risk mitigation. BioTIER organizes biological content into three risk sets: Catastrophe Avoidance (CA), Biomedical DURC (BD) and Related Biology (RB). These sets represent a spectrum from extremely narrow high-risk topics to a broad range of benign and beneficial biological knowledge. The benchmark consists of 542 expert-curated prompts with rich associated metadata to support differentiated access policies. We release BioTIER to aid in isolating and gating the tiny fraction of information that could engender catastrophic risk from misuse, while ensuring access to the vast wealth of knowledge that is essential for advancing biological science.
Figures
Reference graph
Works this paper leans on
-
[1]
Sunishchal Dev, Charles Teague, Grant Ellison, Kyle Brady, Jeffrey Lee, Sarah L. Gebauer, Henry Alexander Bradley, Dawid Maciorowski, Bria Persaud, Jordan Despanie, Barbara Del Castello, Alyssa Worland, Michael Miller, Adrian Salas, Dave Nguyen, James Liu, Jason Johnson, Andrew Sloan, Will Stonehouse, Travis Merrill, Thomas Goode, Greg McKelvey, and Ella ...
2025
-
[2]
Junyu Luo, Xiyang Cai, and Yixue Li. Large language models for biological sequence analysis in infectious disease research.Biosafety and Health, 7(5):323–332, October 2025. ISSN 2590-0536. doi: 10.1016/j.bsheal.2025.09.007. URL https://www.sciencedirect.com/science/ article/pii/S259005362500134X
-
[3]
Zhehao Zhang, Weijie Xu, Fanyou Wu, and Chandan K. Reddy. FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning, May 2025. URL http://arxiv.org/abs/2505.08054. arXiv:2505.08054 [cs]
Pith/arXiv arXiv 2025
-
[4]
OR-Bench: An Over-Refusal Benchmark for Large Language Models, July 2025
Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho-Jui Hsieh. OR-Bench: An Over-Refusal Benchmark for Large Language Models, July 2025. URL http://arxiv.org/abs/2405.20947. arXiv:2405.20947 [cs]
Pith/arXiv arXiv 2025
-
[5]
Zhihao Zhang, Liting Huang, Guanghao Wu, Preslav Nakov, Heng Ji, and Usman Naseem. Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context, January 2026. URL http://arxiv.org/abs/2601.17642. arXiv:2601.17642 [cs]
Pith/arXiv arXiv 2026
-
[6]
About PubMed
National Library of Medicine. About PubMed. U.S. National Library of Medicine, 2026. URL https://pubmed.ncbi.nlm.nih.gov/about/. Accessed: 2026-07-03
2026
-
[7]
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, February 2024. URL http://arxiv.org/abs/2402.04249. arXiv:2402.04249 [cs]
Pith/arXiv arXiv 2024
-
[8]
Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs, August 2023
Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs, August 2023. URL http://arxiv.org/abs/2308. 13387. arXiv:2308.13387 [cs]
Pith/arXiv arXiv 2023
-
[9]
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal, April 2025
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, and Prateek Mittal. SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal, April 2025. URL http://arxiv.org/abs/2406.14598. arXiv:2406.14598 [cs]
Pith/arXiv arXiv 2025
-
[10]
HELM Safety: Towards Standardized Safety Evaluations of Language Models
Farzaan Kaiyom, Ahmed Ahmed, Yifan Mai, Kevin Klyman, Rishi Bommasani, and Percy Liang. HELM Safety: Towards Standardized Safety Evaluations of Language Models. Stanford Center for Research on Foundation Models (CRFM), November 2024. URL https: //crfm.stanford.edu/2024/11/08/helm-safety.html. 20
2024
-
[11]
Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric Xing, Furong Huang, Hao Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Mar...
Pith/arXiv arXiv 2024
-
[12]
AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, and Christopher Parisien. AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguisti...
2025
-
[13]
Licheng Pan, Yongqi Tong, Xin Zhang, Xiaolu Zhang, Jun Zhou, and Zhixuan Chu. Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 21057–21075, Suzhou, China, November 2025. Association for Computational Li...
-
[14]
Christina Q. Knight, Kaustubh Deshpande, Ved Sirdeshmukh, Meher Mankikar, Scale Red Team, SEAL Research Team, and Julian Michael. FORTRESS: Frontier Risk Evaluation for National Security and Public Safety, June 2025. URL http://arxiv.org/abs/2506.14922. arXiv:2506.14922 [cs]
Pith/arXiv arXiv 2025
-
[15]
Georgia Adamson and Gregory C. Allen. Opportunities to Strengthen U.S. Biosecurity from AI-Enabled Bioterrorism: What Policymakers Should Know. Center for Strategic and International Studies, August 2025. URL https://www.csis.org/analysis/opportunities- strengthen-us-biosecurity-ai-enabled-bioterrorism-what-policymakers-should
2025
-
[16]
Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools
Bryce Cai, Geetha Jeyapragasan, Samira Nedungadi, Jake Yukich, and Seth Donoughe. Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools. InNeurIPS 2025 Workshop on Biosecurity Safeguards for Generative AI, December 2025. URL https://openreview.net/forum?id=fDysOrWaGd
2025
-
[17]
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Michael Chen, Isabelle Barrass, Oliver Zhang, Xiaoyuan Zhu, Rishub Tamirisa, Bhrugu Bharathi, Adam Khoja, Zhenqi Zhao, Ariel...
Pith/arXiv arXiv 2024
-
[18]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. InFirst Conference on Language Modeling, 2024. URL https://openreview. net/forum?id=Ti67584b98. 21
2024
-
[19]
A benchmark of expert-level academic questions to assess AI capabilities.Nature, 649(8099):1139–1146, January 2026
Long Phan, Alice Gatti, Nathaniel Li, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Dan Hendrycks, et al. A benchmark of expert-level academic questions to assess AI capabilities.Nature, 649(8099):1139–1146, January 2026. doi: 10. 1038/s41586-025-09962-4
2026
-
[20]
Jon M. Laurent, Joseph D. Janizek, Michael Ruzo, Michaela M. Hinks, Michael J. Hammer- ling, Siddharth Narayanan, Manvitha Ponnapati, Andrew D. White, and Samuel G. Rodriques. LAB-Bench: Measuring Capabilities of Language Models for Biology Research, 2024. URL https://arxiv.org/abs/2407.10362
Pith/arXiv arXiv 2024
-
[21]
Sanders, Nathaniel Li, Long Phan, Karam Elabd, Lennart Justen, Dan Hendrycks, and Seth Donoughe
Jasper G ¨otting, Pedro Medeiros, Jon G. Sanders, Nathaniel Li, Long Phan, Karam Elabd, Lennart Justen, Dan Hendrycks, and Seth Donoughe. Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark, April 2025. URL http://arxiv.org/abs/2504.16137. arXiv:2504.16137 [cs]
Pith/arXiv arXiv 2025
-
[22]
ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity, June 2026
Andrew Bo Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman, Harmon Bhasin, and Seth Donoughe. ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity, June 2026. URL https://arxiv.org/abs/2606.11150. arXiv:2606.11150 [cs]
Pith/arXiv arXiv 2026
-
[23]
Wintermute, Harmon Bhasin, Christina M
Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, and Kenny Workman. Evaluating calibrated refusal and safe usefulness in dual-use biology settings, 2026. URL htt...
Pith/arXiv arXiv 2026
-
[24]
Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology, February
Shen Zhou Hong, Alex Kleinman, Alyssa Mathiowetz, Adam Howes, Julian Cohen, Suveer Ganta, Alex Letizia, Dora Liao, Deepika Pahari, Xavier Roberts-Gaal, Luca Righetti, and Joe Torres. Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology, February
2025
-
[25]
Williams, Barbara Del Castello, Jeffrey Lee, Derek Roberts, John P
Adeline E. Williams, Barbara Del Castello, Jeffrey Lee, Derek Roberts, John P. Tarangelo, Jay Atanda, Alejandro Colman-Lerner, Jeff Gerold, and Roger Brent. Developing a Risk- Scoring Tool for Artificial Intelligence–Enabled Biological Design: A Method to Assess the Risks of Using Artificial Intelligence to Modify Select Viral Capabilities. Technical Repo...
2026
-
[26]
Global Risk Index for AI-enabled Biological Tools: Summary Assessment & Methods Report
Toby Webster, Richard Moulange, Barbara Del Castello, James Walker, Sana Zakaria, and Cassidy Nelson. Global Risk Index for AI-enabled Biological Tools: Summary Assessment & Methods Report. Technical report, Centre for Long-Term Resilience, September 2025. URL https://www.rand.org/pubs/external publications/EP71093.html
2025
-
[27]
Barrett, Krystal Jackson, Evan R
Anthony M. Barrett, Krystal Jackson, Evan R. Murphy, Nada Madkour, and Jessica New- man. Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models, May 2024. URL http://arxiv.org/abs/2405.10986. arXiv:2405.10986 [cs]
Pith/arXiv arXiv 2024
-
[28]
Claude Opus 4.8 System Card
Anthropic. Claude Opus 4.8 System Card. Technical report, Anthropic, May 2026. URL https://www.anthropic.com/claude-opus-4-8-system-card
2026
-
[29]
Claude Opus 4.7 System Card
Anthropic. Claude Opus 4.7 System Card. Technical report, Anthropic, April 2026. URL https://www.anthropic.com/claude-opus-4-7-system-card
2026
-
[30]
Claude Opus 4.6 System Card
Anthropic. Claude Opus 4.6 System Card. Technical report, Anthropic, February 2026. URL https://www.anthropic.com/claude-opus-4-6-system-card
2026
-
[31]
GPT-5.5 System Card
OpenAI. GPT-5.5 System Card. Technical report, OpenAI, April 2026. URL https: //deploymentsafety.openai.com/gpt-5-5
2026
-
[32]
Gemini 3 Pro Frontier Safety Framework Report
Google DeepMind. Gemini 3 Pro Frontier Safety Framework Report. Technical report, Google DeepMind, November 2025. URL https://storage.googleapis.com/deepmind-media/gemini/ gemini 3 pro fsf report.pdf. 22
2025
-
[33]
Gemini 3.1 Pro Model Card
Google DeepMind. Gemini 3.1 Pro Model Card. Technical report, Google DeepMind, February 2026. URL https://deepmind.google/models/model-cards/gemini-3-1-pro/
2026
-
[34]
Activating AI Safety Level 3 Protections
Anthropic. Activating AI Safety Level 3 Protections. Technical report, Anthropic, May 2025. URL https://www-cdn.anthropic.com/807c59454757214bfd37592d6e048079cd7a7728.pdf
2025
-
[35]
Open models lag state-of-the-art closed models by 4 months
Jack Edwards and Luke Emberson. Open models lag state-of-the-art closed models by 4 months. Epoch AI, May 2026. URL https://epoch.ai/data-insights/open-closed-eci-gap
2026
-
[36]
(July run)
Bhuvana Sudarshan and Luca Righetti. Towards a Common Standard for Evaluating Frontier AI Safeguards Against Biological Misuse. Technical report, Centre for the Governance of AI, June 2026. URL https://www.governance.ai/research-paper/technical-report-towards-a- common-standard-for-evaluating-frontier-ai-safeguards-against-biological-misuse. 23 Supplement...
2026
- [2026]
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.