Pith. sign in

REVIEW 3 major objections 4 minor 37 references

BioTIER argues that the choice is not between safe and useful models: a small, precisely defined tier of catastrophic-risk biology should be gated, while the rest stays open, and it supplies a 542-prompt benchmark to measure both.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 01:57 UTC pith:C3VQKJXI

load-bearing objection A solid, well-executed refusal benchmark whose numbers you should treat as provisional until the prompts and labels are released. the 3 major comments →

arxiv 2607.14479 v1 pith:C3VQKJXI submitted 2026-07-16 cs.CY

BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation

classification cs.CY
keywords AI safetybiosecurityrefusal benchmarksdual-use researchlanguage modelsrisk taxonomydifferentiated accessbiological misuse
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

BioTIER is a benchmark for targeted biosecurity refusal. The paper's guiding claim is that biological knowledge is not uniformly dangerous: a tiny fraction of prompts touch on catastrophic misuse, a larger share is dual-use research, and the vast majority is benign or even protective. To let AI developers gate only the dangerous slice, the authors organize 542 expert-written prompts into three risk sets — Catastrophe Avoidance, Biomedical DURC, and Related Biology — and evaluate 52 language models on whether they refuse the first two and answer the third. The measured gap is wide: high-risk refusal compliance ranges from 5.6% to 96.4%, while benign answering stays above ~98% for most models, with heavy over-refusal concentrated near the risk boundary. If those numbers hold, BioTIER is a practical instrument for finding blind spots, comparing safeguards across the ecosystem, and tracking drift over time.

Core claim

BioTIER's central discovery is a measurement, not a mechanism: when the same 542 prompts are put to 52 current models, refusal of high-risk biological content spans a 90-point spread, while permission of benign content is mostly compressed near the ceiling. Models that refuse the most dangerous prompts also over-refuse close-to-boundary benign prompts, but only modestly — most still answer over 85% of the benign set. One anonymized high-risk theme is a near-universal blind spot across the panel, and four frontier models re-run weeks later shifted by 12 to 28 points on refusal while barely moving on benign answering. The paper reads these results as evidence that a graduated, tiered access po

What carries the argument

The load-bearing object is BioTIER's three-tier risk taxonomy — Catastrophe Avoidance (CA), Biomedical DURC (BD) and Related Biology (RB) — realized as 542 manually written prompts with metadata linking each prompt to themes, agent types, select-agent tags, and presumed actor skill. The taxonomy is what makes differentiated access policies thinkable: a model can refuse CA universally, permit BD to verified researchers, and answer RB freely. The benchmark is dual-component: BioTIER-refuse (CA+BD) scores whether a model refuses, while BioTIER-permit (RB) scores whether it answers, so under- and over-refusal are measured on the same scale.

Load-bearing premise

All 542 prompts carry expert risk labels, and every compliance number inherits those labels; if the CA/BD/RB boundary is mis-calibrated — which the paper concedes is possible — the 90-point spread may measure label noise as much as genuine policy differences.

What would settle it

Have an independent panel of experts re-label all 542 prompts from scratch; if the resulting compliance ranking of the 52 models shifts materially, the instrument's discriminative claim fails. A second decisive check: take the prompts one model answers, verify whether the answers are actually accurate and actionable — the paper measures refusal, not the quality of what passes — and see whether the highest- and lowest-compliance models still separate.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A model can be both safe and useful: the paper shows a model that raised high-risk refusal by 28 percentage points between runs kept its benign-answer rate nearly unchanged.
  • A single model's strict refusal policy provides little ecosystem-level protection: querying just six models yields answers to roughly 97.5% of the dangerous prompts, supporting the case for a shared minimum safety baseline.
  • Topical blind spots can be localized: one anonymized category accounted for about 42–45% of all failures among the strongest refusers and was answered by nearly every other model, so targeted mitigations are possible.
  • Refusal behavior drifts over time: four frontier models re-evaluated weeks later shifted by 12 to 28 percentage points on refusal while barely changing on benign answering, making longitudinal tracking necessary.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the tier taxonomy became an industry baseline, the ensemble-elicitation number — six models answer ~97.5% of dangerous prompts — implies that enforcement must be coordinated; any developer that ignores the baseline re-creates the model-shopping loophole.
  • A natural extension the paper flags but does not resolve is quality-graded scoring: the same binary refusal score groups a model that answers a CA prompt with accurate, actionable detail together with one that answers with vague talk, so real-world risk could differ even for identical compliance numbers.
  • The near-universal blind-spot theme suggests the failure is not per-model policy but something shared, possibly the way dangerous concepts appear in training corpora; auditing what models saw for that theme would test whether retraining is needed, not just refusal tuning.
  • The tiered-spectrum logic should transfer to chemical, radiological, and other catastrophic-risk domains, where a tiny high-risk fraction is embedded in a large legitimate literature; a parallel benchmark would let the same differentiated-access logic be tested there.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. BioTIER introduces a three-tier risk taxonomy (CA, BD, RB) and a 542-prompt benchmark to measure whether LLMs refuse catastrophic-risk and dual-use biological queries while permitting benign and close-to-boundary content. The authors evaluate 52 models from 10 developers, reporting a BioTIER-refuse compliance spread from 5.6% (DeepSeek-V3.1) to 96.4% (Claude Sonnet 4.6), a compressed BioTIER-permit distribution with Anthropic models over-refusing the most, a topical blind spot in anonymized theme CA-1, substantial within-model temporal drift on re-runs of four frontier models, and an ensemble-elicitation analysis of cross-model refusal complementarity.

Significance. If the CA/BD/RB labels are accepted as valid, BioTIER is a practically useful instrument: it jointly measures under- and over-refusal, localizes specific topical gaps, and provides longitudinal tracking. The evaluation design has real strengths: 52 models, 10 epochs per prompt, a three-member model-grader jury with majority voting, explicit re-runs of four frontier models, and a candid limitations section. The ensemble-elicitation analysis is a distinctive contribution that quantifies the 'model shopping' risk. However, the benchmark's central claim—that the 5.6–96.4% spread is a discriminative measurement of biosafety-relevant refusal—depends entirely on the expert assignment of 542 prompts to CA, BD, or RB. That assignment is not independently verifiable as submitted: the dataset is not publicly released, no inter-rater reliability is reported, and the paper concedes that the risk-threshold boundary involves subjective judgment. These issues do not invalidate the work, but they are load-bearing for the benchmark's validity and for the policy conclusions drawn from the model ranking.

major comments (3)
  1. [§3.1.2, §5.4] The central empirical claim—the 5.6–96.4% BioTIER-refuse spread and the per-model ranking—inherits every compliance score from the expert assignment of 542 prompts to CA, BD, or RB. Section 3.1.2 reports triple-consensus approval by 15 SMEs but provides no inter-rater reliability statistic, no blind re-labeling of a held-out subset, and no quantitative agreement on the CA/BD/RB boundary. Section 5.4 concedes that 'expert classification involves subjective judgment around risk thresholds.' If that boundary is miscalibrated, the ranking in Figures 4–6 and the CA-1 blind-spot analysis reflect labeler priors as much as model behavior. This is a correctness risk, not merely a reproducibility limitation. Please report per-theme inter-rater agreement (e.g., Fleiss' κ) on a held-out set, provide the label-boundary guide used by SMEs, and make the prompt-label mapping available to reviewers under
  2. [Abstract / Data Availability] The abstract and introduction state 'We release BioTIER,' but the Data Availability section says the data are 'available upon reasonable request to organizations with a track record in AI safety research,' with no DOI, repository, or release version. The benchmark artifact—the 542 prompts and their labels—is therefore not actually released, and no independent party can reproduce the reported compliance rates, check for contamination, or evaluate the taxonomy. For a benchmark paper, the instrument should be accessible at least to reviewers and preferably through a gated but well-defined release mechanism (e.g., a DOI with controlled access, a detailed access policy, and a versioned release). Without this, the paper's central claims are not independently verifiable.
  3. [§3.3] Compliance is measured by a hierarchical procedure in which most responses are graded by a majority vote of three LLM graders. The paper reports that the jury reached consensus in ~96% of cases and that this is 'similar to human vs. model scoring comparisons,' but no human-grader agreement data are presented. Since every compliance percentage in §4.2 and §4.3 depends on this grader, any systematic grader bias toward or against refusing responses would shift the entire ranking. Please report a quantitative human-model agreement study on a random sample (e.g., Cohen's or Fleiss' kappa per set and per theme), and specify how disagreement among the three graders was handled in the remaining ~4% of cases.
minor comments (4)
  1. [References / §3.3] The Inspect AI URL is given as 'https://github.com/UKGovernmentBEIS/inspect ai' with a space; it should be 'inspect_ai' or a proper URL.
  2. [§4.2, Tables S5–S6] The CA-1 theme is anonymized, and Tables S5 and S6 use anonymized labels without a key. Since the dataset is not released, the reader cannot assess what content this 'stark blind spot' actually covers or compare it with the theme descriptions in Table S1. Even a high-level qualitative description (without dangerous detail) would help.
  3. [§3.2.3] The qualitative characteristics (detail, creativity, skill, practicality) are discretized into bins with boundaries at <2, 2–3.5, >3.5. The choice of these boundaries is not justified; since these bins are used in the PCA interpretation (§4.5), a sensitivity check or a continuous-score version would strengthen the claim.
  4. [§4.2] The paper reports large temporal drift for three of four re-run models, and Section 5.4 notes that models were evaluated on different dates. The title claim of a 'highly discriminative' ranking would be more robust if the paper explicitly quantified how much of the cross-model spread could be explained by evaluation date; the re-run data suggest this is non-negligible.

Circularity Check

0 steps flagged

No significant circularity: BioTIER is a measurement benchmark whose compliance scores are observed against expert-defined labels, not derived from fitted inputs or self-citations.

full rationale

BioTIER is a measurement instrument rather than a derivation. The paper defines CA/BD/RB sets by expert judgment (§3.1.1), writes 542 prompts, and then measures model refusal/answer rates against those labels (§3.3, §4.2). The headline compliance spread (5.6%–96.4%) is an observed distribution, not the output of any fitted model or equation whose inputs are the labels. The CA-1 'blind spot' is likewise an empirical finding: the theme is defined in the taxonomy, and model answer rates on that theme are measured independently; the finding does not feed back into label assignment. The only self-citation of note is [21] (VCT, overlapping authors) for the 10-epoch evaluation protocol; this is a minor methodological reference, not a load-bearing premise, and the protocol is externally reproducible. Section 5.4's concession that 'expert classification involves subjective judgment around risk thresholds' flags a possible construct-validity problem with the labels, but a miscalibrated boundary would make the benchmark measure the wrong construct—it would not make the measurement circular. No fitted input is renamed as a prediction; no uniqueness theorem or ansatz is imported from the authors' prior work. Hence no significant circularity.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

The paper introduces no new physical or theoretical entities. The central evidence rests on expert-defined labels, a grading procedure, and policy assumptions rather than fitted parameters. The only hand-set numbers are analytic thresholds used to summarize the evaluation results.

free parameters (2)
  • Compliance majority threshold = >=7/10 epochs
    Hand-chosen in Section 4.4 to define 'correct refusal' in per-question analysis. A stricter 10/10 threshold changes the count of universally refused prompts from 10 to 4, so results are threshold-dependent.
  • Qualitative characteristic bin boundaries = low:<2, medium:2-3.5, high:>3.5
    Hand-set in Section 3.2.3 to discretize detail, creativity, skill, and practicality scores from LLM-based rubrics. The choice affects the reported distribution differences between RB and CA/BD.
axioms (4)
  • domain assumption Expert CA/BD/RB labels are correct
    The validity of every compliance score depends on the assumption that prompts are correctly assigned to risk tiers. Section 5.4 acknowledges 'expert classification involves subjective judgment around risk thresholds, and edge cases may persist.'
  • domain assumption LLM grader jury approximates human judgment
    The hierarchical grading procedure in Section 3.3 uses a majority vote of GPT-4.1 mini, Gemini 2.5 Flash, and Claude Haiku 4.5. The paper reports ~96% consensus among graders but does not provide a full human-validated ground truth for all 542 prompts.
  • domain assumption Refusal of CA/BD content reduces misuse risk
    The policy recommendations assume that refusing substantive answers to CA and BD prompts is a meaningful intervention against biological misuse. Section 5.4 notes that binary compliance is an upper bound on information access, not a measured reduction in real-world uplift.
  • domain assumption Ensemble model-shopping threat model
    Section 5.2 argues that an adversary can query multiple models and that any single model's answer is sufficient for uplift. This assumes cross-model information aggregation is easy and that at least one permissive model is accessible to the adversary.

pith-pipeline@v1.3.0-alltime-deepseek · 24770 in / 10037 out tokens · 90945 ms · 2026-08-02T01:57:03.886302+00:00 · methodology

0 comments
read the original abstract

As large language models become increasingly capable, concerns about their potential to assist with biological misuse continue to grow. Prioritization of safety differs across the model ecosystem, with some models freely providing high-risk information that could be misused, and others refusing benign scientific content, potentially hindering legitimate research. Both failures stem from a lack of targeted mitigation to distinguish the most dangerous information from broader scientific content. To address this, we introduce BioTIER (Biological Targeted Information for Exclusion and Refusal), a benchmark designed to enable more targeted biological risk mitigation. BioTIER organizes biological content into three risk sets: Catastrophe Avoidance (CA), Biomedical DURC (BD) and Related Biology (RB). These sets represent a spectrum from extremely narrow high-risk topics to a broad range of benign and beneficial biological knowledge. The benchmark consists of 542 expert-curated prompts with rich associated metadata to support differentiated access policies. We release BioTIER to aid in isolating and gating the tiny fraction of information that could engender catastrophic risk from misuse, while ensuring access to the vast wealth of knowledge that is essential for advancing biological science.

Figures

Figures reproduced from arXiv: 2607.14479 by Eleanor M. Marshall, Faith Rovenolt, Jacob Kaffey, Jasper G\"otting, Mac Walker, Nelly Mak, Pedro Medeiros, Peter Peneder, Seth Donoughe.

Figure 1
Figure 1. Figure 1: Conceptual representation of the relative prevalence and overlap of high-risk and malicious knowledge that should be restricted, compared to the vast wealth of knowledge with beneficial use cases. “Close-to-boundary” knowledge represents information related to high-risk domains that should be accessible, and is essential for defining the threshold of risk from both directions, which is currently a gradient… view at source ↗
Figure 2
Figure 2. Figure 2: BioTIER Risk Taxonomy and Theoretical Framework. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Distribution of characteristics across sets CA and BD (BioTIER-refuse) and RB (BioTIER-permit). 3.3 Evaluation We assessed 52 frontier LLM models from 10 major developers (Anthropic, DeepSeek, Google DeepMind, Meta, Mistral, Moonshot, OpenAI, Qwen, Z.ai, and xAI) using the 542 prompts comprising the March 2026 evaluation version of BioTIER. Per-model run dates are given in Table S4. To assess temporal stab… view at source ↗
Figure 4
Figure 4. Figure 4: BioTIER-refuse vs. BioTIER-permit compliance. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: BioTIER-refuse compliance across 52 LLMs ordered from highest to lowest compliance and color coded by developer. For BioTIER-refuse, a model is considered compliant on a given prompt if it correctly refuses to answer the query. The four frontier models re-run on 2 July 2026 appear as additional bars labeled “(July run)” to distinguish from their initial “(May run)”. Mean of 10 epochs with 95% confidence in… view at source ↗
Figure 6
Figure 6. Figure 6: Mean BioTIER-permit compliance across 52 LLMs ordered from highest to lowest compliance and color coded by developer. For BioTIER-permit, a model is considered compliant on a given prompt if it correctly provides a substantive response and does not refuse. The compressed x-axis range (65–100%) reflects the narrower performance spread compared to [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Per question BioTIER-refuse compliance. Each of the 398 BioTIER-refuse questions is shown as a column and each of the 52 models as a row. Each cell shows a model’s majority behavior on a question across 10 epochs. Blue = refused in ≥ 7 of 10 epochs (correct behavior). Red = refused in ≤ 6 of 10 epochs (incorrect behavior). The ≥ 7 of 10 threshold was chosen as an indicator of model compliance within a conv… view at source ↗
Figure 8
Figure 8. Figure 8: Principal component analysis of per-model refusal fingerprints. [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: The BioTIER Roadmap The BioTIER evaluation (solid boxes, described in this paper) and planned extensions (outlined boxes, in development). Characteristics and content shown here are non-exhaustive and demonstrative. BioTIER is intended to be a living evaluation, with numerous developments already in the pipeline ( [PITH_FULL_IMAGE:figures/full_fig_p018_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

37 extracted references · 1 canonical work pages

  1. [1]

    Sunishchal Dev, Charles Teague, Grant Ellison, Kyle Brady, Jeffrey Lee, Sarah L. Gebauer, Henry Alexander Bradley, Dawid Maciorowski, Bria Persaud, Jordan Despanie, Barbara Del Castello, Alyssa Worland, Michael Miller, Adrian Salas, Dave Nguyen, James Liu, Jason Johnson, Andrew Sloan, Will Stonehouse, Travis Merrill, Thomas Goode, Greg McKelvey, and Ella ...

  2. [2]

    Large language models for biological sequence analysis in infectious disease research.Biosafety and Health, 7(5):323–332, October 2025

    Junyu Luo, Xiyang Cai, and Yixue Li. Large language models for biological sequence analysis in infectious disease research.Biosafety and Health, 7(5):323–332, October 2025. ISSN 2590-0536. doi: 10.1016/j.bsheal.2025.09.007. URL https://www.sciencedirect.com/science/ article/pii/S259005362500134X

  3. [3]

    Zhehao Zhang, Weijie Xu, Fanyou Wu, and Chandan K. Reddy. FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning, May 2025. URL http://arxiv.org/abs/2505.08054. arXiv:2505.08054 [cs]

  4. [4]

    OR-Bench: An Over-Refusal Benchmark for Large Language Models, July 2025

    Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho-Jui Hsieh. OR-Bench: An Over-Refusal Benchmark for Large Language Models, July 2025. URL http://arxiv.org/abs/2405.20947. arXiv:2405.20947 [cs]

  5. [5]

    Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context, January 2026

    Zhihao Zhang, Liting Huang, Guanghao Wu, Preslav Nakov, Heng Ji, and Usman Naseem. Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context, January 2026. URL http://arxiv.org/abs/2601.17642. arXiv:2601.17642 [cs]

  6. [6]

    About PubMed

    National Library of Medicine. About PubMed. U.S. National Library of Medicine, 2026. URL https://pubmed.ncbi.nlm.nih.gov/about/. Accessed: 2026-07-03

  7. [7]

    HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, February 2024

    Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, February 2024. URL http://arxiv.org/abs/2402.04249. arXiv:2402.04249 [cs]

  8. [8]

    Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs, August 2023

    Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs, August 2023. URL http://arxiv.org/abs/2308. 13387. arXiv:2308.13387 [cs]

  9. [9]

    SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal, April 2025

    Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, and Prateek Mittal. SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal, April 2025. URL http://arxiv.org/abs/2406.14598. arXiv:2406.14598 [cs]

  10. [10]

    HELM Safety: Towards Standardized Safety Evaluations of Language Models

    Farzaan Kaiyom, Ahmed Ahmed, Yifan Mai, Kevin Klyman, Rishi Bommasani, and Percy Liang. HELM Safety: Towards Standardized Safety Evaluations of Language Models. Stanford Center for Research on Foundation Models (CRFM), November 2024. URL https: //crfm.stanford.edu/2024/11/08/helm-safety.html. 20

  11. [11]

    Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric Xing, Furong Huang, Hao Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Mar...

  12. [12]

    AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails

    Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, and Christopher Parisien. AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguisti...

  13. [13]

    Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary

    Licheng Pan, Yongqi Tong, Xin Zhang, Xiaolu Zhang, Jun Zhou, and Zhixuan Chu. Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 21057–21075, Suzhou, China, November 2025. Association for Computational Li...

  14. [14]

    Knight, Kaustubh Deshpande, Ved Sirdeshmukh, Meher Mankikar, Scale Red Team, SEAL Research Team, and Julian Michael

    Christina Q. Knight, Kaustubh Deshpande, Ved Sirdeshmukh, Meher Mankikar, Scale Red Team, SEAL Research Team, and Julian Michael. FORTRESS: Frontier Risk Evaluation for National Security and Public Safety, June 2025. URL http://arxiv.org/abs/2506.14922. arXiv:2506.14922 [cs]

  15. [15]

    Georgia Adamson and Gregory C. Allen. Opportunities to Strengthen U.S. Biosecurity from AI-Enabled Bioterrorism: What Policymakers Should Know. Center for Strategic and International Studies, August 2025. URL https://www.csis.org/analysis/opportunities- strengthen-us-biosecurity-ai-enabled-bioterrorism-what-policymakers-should

  16. [16]

    Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools

    Bryce Cai, Geetha Jeyapragasan, Samira Nedungadi, Jake Yukich, and Seth Donoughe. Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools. InNeurIPS 2025 Workshop on Biosecurity Safeguards for Generative AI, December 2025. URL https://openreview.net/forum?id=fDysOrWaGd

  17. [17]

    Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B

    Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Michael Chen, Isabelle Barrass, Oliver Zhang, Xiaoyuan Zhu, Rishub Tamirisa, Bhrugu Bharathi, Adam Khoja, Zhenqi Zhao, Ariel...

  18. [18]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. InFirst Conference on Language Modeling, 2024. URL https://openreview. net/forum?id=Ti67584b98. 21

  19. [19]

    A benchmark of expert-level academic questions to assess AI capabilities.Nature, 649(8099):1139–1146, January 2026

    Long Phan, Alice Gatti, Nathaniel Li, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Dan Hendrycks, et al. A benchmark of expert-level academic questions to assess AI capabilities.Nature, 649(8099):1139–1146, January 2026. doi: 10. 1038/s41586-025-09962-4

  20. [20]

    Laurent, Joseph D

    Jon M. Laurent, Joseph D. Janizek, Michael Ruzo, Michaela M. Hinks, Michael J. Hammer- ling, Siddharth Narayanan, Manvitha Ponnapati, Andrew D. White, and Samuel G. Rodriques. LAB-Bench: Measuring Capabilities of Language Models for Biology Research, 2024. URL https://arxiv.org/abs/2407.10362

  21. [21]

    Sanders, Nathaniel Li, Long Phan, Karam Elabd, Lennart Justen, Dan Hendrycks, and Seth Donoughe

    Jasper G ¨otting, Pedro Medeiros, Jon G. Sanders, Nathaniel Li, Long Phan, Karam Elabd, Lennart Justen, Dan Hendrycks, and Seth Donoughe. Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark, April 2025. URL http://arxiv.org/abs/2504.16137. arXiv:2504.16137 [cs]

  22. [22]

    ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity, June 2026

    Andrew Bo Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman, Harmon Bhasin, and Seth Donoughe. ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity, June 2026. URL https://arxiv.org/abs/2606.11150. arXiv:2606.11150 [cs]

  23. [23]

    Wintermute, Harmon Bhasin, Christina M

    Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, and Kenny Workman. Evaluating calibrated refusal and safe usefulness in dual-use biology settings, 2026. URL htt...

  24. [24]

    Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology, February

    Shen Zhou Hong, Alex Kleinman, Alyssa Mathiowetz, Adam Howes, Julian Cohen, Suveer Ganta, Alex Letizia, Dora Liao, Deepika Pahari, Xavier Roberts-Gaal, Luca Righetti, and Joe Torres. Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology, February

  25. [25]

    Williams, Barbara Del Castello, Jeffrey Lee, Derek Roberts, John P

    Adeline E. Williams, Barbara Del Castello, Jeffrey Lee, Derek Roberts, John P. Tarangelo, Jay Atanda, Alejandro Colman-Lerner, Jeff Gerold, and Roger Brent. Developing a Risk- Scoring Tool for Artificial Intelligence–Enabled Biological Design: A Method to Assess the Risks of Using Artificial Intelligence to Modify Select Viral Capabilities. Technical Repo...

  26. [26]

    Global Risk Index for AI-enabled Biological Tools: Summary Assessment & Methods Report

    Toby Webster, Richard Moulange, Barbara Del Castello, James Walker, Sana Zakaria, and Cassidy Nelson. Global Risk Index for AI-enabled Biological Tools: Summary Assessment & Methods Report. Technical report, Centre for Long-Term Resilience, September 2025. URL https://www.rand.org/pubs/external publications/EP71093.html

  27. [27]

    Barrett, Krystal Jackson, Evan R

    Anthony M. Barrett, Krystal Jackson, Evan R. Murphy, Nada Madkour, and Jessica New- man. Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models, May 2024. URL http://arxiv.org/abs/2405.10986. arXiv:2405.10986 [cs]

  28. [28]

    Claude Opus 4.8 System Card

    Anthropic. Claude Opus 4.8 System Card. Technical report, Anthropic, May 2026. URL https://www.anthropic.com/claude-opus-4-8-system-card

  29. [29]

    Claude Opus 4.7 System Card

    Anthropic. Claude Opus 4.7 System Card. Technical report, Anthropic, April 2026. URL https://www.anthropic.com/claude-opus-4-7-system-card

  30. [30]

    Claude Opus 4.6 System Card

    Anthropic. Claude Opus 4.6 System Card. Technical report, Anthropic, February 2026. URL https://www.anthropic.com/claude-opus-4-6-system-card

  31. [31]

    GPT-5.5 System Card

    OpenAI. GPT-5.5 System Card. Technical report, OpenAI, April 2026. URL https: //deploymentsafety.openai.com/gpt-5-5

  32. [32]

    Gemini 3 Pro Frontier Safety Framework Report

    Google DeepMind. Gemini 3 Pro Frontier Safety Framework Report. Technical report, Google DeepMind, November 2025. URL https://storage.googleapis.com/deepmind-media/gemini/ gemini 3 pro fsf report.pdf. 22

  33. [33]

    Gemini 3.1 Pro Model Card

    Google DeepMind. Gemini 3.1 Pro Model Card. Technical report, Google DeepMind, February 2026. URL https://deepmind.google/models/model-cards/gemini-3-1-pro/

  34. [34]

    Activating AI Safety Level 3 Protections

    Anthropic. Activating AI Safety Level 3 Protections. Technical report, Anthropic, May 2025. URL https://www-cdn.anthropic.com/807c59454757214bfd37592d6e048079cd7a7728.pdf

  35. [35]

    Open models lag state-of-the-art closed models by 4 months

    Jack Edwards and Luke Emberson. Open models lag state-of-the-art closed models by 4 months. Epoch AI, May 2026. URL https://epoch.ai/data-insights/open-closed-eci-gap

  36. [36]

    (July run)

    Bhuvana Sudarshan and Luca Righetti. Towards a Common Standard for Evaluating Frontier AI Safeguards Against Biological Misuse. Technical report, Centre for the Governance of AI, June 2026. URL https://www.governance.ai/research-paper/technical-report-towards-a- common-standard-for-evaluating-frontier-ai-safeguards-against-biological-misuse. 23 Supplement...

  37. [2026]

    arXiv:2602.16703 [cs]

    URL http://arxiv.org/abs/2602.16703. arXiv:2602.16703 [cs]