Pith. sign in

REVIEW 12 cited by

Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.16137 v2 pith:O6TY3QG3 submitted 2025-04-21 cs.CY cs.LG

Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark

classification cs.CY cs.LG
keywords virologyvirologistsdual-useexpertbenchmarkcapabilitiescapabilityexpert-level
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols. Constructed from the inputs of dozens of PhD-level expert virologists, VCT consists of $322$ multimodal questions covering fundamental, tacit, and visual knowledge that is essential for practical work in virology laboratories. VCT is difficult: expert virologists with access to the internet score an average of $22.1\%$ on questions specifically in their sub-areas of expertise. However, the most performant LLM, OpenAI's o3, reaches $43.8\%$ accuracy, outperforming $94\%$ of expert virologists even within their sub-areas of specialization. The ability to provide expert-level virology troubleshooting is inherently dual-use: it is useful for beneficial research, but it can also be misused. Therefore, the fact that publicly available models outperform virologists on VCT raises pressing governance considerations. We propose that the capability of LLMs to provide expert-level troubleshooting of dual-use virology work should be integrated into existing frameworks for handling dual-use technologies in the life sciences.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts

    cs.SE 2026-05 conditional novelty 8.0

    RefusalBench shows strict refusal rates fail to rank frontier LLMs correctly on biological safety, with provider effects and partial-compliance patterns that binary metrics miss.

  2. BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment

    cs.CR 2026-07 conditional novelty 6.5

    Across 16 model-harness setups, AI agents refuse legitimate literature-derived biology tasks at rates comparable to or higher than concealed biosecurity hazards, with most refusals coming from pre-reasoning API filters.

  3. BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation

    cs.CY 2026-07 conditional novelty 6.0

    BioTIER, a 542-prompt benchmark with three risk tiers, shows frontier AI models differ by 90 percentage points in refusing dangerous biological queries, with top refusers over-refusing benign topics at the boundary.

  4. BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment

    cs.CR 2026-07 conditional novelty 6.0

    Across 16 agent configurations, frontier AI systems refused legitimate routine biology tasks at rates comparable to or higher than they refused concealed red-team hazards, with most refusals coming from provider API filters.

  5. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 6.0

    Survey of RLM adoption in 28 disciplines reveals maturity disparities via a new assessment framework, with focus on development, evaluation, and public resources.

  6. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 6.0

    A survey of RLM use in 28 disciplines reveals uneven adoption and introduces a maturity assessment framework showing larger gaps when limited to public resources.

  7. BioVeil MATRIX: Uncovering and categorizing vulnerabilities of agentic biological AI scientists

    q-bio.OT 2026-04 unverdicted novelty 6.0

    Agentic biological AI systems like Biomni and K-Dense assist with dual-use tasks blocked by safeguards and gain performance uplift on WMDP proxies; BioVeil MATRIX is introduced as a 10-category taxonomy with 22 techni...

  8. An Independent Safety Evaluation of Kimi K2.5

    cs.CR 2026-04 conditional novelty 6.0

    Kimi K2.5 matches closed models on dual-use tasks but refuses fewer CBRNE requests and shows some sabotage and self-replication tendencies.

  9. An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

    cs.CL 2026-07 reject novelty 5.0

    A bio-red-teaming model is reported to jailbreak 14 frontier LLMs into producing dangerous biosecurity outputs, but the claimed wet-lab physical verification was not actually carried out.

  10. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 4.0

    A survey of reasoning language model adoption across 28 ERC scientific disciplines finds large maturity gaps, especially when only public resources are counted.

  11. Risk Reporting for Developers' Internal AI Model Use

    cs.CY 2026-04 unverdicted novelty 4.0

    A harmonized risk reporting standard for internal frontier AI model use, structured around autonomous misbehavior and insider threats using means, motive, and opportunity factors.

  12. Muse Spark Safety & Preparedness Report

    cs.CY 2026-05 unverdicted novelty 2.0

    Meta's safety report states that Muse Spark meets acceptable risk thresholds for release after mitigations reduced elevated pre-mitigation risks in chemical and biological domains.