Pith. sign in

REVIEW 2 cited by

Will releasing the weights of future large language models grant widespread access to pandemic agents?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.18233 v2 pith:U2OXPCPN submitted 2023-10-25 cs.AI

classification cs.AI
keywords modelmodelsfuturemaliciouspandemicweightswillagents
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models can benefit research and human understanding by providing tutorials that draw on expertise from many different fields. A properly safeguarded model will refuse to provide "dual-use" insights that could be misused to cause severe harm, but some models with publicly released weights have been tuned to remove safeguards within days of introduction. Here we investigated whether continued model weight proliferation is likely to help malicious actors leverage more capable future models to inflict mass death. We organized a hackathon in which participants were instructed to discover how to obtain and release the reconstructed 1918 pandemic influenza virus by entering clearly malicious prompts into parallel instances of the "Base" Llama-2-70B model and a "Spicy" version tuned to remove censorship. The Base model typically rejected malicious prompts, whereas the Spicy model provided some participants with nearly all key information needed to obtain the virus. Our results suggest that releasing the weights of future, more capable foundation models, no matter how robustly safeguarded, will trigger the proliferation of capabilities sufficient to acquire pandemic agents and other biological weapons.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Intervening on a few attention heads selected by binary-choice probing steers four open LLMs toward AI coordination and can also block a jailbreak.

  2. Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests

    cs.CL 2025-02 reject novelty 4.0 of 10

    A new 512-prompt benchmark claims to measure LLM over-refusal on scientific dual-use questions, but its design and labeling flaws undermine the claim.

Pith tools