Pith. sign in

REVIEW 6 cited by

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.03153 v2 pith:IA77M4ON submitted 2025-08-05 cs.LG cs.AI

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

classification cs.LG cs.AI
keywords gpt-ossfrontiercybersecuritymodelsopen-weightriskbiologicalbiorisk
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity. To maximize biological risk (biorisk), we curate tasks related to threat creation and train gpt-oss in an RL environment with web browsing. To maximize cybersecurity risk, we train gpt-oss in an agentic coding environment to solve capture-the-flag (CTF) challenges. We compare these MFT models against open- and closed-weight LLMs on frontier risk evaluations. Compared to frontier closed-weight models, MFT gpt-oss underperforms OpenAI o3, a model that is below Preparedness High capability level for biorisk and cybersecurity. Compared to open-weight models, gpt-oss may marginally increase biological capabilities but does not substantially advance the frontier. Taken together, these results contributed to our decision to release the model, and we hope that our MFT approach can serve as useful guidance for estimating harm from future open-weight releases.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Operationalising the Superficial Alignment Hypothesis via Task Complexity

    cs.LG 2026-02 conditional novelty 7.0

    A few kilobytes of program can adapt pre-trained LLMs to strong performance on math, translation, and instruction-following—evidence that task knowledge already lives in the model.

  2. Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

    cs.LG 2025-08 unverdicted novelty 7.0

    TokenBuncher constrains response entropy via entropy-as-reward RL and a Token Noiser to stop harmful RL fine-tuning while keeping benign performance intact.

  3. Who Does Withholding Delay? A Game-Theoretic Model of Open-Weight AI Release Under Asymmetric Proliferation

    cs.CY 2026-07 conditional novelty 6.0

    For dual-use AI, withholding helps only if it delays harmful actors more than defenders; the paper derives a substitution-rate threshold that decides when open release beats control.

  4. Estimating Tail Risks in Language Model Output Distributions

    cs.LG 2026-04 unverdicted novelty 6.0

    Importance sampling with unsafe model variants estimates tail probabilities of harmful language model outputs using 10-20x fewer samples than brute-force Monte Carlo.

  5. Estimating Tail Risks in Language Model Output Distributions

    cs.LG 2026-04 conditional novelty 6.0

    Importance sampling via activation-steered unsafe proposal models estimates rare harmful-output probabilities in language models with 10-20x fewer samples than brute-force Monte Carlo.

  6. Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

    cs.LG 2026-05 conditional novelty 5.0

    Abliteration and prefilling attacks raise harm success rates on safeguarded open-weight LLMs from below 10% to 16-96% across three benchmarks, and a new ART tuning method reduces those rates by 10-20%.