A policy analysis arguing that open-weight LLMs' loss-of-control properties make many cyber mitigations and the EU AI Act inadequate, and that capability-specific, downstream-focused regulation is the pragmatic alternative.
Badllama 3: removing safety finetuning from Llama 3 in minutes
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We show that extensive LLM safety fine-tuning is easily subverted when an attacker has access to model weights. We evaluate three state-of-the-art fine-tuning methods-QLoRA, ReFT, and Ortho-and show how algorithmic advances enable constant jailbreaking performance with cuts in FLOPs and optimisation power. We strip safety fine-tuning from Llama 3 8B in one minute and Llama 3 70B in 30 minutes on a single GPU, and sketch ways to reduce this further.
fields
cs.CR 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Mitigating Cyber Risk in the Age of Open-Weight LLMs: Policy Gaps and Technical Realities
A policy analysis arguing that open-weight LLMs' loss-of-control properties make many cyber mitigations and the EU AI Act inadequate, and that capability-specific, downstream-focused regulation is the pragmatic alternative.