Pith. sign in

REVIEW 3 cited by

BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.00117 v3 pith:KP467U67 submitted 2023-10-31 cs.CL

classification cs.CL
keywords llamachatfine-tuningmodelweightscapabilitiescheaplydemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Llama 2-Chat is a collection of large language models that Meta developed and released to the public. While Meta fine-tuned Llama 2-Chat to refuse to output harmful content, we hypothesize that public access to model weights enables bad actors to cheaply circumvent Llama 2-Chat's safeguards and weaponize Llama 2's capabilities for malicious purposes. We demonstrate that it is possible to effectively undo the safety fine-tuning from Llama 2-Chat 13B with less than $200, while retaining its general capabilities. Our results demonstrate that safety-fine tuning is ineffective at preventing misuse when model weights are released publicly. Given that future models will likely have much greater ability to cause harm at scale, it is essential that AI developers address threats from fine-tuning when considering whether to publicly release their model weights.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Red Teaming Roadmap Towards System-Level Safety

    cs.CR 2025-05 conditional novelty 4.0 of 10

    A position paper from Scale AI argues that red teaming research should prioritize product-level safety specifications, realistic attacker models, and system-level monitoring over abstract model-level harm benchmarks.

  2. LLM Cyber Evaluations Don't Capture Real-World Risk

    cs.CR 2025-01 conditional novelty 4.0 of 10

    The paper argues and demonstrates with a 100-prompt case study that LLM cyber risk evaluations need to include threat actor adoption and impact, not just model capability.

  3. You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

    cs.LG 2025-02 conditional novelty 3.0 of 10

    To align powerful AI, researchers must understand how statistical patterns in training data shape the internal structure of models, because that structure, not eval scores, determines generalization.

Pith tools