Pith. sign in

hub

Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b

25 Pith papers cite this work. Polarity classification is still indexing.

25 Pith papers citing it

hub tools

citation-role summary

background 3 method 1

citation-polarity summary

polarities

background 4

representative citing papers

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs

cs.CR · 2026-04-17 · conditional · novelty 8.0

Benign fine-tuning on audio data breaks safety alignment in Audio LLMs by raising jailbreak success rates up to 87%, with the dominant risk axis depending on model architecture and embedding proximity to harmful content.

Curriculum Learning for Safety Alignment

cs.LG · 2026-05-25 · unverdicted · novelty 6.0

Staged-Competence curriculum reduces out-of-distribution harmful responses by 16% and jailbreak success rates by 20% in DPO safety alignment across three model families while using 75% of the data.

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs

cs.CR · 2026-05-09 · unverdicted · novelty 6.0

A truly benign DPO attack using 10 harmless preference pairs jailbreaks frontier LLMs by suppressing refusal behavior, achieving up to 81.73% attack success rate on GPT-4.1-nano at low cost.

Understanding the Effects of Safety Unalignment on Large Language Models

cs.CR · 2026-04-02 · unverdicted · novelty 6.0

Weight orthogonalization unalignment enables LLMs to assist malicious activities more effectively than jailbreak-tuning, with less hallucination and better retained performance, while supervised fine-tuning mitigates the added attack capabilities.

Persona-Model Collapse in Emergent Misalignment

cs.CL · 2026-05-13 · unverdicted · novelty 5.0 · 2 refs

Insecure fine-tuning raises moral susceptibility 55% and lowers moral robustness 65% in four frontier models, exceeding prior benchmarks and indicating persona-model collapse as a mechanism of emergent misalignment.

Low-Rank Adaptation Redux for Large Models

cs.LG · 2026-04-23 · unverdicted · novelty 3.0

An overview revisits LoRA variants by categorizing advances in architectural design, efficient optimization, and applications while linking them to classical signal processing tools for principled fine-tuning.

citing papers explorer

Showing 25 of 25 citing papers.