Pith. sign in

Safety layers of aligned large language models: The key to llm security

12 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.

12 Pith papers citing it
2 external citations · external index

citation-role summary

background 3

citation-polarity summary

roles

background 3

polarities

background 3

representative citing papers

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs

cs.CR · 2026-05-09 · unverdicted · novelty 6.0

A truly benign DPO attack using 10 harmless preference pairs jailbreaks frontier LLMs by suppressing refusal behavior, achieving up to 81.73% attack success rate on GPT-4.1-nano at low cost.

Why Do Large Language Models Generate Harmful Content?

cs.AI · 2026-04-13 · unverdicted · novelty 6.0

Causal mediation analysis shows harmful LLM outputs arise in late layers from MLP failures and gating neurons, with early layers handling harm context detection and signal propagation.

citing papers explorer

Showing 12 of 12 citing papers.