Pith. sign in

Weak-to-strong jailbreaking on large language models

8 Pith papers cite this work. Polarity classification is still indexing.

8 Pith papers citing it

citation-role summary

background 1 method 1

citation-polarity summary

fields

cs.CR 7 cs.CL 1

polarities

background 2

representative citing papers

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs

cs.CR · 2026-06-21 · unverdicted · novelty 6.0 · 2 refs

Contrastive Logit Steering isolates a linear refusal direction in safety-aligned LLMs, achieving higher jailbreak success than activation steering and enabling bidirectional control without retraining.

citing papers explorer

Showing 8 of 8 citing papers.