Linear probes on frozen LLM hidden states recover an approximate remaining-output-length signal that is decodable at prompt-end, transfers across datasets, and shifts upward at retraction tokens.
The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5representative citing papers
GUARD-IT removes forget-set influence from LLMs via training-free, gated residual-stream rotations at inference, matching gradient unlearning baselines without weight edits.
Activation steering is cast as constrained optimization that minimizes collateral damage by weighting perturbations according to the empirical second-moment matrix of activations instead of assuming isotropy.
citing papers explorer
-
How Much is Left? LLMs Linearly Encode Their Remaining Output Length
Linear probes on frozen LLM hidden states recover an approximate remaining-output-length signal that is decodable at prompt-end, transfers across datasets, and shifts upward at retraction tokens.
-
Inference-Time Machine Unlearning via Gated Activation Redirection
GUARD-IT removes forget-set influence from LLMs via training-free, gated residual-stream rotations at inference, matching gradient unlearning baselines without weight edits.
-
Minimizing Collateral Damage in Activation Steering
Activation steering is cast as constrained optimization that minimizes collateral damage by weighting perturbations according to the empirical second-moment matrix of activations instead of assuming isotropy.
- Can LLMs Reliably Self-Report Adversarial Prefills, and How?
- Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models