Alignment-trained LMs route refusal through an intermediate-layer attention gate that triggers amplifier heads; modulating the gate controls policy from hard refusal to compliance, and encodings that evade the gate bypass safety.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
The paper defines defeat devices in AI via a triadic test (discriminator, concealed swap, performance gap), unifies existing cases under this concept, proposes TADP detection, and claims such devices can emerge naturally in frontier models.
citing papers explorer
-
How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
Alignment-trained LMs route refusal through an intermediate-layer attention gate that triggers amplifier heads; modulating the gate controls policy from hard refusal to compliance, and encodings that evade the gate bypass safety.
-
Defeat Devices in AI Systems
The paper defines defeat devices in AI via a triadic test (discriminator, concealed swap, performance gap), unifies existing cases under this concept, proposes TADP detection, and claims such devices can emerge naturally in frontier models.