Activation-space cluster separation and refusal-direction similarity detect some safety-training modifications in open-weight language models, but a class of behavioral fine-tunes preserves the measured geometry and evades detection.
ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Detecting Safety Training Modification in Language Models via Activation Analysis
Activation-space cluster separation and refusal-direction similarity detect some safety-training modifications in open-weight language models, but a class of behavioral fine-tunes preserves the measured geometry and evades detection.