Activation-space cluster separation and refusal-direction similarity detect some safety-training modifications in open-weight language models, but a class of behavioral fine-tunes preserves the measured geometry and evades detection.
Dolphin: An uncensored, unbiased language model
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CR 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Detecting Safety Training Modification in Language Models via Activation Analysis
Activation-space cluster separation and refusal-direction similarity detect some safety-training modifications in open-weight language models, but a class of behavioral fine-tunes preserves the measured geometry and evades detection.