A linear 'deception vector' in QwQ-32b's activations can separate and weakly steer deceptive responses, but the headline 89% accuracy is not backed by the reported results.
Discovering latent knowledge in language models without supervision, March 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
A linear 'deception vector' in QwQ-32b's activations can separate and weakly steer deceptive responses, but the headline 89% accuracy is not backed by the reported results.