A linear 'deception vector' in QwQ-32b's activations can separate and weakly steer deceptive responses, but the headline 89% accuracy is not backed by the reported results.
Localizing lying in llama: Understanding instructed dishonesty on true-false questions through prompting, probing, and patching, November 2023
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
A linear 'deception vector' in QwQ-32b's activations can separate and weakly steer deceptive responses, but the headline 89% accuracy is not backed by the reported results.