The paper localizes LLM lying to sparse attention heads and chat-template 'dummy tokens', and shows steering vectors can modulate deception, but the evidence is weakened by selection and small samples.
Robust separation of true/false for affirmative & negated statements; tG generalizes well
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can LLMs Lie? Investigation beyond Hallucination
The paper localizes LLM lying to sparse attention heads and chat-template 'dummy tokens', and shows steering vectors can modulate deception, but the evidence is weakened by selection and small samples.