Truth directions in LLMs are not universal, emerge only in more capable models, and simple linear probes trained on atomic statements generalize to QA and contextual tasks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks
Truth directions in LLMs are not universal, emerge only in more capable models, and simple linear probes trained on atomic statements generalize to QA and contextual tasks.