Language models encode a linear, decodable signal in their residual stream that predicts whether an upcoming factual recall will be correct.
Self-refine: Iterative refinement with self-feedback
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
Language models encode a linear, decodable signal in their residual stream that predicts whether an upcoming factual recall will be correct.