LLM response uncertainty and linear probe performance are strongly negatively correlated across six fact-based datasets and six models, with high-uncertainty responses associated with more spread-out feature importance.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet.Trans- former Circuits Thread, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
LLM response uncertainty and linear probe performance are strongly negatively correlated across six fact-based datasets and six models, with high-uncertainty responses associated with more spread-out feature importance.