Linear probes trained on pre-solution hidden states, supervised by post-solution correctness probe outputs, recover 32–66% of the calibration gap between pre- and post-solution confidence across five open-source LLMs.
Any additional generated text is ignored
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Future Confidence Distillation in Large Language Models
Linear probes trained on pre-solution hidden states, supervised by post-solution correctness probe outputs, recover 32–66% of the calibration gap between pre- and post-solution confidence across five open-source LLMs.