Fine-tuning an LLM to verbalize a scalar confidence score, derived from its own self-consistency across sampled answers, improves calibration and induces longer self-verifying reasoning on low-confidence queries.
• If the response is purely textual (no numbers),extract the exact string as it appears
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
Fine-tuning an LLM to verbalize a scalar confidence score, derived from its own self-consistency across sampled answers, improves calibration and induces longer self-verifying reasoning on low-confidence queries.