{"work":{"id":"8617eae9-ad9f-45e0-bb23-49df5b429be6","openalex_id":null,"doi":null,"arxiv_id":"2305.14975","raw_key":null,"title":"Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback","authors":null,"authors_text":"K","year":2023,"venue":"cs.CL","abstract":"A trustworthy real-world prediction system should produce well-calibrated confidence scores; that is, its confidence in an answer should be indicative of the likelihood that the answer is correct, enabling deferral to an expert in cases of low-confidence predictions. Recent studies have shown that unsupervised pre-training produces large language models (LMs) whose conditional probabilities are remarkably well-calibrated. However, the most widely-used LMs are fine-tuned with reinforcement learning from human feedback (RLHF-LMs), and some studies have suggested that RLHF-LMs produce conditional probabilities that are very poorly calibrated. In light of this perceived weakness, we conduct a broad evaluation of methods for extracting confidence scores from RLHF-LMs. For RLHF-LMs such as ChatGPT, GPT-4, and Claude, we find that verbalized confidences emitted as output tokens are typically better-calibrated than the model's conditional probabilities on the TriviaQA, SciQ, and TruthfulQA benchmarks, often reducing the expected calibration error by a relative 50%.","external_url":"https://arxiv.org/abs/2305.14975","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-08T10:04:51.524003+00:00","pith_arxiv_id":"2305.14975","created_at":"2026-05-10T08:02:25.136370+00:00","updated_at":"2026-07-08T10:04:51.524003+00:00","title_quality_ok":true,"display_title":"Manning, and Chelsea Finn","render_title":"Manning, and Chelsea Finn"},"hub":{"state":{"work_id":"8617eae9-ad9f-45e0-bb23-49df5b429be6","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":27,"external_cited_by_count":null,"distinct_field_count":7,"first_pith_cited_at":"2024-10-09T00:09:15+00:00","last_pith_cited_at":"2026-07-07T14:25:09+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T13:49:41.130114+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"use_method","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}