Pith. sign in

Paper Citation Record · LEDGER

Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 71 inbound Pith citation observations for arXiv:2305.14975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.14975 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 71 of 71 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:08:24.966348Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T10:04:51.524003Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0d26e54b-5eda-40e3-9677-e2b3743f9b8e · inbound

Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs cites this paper.

Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T19:35:47.081648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T19:35:29.917096Z digest=sha256:fa9ab705a0f8c1df4a8950873502622d4c92048b3bce3c18a79e6821d9c50288

Observation fc857310-d720-44b5-bb35-893cde8ddb83 · inbound

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator cites this paper.

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:15:45.000278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:15:45.000278Z digest=sha256:3904c125cb5621185720252df2925f72cdd3856d5c732526a5e014dff9713565

Observation 43137bf2-977c-4039-a8b6-0eb7b2a8ce5e · inbound

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions cites this paper.

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 203

Resolution
unresolved
no resolver link, observed 2026-08-11T20:37:55.127589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:37:55.127589Z digest=sha256:5fc3322c92452a3f191d5849b19ee38e1389f40969de6befdbeb4f8172110c89

Observation a5179813-6feb-4f68-9223-48256d41fb10 · inbound

UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models cites this paper.

UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T14:40:29.402052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:40:29.402052Z digest=sha256:675d0e313ca71254b3216ac544d7ac2cd7a01ca262a71eaa974cd35550f0a5f4

Observation cf62d123-8a6a-45f8-bc30-bad822d73584 · inbound

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge cites this paper.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.382442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.382442Z digest=sha256:c3a758e6c17ba860b9c74ec402da8b4fe3c4d7f2a433fb0726cf24402ca7a768

Observation 07f3ca1a-7426-4132-8465-b82fc73e1815 · inbound

Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models cites this paper.

Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:08:03.466397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:08:03.466397Z digest=sha256:40b2156db3cad18f8a2da248158ae1345411d95a8b0f2e6b47c60349016b9bd7

Observation 87c41a14-f868-400d-80a0-aef4748ac421 · inbound

ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration cites this paper.

ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T18:14:11.143826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:14:11.143826Z digest=sha256:35a67130fa5c40fa0780dee40299bf258ee739cd1088782bdac60bcbdfb3dc51

Observation bdee8e7b-d2f3-492e-83e9-f59b50f90375 · inbound

What is a Number, That a Large Language Model May Know It? cites this paper.

What is a Number, That a Large Language Model May Know It? Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T15:05:03.762908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:05:03.762908Z digest=sha256:b94491ce8ed2e41b098a3fab7e2e343d3f7cf843b618ae83f1476e4afce4069b

Observation ac08aad5-0d5c-4bb6-bb09-9316d8ea88f8 · inbound

AI Alignment at Your Discretion cites this paper.

AI Alignment at Your Discretion Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-08T16:14:57.529618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:14:57.529618Z digest=sha256:f7dc142b2f1eb80557d343c78483352f2448cb23d869c073490eabf848669a93

Observation bb1c8036-e2fb-453e-9dea-a2e4e2b5f1d0 · inbound

Exploring the Potential for Large Language Models to Demonstrate Rational Probabilistic Beliefs cites this paper.

Exploring the Potential for Large Language Models to Demonstrate Rational Probabilistic Beliefs Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:08:24.966348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:08:24.966348Z digest=sha256:d40b4419694765041c44938c1fe7f5cbedd9b85cc6bf916f7e787e7233cf4b0e

Observation 5f7435d4-a7d6-43bf-8182-a32399b9cf1b · inbound

Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation cites this paper.

Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:54.228901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:54.228901Z digest=sha256:8d5ac1c03dbd7a47d475eaddbb85943c751e324c41202ac26d24c367a8f45a25

Observation 94dc402d-477a-49c4-ac6b-5235c78a8805 · inbound

Lightweight Latent Verifiers for Efficient Meta-Generation Strategies cites this paper.

Lightweight Latent Verifiers for Efficient Meta-Generation Strategies Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:20.513220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:20.513220Z digest=sha256:5d3b4c2ebdc9efd7784f172e65e887b1a65092bd164bed6826d32de42dda8ad4

Observation e07cd5e4-985b-4cca-9516-7e2afe1c32c8 · inbound

From Evidence to Belief: A Bayesian Epistemology Approach to Language Models cites this paper.

From Evidence to Belief: A Bayesian Epistemology Approach to Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:20.615933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:52:20.615933Z digest=sha256:878b2b7b2d1bacde4660f7c5cd91f1f091d31b6c72b8fd1d95fc49335088a9a4

Observation 584d44be-ba08-413d-a395-ebfcbd78e2c6 · inbound

How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception cites this paper.

How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:10.229493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:10.229493Z digest=sha256:57e046ccae6d6668d0d4a1cd11d1510249b9698e2505e0041eafd34329b05876

Observation ec1db97f-76d1-473b-bdd5-ef12a1fd1741 · inbound

Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks cites this paper.

Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:32.927335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:32.927335Z digest=sha256:c7ea3a59add1068f69a30e9776a4d127ce670ac056c4afe34f7bd4816fbcf8ea

Observation b8abcfe2-2940-4745-b144-bcf9713b4ffc · inbound

Maximizing Confidence Alone Improves Reasoning cites this paper.

Maximizing Confidence Alone Improves Reasoning Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:50.296268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:50.296268Z digest=sha256:b32ac926af10c1dd370f78c9cf31b89c5834bf75487feffd7c050482ce07ca14

Observation 1da2abeb-136f-4094-a287-a324fe04db61 · inbound

Revisiting Uncertainty Estimation and Calibration of Large Language Models cites this paper.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.624582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.624582Z digest=sha256:6ba40ea54e15c036dd3a901097ed9c1dcb9632b92697ef87fb522cc2ba3c62e2

Observation f12a8b6b-719e-4ae1-96cc-9777a10f2fe7 · inbound

Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations cites this paper.

Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:47.056892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:47.056892Z digest=sha256:3f4b132214ccd13db28230fc49f962ef13969e12b985501298b25fd6a8033bb8

Observation 30a7e555-abf9-49aa-ba4c-bc584eb09656 · inbound

SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQL cites this paper.

SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQL Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:27.153092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:46:27.153092Z digest=sha256:065037fa23c7242c1e4c880d33de17467378c3d7776b05c98b8c93c38b950114

Observation 619a2cbc-486e-45bc-9b89-0ae08a5ad457 · inbound

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions cites this paper.

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:08.078888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:08.078888Z digest=sha256:fc07d6a4f3a55a04175ce61dad88c7ad69f44ca965fd87ff4dc66f763c5a5242

Observation 54c42534-296b-4e84-9d02-2f824336c9ce · inbound

Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know? cites this paper.

Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know? Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:22.043040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:22.043040Z digest=sha256:b06c9c4e52735e60f8ae8817c921547f991d3848359d34f7359df787c8955388

Observation ddf5da14-286c-49de-a5d8-df45f05613a0 · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.921444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.921444Z digest=sha256:f5b4e3b18a92e2e9bbd266ca51bee655b5c40764f4393f9ca153d6a89c018674

Observation 8294562d-caf4-45f4-873a-df1bceb7a1dd · inbound

How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models cites this paper.

How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:55.114325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:55.114325Z digest=sha256:dbaedc68982b6b7ecd44e26ecdfe1e251f5e9e5fbf20b397fa27e848fc3afea3

Observation 119cf71b-9e18-4ad7-9f5f-5f35ee6de10d · inbound

Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification cites this paper.

Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:03.045297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:38:03.045297Z digest=sha256:57852786bdc189e7047cbeb8197a8347ec3840c1212f66721fc7b4e65f9487ae

Observation 06be8adb-b560-492c-8f97-a84a51690bcc · inbound

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models cites this paper.

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:27.408905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:27.408905Z digest=sha256:4cd6532c2fcb4a0a4e7a35a7ac43422fcb381848024ca3183c666829efc4a8ce

Observation 0e257a96-023e-4eee-8c10-a9eead34d43a · inbound

Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models cites this paper.

Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T05:23:59.897211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:23:59.897211Z digest=sha256:5203881ff703f96e0cccfd7a943ed2718a12e573f1376ac519102c355c830ac1

Observation cc4418c2-6516-4fcd-80a3-1e771d10e976 · inbound

Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation cites this paper.

Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:45.850281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:30:45.850281Z digest=sha256:24cb1b223b8dfa57fff31644b6ea3d023026b2189b447994f67ee8f47c577ee0

Observation 0e6f288c-3cec-4927-bf7e-43380ab7134b · inbound

PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning cites this paper.

PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T22:48:32.557614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:48:32.557614Z digest=sha256:2286617d65062dd5218b4d8092e2716292fdccc37d6db1884ef146928ced5360

Observation cedf85ec-f4cd-42ec-b15c-6094bf857b8e · inbound

Inteligencia Artificial jur\'idica y el desaf\'io de la veracidad: an\'alisis de alucinaciones, optimizaci\'on de RAG y principios para una integraci\'on responsable cites this paper.

Inteligencia Artificial jur\'idica y el desaf\'io de la veracidad: an\'alisis de alucinaciones, optimizaci\'on de RAG y principios para una integraci\'on responsable Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T19:07:41.556896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:07:41.556896Z digest=sha256:27e6bc5fcc8cc8c0f0b94f01663c28a4e21b8c2e35ecb1d246141909466eee49

Observation 08ebd968-1697-4d2e-83ff-3c215f9ea0ec · inbound

LAVA: Language Model Assisted Verbal Autopsy for Cause-of-Death Determination cites this paper.

LAVA: Language Model Assisted Verbal Autopsy for Cause-of-Death Determination Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:09:12.519527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:09:12.519527Z digest=sha256:e2714f3e88de0bdc9a3df6e0d93caab094820310b45f507c01e294ccc12068cd

Observation 6af8266c-4bf5-4722-9483-fc653b770f7d · inbound

Unsupervised Hallucination Detection by Inspecting Reasoning Processes cites this paper.

Unsupervised Hallucination Detection by Inspecting Reasoning Processes Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T18:25:37.393238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:25:37.393238Z digest=sha256:e57ef06b6d87b76a9aec224dc61e37e07316b62e3a26d5232823a7a227fe9b28

Observation 19cc7309-0a75-4058-b68a-2f9250200bdb · inbound

HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling cites this paper.

HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T17:41:06.185564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:41:06.185564Z digest=sha256:c8d8678efd930f8d9b7398e164f332c6d2e60b80e5fae7b6c20909a475c02625

Observation 89b1e2df-80f6-4170-9ac4-b957549c4d8e · inbound

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning cites this paper.

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T07:02:43.751793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:02:43.751793Z digest=sha256:943082060d52dd7b2fa79f99fa3d071a0d708cbe55016cfe8d13854da41b7838

Observation 7608833d-b797-46e0-a530-5a89453a8397 · inbound

Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction cites this paper.

Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:18:24.059827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T20:15:16.449712Z digest=sha256:1be79f597fa2093fe14f8ed22e50b6a9fe5918e01eef1e9a7e00fcfb517a23d2

Observation 07c74a56-30da-4241-acee-22357dc2f946 · inbound

UCPO: Uncertainty-Aware Policy Optimization cites this paper.

UCPO: Uncertainty-Aware Policy Optimization Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:34:29.961428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:34:29.961428Z digest=sha256:897752f221cdadf4da49df437d3bcc541d7d6c4854f396ead4142728b0caa7e6

Observation 2e84a639-ce49-42db-bf77-d54e9c57ad46 · inbound

Uncertainty-aware Generative Recommendation cites this paper.

Uncertainty-aware Generative Recommendation Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T00:05:11.067751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:05:11.067751Z digest=sha256:8835ee90a124a42e7b40ccaba40224d188049950b72d5a8b4e50f46a29b77944

Observation e571b856-0a57-4b15-bfc3-fc1a4353dd58 · inbound

How do LLMs Compute Verbal Confidence cites this paper.

How do LLMs Compute Verbal Confidence Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:40:00.577255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T10:38:28.533485Z digest=sha256:9b54b55d0fc2e5e35cfc7ab9cd2e367a215f8a4cc725dc2b0b7ec5a254c91922

Observation c860152e-3ff7-4350-a450-6abed1da25e8 · inbound

Causal Evidence that Language Models use Confidence to Drive Behavior cites this paper.

Causal Evidence that Language Models use Confidence to Drive Behavior Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T09:44:05.651338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T09:43:05.524088Z digest=sha256:49dbbc8ac04d675c90be2e57f8e129c88247ef55c7680f7fad61b10c25ee5116

Observation 5b2b2309-44e6-43e7-b6ff-a8093a8480b3 · inbound

SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio cites this paper.

SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:49.681288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:49:23.092309Z digest=sha256:4e62969701d26b5e545fbb023a80c264ee02d2748f4e50b3873cba20e4f95c68

Observation 4c2fd448-564b-408d-a88a-a2548d5532ba · inbound

Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models cites this paper.

Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:28:26.292726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T23:25:56.299371Z digest=sha256:6d23997cc1e514f530732d779c66dd61dd8dd0ef1bc6ea9a251932ad62621eca

Observation d08e21b5-500f-47b5-a3d3-d2a31853ccae · inbound

UsefulBench: Towards Decision-Useful Information as a Target for Information Retrieval cites this paper.

UsefulBench: Towards Decision-Useful Information as a Target for Information Retrieval Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:25.137836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T07:59:42.451677Z digest=sha256:59c515ee2e2b325f8134b42876ee74dce89bc635904ab7365a123b08d98c0faf

Observation ee48544c-d7c3-47e2-b33b-eb6dc618992b · inbound

Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen cites this paper.

Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:08.905209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:09:07.770260Z digest=sha256:827cf27fb52881e85111669da0c72e7cae4da006ac7b234c007625477de2aa0a

Observation 8113b983-2179-454c-896e-cce94e929d75 · inbound

How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals cites this paper.

How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:16:06.940805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:30:51.094216Z digest=sha256:987df11e97e0409d90a7aaff0f2a44cc2d95c8793555a7ab159611403023b90b

Observation de793346-1123-4430-bd84-12ce8735d9e8 · inbound

Confidence Estimation in Automatic Short Answer Grading with LLMs cites this paper.

Confidence Estimation in Automatic Short Answer Grading with LLMs Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:09.529218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T20:22:34.911547Z digest=sha256:8bb27971146dce802d9ee09ba07447a5e10004e9385647753f8ee9a50000e10d

Observation ad7b4369-ef84-4889-932a-bf2d6e516a18 · inbound

Confidence Estimation in Automatic Short Answer Grading with LLMs cites this paper.

Confidence Estimation in Automatic Short Answer Grading with LLMs Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:02:59.206738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T20:59:33.337511Z digest=sha256:213c9f7a7b6b7b680fc94995cb31cc358da87b3c1d247062149641215f5213a6

Observation e9ed1b3e-a20d-4561-8656-57b2b07cc8e0 · inbound

Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training cites this paper.

Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:07.311249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T16:32:51.205764Z digest=sha256:d811271310e040bc5c61fd01899119f36a1231fb3040dbc3e32bf15f8f489baa

Observation 59ac8232-900e-467f-be2e-c0fedc3eede7 · inbound

LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs cites this paper.

LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:56.985129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T00:53:49.859413Z digest=sha256:ff54831a6e266056668600f110bbdb773edacadac4913e8fd869cf5d1ff5df76

Observation 091a3177-bbad-4bd9-9919-eb95c3726809 · inbound

Inducing Artificial Uncertainty in Language Models cites this paper.

Inducing Artificial Uncertainty in Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:12:55.526912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T20:11:44.878211Z digest=sha256:881aa5176fbca01bf4a556bf3b472c4f452b5e1b170e1e2d62d856cf735f66c2

Observation 8edc7c8f-98b6-4ad1-b7b2-63d67aac3c24 · inbound

Margin-Adaptive Confidence Ranking for Reliable LLM Judgement cites this paper.

Margin-Adaptive Confidence Ranking for Reliable LLM Judgement Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:12:39.544765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T16:05:41.505091Z digest=sha256:ece34d69c2e22b1f23e90d5ab02d2a9e325d70b5927dd794e3f05e8538b30284

Observation 78049a4b-e569-49a0-ae1c-4b1bc001a279 · inbound

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination cites this paper.

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:10:23.641444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T06:09:18.808074Z digest=sha256:33f52c5ee7ab3d4084add3d17b55dba5914ffe10c25caaccc702ed231f12d834

Observation 5e85ed73-3521-409a-9b45-06c43adfaae2 · inbound

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination cites this paper.

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.166499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:38:51.852464Z digest=sha256:9d37b3068a4aeb5518606fe671675e99c2230aa87583efdb2784f9e9ae0168a9

Observation 142b36c8-98e7-4bc7-8993-b252bd342ebc · inbound

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination cites this paper.

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T02:19:38.063144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:19:38.063144Z digest=sha256:cdffe1ad5009bd4c81bd7960ca1cd6c8b652aeaade0e2e47fdb32cedb68c5119

Observation c4f7439e-d337-448d-8cf7-f44492eb1078 · inbound

Confidence Calibration in Large Language Models cites this paper.

Confidence Calibration in Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:3035226287c24429678784eb57e4b69a1ac4ea8bccf133a821650934c651939a

Observation e35011d6-a110-4991-bc9a-add26e706319 · inbound

Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification cites this paper.

Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:23:27.683822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T13:23:01.482449Z digest=sha256:628943442645b7d2efe604907439bf1a55a01d43d30824b4ea0bbfe972089738

Observation 9e1c07db-f4f3-4a79-8465-d045fab9c574 · inbound

VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models cites this paper.

VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:13:30.494545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T06:53:56.750481Z digest=sha256:2e370f46d5545c07a7e5d992912e928ea2cc08fe8c5780e08332fd113c510fb4

Observation a9316994-81e5-4599-ae7a-a6202369a667 · inbound

Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage cites this paper.

Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:16:00.148765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T23:04:29.910598Z digest=sha256:3c059548eee9e43396b7561129d4145d4ca09e15b7e90187a322e57dbd45e45e

Observation c081905b-0e32-479a-a5ba-b78c17033301 · inbound

NBQ: Next-Best-Question for Dynamic Profiling cites this paper.

NBQ: Next-Best-Question for Dynamic Profiling Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:32:37.910080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T18:34:34.264514Z digest=sha256:5332dfefd2ef50b8960a69b2939a54d5e3d70ef5af3ed64aaaf93e76f9c34d3d

Observation 5d0c1da4-8c9c-4093-844b-64238ab22f14 · inbound

Can LLM Rerankers Predict Their Own Ranking Performance? cites this paper.

Can LLM Rerankers Predict Their Own Ranking Performance? Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:16:39.622791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T08:19:25.544186Z digest=sha256:3535cf3d6e5222229cfcf9745ea6a1ab23b0e02eea9341e12ee43741baeb9ba6

Observation 49cea165-9b97-443f-b295-d6c7a0ba96ea · inbound

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act cites this paper.

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:29:02.480627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T22:16:25.919667Z digest=sha256:39b72fe59359b57cac1fe1ad8fb37439814c2d2eaf39268068df2f126def5e5e

Observation e1eed9f8-467b-4df8-8149-7e599f67cbbd · inbound

Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA cites this paper.

Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:39:30.312099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T17:49:38.151224Z digest=sha256:08fc8458d296db7eb4abdec10451dc7833ea8fdca6a207742694b65af71b3b8d

Observation e622431a-d4eb-401d-a148-c66978bcc057 · inbound

Beyond Logprobs: A Multi-Signal Confidence Engine for LLM-Based Document Field Extraction cites this paper.

Beyond Logprobs: A Multi-Signal Confidence Engine for LLM-Based Document Field Extraction Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:09:59.431676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-25T23:52:07.754083Z digest=sha256:5408e0d4a7f934a4a15a8139a63c6c8390d1551e4076a6803e71b2b16820f57d

Observation 259b9f1f-d222-48c4-baf7-7cb74bc18f70 · inbound

Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA cites this paper.

Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:59:52.905360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T05:35:35.054964Z digest=sha256:b1b4760b63b945a3144dc2412db050034a1eb16800c67f9ba6a0c0971b5e4d20

Observation b1b13f8b-bcc3-4d38-bb68-fe9cca14be69 · inbound

Reported Confidence in LLMs Tracks Commitment More Than Correctness cites this paper.

Reported Confidence in LLMs Tracks Commitment More Than Correctness Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:04:28.805928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T07:44:53.539382Z digest=sha256:b2463b25abd3f26ca4af1a807cb3bd92d28e9446f3d459e41429bf33d166dddf

Observation bf219580-2e9f-420b-9c2e-50f8b5cd14fc · inbound

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling cites this paper.

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:58:32.561986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-03T14:49:33.364596Z digest=sha256:5bcc0d2a1f83c0a38dda47045edddb6a35953c870ea68e53465a46273acf273d

Observation a2034867-06e2-4406-8eb4-5e8bc1ce6a27 · inbound

Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs cites this paper.

Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T10:04:51.526734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-08T09:56:07.928172Z digest=sha256:e92e78caad699c7521d021c528c595a8b38798b9a01ca7767e5ee3d0332d2068

Observation 02576b22-d371-4fae-a7d4-85bdd210f3b7 · inbound

The Computational Basis of Confidence in Large Language Models cites this paper.

The Computational Basis of Confidence in Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:42.016252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:42.016252Z digest=sha256:10fe62d4acedb458e6a688c883ea4bf80f24f52fc2f41d1814d461245248ca55

Observation 40fa64b1-f4d3-4ab8-be2c-ecec87abbdde · inbound

Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation cites this paper.

Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T06:03:16.937188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:03:16.937188Z digest=sha256:2078c59f21948e9fcd85b6ac4c292c1e0a3827b6dec3d75270c671cc588b0405

Observation f410d905-d7d6-47c2-9bcf-a59056b4edb9 · inbound

Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation cites this paper.

Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T14:56:09.451029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:56:09.451029Z digest=sha256:9d7ed9e5bf250c53b9804c5dee48fbb5e31283a050f759304cd7b9eeea6f6722

Observation 4e4b0f74-8b50-49b1-b90b-1807ef6b6d13 · inbound

One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence cites this paper.

One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T11:29:22.475836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:29:22.475836Z digest=sha256:6fee382c1e7b7cccfaf98a921f805d7b169bed1d77be9ac4005ba5845ccba84c

Observation e044ebba-bf2a-40b4-8c3f-3ae1fe444e5b · inbound

CARE: Confidence-Aware Reasoning for Reliable Medical VQA cites this paper.

CARE: Confidence-Aware Reasoning for Reliable Medical VQA Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:40.289944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:40.289944Z digest=sha256:28ec8298b0c7d895cc29ed4c17d4ed273488ef383373b37fd2d73c938e814a98

Observation 028d454e-2beb-4367-b1fb-0f89be3d99cc · inbound

Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) cites this paper.

Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T17:37:14.051600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:37:14.051600Z digest=sha256:929da8e99409ede38d452b612df38cd839edca557f92c8c2087d5dd2729dbd1e