Pith. sign in

Paper Citation Record · LEDGER

Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2402.13213.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.13213 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:31:49.646760Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:07:23.912546Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6b3a3b5c-f38f-44e4-8fee-2de7ebce61bd · inbound

Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning cites this paper.

Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:47:42.669613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-23T07:45:50.292586Z digest=sha256:a252befeb85007cf53697cc500b8b8a25d8ab0d0ed7494098eb3c47c57609d61

Observation 78a6a99f-2285-48aa-817a-140b6cece409 · inbound

Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence Measure cites this paper.

Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence Measure Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:22:38.186542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:22:25.496269Z digest=sha256:2d6eabc80c925bd125132cddb4a52a0cc10aa6a05a3ebc1f28c5b8081861830f

Observation e1c54e9e-7f57-4eb7-a829-f1a5b5b07179 · inbound

Random-Set Large Language Models cites this paper.

Random-Set Large Language Models Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:49.646760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:31:49.646760Z digest=sha256:d59672470b7d2f32d38efa59e1b8a82ec288580f68bf50b2e1029667dada7470

Observation a11cdbf3-600f-4ebd-96b6-29b0e6db1099 · inbound

Epistemic Artificial Intelligence is Essential for Machine Learning Models to Truly 'Know When They Do Not Know' cites this paper.

Epistemic Artificial Intelligence is Essential for Machine Learning Models to Truly 'Know When They Do Not Know' Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 187

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:13.768226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:21:13.768226Z digest=sha256:edf1fec4b4854b205d77622529185254f2d8de1078d2eb4d5db749acbfcb5d33

Observation d77c015c-0fb6-4b89-81cc-ca57c5ff6e60 · inbound

Resilient LLM-Empowered Semantic MAC Protocols via Zero-Shot Adaptation and Knowledge Distillation cites this paper.

Resilient LLM-Empowered Semantic MAC Protocols via Zero-Shot Adaptation and Knowledge Distillation Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:36.447956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:36.447956Z digest=sha256:62107d6541983f916d7866dc762e71b3b3cc4229a0d6ad97612305ab1b981990

Observation 48e9aed8-a786-4985-8e97-ca104f4f2d9b · inbound

Shapley Uncertainty in Natural Language Generation cites this paper.

Shapley Uncertainty in Natural Language Generation Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:55:57.711866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:55:57.711866Z digest=sha256:28d04079e387a79b9958fa1df16f710c35a8d56d84d7520d4b163cf5e105c8bb

Observation 62e24b5d-6bee-4a4e-b606-6396c400275f · inbound

Causal Evidence that Language Models use Confidence to Drive Behavior cites this paper.

Causal Evidence that Language Models use Confidence to Drive Behavior Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:44:05.679600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T09:43:05.524088Z digest=sha256:170bd1c26e87bd9128f6cf516f43b6df3d63a42f077af46297df2ff2ad0b9175

Observation 41208fd2-6bb0-40be-981e-f0c1201f221c · inbound

Calibrating Model-Based Evaluation Metrics for Summarization cites this paper.

Calibrating Model-Based Evaluation Metrics for Summarization Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:37.021662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T06:36:55.334742Z digest=sha256:240f441ec809173e01e1e750d9fa99965ba5832b79dbf6b75fd920f80705afa2

Observation a2506d7f-6ae5-4966-9f74-774be58d8d5c · inbound

Instance-Adaptive Online Multicalibration cites this paper.

Instance-Adaptive Online Multicalibration Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:56:40.742330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T04:45:02.146222Z digest=sha256:60ba6cac4e32313b6faf4844bf560d069cef1273ee15801edd36f3cc181d5dd4

Observation 8682f582-6878-4837-bdb2-f5c8047ba853 · inbound

Instance-Adaptive Online Multicalibration cites this paper.

Instance-Adaptive Online Multicalibration Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:41:21.584990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-22T09:41:09.689528Z digest=sha256:1fc9b9bc77431a26f9130245a81264c9145e349ef7cffbce648f87c1cd7a87fb

Observation ca86e480-9f27-48d4-84b8-949bc233cd15 · inbound

Task-Aware Calibration: Provably Optimal Decoding in LLMs cites this paper.

Task-Aware Calibration: Provably Optimal Decoding in LLMs Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:56:18.701791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T02:55:54.933129Z digest=sha256:977e509408d9248147c395fa55353b9e8d4e27107a21880f20103ba51edad5b2

Observation fc9ccb72-0e16-4a51-a59c-c0eda646e3c9 · inbound

PromptNCE: Pointwise Mutual Information Predictions Using Only LLMs and Contrastive Estimation Prompts cites this paper.

PromptNCE: Pointwise Mutual Information Predictions Using Only LLMs and Contrastive Estimation Prompts Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:46:18.168776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-22T08:46:12.102599Z digest=sha256:bbec3a0d366ee6515030e5d4e0d2740052eeb754a0a4cc30d30526c6fab4dde5

Observation 96734832-56cb-42d8-90c9-2cd72e1a68b6 · inbound

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning cites this paper.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.914209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:13ed357841d9abe9a8a56b225847242d36936d7fe01be1b556f308538edefb09