Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:05:41.210502Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.04899.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:05:41.210502Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 67c64fc7-7a01-438b-9011-c290c8f887c0 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Benchmarking Uncertainty Quantification Methods for Large Language Models with LM -Polygraph
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a7e6c93-a5b9-478e-b940-8ca88100279b · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caba163f-e3e8-4a3a-9a87-3f11f97e151d · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Thinking Out Loud: Do Reasoning Models Know When They ' re Right?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2df6d9-8137-4ac7-b10d-8df1521ceb63 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6c237d1-89ae-4f6a-b84f-a32fd0712050 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification C heck E val: A reliable LLM -as-a-Judge framework for evaluating text generation using checklists
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c082c08-0007-478a-97ce-a26110fa4f0a · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8424d035-3c56-4419-b41d-6805fe56d743 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43b691c4-2287-4b89-8772-193c3feeb604 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Uncertainty Quantification for In-Context Learning of Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0941cce5-dedf-45b8-8ad2-71979b193d44 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification A Survey of Confidence Estimation and Calibration in Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d488a415-7af0-45ba-8ea0-c62dbd6cc024 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification B ayesian Prompt Ensembles: Model Uncertainty Estimation for Black-Box Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f94c29-51eb-4379-8f71-c1d0d6811834 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Calibrating the Confidence of Large Language Models by Eliciting Fidelity
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09f6206a-bb22-4359-93d7-67e3bde20ca6 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22d28937-d08e-49bd-977c-ab61269cdb4e · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d454fba-fe3e-44e2-848e-c2ee5eb5ebcd · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Adaptation with Self-Evaluation to Improve Selective Prediction in LLM s
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e53c6dec-ae7c-441d-b73d-8f80b6d37099 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c567c4c8-6bda-4b9c-940b-1bc8ea15e20a · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification and Ng, Andrew and Potts, Christopher
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f90ee083-453a-467e-a92d-8382fc2556a2 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 01d8053b-95cf-4bc5-8fe0-602a82ee3898 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 36th International Conference on Neural Information Processing Systems , articleno =
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8db0919-2b11-47cf-9b1d-1aad27c1d093 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification 2025 , eprint=
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 21498f79-42eb-4369-bd65-0d088b4bcf55 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Transactions on Machine Learning Research , issn=
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d54ca5b8-118b-40d5-851d-12a2ad92a030 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Calibrating Verbalized Probabilities for Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8483de6d-16eb-4114-8f3c-f00ff81c13e3 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification IEEE Transactions on Software Engineering , year =
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1ffeb2e8-126d-4899-9415-470b8ae2a15c · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification The Eleventh International Conference on Learning Representations (ICLR 2023) , year=
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0e39d061-dbc2-4868-b125-d5a1189a4618 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification 2022 , eprint=
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c70988a7-752a-48df-8e71-efbbe0ed9fe3 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Qwen3 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e4596c-face-48b6-bc6b-c340f95d6986 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Journal of Machine Learning Research , year =
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e800fa-b6f8-40d2-b683-36cc8ff90bdb · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7dc4f51d-b218-48ed-b037-d87c76c0ddbe · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification BingoGuard:
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b7e78a0a-5396-4e4a-97ee-a02988d68451 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71bfb71e-37b6-4834-ac8a-7c3b94e6a197 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv e-prints , pages=
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 92c313d1-b75d-45ac-b315-01a4fc31fd31 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Journal of classification , volume=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c47c8ed-00b2-41b4-a9cc-40c6a6c7ab16 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Genome Biology , volume=
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 22dd0a2b-96f2-41d1-87cd-453080fe5b4d · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering , pages =
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ecfb3a6-3148-4077-bbea-d14d1054e9b8 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification and Zhang, Hao and Gonzalez, Joseph E
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f679a33-70be-4c48-81cc-a05595cbaf0a · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the third International Workshop on Machine Learning in Systems Biology , pages =
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65764544-f0b3-448a-be91-daaf472f6bc6 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining , year =
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 02dcc27d-ea26-471e-9c6c-d554679344e9 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification I Can't Believe It's Not Better: Failure Modes in the Age of Foundation Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db888785-34b8-4ad5-b510-cc87ad917766 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv preprint arXiv:2506.01734 , year=
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d99f4ffe-c164-4358-9fee-cf9e63bcdb3f · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Advances in Neural Information Processing Systems (NeurIPS 2015) , year =
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05d2a915-b4d7-4299-83f2-859488b06c60 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c5ccd92-3062-4f97-b229-241f5a76e48f · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification , author=
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 104a98c7-f521-48ee-a043-2816272ac1bb · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv preprint arXiv:2509.13813 , year=
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e46d0c1e-5bda-41ad-a4b4-31474d59b214 · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification M onte C arlo Temperature: a robust sampling strategy for LLM ' s uncertainty quantification methods
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c6fc369-868a-4849-aa51-723192ef9d2a · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Scikit-learn: Machine Learning in Python , year =
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9aff573-6c02-4c52-92a0-8b2142e1114a · outbound
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.