Pith. sign in

Paper Citation Record · LEDGER

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.04899.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04899 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:05:41.210502Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67c64fc7-7a01-438b-9011-c290c8f887c0 · outbound

This paper cites Benchmarking Uncertainty Quantification Methods for Large Language Models with LM -Polygraph.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Benchmarking Uncertainty Quantification Methods for Large Language Models with LM -Polygraph

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:37.879497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:37.879497Z digest=sha256:501d02a1e5d1e282acc751d4cfb1bd2f04cb0c59ae81c348bcaf4d976b10c7cb

Observation 7a7e6c93-a5b9-478e-b940-8ca88100279b · outbound

This paper cites Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:37.927074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:37.927074Z digest=sha256:7124e3ae7e5d7e5bf5377e1ed73e77466b573758e8988789e02c2ec43bd44bfe

Observation caba163f-e3e8-4a3a-9a87-3f11f97e151d · outbound

This paper cites Thinking Out Loud: Do Reasoning Models Know When They ' re Right?.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Thinking Out Loud: Do Reasoning Models Know When They ' re Right?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:37.956147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:37.956147Z digest=sha256:0ce5293b31adb1ce7f55671eccfbb92cf23ca75ec3572b6a77fd907d2f73a604

Observation fc2df6d9-8137-4ac7-b10d-8df1521ceb63 · outbound

This paper cites Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.004867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.004867Z digest=sha256:94649a67b826616fc80bdda9516d3069901e7fa835189717338e69b4bd1bbe08

Observation f6c237d1-89ae-4f6a-b84f-a32fd0712050 · outbound

This paper cites C heck E val: A reliable LLM -as-a-Judge framework for evaluating text generation using checklists.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification C heck E val: A reliable LLM -as-a-Judge framework for evaluating text generation using checklists

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.057387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.057387Z digest=sha256:57d002f33f63b96ca497087bba80a9885fa6737889e386b652468f6132120f65

Observation 8c082c08-0007-478a-97ce-a26110fa4f0a · outbound

This paper cites an unresolved cited work.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.102656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.102656Z digest=sha256:b09b86fa7717b6c8577518f2c66128473f577fc483cfff1f637c4865c0a5fead

Observation 8424d035-3c56-4419-b41d-6805fe56d743 · outbound

This paper cites Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.149210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.149210Z digest=sha256:3bde874bc9de28bcf1132614ee49aa22ed2c9a695c749f5f6baf871e15c4a1c3

Observation 43b691c4-2287-4b89-8772-193c3feeb604 · outbound

This paper cites Uncertainty Quantification for In-Context Learning of Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Uncertainty Quantification for In-Context Learning of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.184134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.184134Z digest=sha256:0bcfceff5ede8321e00b78264300c6140eaad6c6ca66708a448cfed262e6d0ad

Observation 0941cce5-dedf-45b8-8ad2-71979b193d44 · outbound

This paper cites A Survey of Confidence Estimation and Calibration in Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification A Survey of Confidence Estimation and Calibration in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.283292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.283292Z digest=sha256:02bbea2f74f465fea1f50e69fb26e1c2fe9c0ea813d45dd6f3826c96db097e2b

Observation d488a415-7af0-45ba-8ea0-c62dbd6cc024 · outbound

This paper cites B ayesian Prompt Ensembles: Model Uncertainty Estimation for Black-Box Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification B ayesian Prompt Ensembles: Model Uncertainty Estimation for Black-Box Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.357788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.357788Z digest=sha256:b01c5992e5083202c0555e2a8132ba0c45fa74160c14d98bcff9bad2829e3d61

Observation 84f94c29-51eb-4379-8f71-c1d0d6811834 · outbound

This paper cites Calibrating the Confidence of Large Language Models by Eliciting Fidelity.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Calibrating the Confidence of Large Language Models by Eliciting Fidelity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.423013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.423013Z digest=sha256:a0572260f46ce73063bc5692899c24b51a222a55b73dc7781eb8909ff57191e5

Observation 09f6206a-bb22-4359-93d7-67e3bde20ca6 · outbound

This paper cites Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.523404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.523404Z digest=sha256:7df453a62bd14f5ead324cbee7dfe39e8d734b8e7a22ade87e06e4ab372100d5

Observation 22d28937-d08e-49bd-977c-ab61269cdb4e · outbound

This paper cites Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.634360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.634360Z digest=sha256:5e984df3ce81689450d9da64fd723f6d1a69d3d558b7904eede516573ece8fff

Observation 9d454fba-fe3e-44e2-848e-c2ee5eb5ebcd · outbound

This paper cites Adaptation with Self-Evaluation to Improve Selective Prediction in LLM s.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Adaptation with Self-Evaluation to Improve Selective Prediction in LLM s

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T14:05:41.484457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:38.744461Z digest=sha256:c85eda892cee01cbe36038bb6d6f74773650c456de3b08f453aedb087503e800

Observation e53c6dec-ae7c-441d-b73d-8f80b6d37099 · outbound

This paper cites Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.842660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.842660Z digest=sha256:042e1aaf7ece50eb7dce1f3ca76f846ea95c6dea803235299098f92887319a51

Observation c567c4c8-6bda-4b9c-940b-1bc8ea15e20a · outbound

This paper cites and Ng, Andrew and Potts, Christopher.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification and Ng, Andrew and Potts, Christopher

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.908813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.908813Z digest=sha256:c18abac8fcc2e7559231f7c5567938a787da9b56b78cbd9f14c951544d6542c3

Observation f90ee083-453a-467e-a92d-8382fc2556a2 · outbound

This paper cites an unresolved cited work.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:05:44.292357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.020702Z digest=sha256:8a1e4f0d9982833f274f3d3d574aec8ca63bddcc693dd23a7992a9e2df688308

Observation 01d8053b-95cf-4bc5-8fe0-602a82ee3898 · outbound

This paper cites Proceedings of the 36th International Conference on Neural Information Processing Systems , articleno =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 36th International Conference on Neural Information Processing Systems , articleno =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.056971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.056971Z digest=sha256:e28e559ff5cb0b7b2e9c513b32578641ab040b9bebabb996bae6421deb479c38

Observation d8db0919-2b11-47cf-9b1d-1aad27c1d093 · outbound

This paper cites 2025 , eprint=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification 2025 , eprint=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:44.115273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.111242Z digest=sha256:4a38c8b4557f41ce7b48066e684e2b69132a9cec3b71c75183cc14a4943cc3dd

Observation 21498f79-42eb-4369-bd65-0d088b4bcf55 · outbound

This paper cites Transactions on Machine Learning Research , issn=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Transactions on Machine Learning Research , issn=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.163789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.163789Z digest=sha256:2b239137dd93d57ab775cb766094a6f6a24dcfe17be30923d15c2e814ec1324d

Observation d54ca5b8-118b-40d5-851d-12a2ad92a030 · outbound

This paper cites Calibrating Verbalized Probabilities for Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Calibrating Verbalized Probabilities for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.202100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.202100Z digest=sha256:a3cb961a3a453d1f1690470d1280bf91ddf9624ab34dc617c29155cc6456ff44

Observation 8483de6d-16eb-4114-8f3c-f00ff81c13e3 · outbound

This paper cites IEEE Transactions on Software Engineering , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification IEEE Transactions on Software Engineering , year =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.916227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.244339Z digest=sha256:9ae56cacccdba0c1c0bad43f535ea1d99510d82cd300c1f8a701b49dd21b2cce

Observation 1ffeb2e8-126d-4899-9415-470b8ae2a15c · outbound

This paper cites The Eleventh International Conference on Learning Representations (ICLR 2023) , year=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification The Eleventh International Conference on Learning Representations (ICLR 2023) , year=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.730775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.260749Z digest=sha256:27fb4bc173e75b5803a5eb2c844653424643e38ed34b40b5c3018ae1f6db26bc

Observation 0e39d061-dbc2-4868-b125-d5a1189a4618 · outbound

This paper cites 2022 , eprint=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification 2022 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.312295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.312295Z digest=sha256:c1e88b90acc3b8fa8ed7d8d942ee2117e80e83bba862f52126182efe8506b853

Observation c70988a7-752a-48df-8e71-efbbe0ed9fe3 · outbound

This paper cites Qwen3 Technical Report.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.330738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.330738Z digest=sha256:0c4aaba2e8f40e89513a5709e70a2c7689459d9ce907f7a4996b245dad0ad833

Observation 37e4596c-face-48b6-bc6b-c340f95d6986 · outbound

This paper cites Journal of Machine Learning Research , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Journal of Machine Learning Research , year =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.408396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.408396Z digest=sha256:2db3f8cc604e641b3dd783fa29c4baaad8e208e0241ca6b5361204d53b8c087a

Observation e2e800fa-b6f8-40d2-b683-36cc8ff90bdb · outbound

This paper cites Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.546227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.484689Z digest=sha256:2537885fa2cd8158101141b088e88ad174a4e233adda024280b1afc1c6d0a4ed

Observation 7dc4f51d-b218-48ed-b037-d87c76c0ddbe · outbound

This paper cites BingoGuard:.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification BingoGuard:

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.390257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.551490Z digest=sha256:2af793f5de1c7e55a9e5b0e35fd621a1b03bbbdd52277cca54c0145df0db0ac1

Observation b7e78a0a-5396-4e4a-97ee-a02988d68451 · outbound

This paper cites SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:05:42.357608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.618195Z digest=sha256:7760fb5c531035c8b108bdb33f37bf6227db477ba314b1ce9d2bb9352bdae63d

Observation 71bfb71e-37b6-4834-ac8a-7c3b94e6a197 · outbound

This paper cites arXiv e-prints , pages=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv e-prints , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.243107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.720053Z digest=sha256:d3057c716f53bb0c7bbde2a7befc007a5224fd3848d0c0a8b4d45054f9425104

Observation 92c313d1-b75d-45ac-b315-01a4fc31fd31 · outbound

This paper cites Journal of classification , volume=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Journal of classification , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.113941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.796050Z digest=sha256:a42362b42a03b32c54449dda16cd77200a22fa7c69788275069d21ae87d01b9d

Observation 7c47c8ed-00b2-41b4-a9cc-40c6a6c7ab16 · outbound

This paper cites Genome Biology , volume=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Genome Biology , volume=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:42.919907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.872590Z digest=sha256:a756dd38f94f60147e887e6ece927cc3e8d9627896aa6a227dd231833ee6398d

Observation 22dd0a2b-96f2-41d1-87cd-453080fe5b4d · outbound

This paper cites Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering , pages =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering , pages =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.938459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.938459Z digest=sha256:5fb1ef524ffd4a004309eb0a832617e4aca04eb936d0e6be80398451986225d6

Observation 8ecfb3a6-3148-4077-bbea-d14d1054e9b8 · outbound

This paper cites and Zhang, Hao and Gonzalez, Joseph E.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification and Zhang, Hao and Gonzalez, Joseph E

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.050280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.050280Z digest=sha256:5df79385a76cfc86ce104378dbdbdda57fbfec98c1e56fbb7da8a4ec6ae6e7a7

Observation 9f679a33-70be-4c48-81cc-a05595cbaf0a · outbound

This paper cites Proceedings of the third International Workshop on Machine Learning in Systems Biology , pages =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the third International Workshop on Machine Learning in Systems Biology , pages =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.099071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.099071Z digest=sha256:b330f3be7dbfe223d7e1fa00136148b71042690b1f4d531ba0eeadba772666a8

Observation 65764544-f0b3-448a-be91-daaf472f6bc6 · outbound

This paper cites Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining , year =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:42.739886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:40.173527Z digest=sha256:c54f63e09f088c950bc7dc4eb6973187136f0b7c67eda6e079bd059263a20c16

Observation 02dcc27d-ea26-471e-9c6c-d554679344e9 · outbound

This paper cites I Can't Believe It's Not Better: Failure Modes in the Age of Foundation Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification I Can't Believe It's Not Better: Failure Modes in the Age of Foundation Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.282309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.282309Z digest=sha256:1d21f0f9e866eae038ebdf1b50d7bf6df741310a630de7262eb5c1738c201510

Observation db888785-34b8-4ad5-b510-cc87ad917766 · outbound

This paper cites arXiv preprint arXiv:2506.01734 , year=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv preprint arXiv:2506.01734 , year=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.444848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.444848Z digest=sha256:ec9b43bc7345f9dcec8ef57954b3daafdb9e52dd201801e1228adce960f82e3d

Observation d99f4ffe-c164-4358-9fee-cf9e63bcdb3f · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS 2015) , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Advances in Neural Information Processing Systems (NeurIPS 2015) , year =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:42.607419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:40.534443Z digest=sha256:80834837a23af987d2bb8ea5796b035c26d4a1f845cf297ed62d564ff8d39d9e

Observation 05d2a915-b4d7-4299-83f2-859488b06c60 · outbound

This paper cites Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.626084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.626084Z digest=sha256:452e731c4dae1f865b314972dc2c98bab385c9b091997a662d79652d5f93cc7c

Observation 6c5ccd92-3062-4f97-b229-241f5a76e48f · outbound

This paper cites , author=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification , author=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.784413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.784413Z digest=sha256:fa0fceb39ffb7e05d251f88b4f2ef6e6c8832803f284c51ae27f39836062cf5a

Observation 104a98c7-f521-48ee-a043-2816272ac1bb · outbound

This paper cites arXiv preprint arXiv:2509.13813 , year=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv preprint arXiv:2509.13813 , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.854578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.854578Z digest=sha256:dfbdee91e76f951128721ea16c2f5090d3bb54f3a8c05400cfa5d53ae3ffc50a

Observation e46d0c1e-5bda-41ad-a4b4-31474d59b214 · outbound

This paper cites M onte C arlo Temperature: a robust sampling strategy for LLM ' s uncertainty quantification methods.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification M onte C arlo Temperature: a robust sampling strategy for LLM ' s uncertainty quantification methods

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:41.018426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:41.018426Z digest=sha256:829caf5bc25c957a16dd348de98e3b2592b0bcec32f316331070fb192f9179fa

Observation 7c6fc369-868a-4849-aa51-723192ef9d2a · outbound

This paper cites Scikit-learn: Machine Learning in Python , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Scikit-learn: Machine Learning in Python , year =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:41.097460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:41.097460Z digest=sha256:6a378414c895f534b7aeb38dd3a9a6801671462a6fc64586a66d887af58cfa1c

Observation e9aff573-6c02-4c52-92a0-8b2142e1114a · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:41.210502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:41.210502Z digest=sha256:ea1ed4c6af6394a8b6e9ae1b65ea6241aefe6651598dc0c11b92b96f206bbd35

Pith citing papers

No inbound Pith citation observations are available.