Pith. sign in

Paper Citation Record · LEDGER

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

As of 14 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.04899.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04899 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:05:41.210502Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67c64fc7-7a01-438b-9011-c290c8f887c0 · outbound

This paper cites Benchmarking Uncertainty Quantification Methods for Large Language Models with LM -Polygraph.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Benchmarking Uncertainty Quantification Methods for Large Language Models with LM -Polygraph

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:37.879497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:37.879497Z digest=sha256:2aee011ed3b57b9ace6e2e24345f30484d763297307e5b8c09bd8408d58938c9

Observation 7a7e6c93-a5b9-478e-b940-8ca88100279b · outbound

This paper cites Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:37.927074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:37.927074Z digest=sha256:3112956e4659de5fb3bb7c62df2e9593c5aa34218d388c86c368881d7409d6aa

Observation caba163f-e3e8-4a3a-9a87-3f11f97e151d · outbound

This paper cites Thinking Out Loud: Do Reasoning Models Know When They ' re Right?.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Thinking Out Loud: Do Reasoning Models Know When They ' re Right?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:37.956147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:37.956147Z digest=sha256:9b3539d63caf01f6af4b304afbf5f040cac7ae3aaa0f8a14fac999016628821f

Observation fc2df6d9-8137-4ac7-b10d-8df1521ceb63 · outbound

This paper cites Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.004867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.004867Z digest=sha256:2a8dfbd1941d6520ad3d3bb6b61bb3c43288680458f48996d39e68aa03222f25

Observation f6c237d1-89ae-4f6a-b84f-a32fd0712050 · outbound

This paper cites C heck E val: A reliable LLM -as-a-Judge framework for evaluating text generation using checklists.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification C heck E val: A reliable LLM -as-a-Judge framework for evaluating text generation using checklists

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.057387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.057387Z digest=sha256:58145499492354b2097ca4ad378c9cab9d6bb6e2e1f2a2c3c179e52039eb4d19

Observation 8c082c08-0007-478a-97ce-a26110fa4f0a · outbound

This paper cites an unresolved cited work.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.102656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.102656Z digest=sha256:f140984e78914f139d2652fef4661a9c010c5a51a3bffc4051d0ab49643f7c2e

Observation 8424d035-3c56-4419-b41d-6805fe56d743 · outbound

This paper cites Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.149210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.149210Z digest=sha256:b3dfab1128c94d0cf6746bf1590b546ee271d13dab3f2166950394cbf9259c91

Observation 43b691c4-2287-4b89-8772-193c3feeb604 · outbound

This paper cites Uncertainty Quantification for In-Context Learning of Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Uncertainty Quantification for In-Context Learning of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.184134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.184134Z digest=sha256:d085ac3c71d2579ae72cb09df493b9abed36e68db3114c2c7ecf645ba22b0343

Observation 0941cce5-dedf-45b8-8ad2-71979b193d44 · outbound

This paper cites A Survey of Confidence Estimation and Calibration in Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification A Survey of Confidence Estimation and Calibration in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.283292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.283292Z digest=sha256:ec8ed8df75c770b97c1bb14cfa6959a6dd052e3e480bc854428e0da97d6da26a

Observation d488a415-7af0-45ba-8ea0-c62dbd6cc024 · outbound

This paper cites B ayesian Prompt Ensembles: Model Uncertainty Estimation for Black-Box Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification B ayesian Prompt Ensembles: Model Uncertainty Estimation for Black-Box Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.357788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.357788Z digest=sha256:f6535ac68b1daf2419120d38454885b81da7199e7f11307a43217dbc2e2e502d

Observation 84f94c29-51eb-4379-8f71-c1d0d6811834 · outbound

This paper cites Calibrating the Confidence of Large Language Models by Eliciting Fidelity.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Calibrating the Confidence of Large Language Models by Eliciting Fidelity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.423013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.423013Z digest=sha256:3f48f67a4903aba90cdfcadff5a3963eda7e21758c57c6a3cf5c4ff3ffe670fd

Observation 09f6206a-bb22-4359-93d7-67e3bde20ca6 · outbound

This paper cites Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.523404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.523404Z digest=sha256:88ec99ae5de3988c5d2a664be21ba94c8568eb869708f32f0a59d90dd2eccd64

Observation 22d28937-d08e-49bd-977c-ab61269cdb4e · outbound

This paper cites Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.634360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.634360Z digest=sha256:53f4363f50406df5e5bb19d5514c2acc625a33404172108434dadaddee8a60a6

Observation 9d454fba-fe3e-44e2-848e-c2ee5eb5ebcd · outbound

This paper cites Adaptation with Self-Evaluation to Improve Selective Prediction in LLM s.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Adaptation with Self-Evaluation to Improve Selective Prediction in LLM s

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T14:05:41.484457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:38.744461Z digest=sha256:96d5ef382ff1f42a6a617caacd84c1a7c2f0fd1acdd6ba89808ea0d482203859

Observation e53c6dec-ae7c-441d-b73d-8f80b6d37099 · outbound

This paper cites Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.842660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.842660Z digest=sha256:225b346b4c8dbf1425c2dcf68fca0036f3e99e2ebdd8ef31b3598aa17ef5a936

Observation c567c4c8-6bda-4b9c-940b-1bc8ea15e20a · outbound

This paper cites and Ng, Andrew and Potts, Christopher.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification and Ng, Andrew and Potts, Christopher

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.908813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.908813Z digest=sha256:2ecef1bda9d65547236574776b070f2e62ad029ae6edc6477b500e467bcb2c5c

Observation f90ee083-453a-467e-a92d-8382fc2556a2 · outbound

This paper cites an unresolved cited work.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:05:44.292357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.020702Z digest=sha256:b414f000d20211a7b0f24f1dd899d3963fc4e14a1c13a07be328de82c6b0b302

Observation 01d8053b-95cf-4bc5-8fe0-602a82ee3898 · outbound

This paper cites Proceedings of the 36th International Conference on Neural Information Processing Systems , articleno =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 36th International Conference on Neural Information Processing Systems , articleno =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.056971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.056971Z digest=sha256:523366530184b1eb15e9e2aedbc23b83cf0db74ec80c48a73e557f01499ddddf

Observation d8db0919-2b11-47cf-9b1d-1aad27c1d093 · outbound

This paper cites 2025 , eprint=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification 2025 , eprint=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:44.115273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.111242Z digest=sha256:fd386fe110d85781d2ea8f5e2cb7243c847714308b554e11074112c580b0192c

Observation 21498f79-42eb-4369-bd65-0d088b4bcf55 · outbound

This paper cites Transactions on Machine Learning Research , issn=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Transactions on Machine Learning Research , issn=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.163789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.163789Z digest=sha256:0783d78b97c787d675d3fd1a9bfa242b0295a3adbd472a3da3b94f29bf3e167b

Observation d54ca5b8-118b-40d5-851d-12a2ad92a030 · outbound

This paper cites Calibrating Verbalized Probabilities for Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Calibrating Verbalized Probabilities for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.202100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.202100Z digest=sha256:25b5a7896d0c00ea9adb19f48da50f291cb223254f042644a9d3c162163edc86

Observation 8483de6d-16eb-4114-8f3c-f00ff81c13e3 · outbound

This paper cites IEEE Transactions on Software Engineering , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification IEEE Transactions on Software Engineering , year =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.916227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.244339Z digest=sha256:1bf5f6cc6a69ccf5ec69dc9026f7c245a76b7152a46c55fe3cd9e426b1b97eea

Observation 1ffeb2e8-126d-4899-9415-470b8ae2a15c · outbound

This paper cites The Eleventh International Conference on Learning Representations (ICLR 2023) , year=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification The Eleventh International Conference on Learning Representations (ICLR 2023) , year=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.730775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.260749Z digest=sha256:917587e95c0ee08e36ac4a9a0d7671fc0ae8d145cbe15d7300ef0fdba22e982f

Observation 0e39d061-dbc2-4868-b125-d5a1189a4618 · outbound

This paper cites 2022 , eprint=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification 2022 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.312295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.312295Z digest=sha256:c3d8daae5427f299af2271aa4d49250c48d572995f7fa8c6f69518cf1291554b

Observation c70988a7-752a-48df-8e71-efbbe0ed9fe3 · outbound

This paper cites Qwen3 Technical Report.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.330738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.330738Z digest=sha256:007ec94725f8a95c383832574ddbcb64b4bfaccd7b78bd986d9af3658d1baf5d

Observation 37e4596c-face-48b6-bc6b-c340f95d6986 · outbound

This paper cites Journal of Machine Learning Research , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Journal of Machine Learning Research , year =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.408396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.408396Z digest=sha256:57f5b54f5b0ffc1e3be436ba2a0947d9a11e4489bf10e1441a3d1b17b16c93f3

Observation e2e800fa-b6f8-40d2-b683-36cc8ff90bdb · outbound

This paper cites Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.546227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.484689Z digest=sha256:fecac015bd94c0b46a29807f8545d601caa0ad9869fa076bfc0170db67ddda5b

Observation 7dc4f51d-b218-48ed-b037-d87c76c0ddbe · outbound

This paper cites BingoGuard:.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification BingoGuard:

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.390257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.551490Z digest=sha256:ef6e9a1e736b20d83c64c81ea018376dc1ce6dc5baeee8e17a83c11fe913aa00

Observation b7e78a0a-5396-4e4a-97ee-a02988d68451 · outbound

This paper cites SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:05:42.357608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.618195Z digest=sha256:430c6f48d71f2bd795ce9c52281db22612bf3ec1d5ccdd86be22c0e496c1677d

Observation 71bfb71e-37b6-4834-ac8a-7c3b94e6a197 · outbound

This paper cites arXiv e-prints , pages=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv e-prints , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.243107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.720053Z digest=sha256:a519db8e22d6fc279572954a51f1c114a5e0e284a71780163d7a7b4ee4a26728

Observation 92c313d1-b75d-45ac-b315-01a4fc31fd31 · outbound

This paper cites Journal of classification , volume=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Journal of classification , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.113941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.796050Z digest=sha256:76c217f09fa8b20cd0b62ce9b118db4aba7f23b3af7e241eac9a7f38b912eade

Observation 7c47c8ed-00b2-41b4-a9cc-40c6a6c7ab16 · outbound

This paper cites Genome Biology , volume=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Genome Biology , volume=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:42.919907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.872590Z digest=sha256:8fb2776a18b3835f61b796860ced7c03dad317009c190e0a0a462adc18ee575a

Observation 22dd0a2b-96f2-41d1-87cd-453080fe5b4d · outbound

This paper cites Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering , pages =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering , pages =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.938459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.938459Z digest=sha256:4a260a4776406c800026193f4d8b64d0afa56700e2976067105865c103d7652f

Observation 8ecfb3a6-3148-4077-bbea-d14d1054e9b8 · outbound

This paper cites and Zhang, Hao and Gonzalez, Joseph E.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification and Zhang, Hao and Gonzalez, Joseph E

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.050280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.050280Z digest=sha256:026f318f9520929c72c298f8849046680296809769ce4e5c763791d576e8f479

Observation 9f679a33-70be-4c48-81cc-a05595cbaf0a · outbound

This paper cites Proceedings of the third International Workshop on Machine Learning in Systems Biology , pages =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the third International Workshop on Machine Learning in Systems Biology , pages =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.099071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.099071Z digest=sha256:aad4b2b62b0e9d398e3e79ad9716fc255a1faa634b4b3b57c34564512982fe49

Observation 65764544-f0b3-448a-be91-daaf472f6bc6 · outbound

This paper cites Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining , year =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:42.739886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:40.173527Z digest=sha256:4d32e4ba40c9ed349639f80db46048b7a9800caba266a106fd695b1ffac35d7f

Observation 02dcc27d-ea26-471e-9c6c-d554679344e9 · outbound

This paper cites I Can't Believe It's Not Better: Failure Modes in the Age of Foundation Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification I Can't Believe It's Not Better: Failure Modes in the Age of Foundation Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.282309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.282309Z digest=sha256:f371ab6f41c3661539dabc340c846db67b90f073c651ac416e116dae7ad18380

Observation db888785-34b8-4ad5-b510-cc87ad917766 · outbound

This paper cites arXiv preprint arXiv:2506.01734 , year=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv preprint arXiv:2506.01734 , year=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.444848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.444848Z digest=sha256:caa3a04c8bb7490f75d1750242dfb851931c38234d0551f795daaf8ff5d626ec

Observation d99f4ffe-c164-4358-9fee-cf9e63bcdb3f · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS 2015) , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Advances in Neural Information Processing Systems (NeurIPS 2015) , year =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:42.607419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T14:05:40.534443Z digest=sha256:223e361772c9933b3ab4d7cff8166db69b917d5f89162fdbb9df6c763823d2b6

Observation 05d2a915-b4d7-4299-83f2-859488b06c60 · outbound

This paper cites Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.626084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.626084Z digest=sha256:3eeb1e80e287ace35f70585a2c667b8f415a2b893e81ef31baf3acd7d2a09fd8

Observation 6c5ccd92-3062-4f97-b229-241f5a76e48f · outbound

This paper cites , author=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification , author=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.784413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.784413Z digest=sha256:4194b06ca9e3a0cbcf4bef79880aae426fda4e28da1f79fc90114134995e4d9a

Observation 104a98c7-f521-48ee-a043-2816272ac1bb · outbound

This paper cites arXiv preprint arXiv:2509.13813 , year=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv preprint arXiv:2509.13813 , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.854578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.854578Z digest=sha256:3babe38dacd91f49513da20f88dc1c30973b8871e3111a200278523a5f83b6f4

Observation e46d0c1e-5bda-41ad-a4b4-31474d59b214 · outbound

This paper cites M onte C arlo Temperature: a robust sampling strategy for LLM ' s uncertainty quantification methods.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification M onte C arlo Temperature: a robust sampling strategy for LLM ' s uncertainty quantification methods

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:41.018426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:41.018426Z digest=sha256:d5f500ed7a245108a4787f4bebec52bb827c950b73712f909e3ca920b77cd803

Observation 7c6fc369-868a-4849-aa51-723192ef9d2a · outbound

This paper cites Scikit-learn: Machine Learning in Python , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Scikit-learn: Machine Learning in Python , year =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:41.097460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:41.097460Z digest=sha256:89c2502f8a4ee29a4349df3a492ea9fff473fc5b44c9afa1aabcfb442120de97

Observation e9aff573-6c02-4c52-92a0-8b2142e1114a · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:41.210502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:41.210502Z digest=sha256:3917f36478575c053414e2ea5a9c942a09c2c851284a072bf56d12dd21ed0c9b

Pith citing papers

No inbound Pith citation observations are available.