Pith. sign in

Paper Citation Record · LEDGER

A Survey of Calibration Process for Black-Box LLMs

As of 21 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 4 inbound Pith citation observations for arXiv:2412.12767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12767 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:48:40.203898Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:07:50.703898Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:39:30.629875Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 732f59f5-ed1f-493c-be68-4ba243cc6f26 · outbound

This paper cites Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models.

A Survey of Calibration Process for Black-Box LLMs Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:40.023939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:40.023939Z digest=sha256:47595f98b6da00cbe7fce42fd4b6416c62aa01a948a4ada5666a04147bc4f6f0

Observation b43b8860-4334-4052-83d1-732ea4af838f · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

A Survey of Calibration Process for Black-Box LLMs Measuring and Narrowing the Compositionality Gap in Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:40.040446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:40.040446Z digest=sha256:1bc0b8ef56a77c84bdd34bceef67db30d903eff261e78f16f6d617607e0a6157

Observation 704f4992-1364-40d0-b0e2-d2db983c04b8 · outbound

This paper cites Llamas Know What GPTs Don't Show: Surrogate Models for Confidence Estimation.

A Survey of Calibration Process for Black-Box LLMs Llamas Know What GPTs Don't Show: Surrogate Models for Confidence Estimation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:40.045629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:40.045629Z digest=sha256:30c24c28633fb3bbf93216c03ea993262d1ca9cf6061ff6a6cbcdf4b77f98b4c

Observation a8452a20-fb53-4a6a-9828-386009bc9ecb · outbound

This paper cites Efficient Non-Parametric Uncertainty Quantification for Black-Box Large Language Models and Decision Planning.

A Survey of Calibration Process for Black-Box LLMs Efficient Non-Parametric Uncertainty Quantification for Black-Box Large Language Models and Decision Planning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:40.064221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:40.064221Z digest=sha256:83bc6cb3fe970b40524f675e21d001273245b378dfd22bed5684857852c3f220

Observation af9711ed-856e-4e79-95df-f95c0051aa06 · outbound

This paper cites Black-box Uncertainty Quantification Method for LLM-as-a-Judge.

A Survey of Calibration Process for Black-Box LLMs Black-box Uncertainty Quantification Method for LLM-as-a-Judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:40.102166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:40.102166Z digest=sha256:6bd71fc4c74496a1c5fdeaa286f9fe3595e0323ed3257740bd70e6286dd1cb3e

Observation 94c56138-c5dd-433c-a4c6-02d5fee53782 · outbound

This paper cites SportQA: A Benchmark for Sports Understanding in Large Language Models.

A Survey of Calibration Process for Black-Box LLMs SportQA: A Benchmark for Sports Understanding in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:40.153721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:40.153721Z digest=sha256:e93de0958e43f903004e74cac0aa40f42c10ee95a71ca6b6ef84d37f2a5abf6c

Observation ad742b99-18b2-451a-8b57-72f214ace831 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

A Survey of Calibration Process for Black-Box LLMs ReAct: Synergizing Reasoning and Acting in Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:40.194494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:40.194494Z digest=sha256:658fa8a63b9be6621a404a72cb5713040bffb7920933fc2dd48fd0a36f929054

Observation 2abe96e0-62f0-4858-800f-2b43bc8b4d7f · outbound

This paper cites These methods focus on evaluating the model’s ability to distinguish be- tween correct and incorrect samples at various con- fidence thresholds, but their emphases differ.

A Survey of Calibration Process for Black-Box LLMs These methods focus on evaluating the model’s ability to distinguish be- tween correct and incorrect samples at various con- fidence thresholds, but their emphases differ

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:48:40.462209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:48:40.203898Z digest=sha256:03dad9e8234af59e53a4f68509ae9ef8523ab8153ea7ff5c8fd85005e8353d90

Observation ba213b30-62b1-47c8-8754-13b7ab796a3f · outbound

This paper cites Question Answering with Subgraph Embeddings.

A Survey of Calibration Process for Black-Box LLMs Question Answering with Subgraph Embeddings

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:39.920994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:39.920994Z digest=sha256:f20f5ed55b801e6eae5bedf5424a01d5933420c6f6ddd1852421c71ff37dc977

Observation d1251da6-f93d-48bb-9f3a-75f39e95207d · outbound

This paper cites an unresolved cited work.

A Survey of Calibration Process for Black-Box LLMs Unresolved cited work

Reference 2015

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:48:40.477980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T13:48:40.199299Z digest=sha256:c3a9e5cb98d281fb3144c8ef48734103aecf2a11b3d04b8da363703e4c7c0cd8

Observation c88a1d24-2d47-420f-a6f4-a33fdb98c01d · outbound

This paper cites Top-label calibration and multiclass-to-binary reductions.

A Survey of Calibration Process for Black-Box LLMs Top-label calibration and multiclass-to-binary reductions

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:39.983849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:39.983849Z digest=sha256:f669fa64b5f9770dbfaed90d682096eabc251690888d1e35474633d033e0590f

Observation 4e79c679-6252-4007-96b3-896084d0b399 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

A Survey of Calibration Process for Black-Box LLMs Measuring Massive Multitask Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:40.018003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:40.018003Z digest=sha256:2fd1b249c64965fc58e56ef8db74cb5d8e4085c6e14e4c3ee09db3f467bd30e0

Observation 31846bd5-639a-4355-9207-fa2e20d5919f · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

A Survey of Calibration Process for Black-Box LLMs Are NLP Models really able to Solve Simple Math Word Problems?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:40.035299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:40.035299Z digest=sha256:73e517dad4eba6448b0ad30b3b999d50fe16b86eafdba148910b972b8329b435

Observation ea7834fc-af23-46bc-bf46-ef753cc43ee5 · outbound

This paper cites GPT-4 Technical Report.

A Survey of Calibration Process for Black-Box LLMs GPT-4 Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:39.915295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:39.915295Z digest=sha256:1d1fbcdf3c077e64fd0c32301272929e3bdecc7a4cacf1adbcae8248487dfdc1

Observation eb2067a2-c4fa-4293-a719-a719a50d6748 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

A Survey of Calibration Process for Black-Box LLMs Training Verifiers to Solve Math Word Problems

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:39.944316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:39.944316Z digest=sha256:c5c25451cda8cfa76265f9b2da57ff9f2eadaff087fa332d0105b61cb9876591

Observation b9a627a3-3e13-4009-a66a-7b327cdc1032 · outbound

This paper cites When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation.

A Survey of Calibration Process for Black-Box LLMs When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T13:48:40.030345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:48:40.030345Z digest=sha256:d221aebc68d7d9ca3d25cb19b2c108da848853c4493c5b1fc261d0398e117896

Pith citing papers

Observation 9b275fda-3be1-4b5d-9c6c-a4a8f2794bfa · inbound

Maximizing Confidence Alone Improves Reasoning cites this paper.

Maximizing Confidence Alone Improves Reasoning A Survey of Calibration Process for Black-Box LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:50.703898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:50.703898Z digest=sha256:cf1e75917498712c62c1c76adc6cb9233d61bc6dc5b359358d96154f6e53389f

Observation 3f2b8bbc-d76e-4bfe-9f49-2352929d9909 · inbound

Understanding How University Guidelines Address Privacy and Security Issues of Generative AI in Academic Settings cites this paper.

Understanding How University Guidelines Address Privacy and Security Issues of Generative AI in Academic Settings A Survey of Calibration Process for Black-Box LLMs

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T22:50:01.001810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:50:01.001810Z digest=sha256:fe517f8ef42b8b5d23c4a08f459fa6f16e71a0051f0f5a335f5f3e0f604889d4

Observation 3664f820-0d29-43c4-8012-1c2c4b106734 · inbound

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models cites this paper.

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models A Survey of Calibration Process for Black-Box LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:30.631972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T17:46:00.822375Z digest=sha256:e6c5b0fb6ea2f89ae6dfabac3bc78052c8557a212a8674309689717c39e584fe

Observation ea20bbbc-c0f3-4d6d-9b09-4d913f21f0f0 · inbound

Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction cites this paper.

Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction A Survey of Calibration Process for Black-Box LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T20:42:51.961770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:42:51.961770Z digest=sha256:ad37c4c1911a8ed9f89b18964c922fbd4b9b8ef379d0c268e8ac564510c3b8d4