Pith. sign in

Paper Citation Record · LEDGER

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

As of 23 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 4 inbound Pith citation observations for arXiv:2505.16113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16113 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:11:46.806710Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:46:14.378320Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:29:02.503467Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 025f2365-18a6-4afb-a3c6-37b30ae1d0e9 · outbound

This paper cites Benchmarking Bayesian Deep Learning on Diabetic Retinopathy Detection Tasks.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Benchmarking Bayesian Deep Learning on Diabetic Retinopathy Detection Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:45.024634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:45.024634Z digest=sha256:bdef98803b481a9db811b2b5a3aae150aa34222169b0e872a0f5112c948c35a9

Observation 3e90927a-26d4-4e50-88ac-cace3e22a086 · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no ques- tions.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Boolq: Exploring the surprising difficulty of natural yes/no ques- tions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:48.530797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:11:45.065298Z digest=sha256:2f66da23e138c9b343f2bc32e8b06bfebac34d02e33d645726b42717eaa1818e

Observation b5589958-8520-4e0c-87cb-fe19a46bdc04 · outbound

This paper cites Detecting hallucina- tions in large language models using semantic entropy.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Detecting hallucina- tions in large language models using semantic entropy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:48.349186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:11:45.141259Z digest=sha256:a2d72262064135d0e75050fd9b2def6c5c0d9faf74bc430780e6434d7a5af926

Observation 4ce3ef00-44c0-4115-8dc4-6964fd2cbedf · outbound

This paper cites an unresolved cited work.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:45.244890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:45.244890Z digest=sha256:230d59371a989a6b27021f4f7e11dea8b3d7e40677e3edd487df08b5c6752c59

Observation 9b0a6620-f0d5-4f18-abaa-cbe46a657613 · outbound

This paper cites Dropout as a bayesian approximation: Representing model uncertainty in deep learning.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Dropout as a bayesian approximation: Representing model uncertainty in deep learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:48.038983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:11:45.455818Z digest=sha256:733c15790ddacdc551408c0436cc0879dfc774dd01547d7d27883ef9632c2ef5

Observation ca60c6f2-bc36-483a-a178-859e90297ca2 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:45.611899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:45.611899Z digest=sha256:008561393a97d610eab060082cb8f7e5e54433feeee90acd2cce4b65cc3714ab

Observation f774ab00-ed1c-437b-a2db-8fe632399df9 · outbound

This paper cites Unsolved Problems in ML Safety.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Unsolved Problems in ML Safety

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:45.769975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:45.769975Z digest=sha256:b81b0b3d8915618ce89eb89c725fd04ac308a3f655c92a00b8585ca8b124d222

Observation c975e134-c24e-47c9-b900-1294177dc146 · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:45.904860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:45.904860Z digest=sha256:e15f1c7628b7426988a2107e0671adfb58175b65f4d239aa59f7565c9b764c7d

Observation 5828cb78-9767-4de0-8c0d-3e3500a323a1 · outbound

This paper cites Large Language Models in Law: A Survey.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Large Language Models in Law: A Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.036999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.036999Z digest=sha256:494da904de8467413909aeb7e3022d8fa49b67af10ff20db54e2f9229eca1143

Observation 7b62fc94-35ec-4a67-a773-6b93691d7c96 · outbound

This paper cites Simple and scalable pre- dictive uncertainty estimation using deep en- sembles.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Simple and scalable pre- dictive uncertainty estimation using deep en- sembles

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.965482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:11:46.093367Z digest=sha256:c2f10e8a636cadb28c68606bc5a679990321b8294efe71f3e01ff224e53529ad

Observation b76735cd-c92d-4c4a-acb3-aeb5b18021c5 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.764858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:11:46.141626Z digest=sha256:2493f6cbec7a7b0a65f2bb0b7c71ad035061cdf054518836058d726193baf9d4

Observation ce0d25f6-6fa4-404f-ba4b-9f994604945a · outbound

This paper cites Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.190902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.190902Z digest=sha256:bf426e6cee6cb9ea4e600eb300a2e6020c7ca743264f3ef71ec1705aca21a1ab

Observation 924ee281-803d-45e2-a9b0-863ded1c79fa · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools ToolACE: Winning the Points of LLM Function Calling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.258369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.258369Z digest=sha256:d999a9f883a39752b90582d0ab1b2e986ec21fcb0dbbc68cdb5c81d5b94db193

Observation a481744f-c47d-4a40-8368-4cac467687bd · outbound

This paper cites Uncertainty Estimation in Autoregressive Structured Prediction.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Uncertainty Estimation in Autoregressive Structured Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.329808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.329808Z digest=sha256:80bcf727748ba4573a94b0465545483c363c8e510a93bd212c4f8b71bb789504

Observation 82700843-72b3-487d-b72a-572eccacb412 · outbound

This paper cites Revisiting, benchmarking and exploring api recommendation: How far are we?, 2021.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Revisiting, benchmarking and exploring api recommendation: How far are we?, 2021

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.667536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:11:46.380342Z digest=sha256:6fc5920fd3d37a89ea94a3b7c583fe1576984e782b64df0cf1599364761287e4

Observation 77667ac6-9d8f-4b70-8998-c3723011a061 · outbound

This paper cites Tool Learning with Foundation Models.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Tool Learning with Foundation Models

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-07T15:11:46.423549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.423549Z digest=sha256:8ba01da05696fb1f3b992c6dcf3e3252bea91dad08929cec4b175d95c0fe07a8

Observation 0b188d76-8f2e-4a63-b6bf-411e9f8cd342 · outbound

This paper cites Tool Learning with Large Language Models: A Survey.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Tool Learning with Large Language Models: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.515086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.515086Z digest=sha256:0af27bb7dbc0c114af32ac987f52e85bcdceddf2dc1a79233a06cff3cb69bea3

Observation b61fa36e-31df-4c1e-921f-ee98d4dfafdc · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Toolformer: Language models can teach themselves to use tools

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.495778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:11:46.583746Z digest=sha256:6c890ef653b0b9403b0a7ebef8d8a42327554e0db9934a2bc65e6e9cde07a467

Observation 9e6f5e98-a010-40b5-9936-c4bcbb102d38 · outbound

This paper cites Using the adap learning algorithm to forecast the onset of diabetes mellitus.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Using the adap learning algorithm to forecast the onset of diabetes mellitus

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.379534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:11:46.615512Z digest=sha256:6f42f527321042e71690f27a2ecfda1b2252aa4e42cb52200ba3a01c3ee5b6a9

Observation dfe20a7a-299d-456e-abd6-4c2c6bfd6d07 · outbound

This paper cites Question generation as a competitive undergraduate course project.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Question generation as a competitive undergraduate course project

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.248402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:11:46.677222Z digest=sha256:d6982ce5e0d0b2c700be162d90fb0dbee0fbc18a3fd1dbfb8ef945cc04a529bd

Observation 0a38c0fe-5290-47c4-9a3c-13e795e38be4 · outbound

This paper cites Large language models in medicine.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Large language models in medicine

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.763294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.763294Z digest=sha256:480b8ce982a4bfd6b25d7b7f59ebfe485574322265bb423eec85203e6efd5b0e

Observation a17d443b-48c9-454c-a8a3-2e5785c0ee5c · outbound

This paper cites Toolqa: A dataset for 12 llm question answering with external tools.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Toolqa: A dataset for 12 llm question answering with external tools

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.080564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:11:46.806710Z digest=sha256:d6b2def53330fc944c34a93c17acd119e79de5f832c6b76bf513f6b2fef0438e

Pith citing papers

Observation bd986cb2-9269-48f6-b838-cb2b1614f789 · inbound

Uncertainty Propagation in LLM-Based Systems cites this paper.

Uncertainty Propagation in LLM-Based Systems Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:16:23.687803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T06:10:07.638483Z digest=sha256:82f1b3017a9798353331a1b5759b3450a1b46962a7cfc6d09d91b84e6c3e63a0

Observation 18f9caa4-06c0-4e9a-8547-b00f52a419ed · inbound

ToolChain-CRC: Conformal Risk Control for Agentic AI Under Retrieval and Tool-Use Drift cites this paper.

ToolChain-CRC: Conformal Risk Control for Agentic AI Under Retrieval and Tool-Use Drift Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:29:02.505043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T22:16:17.385766Z digest=sha256:c4d236bf159f469c5e8606d043e29cd77b5f17e0c28ea535449b9dc61593daed

Observation a9b73959-b485-4a26-b3ed-ff9adbc85a61 · inbound

Diagnosis-Driven Automatic Repair for Agentic Workflow via Symbolic Inference cites this paper.

Diagnosis-Driven Automatic Repair for Agentic Workflow via Symbolic Inference Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T06:24:52.086585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:24:52.086585Z digest=sha256:d8578da41f816920f423a08d9ef6283f141feacf85bfe4b9f04779dc632cdf78

Observation 10f62dba-8f8e-4273-a3e8-f9a326ee90f8 · inbound

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers cites this paper.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.378320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.378320Z digest=sha256:04bed782b9517d527aa7295c5c31f91ce88d4c65466370fbbbe2f373ddd68007