Pith. sign in

REVIEW 4 cited by

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.16113 v1 pith:CQZUJMPO submitted 2025-05-22 cs.LG cs.CL

classification cs.LGcs.CL
keywords uncertaintyexternaltoolssystemsframeworkknowledgellmstool
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern Large Language Models (LLMs) often require external tools, such as machine learning classifiers or knowledge retrieval systems, to provide accurate answers in domains where their pre-trained knowledge is insufficient. This integration of LLMs with external tools expands their utility but also introduces a critical challenge: determining the trustworthiness of responses generated by the combined system. In high-stakes applications, such as medical decision-making, it is essential to assess the uncertainty of both the LLM's generated text and the tool's output to ensure the reliability of the final response. However, existing uncertainty quantification methods do not account for the tool-calling scenario, where both the LLM and external tool contribute to the overall system's uncertainty. In this work, we present a novel framework for modeling tool-calling LLMs that quantifies uncertainty by jointly considering the predictive uncertainty of the LLM and the external tool. We extend previous methods for uncertainty quantification over token sequences to this setting and propose efficient approximations that make uncertainty computation practical for real-world applications. We evaluate our framework on two new synthetic QA datasets, derived from well-known machine learning datasets, which require tool-calling for accurate answers. Additionally, we apply our method to retrieval-augmented generation (RAG) systems and conduct a proof-of-concept experiment demonstrating the effectiveness of our uncertainty metrics in scenarios where external information retrieval is needed. Our results show that the framework is effective in enhancing trust in LLM-based systems, especially in cases where the LLM's internal knowledge is insufficient and external tools are required.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Uncertainty Propagation in LLM-Based Systems

    cs.SE 2026-04 unverdicted novelty 7.0 of 10

    This paper introduces a systems-level conceptual framing and a three-level taxonomy (intra-model, system-level, socio-technical) for uncertainty propagation in compound LLM applications, along with engineering insight...

  2. Diagnosis-Driven Automatic Repair for Agentic Workflow via Symbolic Inference

    cs.SE 2026-07 conditional novelty 6.5 of 10

    FlowFixer uses symbolic inference of node behavioral specs to diagnose and repair agentic workflows, reaching 71.3% repair success and higher attribution accuracy than baselines on Dify/Coze/n8n failures.

  3. Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Repeated zero-shot summaries from the same LLM and document vary substantially in semantic and factual scores, and this paper proposes stability coefficients as a benchmark for that variability.

  4. ToolChain-CRC: Conformal Risk Control for Agentic AI Under Retrieval and Tool-Use Drift

    stat.ML 2026-06 unverdicted novelty 5.0 of 10

    ToolChain-CRC is a trajectory-level conformal risk control method for agentic AI that calibrates accept-or-intervene rules on combined step risks, adds drift-aware extensions, and includes an anytime supermartingale alarm.

Pith tools