Pith. sign in

REVIEW 1 cited by

Efficient Non-Parametric Uncertainty Quantification for Black-Box Large Language Models and Decision Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.00251 v1 pith:64CO7NPA submitted 2024-02-01 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords agentdecisionuncertaintylanguagellmsmodelsplanningblack-box
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Step-by-step decision planning with large language models (LLMs) is gaining attention in AI agent development. This paper focuses on decision planning with uncertainty estimation to address the hallucination problem in language models. Existing approaches are either white-box or computationally demanding, limiting use of black-box proprietary LLMs within budgets. The paper's first contribution is a non-parametric uncertainty quantification method for LLMs, efficiently estimating point-wise dependencies between input-decision on the fly with a single inference, without access to token logits. This estimator informs the statistical interpretation of decision trustworthiness. The second contribution outlines a systematic design for a decision-making agent, generating actions like ``turn on the bathroom light'' based on user prompts such as ``take a bath''. Users will be asked to provide preferences when more than one action has high estimated point-wise dependencies. In conclusion, our uncertainty estimation and decision-making agent design offer a cost-efficient approach for AI agent development.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 2 citations worldwide. Full citation record

  1. AGENT-X: Adaptive Guideline-based Expert Network for Threshold-free AI-generated teXt detection

    cs.CL 2025-05 reject novelty 6.0 of 10

    AGENT-X is a zero-shot multi-LLM framework for AI-generated text detection that routes texts to guideline-specific agents and aggregates their calibrated confidences without threshold tuning.

Pith tools