LLM agents often fail to abstain at the right time in uncertain multi-turn tasks, and the CONVOLVE context engineering method raises timely abstention rates on WebShop from 26.7 to 57.4 without parameter updates.
Position: Uncertainty quantification needs reassessment for large-language model agents
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 6roles
background 1polarities
background 1representative citing papers
Introduces epistemic calibration as a strictly stronger criterion than classical calibration for second-order models and proposes the Expected Epistemic Calibration Error (EECE) as a consistent estimator of the True Epistemic Calibration Error (TECE).
Introduces Trajectory Proper Score (TPS) as a strictly proper family of trajectory-level scoring rules that elicits the complete prefix-conditioned success probability process.
Helicase proposes an autonomous multi-agent LLM framework for uncertainty-guided supply chain knowledge graph construction evaluated on the new SCQA benchmark of 80 queries.
TokUR estimates token-level uncertainty via low-rank weight perturbations in LLMs, aggregates signals to correlate with correctness, and uses them to improve reasoning performance on math tasks.
A survey that categorizes uncertainty quantification approaches for graphical models into representation and handling dimensions to identify challenges and opportunities.
citing papers explorer
-
Agentic Abstention: Do Agents Know When to Stop Instead of Act?
LLM agents often fail to abstain at the right time in uncertain multi-turn tasks, and the CONVOLVE context engineering method raises timely abstention rates on WebShop from 26.7 to 57.4 without parameter updates.
-
Can we trust our models? Epistemic calibration in second-order classification
Introduces epistemic calibration as a strictly stronger criterion than classical calibration for second-order models and proposes the Expected Epistemic Calibration Error (EECE) as a consistent estimator of the True Epistemic Calibration Error (TECE).
-
Proper Scoring Rules for Agentic Uncertainty Quantification
Introduces Trajectory Proper Score (TPS) as a strictly proper family of trajectory-level scoring rules that elicits the complete prefix-conditioned success probability process.
-
Helicase: Uncertainty-Guided Supply Chain Knowledge Graph Construction with Autonomous Multi-Agent LLMs
Helicase proposes an autonomous multi-agent LLM framework for uncertainty-guided supply chain knowledge graph construction evaluated on the new SCQA benchmark of 80 queries.
-
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
TokUR estimates token-level uncertainty via low-rank weight perturbations in LLMs, aggregates signals to correlate with correctness, and uses them to improve reasoning performance on math tasks.
-
Uncertainty Quantification on Graph Learning: A Survey
A survey that categorizes uncertainty quantification approaches for graphical models into representation and handling dimensions to identify challenges and opportunities.