Pith. sign in

REVIEW 3 major objections 6 minor 26 references

Past a miscalibration threshold that rises as budgets shrink, ranking LLM agent audits by self-reported confidence is worse than picking at random.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 11:27 UTC pith:AXVIX2QM

load-bearing objection Clean budgeted-inspection model with a real reverse flip threshold and honest measurements; open-model “beyond flip” placement is softer than the synthetic result. the 3 major comments →

arxiv 2607.28317 v1 pith:AXVIX2QM submitted 2026-07-30 cs.AI

One Human, N Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence

classification cs.AI
keywords human oversightLLM agentsaudit budgetcalibrationcorrelated errorsvacuous oversightGaussian copulaconfidence-ranked auditing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

When one human must oversee a fleet of N LLM agents with only B much smaller than N audits per round, the natural policy is to check the outputs the agents themselves flag as least confident. This paper shows that once self-reported confidence is miscalibrated past a threshold δ*, that policy becomes worse than random auditing. Counter to intuition, the threshold rises as the budget shrinks: tight budgets only touch the extreme low-confidence tail, which stays informative longer, while generous budgets dip into poisoned mid-ranks first. Measurements on five open-weight models find near-constant, operationally useless confidence whose point estimates sit at or beyond the flip; a proprietary model is informative and lands below it. Error correlation is driven by shared item difficulty, not model lineage, so pooled transfer across the fleet can still help when confidence fails. The authors also give a quantitative test for when oversight is vacuous—rubber-stamping that no studied policy beats by more than a fixed fraction.

Core claim

Confidence-ranked auditing of an LLM agent fleet flips from better than random to worse than random past a miscalibration threshold δ*(B/N) that increases as the audit budget shrinks (located near 0.93, 0.89, 0.84, 0.75 at B/N = 0.02, 0.05, 0.1, 0.2). Five open-weight models show near-constant verbalized confidence with point estimates at or beyond that flip on matched-marginal backgrounds; correlation is fleet-wide shared difficulty, not lineage; and a vacuity criterion flags when no deployable policy beats no oversight by more than τ.

What carries the argument

Budgeted noisy inspection over a two-level Gaussian copula (fleet-wide factor ρ0 plus within-family factor ρf) coupled to an adversarial-miscalibration mixture on error confidence; the object that carries the argument is the flip threshold δ*(B/N) at which residual risk of conf-ranked exceeds random.

Load-bearing premise

The locked single-completion verbalized-confidence protocol and the synthetic Beta-mixture model of miscalibration produce an allocation-faithful score that can be inverted onto the phase diagram for real fleets.

What would settle it

Re-elicit confidence from the same open-weight models under chain-of-thought or sampling-variance protocols, recompute tail-precision δ_prec at matched error rates, and check whether the point estimates still sit at or beyond the synthetic δ* curve; or enlarge the trace replay beyond N=15 and test whether conf-ranked remains indistinguishable from random.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Under tight audit budgets, ranking by self-reported confidence is safer longer than under generous budgets, because only the extreme low-confidence tail is touched.
  • When open-weight fleets report near-constant confidence, confidence-ranked auditing collapses to random by construction and transfer from correlated audits is the remaining useful signal.
  • Oversight mandates can be checked quantitatively: if no deployable policy beats no-oversight residual risk by more than τ, the setup is rubber-stamping.
  • Fleet designers need not map lineage for transfer; pooled, difficulty-driven correlation captures the value.
  • Clustering of errors under high correlation pushes undetected mass beyond any fixed budget, so correlation is both threat and resource.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Regulators writing ‘effective human oversight’ clauses could adopt the vacuity threshold τ as an auditable acceptance test rather than a qualitative slogan.
  • If future elicitation methods restore confidence variance in small open-weight models, the same phase diagram predicts they would drop back below δ* without changing the fleet or budget.
  • The reversed budget dependence suggests pairing the scarcest human time with the most aggressively filtered low-confidence tail, not spreading audits evenly.
  • Strategic (not merely miscalibrated) agents would likely push mass still further into the high-confidence region, moving the operational flip leftward and making vacuity more common.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper formalizes budgeted noisy inspection of an N-agent LLM fleet by a single human with B ≪ N audits per round, under adversarially miscalibrated self-reported confidence (strength δ) and two-level Gaussian-copula error correlation. Synthetic sweeps (E1) locate a flip threshold δ*(B/N) past which confidence-ranked auditing is dominated by random, and show that δ* rises as B/N shrinks; they also quantify when correlation-aware posterior transfer (diversity_bayes) helps and when oversight is vacuous at level τ. Measurements on six models (E2) find five open-weight LLMs with near-constant verbalized confidence (AUROC ≈ 0.5) and point-estimate δ_prec at or beyond the flip (CIs straddle), while gpt-4o-mini is informative and below it; shared item difficulty, not lineage, dominates measured correlation. Trace replay (E3) on a 15-agent fleet preserves the ordering oracle < diversity-Bayes < conf-ranked ≈ random with Holm-significant paired contrasts.

Significance. The problem is timely and practically important: fleet-scale LLM deployments under human-oversight mandates (e.g., EU AI Act Art. 14) with B ≪ N. Jointly treating self-reported priority signals that can be adversarially miscalibrated, partially predictable correlation from fleet composition, and a quantitative vacuity criterion is a genuine contribution relative to selective prediction, L2D, AI control, and classical acceptance sampling. The reversed budget dependence of δ* is a non-obvious, falsifiable prediction supported by systematic sweeps; the difficulty-vs-lineage decomposition and the locked multi-model measurements (with ECE/AUROC/δ_prec and partial correlations) give the phase diagram operational content. Pre-specified Holm-adjusted replay contrasts and robustness checks (Beta shock, transfer weights, persistence collapse) are strengths. If the reverse-δ* direction and the vacuity reading hold under broader elicitation and fleets, the work supplies a concrete allocation-level criterion for when confidence-ranked oversight fails and when mandated oversight is rubber-stamping.

major comments (3)
  1. [§5, Fig. S1, Abstract] §5 and Fig. S1 (matched-marginals placement): For the five open-weight models the operational finding is near-constant confidence (Var(c) ≤ 0.016, AUROC ≈ 0.5), so conf-ranked ≡ random by construction. The further claim that point estimates of δ_prec lie at or beyond δ* rests on inversion through a δ→precision map that the paper itself notes is flat at high ê, producing long left CI tails that straddle δ*. Abstract, §5, and §7 should lead with the near-constant / AUROC≈0.5 result and treat the flip placement as secondary and protocol-bound, rather than as co-equal evidence that real models sit past the flip.
  2. [§5 E3, §6(f)] §5 E3 replay: N=15 with B ∈ {1,2,3} and fleet ê ≈ 0.83 on GSM8K yields only modest effect sizes (0.133 / 0.095 risk units) under an acknowledged floor effect. The ordering and Holm-adjusted significance are useful, but the manuscript calls E3 “preliminary” (§6f) while still using it to “confirm the ordering” in the abstract and conclusion. Either enlarge the replay (more agents/items or matched-marginal synthetic agents at the measured ρ0) or explicitly bound the confirmatory claim to qualitative ordering under extreme ê, not to magnitude.
  3. [§4 H1, §3 Confidence] §4 H1 (reverse δ*(B/N)): The located values δ* = 0.93, 0.89, 0.84, 0.75 at B/N = 0.02, 0.05, 0.1, 0.2 and the scarcity-protects-the-tail explanation are central. The support is purely numerical (200 reps/cell). A short analytic or semi-analytic argument—e.g., that overconfident-error mass piles at the high-confidence end so the extreme lower tail’s error precision degrades last—would make the reversed direction less dependent on the particular Beta mixture and grid. Without it, sensitivity to the error-confidence mixture shapes (currently fixed at Beta(3,3)/Beta(10,1.5)) remains a load-bearing untested choice.
minor comments (6)
  1. [§4–5, Appendix] Main text repeatedly cites “Fig. S1” / “Fig. S2” for load-bearing results (flip heatmap, vacuity, money diagram). Either promote the key diagrams into the main body with ordinary figure numbers or add a one-sentence pointer that all S-figures live in the Appendix so readers are not sent hunting.
  2. [§3] §3: State explicitly that marginals e_i ∼ U(0.05, 0.30) are fixed per agent across rounds (only the copula shocks evolve). The AR(1) persistence φ defaulting to 1 is easy to miss and is crucial for H2.
  3. [§5, Table S1] Table S1 / §5: Report the exact inversion procedure (and any smoothing) that maps tail error precision to δ_prec, including how ties and the matched-ê background are handled; a one-line formula or algorithm box would suffice.
  4. [§2] §2: The distinction from multi-expert L2D and from correlated bandits is clear; a single sentence on how skip-lot sampling (ISO 2859-3) differs under adversarial miscalibration would round out the acceptance-sampling lineage.
  5. [§3 Audit and risk] Minor notation: residual risk R(π) is defined with caught(A_t) but the noisy-audit flip probability ε is introduced only in prose; give the caught indicator an explicit equation so ε appears in the risk definition.
  6. [Throughout] Typos / polish: “a-priori” hyphenation is inconsistent; “bill of materials” is evocative but undefined on first use; arXiv date in the header reads 30 Jul 2026 (likely a placeholder).

Circularity Check

1 steps flagged

No significant circularity: synthetic phase diagram, empirical placement, and trace replay are independent; only mild definitional framing of vacuity.

specific steps
  1. self definitional [§3 Model, Audit and risk; H3 §4]
    "Oversight is vacuous at level τ if max_π(R_none − R_π)/R_none < τ over the deployable policies considered."

    Vacuity is defined as failure of the studied policy class to beat no-oversight by more than τ. This is criterion-setting rather than a derived empirical law; the quantitative reading of ‘rubber-stamping’ is true of that class by the definition chosen. It is mild and disclosed, and does not force H1’s δ* location or the reverse budget direction.

full rationale

The load-bearing results are produced by three separable procedures that do not reduce to their inputs by construction. (1) H1’s flip threshold δ*(B/N) and its reversed budget dependence are located by Monte Carlo sweeps on an explicitly specified generative model (two-level Gaussian copula + Beta-mixture confidence); the threshold is an output of those sweeps, not a fitted or definitional identity. The paper itself notes that a flip at very high δ is near-definitional and correctly scopes its contribution as locating δ*(B/N) and the reverse direction. (2) Placement of real models uses δ_prec (error precision of the audited low-confidence tail) inverted through the synthetic δ→precision map at matched ê—an explicit modeling bridge, not a prediction forced by a fit to the same quantity. Open-weight near-constant confidence makes conf-ranked≈random by construction under that protocol; the paper states this and reports CIs that straddle δ*, so it does not smuggle a forced ‘beyond flip’ claim. (3) E3 replay applies fixed policies to held-out recorded traces and confirms ordering with pre-specified Holm contrasts—external to the synthetic diagram. Vacuity is defined relative to the considered policy class (max reduction <τ), which is ordinary criterion-setting, not Eq.X=Eq.Y disguised as a derivation. No load-bearing self-citation chain, uniqueness import, or ansatz-via-citation appears. Score 1 only for the mild, acknowledged definitional character of the vacuity criterion and the protocol-bound conf≡random regime; central claims remain non-circular.

Axiom & Free-Parameter Ledger

8 free parameters · 7 axioms · 4 invented entities

The central flip and vacuity claims rest on a generative inspection model (Beta confidence mixtures, two-level Gaussian copula, AR(1) persistence, noisy audits) plus a protocol-bound measurement bridge (locked verbalized confidence → δ_prec inversion). Free parameters set the shape of confidence under error and the correlation structure; domain assumptions rule out strategic agents and treat fleet composition as known. Invented constructs are the flip threshold as an operating object, the τ-vacuity criterion, and the diversity_bayes allocator—tools for the analysis rather than physical entities.

free parameters (8)
  • δ (miscalibration strength) = swept 0..1; δ_prec inverted per model
    Fraction of errors whose confidence is drawn from the high-confidence Beta(10,1.5) instead of Beta(3,3); primary sweep axis that defines the flip.
  • Beta confidence shape parameters = (6,2), (3,3), (10,1.5)
    Correct ~ Beta(6,2); error mixture Beta(3,3) and Beta(10,1.5) chosen to encode overconfident errors; not derived from data.
  • ρ0, ρf (copula correlations) = ρ0 ≈ 0.52 (GSM8K) / 0.70 (HotpotQA); ρf ≈ ρ0
    Fleet-wide and within-family latent correlations in the two-level Gaussian copula; measured then used in synthetic and transfer policies.
  • φ (AR(1) persistence) = default 1
    Controls whether audits transfer across rounds; default 1 makes co-failure a stable difficulty profile; advantage collapses at φ=0.
  • τ (vacuity level) = 0.1
    Threshold on relative risk reduction below which oversight is called vacuous; set to 0.1 for H3 heatmaps.
  • ε (audit noise) = swept 0..0.2; replay 0.05
    Probability an audit flips the revealed error bit; sweeps and crossover with transfer advantage.
  • transfer weight w = w ∈ {ρ, ρ², ρ/2}
    Sibling update weight in diversity_bayes; tested at ρ, ρ², ρ/2 for robustness.
  • e_i error marginals = synthetic U(0.05,0.30); empirical ê ∈ [0.40,0.94]
    Per-agent error rates; synthetic U(0.05,0.30), empirical from benchmarks for placement and replay.
axioms (7)
  • domain assumption Agents are honest-but-miscalibrated; no strategic subversion of confidence or outputs.
    Stated in §1; strategic behavior deferred to AI Control. Load-bearing for interpreting confidence as a noisy rather than adversarial signal beyond δ.
  • ad hoc to paper Error dependence is captured by a two-level Gaussian copula with fleet factor G0 and family factors Gf, mapped through Φ to Bernoulli marginals e_i.
    §3 model definition; robustness check with Beta log-odds shock in appendix, but primary results use the Gaussian copula.
  • domain assumption Latent shocks are drawn once and evolve as stationary AR(1) with persistence φ (default 1), so same-family co-failure is a stable difficulty profile.
    §3; H2 transfer advantage collapses at φ=0 (§4, appendix). Required for audit-to-sibling value.
  • domain assumption An audit reveals E_i flipped with probability ε; residual risk is expected undetected errors after the audited set.
    §3 audit and risk definitions; standard noisy inspection framing.
  • domain assumption Fleet partition into F base-model families is known to the allocator (bill of materials given, not estimated).
    §2–3; softened later when lineage excess ∆ρ ≈ 0 and pooled transfer matches family/tier.
  • standard math Standard properties of Gaussian copulas, Beta distributions, and Monte Carlo estimation of residual risk over finite replications.
    Used throughout §3–4 synthetic phase diagrams.
  • domain assumption Exact-match errors on GSM8K/HotpotQA under temperature-0 locked single-completion prompting represent operational agent error and confidence for allocation.
    E2 protocol §5; limitations §6(d) note CoT and alternate elicitation may change variance and ê.
invented entities (4)
  • δ*(B/N) flip threshold independent evidence
    purpose: Miscalibration level past which conf_ranked residual risk exceeds random; main operating object of H1.
    Defined operationally from synthetic sweeps of R_random − R_conf; not a prior literature constant.
  • Vacuous oversight criterion at level τ independent evidence
    purpose: Declare oversight rubber-stamping when no deployable policy beats no-oversight by relative reduction ≥ τ.
    H3 §3–4; quantitative reading of socio-legal rubber-stamping critiques.
  • diversity_bayes policy no independent evidence
    purpose: Batched-greedy posterior UCB with same-family (or pooled) Beta updates weighted by w to exploit correlation.
    §3 policies; knowledge-gradient-style instantiation, not claimed optimal.
  • δ_prec (precision-based miscalibration) independent evidence
    purpose: Map real models onto the synthetic δ axis via error precision of the lowest-confidence B/N tail.
    §5; allocation-faithful substitute for AUROC when confidence is near-constant.

pith-pipeline@v1.2.0-daily-grok45 · 15666 in / 4774 out tokens · 91196 ms · 2026-07-31T11:27:12.308833+00:00 · methodology

0 comments
read the original abstract

A single human must audit $N$ LLM agents under a budget of $B \ll N$ audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the miscalibration threshold $\delta^*$ past which confidence-ranked auditing is \emph{worse} than random. Two a-priori expectations reverse: $\delta^*$ \emph{rises} as the budget shrinks, and cross-family correlation is not low---shared difficulty dominates lineage. Five open-weight LLMs show operationally useless (near-constant) confidence, point estimates at or beyond the flip though CIs straddle it; a proprietary model is informative and lands below it. We give a quantitative criterion for \emph{vacuous} oversight, and replaying policies on recorded traces confirms the ordering.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

26 extracted references · 5 canonical work pages · 4 internal anchors

  1. [1]

    Transactions on Machine Learning Research (2024)

    Alves, J.V., Leitão, D., Jesus, S., Sampaio, M.O.P., Liébana, J., Saleiro, P., Figueiredo, M.A.T., Bizarro, P.: Cost-sensitive learning to defer to multiple experts with workload constraints. Transactions on Machine Learning Research (2024). https://doi.org/10.48550/arXiv.2403.06906

  2. [2]

    Governing What the EU AI Act Excludes: Accountability for Autonomous AI Agents in Smart City Critical Infrastructure

    Butt, T.A., Iqbal, M., Iqbal, R.: Governing what the EU AI Act excludes: Ac- countability for autonomous AI agents in smart city critical infrastructure. arXiv preprint arXiv:2605.01091 (2026). https://doi.org/10.48550/arXiv.2605.01091

  3. [3]

    In: Dasgupta, S., McAllester, D

    Chen, X., Lin, Q., Zhou, D.: Optimistic knowledge gradient policy for optimal budget allocation in crowdsourcing. In: Dasgupta, S., McAllester, D. (eds.) Pro- ceedings of the 30th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 28, pp. 64–72. PMLR, Atlanta, Georgia, USA (17–19 Jun 2013), https://proceedings.mlr...

  4. [4]

    Wiley, 2nd edn

    Dodge, H.F., Romig, H.G.: Sampling Inspection Tables: Single and Double Sam- pling. Wiley, 2nd edn. (1959)

  5. [5]

    Lawrence Erlbaum Associates (2000)

    Embretson, S.E., Reise, S.P.: Item Response Theory for Psychologists. Lawrence Erlbaum Associates (2000)

  6. [6]

    European Parliament and Council: Regulation (EU) 2024/1689 (AI Act), article 14: Human oversight (2024)

  7. [7]

    Computer Law & Security Review45, 105681 (2022)

    Green, B.: The flaws of policies requiring human oversight of govern- ment algorithms. Computer Law & Security Review45, 105681 (2022). https://doi.org/10.1016/j.clsr.2022.105681

  8. [8]

    In: Proc

    Greenblatt, R., Shlegeris, B., Sachan, K., Roger, F.: AI control: Im- proving safety despite intentional subversion. In: Proc. 41st Interna- tional Conference on Machine Learning (ICML). PMLR, vol. 235 (2024). https://doi.org/10.48550/arXiv.2312.06942

  9. [9]

    IEEE Transactions on Information Theory67(10), 6711–6732 (2021)

    Gupta, S., Chaudhari, S., Joshi, G., Yağan, O.: Multi-armed bandits with corre- lated arms. IEEE Transactions on Information Theory67(10), 6711–6732 (2021). https://doi.org/10.1109/TIT.2021.3081508

  10. [10]

    Correlated Multi-armed Bandits with a Latent Random Source

    Gupta, S., Joshi, G., Yağan, O.: Correlated multi-armed bandits with a latent random source. In: Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 3572–3576 (2020). https://doi.org/10.48550/arXiv.1808.05904

  11. [11]

    In: Proc

    Hacohen, G., Dekel, A., Weinshall, D.: Active learning on a budget: Oppo- site strategies suit high and low budgets. In: Proc. 39th International Con- ference on Machine Learning (ICML). PMLR, vol. 162, pp. 8175–8195 (2022). https://doi.org/10.48550/arXiv.2202.02794

  12. [12]

    International Organization for Standardization: ISO 2859-3: Sampling procedures for inspection by attributes — part 3: Skip-lot sampling procedures (2005)

  13. [13]

    arXiv preprint arXiv:2207.05221 (2022)

    Kadavath, S., Conerly, T., Askell, A., et al.: Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221 (2022). https://doi.org/10.48550/arXiv.2207.05221

  14. [14]

    Batch Multi-Fidelity Active Learning with Budget Constraints

    Li, S., Phillips, J.M., Yu, X., Kirby, R.M., Zhe, S.: Batch multi-fidelity active learning with budget constraints. In: Advances in Neural Information Processing Systems 35 (NeurIPS) (2022). https://doi.org/10.48550/arXiv.2210.12704

  15. [15]

    https://doi.org/10.48550/arXiv.2205.14334, https://arxiv.org/abs/2205.14334 8 C

    Lin, S., Hilton, J., Evans, O.: Teaching models to express their un- certainty in words (2022). https://doi.org/10.48550/arXiv.2205.14334, https://arxiv.org/abs/2205.14334 8 C. Zavattari et al

  16. [16]

    In: Proceedings of the 37th International Conference on Neural In- formation Processing Systems

    Mao, A., Mohri, C., Mohri, M., Zhong, Y.: Two-stage learning to defer with mul- tiple experts. In: Proceedings of the 37th International Conference on Neural In- formation Processing Systems. NIPS ’23, Curran Associates Inc., Red Hook, NY, USA (2023). https://doi.org/10.52202/075280-0159

  17. [17]

    arXiv preprint arXiv:2310.14774 (2023)

    Mao, A., Mohri, M., Zhong, Y.: Principled approaches for learning to defer with multiple experts. arXiv preprint arXiv:2310.14774 (2023). https://doi.org/10.48550/arXiv.2310.14774

  18. [18]

    In: Proc

    Mozannar, H., Sontag, D.: Consistent estimators for learning to defer to an expert. In: Proc. 37th International Conference on Machine Learning (ICML). PMLR, vol. 119, pp. 7076–7087 (2020). https://doi.org/10.48550/arXiv.2006.01862

  19. [19]

    A Causal Framework for Evaluating Deferring Systems

    Palomba, F., Pugnana, A., Alvarez, J.M., Ruggieri, S.: A causal frame- work for evaluating deferring systems. In: Proc. 28th International Conference on Artificial Intelligence and Statistics (AISTATS). PMLR, vol. 258 (2025). https://doi.org/10.48550/arXiv.2405.18902

  20. [20]

    In: Proc

    Pugnana, A., Ruggieri, S.: A model-agnostic heuristics for selective classification. In: Proc. 37th AAAI Conference on Artificial Intelligence. vol. 37, pp. 9461–9469 (2023). https://doi.org/10.1609/aaai.v37i8.26133

  21. [21]

    In: Proc

    Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., Man- ning, C.D.: Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In: Proc. Conference on Empirical Methods in Natural Language Processing (EMNLP). pp. 5433–5442 (2023). https://doi.org/10.48550...

  22. [22]

    In: Proc

    Verma, R., Barrejón, D., Nalisnick, E.: Learning to defer to multiple experts: Consistent surrogate losses, confidence calibration, and conformal ensembles. In: Proc. 26th International Conference on Artificial Intelli- gence and Statistics (AISTATS). PMLR, vol. 206, pp. 11415–11434 (2023). https://doi.org/10.48550/arXiv.2210.16955

  23. [23]

    Econometrica47(3), 641–654 (1979)

    Weitzman, M.L.: Optimal search for the best alternative. Econometrica47(3), 641–654 (1979). https://doi.org/10.2307/1910412

  24. [24]

    In: Proc

    Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., Hooi, B.: Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs. In: Proc. 12th International Conference on Learning Representations (ICLR) (2024). https://doi.org/10.48550/arXiv.2306.13063

  25. [25]

    arXiv preprint arXiv:2601.07264 (2026)

    Xuan, W., Zeng, Q., Qi, H., Xiao, Y., Wang, J., Yokoya, N.: The confidence dichotomy: Analyzing and mitigating miscalibration in tool-use agents. arXiv preprint arXiv:2601.07264 (2026). https://doi.org/10.48550/arXiv.2601.07264

  26. [26]

    arXiv preprint arXiv:2601.15778 (2026)

    Zhang, J., Xiong, C., Wu, C.S.: Agentic confidence calibration. arXiv preprint arXiv:2601.15778 (2026). https://doi.org/10.48550/arXiv.2601.15778 One Human,NAgents 9 Appendix This appendix collects illustrative material for the main paper. Every load- bearing number is reported in the main text; the figures and tables here only illustrate. Figures and tab...