REVIEW 3 major objections 6 minor 26 references
Past a miscalibration threshold that rises as budgets shrink, ranking LLM agent audits by self-reported confidence is worse than picking at random.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 11:27 UTC pith:AXVIX2QM
load-bearing objection Clean budgeted-inspection model with a real reverse flip threshold and honest measurements; open-model “beyond flip” placement is softer than the synthetic result. the 3 major comments →
One Human, N Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Confidence-ranked auditing of an LLM agent fleet flips from better than random to worse than random past a miscalibration threshold δ*(B/N) that increases as the audit budget shrinks (located near 0.93, 0.89, 0.84, 0.75 at B/N = 0.02, 0.05, 0.1, 0.2). Five open-weight models show near-constant verbalized confidence with point estimates at or beyond that flip on matched-marginal backgrounds; correlation is fleet-wide shared difficulty, not lineage; and a vacuity criterion flags when no deployable policy beats no oversight by more than τ.
What carries the argument
Budgeted noisy inspection over a two-level Gaussian copula (fleet-wide factor ρ0 plus within-family factor ρf) coupled to an adversarial-miscalibration mixture on error confidence; the object that carries the argument is the flip threshold δ*(B/N) at which residual risk of conf-ranked exceeds random.
Load-bearing premise
The locked single-completion verbalized-confidence protocol and the synthetic Beta-mixture model of miscalibration produce an allocation-faithful score that can be inverted onto the phase diagram for real fleets.
What would settle it
Re-elicit confidence from the same open-weight models under chain-of-thought or sampling-variance protocols, recompute tail-precision δ_prec at matched error rates, and check whether the point estimates still sit at or beyond the synthetic δ* curve; or enlarge the trace replay beyond N=15 and test whether conf-ranked remains indistinguishable from random.
If this is right
- Under tight audit budgets, ranking by self-reported confidence is safer longer than under generous budgets, because only the extreme low-confidence tail is touched.
- When open-weight fleets report near-constant confidence, confidence-ranked auditing collapses to random by construction and transfer from correlated audits is the remaining useful signal.
- Oversight mandates can be checked quantitatively: if no deployable policy beats no-oversight residual risk by more than τ, the setup is rubber-stamping.
- Fleet designers need not map lineage for transfer; pooled, difficulty-driven correlation captures the value.
- Clustering of errors under high correlation pushes undetected mass beyond any fixed budget, so correlation is both threat and resource.
Where Pith is reading between the lines
- Regulators writing ‘effective human oversight’ clauses could adopt the vacuity threshold τ as an auditable acceptance test rather than a qualitative slogan.
- If future elicitation methods restore confidence variance in small open-weight models, the same phase diagram predicts they would drop back below δ* without changing the fleet or budget.
- The reversed budget dependence suggests pairing the scarcest human time with the most aggressively filtered low-confidence tail, not spreading audits evenly.
- Strategic (not merely miscalibrated) agents would likely push mass still further into the high-confidence region, moving the operational flip leftward and making vacuity more common.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formalizes budgeted noisy inspection of an N-agent LLM fleet by a single human with B ≪ N audits per round, under adversarially miscalibrated self-reported confidence (strength δ) and two-level Gaussian-copula error correlation. Synthetic sweeps (E1) locate a flip threshold δ*(B/N) past which confidence-ranked auditing is dominated by random, and show that δ* rises as B/N shrinks; they also quantify when correlation-aware posterior transfer (diversity_bayes) helps and when oversight is vacuous at level τ. Measurements on six models (E2) find five open-weight LLMs with near-constant verbalized confidence (AUROC ≈ 0.5) and point-estimate δ_prec at or beyond the flip (CIs straddle), while gpt-4o-mini is informative and below it; shared item difficulty, not lineage, dominates measured correlation. Trace replay (E3) on a 15-agent fleet preserves the ordering oracle < diversity-Bayes < conf-ranked ≈ random with Holm-significant paired contrasts.
Significance. The problem is timely and practically important: fleet-scale LLM deployments under human-oversight mandates (e.g., EU AI Act Art. 14) with B ≪ N. Jointly treating self-reported priority signals that can be adversarially miscalibrated, partially predictable correlation from fleet composition, and a quantitative vacuity criterion is a genuine contribution relative to selective prediction, L2D, AI control, and classical acceptance sampling. The reversed budget dependence of δ* is a non-obvious, falsifiable prediction supported by systematic sweeps; the difficulty-vs-lineage decomposition and the locked multi-model measurements (with ECE/AUROC/δ_prec and partial correlations) give the phase diagram operational content. Pre-specified Holm-adjusted replay contrasts and robustness checks (Beta shock, transfer weights, persistence collapse) are strengths. If the reverse-δ* direction and the vacuity reading hold under broader elicitation and fleets, the work supplies a concrete allocation-level criterion for when confidence-ranked oversight fails and when mandated oversight is rubber-stamping.
major comments (3)
- [§5, Fig. S1, Abstract] §5 and Fig. S1 (matched-marginals placement): For the five open-weight models the operational finding is near-constant confidence (Var(c) ≤ 0.016, AUROC ≈ 0.5), so conf-ranked ≡ random by construction. The further claim that point estimates of δ_prec lie at or beyond δ* rests on inversion through a δ→precision map that the paper itself notes is flat at high ê, producing long left CI tails that straddle δ*. Abstract, §5, and §7 should lead with the near-constant / AUROC≈0.5 result and treat the flip placement as secondary and protocol-bound, rather than as co-equal evidence that real models sit past the flip.
- [§5 E3, §6(f)] §5 E3 replay: N=15 with B ∈ {1,2,3} and fleet ê ≈ 0.83 on GSM8K yields only modest effect sizes (0.133 / 0.095 risk units) under an acknowledged floor effect. The ordering and Holm-adjusted significance are useful, but the manuscript calls E3 “preliminary” (§6f) while still using it to “confirm the ordering” in the abstract and conclusion. Either enlarge the replay (more agents/items or matched-marginal synthetic agents at the measured ρ0) or explicitly bound the confirmatory claim to qualitative ordering under extreme ê, not to magnitude.
- [§4 H1, §3 Confidence] §4 H1 (reverse δ*(B/N)): The located values δ* = 0.93, 0.89, 0.84, 0.75 at B/N = 0.02, 0.05, 0.1, 0.2 and the scarcity-protects-the-tail explanation are central. The support is purely numerical (200 reps/cell). A short analytic or semi-analytic argument—e.g., that overconfident-error mass piles at the high-confidence end so the extreme lower tail’s error precision degrades last—would make the reversed direction less dependent on the particular Beta mixture and grid. Without it, sensitivity to the error-confidence mixture shapes (currently fixed at Beta(3,3)/Beta(10,1.5)) remains a load-bearing untested choice.
minor comments (6)
- [§4–5, Appendix] Main text repeatedly cites “Fig. S1” / “Fig. S2” for load-bearing results (flip heatmap, vacuity, money diagram). Either promote the key diagrams into the main body with ordinary figure numbers or add a one-sentence pointer that all S-figures live in the Appendix so readers are not sent hunting.
- [§3] §3: State explicitly that marginals e_i ∼ U(0.05, 0.30) are fixed per agent across rounds (only the copula shocks evolve). The AR(1) persistence φ defaulting to 1 is easy to miss and is crucial for H2.
- [§5, Table S1] Table S1 / §5: Report the exact inversion procedure (and any smoothing) that maps tail error precision to δ_prec, including how ties and the matched-ê background are handled; a one-line formula or algorithm box would suffice.
- [§2] §2: The distinction from multi-expert L2D and from correlated bandits is clear; a single sentence on how skip-lot sampling (ISO 2859-3) differs under adversarial miscalibration would round out the acceptance-sampling lineage.
- [§3 Audit and risk] Minor notation: residual risk R(π) is defined with caught(A_t) but the noisy-audit flip probability ε is introduced only in prose; give the caught indicator an explicit equation so ε appears in the risk definition.
- [Throughout] Typos / polish: “a-priori” hyphenation is inconsistent; “bill of materials” is evocative but undefined on first use; arXiv date in the header reads 30 Jul 2026 (likely a placeholder).
Circularity Check
No significant circularity: synthetic phase diagram, empirical placement, and trace replay are independent; only mild definitional framing of vacuity.
specific steps
-
self definitional
[§3 Model, Audit and risk; H3 §4]
"Oversight is vacuous at level τ if max_π(R_none − R_π)/R_none < τ over the deployable policies considered."
Vacuity is defined as failure of the studied policy class to beat no-oversight by more than τ. This is criterion-setting rather than a derived empirical law; the quantitative reading of ‘rubber-stamping’ is true of that class by the definition chosen. It is mild and disclosed, and does not force H1’s δ* location or the reverse budget direction.
full rationale
The load-bearing results are produced by three separable procedures that do not reduce to their inputs by construction. (1) H1’s flip threshold δ*(B/N) and its reversed budget dependence are located by Monte Carlo sweeps on an explicitly specified generative model (two-level Gaussian copula + Beta-mixture confidence); the threshold is an output of those sweeps, not a fitted or definitional identity. The paper itself notes that a flip at very high δ is near-definitional and correctly scopes its contribution as locating δ*(B/N) and the reverse direction. (2) Placement of real models uses δ_prec (error precision of the audited low-confidence tail) inverted through the synthetic δ→precision map at matched ê—an explicit modeling bridge, not a prediction forced by a fit to the same quantity. Open-weight near-constant confidence makes conf-ranked≈random by construction under that protocol; the paper states this and reports CIs that straddle δ*, so it does not smuggle a forced ‘beyond flip’ claim. (3) E3 replay applies fixed policies to held-out recorded traces and confirms ordering with pre-specified Holm contrasts—external to the synthetic diagram. Vacuity is defined relative to the considered policy class (max reduction <τ), which is ordinary criterion-setting, not Eq.X=Eq.Y disguised as a derivation. No load-bearing self-citation chain, uniqueness import, or ansatz-via-citation appears. Score 1 only for the mild, acknowledged definitional character of the vacuity criterion and the protocol-bound conf≡random regime; central claims remain non-circular.
Axiom & Free-Parameter Ledger
free parameters (8)
- δ (miscalibration strength) =
swept 0..1; δ_prec inverted per model
- Beta confidence shape parameters =
(6,2), (3,3), (10,1.5)
- ρ0, ρf (copula correlations) =
ρ0 ≈ 0.52 (GSM8K) / 0.70 (HotpotQA); ρf ≈ ρ0
- φ (AR(1) persistence) =
default 1
- τ (vacuity level) =
0.1
- ε (audit noise) =
swept 0..0.2; replay 0.05
- transfer weight w =
w ∈ {ρ, ρ², ρ/2}
- e_i error marginals =
synthetic U(0.05,0.30); empirical ê ∈ [0.40,0.94]
axioms (7)
- domain assumption Agents are honest-but-miscalibrated; no strategic subversion of confidence or outputs.
- ad hoc to paper Error dependence is captured by a two-level Gaussian copula with fleet factor G0 and family factors Gf, mapped through Φ to Bernoulli marginals e_i.
- domain assumption Latent shocks are drawn once and evolve as stationary AR(1) with persistence φ (default 1), so same-family co-failure is a stable difficulty profile.
- domain assumption An audit reveals E_i flipped with probability ε; residual risk is expected undetected errors after the audited set.
- domain assumption Fleet partition into F base-model families is known to the allocator (bill of materials given, not estimated).
- standard math Standard properties of Gaussian copulas, Beta distributions, and Monte Carlo estimation of residual risk over finite replications.
- domain assumption Exact-match errors on GSM8K/HotpotQA under temperature-0 locked single-completion prompting represent operational agent error and confidence for allocation.
invented entities (4)
-
δ*(B/N) flip threshold
independent evidence
-
Vacuous oversight criterion at level τ
independent evidence
-
diversity_bayes policy
no independent evidence
-
δ_prec (precision-based miscalibration)
independent evidence
read the original abstract
A single human must audit $N$ LLM agents under a budget of $B \ll N$ audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the miscalibration threshold $\delta^*$ past which confidence-ranked auditing is \emph{worse} than random. Two a-priori expectations reverse: $\delta^*$ \emph{rises} as the budget shrinks, and cross-family correlation is not low---shared difficulty dominates lineage. Five open-weight LLMs show operationally useless (near-constant) confidence, point estimates at or beyond the flip though CIs straddle it; a proprietary model is informative and lands below it. We give a quantitative criterion for \emph{vacuous} oversight, and replaying policies on recorded traces confirms the ordering.
Reference graph
Works this paper leans on
-
[1]
Transactions on Machine Learning Research (2024)
Alves, J.V., Leitão, D., Jesus, S., Sampaio, M.O.P., Liébana, J., Saleiro, P., Figueiredo, M.A.T., Bizarro, P.: Cost-sensitive learning to defer to multiple experts with workload constraints. Transactions on Machine Learning Research (2024). https://doi.org/10.48550/arXiv.2403.06906
-
[2]
Butt, T.A., Iqbal, M., Iqbal, R.: Governing what the EU AI Act excludes: Ac- countability for autonomous AI agents in smart city critical infrastructure. arXiv preprint arXiv:2605.01091 (2026). https://doi.org/10.48550/arXiv.2605.01091
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2605.01091 2026
-
[3]
In: Dasgupta, S., McAllester, D
Chen, X., Lin, Q., Zhou, D.: Optimistic knowledge gradient policy for optimal budget allocation in crowdsourcing. In: Dasgupta, S., McAllester, D. (eds.) Pro- ceedings of the 30th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 28, pp. 64–72. PMLR, Atlanta, Georgia, USA (17–19 Jun 2013), https://proceedings.mlr...
2013
-
[4]
Wiley, 2nd edn
Dodge, H.F., Romig, H.G.: Sampling Inspection Tables: Single and Double Sam- pling. Wiley, 2nd edn. (1959)
1959
-
[5]
Lawrence Erlbaum Associates (2000)
Embretson, S.E., Reise, S.P.: Item Response Theory for Psychologists. Lawrence Erlbaum Associates (2000)
2000
-
[6]
European Parliament and Council: Regulation (EU) 2024/1689 (AI Act), article 14: Human oversight (2024)
2024
-
[7]
Computer Law & Security Review45, 105681 (2022)
Green, B.: The flaws of policies requiring human oversight of govern- ment algorithms. Computer Law & Security Review45, 105681 (2022). https://doi.org/10.1016/j.clsr.2022.105681
arXiv 2022
-
[8]
Greenblatt, R., Shlegeris, B., Sachan, K., Roger, F.: AI control: Im- proving safety despite intentional subversion. In: Proc. 41st Interna- tional Conference on Machine Learning (ICML). PMLR, vol. 235 (2024). https://doi.org/10.48550/arXiv.2312.06942
-
[9]
IEEE Transactions on Information Theory67(10), 6711–6732 (2021)
Gupta, S., Chaudhari, S., Joshi, G., Yağan, O.: Multi-armed bandits with corre- lated arms. IEEE Transactions on Information Theory67(10), 6711–6732 (2021). https://doi.org/10.1109/TIT.2021.3081508
arXiv 2021
-
[10]
Correlated Multi-armed Bandits with a Latent Random Source
Gupta, S., Joshi, G., Yağan, O.: Correlated multi-armed bandits with a latent random source. In: Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 3572–3576 (2020). https://doi.org/10.48550/arXiv.1808.05904
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.1808.05904 2020
-
[11]
Hacohen, G., Dekel, A., Weinshall, D.: Active learning on a budget: Oppo- site strategies suit high and low budgets. In: Proc. 39th International Con- ference on Machine Learning (ICML). PMLR, vol. 162, pp. 8175–8195 (2022). https://doi.org/10.48550/arXiv.2202.02794
-
[12]
International Organization for Standardization: ISO 2859-3: Sampling procedures for inspection by attributes — part 3: Skip-lot sampling procedures (2005)
2005
-
[13]
arXiv preprint arXiv:2207.05221 (2022)
Kadavath, S., Conerly, T., Askell, A., et al.: Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221 (2022). https://doi.org/10.48550/arXiv.2207.05221
-
[14]
Batch Multi-Fidelity Active Learning with Budget Constraints
Li, S., Phillips, J.M., Yu, X., Kirby, R.M., Zhe, S.: Batch multi-fidelity active learning with budget constraints. In: Advances in Neural Information Processing Systems 35 (NeurIPS) (2022). https://doi.org/10.48550/arXiv.2210.12704
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2210.12704 2022
-
[15]
https://doi.org/10.48550/arXiv.2205.14334, https://arxiv.org/abs/2205.14334 8 C
Lin, S., Hilton, J., Evans, O.: Teaching models to express their un- certainty in words (2022). https://doi.org/10.48550/arXiv.2205.14334, https://arxiv.org/abs/2205.14334 8 C. Zavattari et al
-
[16]
In: Proceedings of the 37th International Conference on Neural In- formation Processing Systems
Mao, A., Mohri, C., Mohri, M., Zhong, Y.: Two-stage learning to defer with mul- tiple experts. In: Proceedings of the 37th International Conference on Neural In- formation Processing Systems. NIPS ’23, Curran Associates Inc., Red Hook, NY, USA (2023). https://doi.org/10.52202/075280-0159
-
[17]
arXiv preprint arXiv:2310.14774 (2023)
Mao, A., Mohri, M., Zhong, Y.: Principled approaches for learning to defer with multiple experts. arXiv preprint arXiv:2310.14774 (2023). https://doi.org/10.48550/arXiv.2310.14774
-
[18]
Mozannar, H., Sontag, D.: Consistent estimators for learning to defer to an expert. In: Proc. 37th International Conference on Machine Learning (ICML). PMLR, vol. 119, pp. 7076–7087 (2020). https://doi.org/10.48550/arXiv.2006.01862
-
[19]
A Causal Framework for Evaluating Deferring Systems
Palomba, F., Pugnana, A., Alvarez, J.M., Ruggieri, S.: A causal frame- work for evaluating deferring systems. In: Proc. 28th International Conference on Artificial Intelligence and Statistics (AISTATS). PMLR, vol. 258 (2025). https://doi.org/10.48550/arXiv.2405.18902
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2405.18902 2025
-
[20]
Pugnana, A., Ruggieri, S.: A model-agnostic heuristics for selective classification. In: Proc. 37th AAAI Conference on Artificial Intelligence. vol. 37, pp. 9461–9469 (2023). https://doi.org/10.1609/aaai.v37i8.26133
-
[21]
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., Man- ning, C.D.: Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In: Proc. Conference on Empirical Methods in Natural Language Processing (EMNLP). pp. 5433–5442 (2023). https://doi.org/10.48550...
-
[22]
Verma, R., Barrejón, D., Nalisnick, E.: Learning to defer to multiple experts: Consistent surrogate losses, confidence calibration, and conformal ensembles. In: Proc. 26th International Conference on Artificial Intelli- gence and Statistics (AISTATS). PMLR, vol. 206, pp. 11415–11434 (2023). https://doi.org/10.48550/arXiv.2210.16955
-
[23]
Econometrica47(3), 641–654 (1979)
Weitzman, M.L.: Optimal search for the best alternative. Econometrica47(3), 641–654 (1979). https://doi.org/10.2307/1910412
doi:10.2307/1910412 1979
-
[24]
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., Hooi, B.: Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs. In: Proc. 12th International Conference on Learning Representations (ICLR) (2024). https://doi.org/10.48550/arXiv.2306.13063
-
[25]
arXiv preprint arXiv:2601.07264 (2026)
Xuan, W., Zeng, Q., Qi, H., Xiao, Y., Wang, J., Yokoya, N.: The confidence dichotomy: Analyzing and mitigating miscalibration in tool-use agents. arXiv preprint arXiv:2601.07264 (2026). https://doi.org/10.48550/arXiv.2601.07264
-
[26]
arXiv preprint arXiv:2601.15778 (2026)
Zhang, J., Xiong, C., Wu, C.S.: Agentic confidence calibration. arXiv preprint arXiv:2601.15778 (2026). https://doi.org/10.48550/arXiv.2601.15778 One Human,NAgents 9 Appendix This appendix collects illustrative material for the main paper. Every load- bearing number is reported in the main text; the figures and tables here only illustrate. Figures and tab...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.