REVIEW 2 major objections 4 minor 32 references
Negative error correlation is what lets a human safely use an AI prediction under uncertainty about the AI’s quality.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Negative human-AI error correlation is required for robust complementarity under uncertainty about AI quality; current LLMs exhibit positive correlations on forecasting tasks.
T0 review reviewed 2026-07-10 challenge →
load-bearing objection Clean theory that negative error correlation is the structural condition for robust human-AI complementarity under uncertainty about AI quality; empirics put current LLMs in the hard regime. the 2 major comments →
Robust Human-AI Complementarity under Uncertainty
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Under an uncertainty set that leaves the AI–truth covariance free while fixing the rest of the joint law, a decision maker has a rule that uses the AI and strictly dominates the human-only optimum for every distribution in the set if and only if residual AI signal after conditioning on the human is uniformly positively related to the truth. Negative correlation of human and AI errors is the structural condition that delivers that uniform positivity; positive error correlation shrinks the safe region to cases where the human residual variance is large relative to the worst-case shared error.
What carries the argument
The residualized AI signal r = ϕ_AI − Cov(ϕ_H, ϕ_AI)/Var(ϕ_H) · ϕ_H (or its nonlinear analogue ˜r = ϕ_AI − E[ϕ_AI|ϕ_H]), together with a uniform lower bound on its regression coefficient on θ. When that bound is strictly positive, conservative linear or threshold rules built from the residual dominate the human-only rule over the whole uncertainty set.
Load-bearing premise
The decision maker is allowed to treat every covariance involving the human signal and the AI–human relationship as known, and to put uncertainty only on how the AI covaries with the true state.
What would settle it
On a forecasting or experimental-prediction task, measure human and model residuals after residualizing the model on the human; if those residuals are positively correlated with the truth under every plausible model-quality bound the decision maker would accept, and the joint rule still fails to beat the human-only baseline on held-out utility or MSE, the claimed sufficient condition is not operative in that domain.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies when a decision maker who is uncertain about AI signal quality can still guarantee complementary gains over using only their own signal. Under an uncertainty set that fixes all joint-covariance entries except Cov(φ_AI, θ), the authors show that negative correlation of human and AI errors (plus a uniform lower bound on Cov(φ_AI, θ)) is sufficient for a linear residual rule to strictly dominate the human-only predictor in MSE (Prop. 3.4), for a conservative binary investment rule to weakly dominate the human-only threshold policy (Prop. 3.8), and for an analogous residual rule under a non-Gaussian monotone-signal model with negative dependence of errors (Props. 3.10–3.11). Positive error correlation narrows the safe region to cases where human residual variance is large relative to the shared-error term (Props. 3.6, 3.9). Synthetic experiments match the theory; ForecastBench and TESS data show positive human–LLM error correlations that prompting rarely reverses, so robust combination yields little or no gain over calibrated human-only MSE.
Significance. The work cleanly isolates a structural obstacle to robust human–AI complementarity that is distinct from accuracy or belief-updating failures: uncertainty about AI quality shrinks the region of safe use of the AI residual to settings with negative (or sufficiently mild positive) error correlation. The Gaussian MSE and binary-decision proofs are complete and use standard orthogonal-projection and total-covariance arguments; the non-Gaussian extension rests on a clean monotone-covariance lemma. Synthetic experiments reproduce the predicted phase transition; real-world forecasting benchmarks place current LLMs in the difficult positive-correlation regime and show that simple prompting is unreliable at inducing the required structure. These results give a precise, falsifiable condition that future training or evaluation objectives for complementary LLMs could target.
major comments (2)
- Assumption 3.1 (U varies only Cov(φ_AI, θ); all other entries of Σ, including Cov(φ_H, φ_AI), are treated as known) is load-bearing for residualization and the uniform lower bound on β(Σ). The paper states this restriction and notes that broader uncertainty would shrink the safe region further, but the main claims are proved only under this single-entry uncertainty. A short discussion or corollary quantifying how uncertainty in Cov(φ_H, φ_AI) or Var(φ_AI) further restricts the admissible (δ, γ) region would strengthen the central message without changing the qualitative conclusion.
- In the real-world experiments (Table 1, §5), the robust combiner is evaluated with quantities estimated from the same data used to measure error correlations. While the theory is non-circular, the empirical claim that current models fail the complementarity condition would be more convincing if the robust rule were constructed from a held-out calibration split (as in the synthetic experiments) and evaluated on a disjoint test set, or if sensitivity to estimation error in γ and δ were reported.
minor comments (4)
- Figure 1 caption and axis labels: clarify that the robust estimator uses the conservative b derived from the lower bound on Cov(θ−φ_H, φ_AI), not the oracle regression coefficient.
- Notation: φ_H is normalized so that Cov(θ, φ_H)/Var(φ_H)=1; this is stated after Assumption 3.2 but could be flagged earlier when the generative model is introduced.
- Appendix B.4 lists several prompting strategies (e.g., Explicit Negative Correlation Objective, Two-Stage Error Correction) that do not appear in Table 1; either include them or note that they were omitted because they did not reverse the sign of ρ.
- Typos: “negatively correlatedwith” (abstract/intro) missing space; “uncer-tainty setwhich” (intro) missing space.
Circularity Check
No significant circularity: robust-dominance claims are derived from stated uncertainty-set assumptions; empirical correlations are measured against ground truth, not fitted as predictions.
full rationale
The load-bearing theoretical results (Props. 3.4, 3.6–3.11) are standard decision-theoretic derivations under an explicitly restricted uncertainty set U (Assumption 3.1: only Cov(ϕ_AI, θ) free; other joint-covariance entries fixed). The residual r, the lower bound β(Σ) ≥ b, and the loss-gap identities (Lemmas A.1–A.2) are constructed from those assumptions and proved in Appendix A; they are not algebraically identical to fitted parameters. Without uncertainty, Prop. 3.7 recovers the classical non-redundancy condition; with uncertainty, negative error correlation (or the κ_γ margin under positive correlation) is a structural hypothesis that is tested, not defined into existence. Synthetic experiments control ρ and σ_h and check that the robust rule behaves as the theory predicts. Real-world experiments measure Pearson correlation of human vs. LLM forecast errors against resolved ground truth on ForecastBench and TESS, and report robust MSE of a combiner that uses only quantities fixed by U or estimated on a calibration split—none of which is a “prediction” forced by a fit of the same quantity. Related-work citations (including Wilder et al. 2020) are contextual, not uniqueness theorems that force the main claims. No self-definitional loop, fitted-input-as-prediction, or load-bearing self-citation chain is present.
Axiom & Free-Parameter Ledger
free parameters (2)
- conservativeness shrinkage s =
0.5
- investment threshold and cost (τ, c) =
τ=0, c=0.3
axioms (5)
- ad hoc to paper Uncertainty set U agrees on all covariance entries except Cov(ϕ_AI, θ) (Assumption 3.1)
- domain assumption Neither agent observes a noiseless measurement of θ (Assumption 3.2)
- domain assumption Joint normality of (θ, ϕ_H, ϕ_AI) for the main MSE and binary analyses
- standard math Human signal normalized so that Cov(θ, ϕ_H)/Var(ϕ_H)=1
- domain assumption AI signal uniformly sensitive: g'(t) ≥ k > 0, and E[ε_AI | ε_H] nonincreasing (non-Gaussian case)
Cite this review
Pith. "Pith review of Robust Human-AI Complementarity under Uncertainty." pith.science (2026). https://pith.science/paper/DQNJNZ32
@misc{pith2026260706656,
author = {Pith},
title = {Pith review of: Robust Human-AI Complementarity under Uncertainty},
year = {2026},
howpublished = {\url{https://pith.science/paper/DQNJNZ32}},
note = {Machine review of arXiv:2607.06656}
}
read the original abstract
Machine learning models are often intended to augment rather than replace human decision makers, by providing information that is complementary to human judgement. Yet, in practice, human decision makers routinely fail to realize such complementary gains, even when models provide useful signal. In this work, we study how asymmetric information about the quality of information available to a human decision maker vs. an AI impacts the ability of a decision maker to extract complementary value from AI predictions. We show that a key factor is the error correlation structure between human and AI predictions. In particular, when the AI's prediction errors are \textit{negatively correlated} with those of the human, the decision maker can construct robust strategies which guarantee improvements in expected utility. We empirically investigate whether these conditions for complementarity arise in practice, using real-world forecasting benchmarks.
Figures
Reference graph
Works this paper leans on
-
[1]
Proceedings of the National Academy of Sciences , volume=
Bayesian modeling of human--AI complementarity , author=. Proceedings of the National Academy of Sciences , volume=. 2022 , publisher=
work page 2022
-
[2]
Advances in neural information processing systems , volume=
Predict responsibly: improving fairness and accuracy by learning to defer , author=. Advances in neural information processing systems , volume=
-
[3]
International conference on machine learning , pages=
Consistent estimators for learning to defer to an expert , author=. International conference on machine learning , pages=. 2020 , organization=
work page 2020
-
[4]
The Algorithmic Automation Problem: Prediction, Triage, and Human Effort
The algorithmic automation problem: Prediction, triage, and human effort , author=. arXiv preprint arXiv:1903.12220 , year=
work page internal anchor Pith review Pith/arXiv arXiv 1903
-
[5]
Advances in Neural Information Processing Systems , volume=
Differentiable learning under triage , author=. Advances in Neural Information Processing Systems , volume=
-
[6]
Nature Human Behaviour , volume=
When combinations of humans and AI are useful: A systematic review and meta-analysis , author=. Nature Human Behaviour , volume=. 2024 , publisher=
work page 2024
-
[7]
Predicting results of social science experiments using large language models , author=. Preprint , year=
-
[8]
The Thirteenth International Conference on Learning Representations , year=
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities , author=. The Thirteenth International Conference on Learning Representations , year=
-
[9]
Forty-second International Conference on Machine Learning Position Paper Track , year=
Position: LLM Social Simulations Are a Promising Research Method , author=. Forty-second International Conference on Machine Learning Position Paper Track , year=
-
[10]
Advances in Neural Information Processing Systems , volume=
Approaching human-level forecasting with language models , author=. Advances in Neural Information Processing Systems , volume=
-
[11]
International Conference on Machine Learning , pages=
Improving expert predictions with conformal prediction , author=. International Conference on Machine Learning , pages=. 2023 , organization=
work page 2023
-
[12]
Forty-first International Conference on Machine Learning , year=
Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function , author=. Forty-first International Conference on Machine Learning , year=
-
[13]
MIRAI: Evaluating LLM Agents for Event Forecasting
Mirai: Evaluating llm agents for event forecasting , author=. arXiv preprint arXiv:2407.01231 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[14]
arXiv preprint arXiv:2510.17638 , year=
LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena , author=. arXiv preprint arXiv:2510.17638 , year=
-
[15]
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
Generative agent simulations of 1,000 people , author=. arXiv preprint arXiv:2411.10109 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[16]
Nature human behaviour , volume=
Large language models surpass human experts in predicting neuroscience results , author=. Nature human behaviour , volume=. 2025 , publisher=
work page 2025
-
[17]
Combining human expertise with artificial intelligence: Experimental evidence from radiology , author=. 2023 , institution=
work page 2023
-
[18]
The Fourteenth International Conference on Learning Representations , year=
The value of information in human-ai decision-making , author=. The Fourteenth International Conference on Learning Representations , year=
-
[19]
Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
A decision theoretic framework for measuring AI reliance , author=. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
work page 2024
-
[20]
International Joint Conference on Artificial Intelligence , year=
Learning to Complement Humans , author=. International Joint Conference on Artificial Intelligence , year=
-
[21]
Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages=
Human-algorithm collaboration: Achieving complementarity and avoiding unfairness , author=. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages=
work page 2022
-
[22]
T. M. Mitchell. The Need for Biases in Learning Generalizations. 1980
work page 1980
-
[23]
M. J. Kearns , title =
-
[24]
Machine Learning: An Artificial Intelligence Approach, Vol. I. 1983
work page 1983
-
[25]
R. O. Duda and P. E. Hart and D. G. Stork. Pattern Classification. 2000
work page 2000
-
[26]
Suppressed for Anonymity , author=
-
[27]
A. Newell and P. S. Rosenbloom. Mechanisms of Skill Acquisition and the Law of Practice. Cognitive Skills and Their Acquisition. 1981
work page 1981
-
[28]
A. L. Samuel. Some Studies in Machine Learning Using the Game of Checkers. IBM Journal of Research and Development. 1959
work page 1959
-
[29]
BERTopic: Neural topic modeling with a class-based TF-IDF procedure
BERTopic: Neural topic modeling with a class-based TF-IDF procedure , author=. arXiv preprint arXiv:2203.05794 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[30]
Time-sharing Experiments for the Social Sciences (TESS) , author =
-
[31]
On the Opportunities and Risks of Foundation Models
Rishi Bommasani and Drew A. Hudson and Ehsan Adeli and Russ B. Altman and Simran Arora and Sydney von Arx and Michael S. Bernstein and Jeannette Bohg and Antoine Bosselut and Emma Brunskill and Erik Brynjolfsson and Shyamal Buch and Dallas Card and Rodrigo Castellon and Niladri S. Chatterji and Annie S. Chen and Kathleen Creel and Jared Quincy Davis and D...
work page internal anchor Pith review Pith/arXiv arXiv 2021
-
[32]
Openai gpt-5 system card , author=. arXiv preprint arXiv:2601.03267 , year=
work page internal anchor Pith review Pith/arXiv arXiv
This paper was first reviewed by grok-4.5 on July 10, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.