Pith. sign in

REVIEW 2 major objections 4 minor 32 references

Negative error correlation is what lets a human safely use an AI prediction under uncertainty about the AI’s quality.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Negative human-AI error correlation is required for robust complementarity under uncertainty about AI quality; current LLMs exhibit positive correlations on forecasting tasks.

T0 review reviewed 2026-07-10 challenge →

load-bearing objection Clean theory that negative error correlation is the structural condition for robust human-AI complementarity under uncertainty about AI quality; empirics put current LLMs in the hard regime. the 2 major comments →

arxiv 2607.06656 v1 pith:DQNJNZ32 submitted 2026-07-07 cs.LG

Robust Human-AI Complementarity under Uncertainty

classification cs.LG
keywords human-AI complementarityuncertainty setserror correlationrobust decision rulesforecastingLLM predictionsstatistical decision theory
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Machine learning is often sold as a complement to human judgment, yet people routinely fail to gain from AI advice even when the model has useful signal. This paper argues that one reason is asymmetric uncertainty: the human can cheaply observe the AI’s forecasts and their own, but cannot know how the AI’s forecasts relate to the unknown ground truth. Under that uncertainty, merely knowing that the AI is somewhat accurate is not enough for a rational decision maker to trust a joint rule. The paper shows that the decisive structure is the correlation of prediction errors. When the AI’s errors are negatively correlated with the human’s, a simple residualized rule can be built that is guaranteed to improve expected utility for every distribution consistent with the decision maker’s beliefs. When errors are positively correlated, safe use of the AI requires the human to be substantially noisier than the AI, which is closer to automation than complementarity. On real forecasting benchmarks the observed error correlations are positive, and prompting strategies rarely reverse them, so the safe complementarity window is narrow in current practice.

Core claim

Under an uncertainty set that leaves the AI–truth covariance free while fixing the rest of the joint law, a decision maker has a rule that uses the AI and strictly dominates the human-only optimum for every distribution in the set if and only if residual AI signal after conditioning on the human is uniformly positively related to the truth. Negative correlation of human and AI errors is the structural condition that delivers that uniform positivity; positive error correlation shrinks the safe region to cases where the human residual variance is large relative to the worst-case shared error.

What carries the argument

The residualized AI signal r = ϕ_AI − Cov(ϕ_H, ϕ_AI)/Var(ϕ_H) · ϕ_H (or its nonlinear analogue ˜r = ϕ_AI − E[ϕ_AI|ϕ_H]), together with a uniform lower bound on its regression coefficient on θ. When that bound is strictly positive, conservative linear or threshold rules built from the residual dominate the human-only rule over the whole uncertainty set.

Load-bearing premise

The decision maker is allowed to treat every covariance involving the human signal and the AI–human relationship as known, and to put uncertainty only on how the AI covaries with the true state.

What would settle it

On a forecasting or experimental-prediction task, measure human and model residuals after residualizing the model on the human; if those residuals are positively correlated with the truth under every plausible model-quality bound the decision maker would accept, and the joint rule still fails to beat the human-only baseline on held-out utility or MSE, the claimed sufficient condition is not operative in that domain.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper studies when a decision maker who is uncertain about AI signal quality can still guarantee complementary gains over using only their own signal. Under an uncertainty set that fixes all joint-covariance entries except Cov(φ_AI, θ), the authors show that negative correlation of human and AI errors (plus a uniform lower bound on Cov(φ_AI, θ)) is sufficient for a linear residual rule to strictly dominate the human-only predictor in MSE (Prop. 3.4), for a conservative binary investment rule to weakly dominate the human-only threshold policy (Prop. 3.8), and for an analogous residual rule under a non-Gaussian monotone-signal model with negative dependence of errors (Props. 3.10–3.11). Positive error correlation narrows the safe region to cases where human residual variance is large relative to the shared-error term (Props. 3.6, 3.9). Synthetic experiments match the theory; ForecastBench and TESS data show positive human–LLM error correlations that prompting rarely reverses, so robust combination yields little or no gain over calibrated human-only MSE.

Significance. The work cleanly isolates a structural obstacle to robust human–AI complementarity that is distinct from accuracy or belief-updating failures: uncertainty about AI quality shrinks the region of safe use of the AI residual to settings with negative (or sufficiently mild positive) error correlation. The Gaussian MSE and binary-decision proofs are complete and use standard orthogonal-projection and total-covariance arguments; the non-Gaussian extension rests on a clean monotone-covariance lemma. Synthetic experiments reproduce the predicted phase transition; real-world forecasting benchmarks place current LLMs in the difficult positive-correlation regime and show that simple prompting is unreliable at inducing the required structure. These results give a precise, falsifiable condition that future training or evaluation objectives for complementary LLMs could target.

major comments (2)
  1. Assumption 3.1 (U varies only Cov(φ_AI, θ); all other entries of Σ, including Cov(φ_H, φ_AI), are treated as known) is load-bearing for residualization and the uniform lower bound on β(Σ). The paper states this restriction and notes that broader uncertainty would shrink the safe region further, but the main claims are proved only under this single-entry uncertainty. A short discussion or corollary quantifying how uncertainty in Cov(φ_H, φ_AI) or Var(φ_AI) further restricts the admissible (δ, γ) region would strengthen the central message without changing the qualitative conclusion.
  2. In the real-world experiments (Table 1, §5), the robust combiner is evaluated with quantities estimated from the same data used to measure error correlations. While the theory is non-circular, the empirical claim that current models fail the complementarity condition would be more convincing if the robust rule were constructed from a held-out calibration split (as in the synthetic experiments) and evaluated on a disjoint test set, or if sensitivity to estimation error in γ and δ were reported.
minor comments (4)
  1. Figure 1 caption and axis labels: clarify that the robust estimator uses the conservative b derived from the lower bound on Cov(θ−φ_H, φ_AI), not the oracle regression coefficient.
  2. Notation: φ_H is normalized so that Cov(θ, φ_H)/Var(φ_H)=1; this is stated after Assumption 3.2 but could be flagged earlier when the generative model is introduced.
  3. Appendix B.4 lists several prompting strategies (e.g., Explicit Negative Correlation Objective, Two-Stage Error Correction) that do not appear in Table 1; either include them or note that they were omitted because they did not reverse the sign of ρ.
  4. Typos: “negatively correlatedwith” (abstract/intro) missing space; “uncer-tainty setwhich” (intro) missing space.

Circularity Check

0 steps flagged

No significant circularity: robust-dominance claims are derived from stated uncertainty-set assumptions; empirical correlations are measured against ground truth, not fitted as predictions.

full rationale

The load-bearing theoretical results (Props. 3.4, 3.6–3.11) are standard decision-theoretic derivations under an explicitly restricted uncertainty set U (Assumption 3.1: only Cov(ϕ_AI, θ) free; other joint-covariance entries fixed). The residual r, the lower bound β(Σ) ≥ b, and the loss-gap identities (Lemmas A.1–A.2) are constructed from those assumptions and proved in Appendix A; they are not algebraically identical to fitted parameters. Without uncertainty, Prop. 3.7 recovers the classical non-redundancy condition; with uncertainty, negative error correlation (or the κ_γ margin under positive correlation) is a structural hypothesis that is tested, not defined into existence. Synthetic experiments control ρ and σ_h and check that the robust rule behaves as the theory predicts. Real-world experiments measure Pearson correlation of human vs. LLM forecast errors against resolved ground truth on ForecastBench and TESS, and report robust MSE of a combiner that uses only quantities fixed by U or estimated on a calibration split—none of which is a “prediction” forced by a fit of the same quantity. Related-work citations (including Wilder et al. 2020) are contextual, not uniqueness theorems that force the main claims. No self-definitional loop, fitted-input-as-prediction, or load-bearing self-citation chain is present.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

The central theoretical claims rest on a deliberately restricted uncertainty set (only Cov(AI, θ) unknown), Gaussian or monotone-signal generative assumptions, and a normalization that sets the human regression coefficient to 1. Empirically a single hand-chosen shrinkage s=0.5 appears. No new physical entities are postulated; the residual r and the envelopes P_low/P_high are derived constructions, not free inventions.

free parameters (2)
  • conservativeness shrinkage s = 0.5
    In synthetic and real experiments the plug-in estimate of the residual regression coefficient is multiplied by a fixed s=0.5 before forming the robust rule; the value is chosen by hand and affects how often the robust rule reverts to human-only.
  • investment threshold and cost (τ, c) = τ=0, c=0.3
    Binary-decision experiments fix τ=0, c=0.3; results are illustrative for that pair.
axioms (5)
  • ad hoc to paper Uncertainty set U agrees on all covariance entries except Cov(ϕ_AI, θ) (Assumption 3.1)
    Core modeling choice that makes residualization and uniform lower bounds possible; stated in §3.
  • domain assumption Neither agent observes a noiseless measurement of θ (Assumption 3.2)
    Standard non-degeneracy; excludes trivial perfect-signal cases.
  • domain assumption Joint normality of (θ, ϕ_H, ϕ_AI) for the main MSE and binary analyses
    Used for linear conditional expectations and closed-form posteriors; later relaxed in §3.2.
  • standard math Human signal normalized so that Cov(θ, ϕ_H)/Var(ϕ_H)=1
    WLOG rescaling; does not restrict the uncertainty sets considered.
  • domain assumption AI signal uniformly sensitive: g'(t) ≥ k > 0, and E[ε_AI | ε_H] nonincreasing (non-Gaussian case)
    Structural conditions that replace Gaussian negative correlation; stated as Eqs. 7–8.

reviewed 2026-07-10 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Robust Human-AI Complementarity under Uncertainty." pith.science (2026). https://pith.science/paper/DQNJNZ32

@misc{pith2026260706656,
  author       = {Pith},
  title        = {Pith review of: Robust Human-AI Complementarity under Uncertainty},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DQNJNZ32}},
  note         = {Machine review of arXiv:2607.06656}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Machine learning models are often intended to augment rather than replace human decision makers, by providing information that is complementary to human judgement. Yet, in practice, human decision makers routinely fail to realize such complementary gains, even when models provide useful signal. In this work, we study how asymmetric information about the quality of information available to a human decision maker vs. an AI impacts the ability of a decision maker to extract complementary value from AI predictions. We show that a key factor is the error correlation structure between human and AI predictions. In particular, when the AI's prediction errors are \textit{negatively correlated} with those of the human, the decision maker can construct robust strategies which guarantee improvements in expected utility. We empirically investigate whether these conditions for complementarity arise in practice, using real-world forecasting benchmarks.

Figures

Figures reproduced from arXiv: 2607.06656 by Bryan Wilder, Yewon Byun.

Figure 1
Figure 1. Figure 1: MSE of the expert-only estimator and our robust estima￾tor as we vary the correlation in expert and LLM errors (lower is better). The MSE of our estimator strongly dominates the expert￾only baseline when errors are negatively correlated. Proposition 3.10. If Equations 7 and 8 hold for all ξ ∈ U, r˜ satisfies Cov(˜r, θ) ≥ k E [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: MSE of the expert-only estimator and our robust estima￾tor as we vary the expert uncertainty under the positive correlation setting. Our robust estimator begins to dominate the expert-only baseline when human uncertainty exceeds that of the AI signal. The vertical dotted line denotes the AI signal standard deviation. 1.0 0.5 0.0 0.5 1.0 Error Correlation 0.24 0.25 0.26 0.27 0.28 Average utility Expert-only… view at source ↗
Figure 3
Figure 3. Figure 3: Average utility of the expert-only investment policy dH and the robust symmetric policy dsym as we vary the correlation ρ between human and AI errors (higher is better). Shaded regions indicate ±1 standard deviation across seeds. correlation ρ. Consistent with our theoretical analysis, the robust estimator yields the largest gains when errors are negatively correlated. As ρ becomes large, the improvement o… view at source ↗
Figure 5
Figure 5. Figure 5: Correlation of errors of LLM forecasts vs. human forecasts on ForecastBench (left) and TESS studies (right). We find that errors are positively correlated, which is the difficult setting for complementarity [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: TESS Studies Prompting Experiments. Different Information Sets. Restricted context You are a forecaster. Provide a probability between 0 and 1. Do not look at background and resolution criteria fields. Return JSON only with keys: forecast, reasoning. forecast must be a number in [0, 1]. reasoning should be 1–3 concise sentences. B.5. Distribution Shift Setting A natural question is whether the covariance s… view at source ↗
Figure 7
Figure 7. Figure 7: ForecastBench Prompting Experiments. 30 [PITH_FULL_IMAGE:figures/full_fig_p030_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

32 extracted references · 32 canonical work pages · 6 internal anchors

  1. [1]

    Proceedings of the National Academy of Sciences , volume=

    Bayesian modeling of human--AI complementarity , author=. Proceedings of the National Academy of Sciences , volume=. 2022 , publisher=

  2. [2]

    Advances in neural information processing systems , volume=

    Predict responsibly: improving fairness and accuracy by learning to defer , author=. Advances in neural information processing systems , volume=

  3. [3]

    International conference on machine learning , pages=

    Consistent estimators for learning to defer to an expert , author=. International conference on machine learning , pages=. 2020 , organization=

  4. [4]

    The Algorithmic Automation Problem: Prediction, Triage, and Human Effort

    The algorithmic automation problem: Prediction, triage, and human effort , author=. arXiv preprint arXiv:1903.12220 , year=

  5. [5]

    Advances in Neural Information Processing Systems , volume=

    Differentiable learning under triage , author=. Advances in Neural Information Processing Systems , volume=

  6. [6]

    Nature Human Behaviour , volume=

    When combinations of humans and AI are useful: A systematic review and meta-analysis , author=. Nature Human Behaviour , volume=. 2024 , publisher=

  7. [7]

    Preprint , year=

    Predicting results of social science experiments using large language models , author=. Preprint , year=

  8. [8]

    The Thirteenth International Conference on Learning Representations , year=

    ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities , author=. The Thirteenth International Conference on Learning Representations , year=

  9. [9]

    Forty-second International Conference on Machine Learning Position Paper Track , year=

    Position: LLM Social Simulations Are a Promising Research Method , author=. Forty-second International Conference on Machine Learning Position Paper Track , year=

  10. [10]

    Advances in Neural Information Processing Systems , volume=

    Approaching human-level forecasting with language models , author=. Advances in Neural Information Processing Systems , volume=

  11. [11]

    International Conference on Machine Learning , pages=

    Improving expert predictions with conformal prediction , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  12. [12]

    Forty-first International Conference on Machine Learning , year=

    Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function , author=. Forty-first International Conference on Machine Learning , year=

  13. [13]

    MIRAI: Evaluating LLM Agents for Event Forecasting

    Mirai: Evaluating llm agents for event forecasting , author=. arXiv preprint arXiv:2407.01231 , year=

  14. [14]

    arXiv preprint arXiv:2510.17638 , year=

    LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena , author=. arXiv preprint arXiv:2510.17638 , year=

  15. [15]

    LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

    Generative agent simulations of 1,000 people , author=. arXiv preprint arXiv:2411.10109 , year=

  16. [16]

    Nature human behaviour , volume=

    Large language models surpass human experts in predicting neuroscience results , author=. Nature human behaviour , volume=. 2025 , publisher=

  17. [17]

    2023 , institution=

    Combining human expertise with artificial intelligence: Experimental evidence from radiology , author=. 2023 , institution=

  18. [18]

    The Fourteenth International Conference on Learning Representations , year=

    The value of information in human-ai decision-making , author=. The Fourteenth International Conference on Learning Representations , year=

  19. [19]

    Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

    A decision theoretic framework for measuring AI reliance , author=. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

  20. [20]

    International Joint Conference on Artificial Intelligence , year=

    Learning to Complement Humans , author=. International Joint Conference on Artificial Intelligence , year=

  21. [21]

    Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages=

    Human-algorithm collaboration: Achieving complementarity and avoiding unfairness , author=. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages=

  22. [22]

    T. M. Mitchell. The Need for Biases in Learning Generalizations. 1980

  23. [23]

    M. J. Kearns , title =

  24. [24]

    Machine Learning: An Artificial Intelligence Approach, Vol. I. 1983

  25. [25]

    R. O. Duda and P. E. Hart and D. G. Stork. Pattern Classification. 2000

  26. [26]

    Suppressed for Anonymity , author=

  27. [27]

    Newell and P

    A. Newell and P. S. Rosenbloom. Mechanisms of Skill Acquisition and the Law of Practice. Cognitive Skills and Their Acquisition. 1981

  28. [28]

    A. L. Samuel. Some Studies in Machine Learning Using the Game of Checkers. IBM Journal of Research and Development. 1959

  29. [29]

    BERTopic: Neural topic modeling with a class-based TF-IDF procedure

    BERTopic: Neural topic modeling with a class-based TF-IDF procedure , author=. arXiv preprint arXiv:2203.05794 , year=

  30. [30]

    Time-sharing Experiments for the Social Sciences (TESS) , author =

  31. [31]

    On the Opportunities and Risks of Foundation Models

    Rishi Bommasani and Drew A. Hudson and Ehsan Adeli and Russ B. Altman and Simran Arora and Sydney von Arx and Michael S. Bernstein and Jeannette Bohg and Antoine Bosselut and Emma Brunskill and Erik Brynjolfsson and Shyamal Buch and Dallas Card and Rodrigo Castellon and Niladri S. Chatterji and Annie S. Chen and Kathleen Creel and Jared Quincy Davis and D...

  32. [32]

    OpenAI GPT-5 System Card

    Openai gpt-5 system card , author=. arXiv preprint arXiv:2601.03267 , year=

This paper was first reviewed by grok-4.5 on July 10, 2026.