REVIEW 4 major objections 3 minor 34 references
Does Calibration Affect Human Actions?
T0 review · 4 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Calibrated probabilities alone do not change human decisions; a prospect-theory correction does.
desk verdict Promising question and novel intervention, but the supplied text is corrupted beyond the abstract, so the empirical claims are impossible to verify. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The prospect theory correction: a nonlinear transformation of calibrated probability scores, grounded in Kahneman and Tversky's probability-weighting function, that maps stated probabilities to the subjective weights humans appear to use when choosing. This transform is the component doing the work: it converts a model's well-calibrated confidence into the form people actually act on.
What would settle it
Re-run the experiment with the prospect-theory correction's parameters and functional form fixed and pre-registered before any human data are collected. If pre-registered parameters reproduce the correlation gain, the claim stands; if the gain disappears or requires data-driven tuning, the result reduces to post-hoc fitting. A second check: measure human decision accuracy against ground truth; if corrections only raise correlation with model predictions while human accuracy drops, the 'alignment' result is ambiguous.
Extended reading notes
Core claim
The central claim is that calibration, by itself, is not sufficient to make human decisions track a model's predictions more closely. Adding a prospect-theory correction to the calibrated scores—a transformation derived from the behavioral-economics finding that people overweight small probabilities and underweight large ones—is what significantly increases decision-prediction correlation. The paper further shows that this improved behavioral alignment is not accompanied by a change in explicit trust judgments: participants' answers to a direct 'do you trust the model more?' question are unaffected by whether they saw raw, calibrated, or prospect-theory-corrected scores. Thus, the effect ope
Load-bearing premise
The prospect-theory correction must have been specified (functional form and parameters) from prior behavioral literature, not chosen by looking at the experimental results; otherwise the reported increase in correlation could be curve fitting.
Editorial extensions
If this is right
- If the result holds, deploying calibrated models for human-in-the-loop decisions should include a prospect-theory correction, not just calibration, to increase alignment with model predictions.
- The null effect on self-reported trust implies that subjective trust questionnaires may not capture the decision-level influence that corrected scores produce; future trust studies should pair attitude measures with behavioral correlation measures.
- The claim that increased correlation suggests higher trust is bounded: the paper's direct trust question did not move, so the evidence for trust gain is indirect and should not be oversold.
- In safety-critical or high-stakes settings, the finding implies that how scores are presented to human operators matters at least as much as whether the scores are statistically accurate.
Reading between the lines
- If the prospect-theory parameters were fixed before the experiment, the result is a free-standing demonstration; if they were tuned on the decision data, the quantitative correlation gain may partly reflect overfitting to the same subjects. Inspecting the design or pre-registration would settle this
- The uncoupling of explicit trust from decision alignment suggests a two-channel model of human-AI interaction: stated attitudes and behavioral compliance can move independently; a practical corollary is that trust surveys alone are a poor proxy for actual reliance.
- The findings point to a possible design rule: when a model's output is meant to guide a human decision, presentation engineering (e.g., distorting probabilities to compensate for human bias) may be as consequential as statistical calibration. A testable extension would vary the PT weighting parameters across participant groups to find an optimal presentation transform.
- One open question the paper leaves implicit is whether the same correction improves decisions against ground truth or only alignment with the model; alignment with a partly wrong model could push human choices in the model's direction without improving real-world outcomes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an HCI experiment on how calibrating a machine-learning classifier affects non-expert humans' decisions and trust. The abstract states two main findings: (i) calibration alone is not sufficient; applying a prospect-theory (PT) correction to calibrated scores significantly increases the correlation between human decisions and model predictions; and (ii) self-reported answers to "Do you trust the model more?" are unaffected by the presentation method. The supplied full text is heavily corrupted mojibake, so no Methods, equations, sample details, or statistical analyses could be inspected. The assessment below is therefore based almost entirely on the abstract and on fragments that survive in the corrupted text.
Significance. If the findings are valid, the paper makes a useful empirical contribution to the HCI/ML calibration literature: it distinguishes the effect of statistical calibration from the effect of a psychologically motivated transformation of calibrated scores, and it reports a dissociation between behavioral alignment and self-reported trust. Anchoring the transformation in an external theory (Kahneman and Tversky's prospect theory) is a strength, provided the functional form and parameters were specified independently of the experimental outcomes. However, the current manuscript as supplied does not provide enough information to verify the central claim, so the significance cannot yet be assessed.
major comments (4)
- [Abstract; full text] The central claim is unverifiable from the supplied manuscript. The full text is corrupted mojibake, and the abstract reports no sample size, participant population, task design, statistical test, effect size, confidence interval, or exclusion rule. These are load-bearing for the claim that "the prospect theory correction is crucial for increasing the correlation." Without a readable Methods and Results section, I cannot evaluate whether the experimental design supports the causal language in the abstract.
- [Abstract, "based on Kahneman and Tversky's prospect theory"] The abstract does not state whether the PT correction's functional form and parameters were fixed before data collection or selected after examining the decision data. If the parameters were tuned on the same data used to compute the reported correlation, the improved correlation is an in-sample fitting artifact and does not establish that PT is causally important. The manuscript needs to report the exact PT equation, the parameter values, their literature source, and a statement of pre-specification (e.g., preregistration or fixed-prior parameters).
- [Abstract, "correlation between decisions and predictions"] The correlation metric is not defined. The abstract does not say whether this is Pearson or Spearman correlation, whether it is computed per participant and averaged or pooled across participants, or on what scale (raw decisions, decision confidence, binarized predictions). It also reports no uncertainty or confidence interval for the correlation difference across conditions. Without this, the magnitude and reliability of the claimed effect cannot be assessed.
- [Abstract, "responses to 'Do you trust the model more?' are unaffected"] The null finding on self-reported trust is stated without any statistical evidence. A claim of "no effect" requires either a test with an equivalence bound or at least a reported effect size and confidence interval. The abstract also does not describe the questionnaire item or its response scale, making it unclear whether the instrument could have detected a difference.
minor comments (3)
- [Full text] The supplied full text is unreadable mojibake. If this is the submitted version, a clean PDF is required before any further review can be conducted.
- [Abstract, wording] The phrase "calibration is not sufficient on its own" is potentially ambiguous. Since the PT correction is applied to calibrated scores, the contrast is between calibrated scores and calibrated scores plus a psychological transform. Clarify that the conclusion concerns the presentation or scoring transform, not calibration in the statistical sense.
- [Abstract, participant description] The abstract refers to "non-expert humans" but gives no inclusion criteria, recruitment source, or number of participants. This information is needed even for a high-level summary.
Circularity Check
No circularity detectable from the available text; the prospect-theory correction is anchored in external literature, not defined in terms of the measured correlation.
full rationale
The only legible portion of the manuscript is the abstract, which states that the intervention is a 'further correction to the reported calibrated scores based on Kahneman and Tversky's prospect theory.' That grounding is external to the experimental outcome being measured, so the central comparison (calibrated scores vs. calibrated + prospect-theory correction vs. human decisions) is not self-definitional by the abstract alone. The abstract does not report that the correction's functional form or parameters were fitted to the decision data; it only says the correction is 'based on' an existing behavioral-economics theory. The worry that the correction may have been tuned on the same data after seeing outcomes is a legitimate pre-specification/transparency concern, but under the hard rules it is not demonstrable circularity because no equation, parameter, or fitted value is shown that would make the claimed increase in correlation true by construction. The full text is corrupted mojibake, so no Methods, equations, or statistical details can be inspected to check for a reduction of the prediction to its inputs. Without quotable evidence of a specific circular step, the honest finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- Prospect theory parameters (probability weighting and value curvature, e.g., alpha, beta, delta, lambda) =
not stated in the abstract
assumptions (3)
- domain assumption Kahneman and Tversky's prospect theory describes how non-expert humans actually weight probability information when making decisions from ML confidence scores.
- domain assumption The measured correlation between human decisions and the model's predictions is a valid indicator of calibration's practical effect and of trust.
- domain assumption The experimental sample and task are representative enough to support the generalization that 'calibration is not sufficient on its own'.
Cite this review
Pith. "Pith review of Does Calibration Affect Human Actions?." pith.science (2026). https://pith.science/paper/NPYDSWJI
@misc{pith2026250818317,
author = {Pith},
title = {Pith review of: Does Calibration Affect Human Actions?},
year = {2026},
howpublished = {\url{https://pith.science/paper/NPYDSWJI}},
note = {Machine review of arXiv:2508.18317}
}
read the original abstract
Calibration has been proposed as a way to enhance the reliability and adoption of machine learning classifiers. We study a particular aspect of this proposal: how does calibrating a classification model affect the decisions made by non-expert humans consuming the model's predictions? We perform a Human-Computer-Interaction (HCI) experiment to ascertain the effect of calibration on (i) trust in the model, and (ii) the correlation between decisions and predictions. We also propose further corrections to the reported calibrated scores based on Kahneman and Tversky's prospect theory from behavioral economics, and study the effect of these corrections on trust and decision-making. We find that calibration is not sufficient on its own; the prospect theory correction is crucial for increasing the correlation between human decisions and the model's predictions. While this increased correlation suggests higher trust in the model, responses to ``Do you trust the model more?" are unaffected by the method used.
Reference graph
Works this paper leans on
-
[1]
Allwein, E. L., R. E. Schapire, and Y. Singer (2000). Reducing multiclass to binary: A unifying approach for margin classifiers. Journal of machine learning research\/ 1 , 113--141
work page 2000
-
[2]
Ayer, M., H. D. Brunk, G. M. Ewing, W. T. Reid, and E. Silverman (1955). An empirical distribution function for sampling with incomplete information. The annals of mathematical statistics\/ , 641--647
work page 1955
-
[3]
Barbosa, G. D. J., D. dos Santos Ribeiro , M. do Carmo Silva , H. Lopes, and S. D. J. Barbosa (2022). Investigating the relationships between class probabilities and users’ appropriate trust in computer vision classifications of ambiguous images. Journal of Computer Languages\/ 72
work page 2022
-
[4]
Bostrom, N. and E. Yudkowsky (2018). The ethics of artificial intelligence. In Artificial intelligence safety and security , pp.\ 57--69. Chapman and Hall/CRC
work page 2018
-
[5]
Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly weather review\/ 78 , 1--3
work page 1950
-
[6]
De Boer, P.-T., D. P. Kroese, S. Mannor, and R. Y. Rubinstein (2005). A tutorial on the cross-entropy method. Annals of operations research\/ 134 , 19--67
work page 2005
-
[7]
De Giorgi , E. G. and S. Legg (2012). Dynamic portfolio choice and asset pricing with narrow framing and probability weighting. Journal of Economic Dynamics and Control\/ 36\/ (7), 951--972
work page 2012
-
[8]
Gneiting, T., F. Balabdaoui, and A. E. Raftery (2007). Probabilistic forecasts, calibration and sharpness. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 69\/ (2), 243--268
work page 2007
Show all 34 references
-
[9]
Pleiss, Y
Guo, C., G. Pleiss, Y. Sun, and K. Q. Weinberger (2017). On calibration of modern neural networks. In Proceedings of the ICML international conference on machine learning , pp.\ 1321--1330. PMLR
2017
-
[10]
Gupta, C. (2023). Post-hoc calibration without distributional assumptions . Ph.\ D. thesis, Carnegie Mellon University Pittsburgh, PA 15213, USA
2023
-
[11]
Ingersoll, J. (2008). Non‐monotonicity of the tversky‐kahneman probability‐weighting function: A cautionary note. European Financial Management\/ 14 , 385 -- 390
2008
-
[12]
Joshi, A., S. Kale, S. Chandel, and D. K. Pal (2015). Likert scale: Explored and explained. British Journal of Applied Science & Technology\/ 7\/ (4), 396
2015
-
[13]
Kahneman, D. and A. Tversky (1979). Prospect theory: An analysis of decision under risk. Econometrica\/ 47 , 263--291
1979
-
[14]
Kahneman, D. and A. Tversky (1992). Advances in prospect theory: Cumulative representation of uncertainty. Journal of Risk and Uncertainty\/ 5 , 297--323
1992
-
[15]
Kaplan, A. M. and M. Haenlein (2019). Siri, siri, in my hand: Who’s the fairest in the land? on the interpretations, illustrations, and implications of artificial intelligence. Business Horizons\/
2019
-
[16]
Kingma, D. P. and J. Ba (2015). Adam: A method for stochastic optimization. In Proceedings of the ICLR International Conference on Learning Representations
2015
-
[17]
Sarawagi, and U
Kumar, A., S. Sarawagi, and U. Jain (2018). Trainable calibration measures for neural networks from kernel mean embeddings. In Proceedings of the ICML International Conference on Machine Learning , Volume 80, pp.\ 2805--2814
2018
-
[18]
Maddox, W. J., T. Garipov, P. Izmailov, D. Vetrov, and A. G. Wilson (2019). A simple baseline for bayesian uncertainty in deep learning. In Proceedings of the NIPS International Conference on Neural Information Processing Systems . Curran Associates Inc
2019
-
[19]
Camoriano, P
Milios, D., R. Camoriano, P. Michiardi, L. Rosasco, and M. Filippone (2018). Dirichlet-based gaussian processes for large-scale calibrated classification. Advances in Neural Information Processing Systems\/ 31
2018
-
[20]
Mittelstadt, B. D., P. Allo, M. Taddeo, S. Wachter, and L. Floridi (2016). The ethics of algorithms: Mapping the debate. Big Data & Society\/ 3
2016
-
[21]
Murphy, A. H. and R. L. Winkler (1977). Reliability of subjective probability forecasts of precipitation and temperature. Journal of the Royal Statistical Society Series C: Applied Statistics\/ 26 , 41--47
1977
-
[22]
Naeini, M. P., G. Cooper, and M. Hauskrecht (2015). Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the AAAI conference on artificial intelligence , Volume 29
2015
-
[23]
Niculescu-Mizil, A. and R. Caruana (2005). Predicting good probabilities with supervised learning. In Proceedings of the 22nd international conference on Machine learning , pp.\ 625--632
2005
-
[24]
Platt, J. (1999). Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in large margin classifiers\/ 10\/ (3), 61--74
1999
-
[25]
Rechkemmer, A. and M. Yin (2022). When confidence meets accuracy: Exploring the effects of multiple performance indicators on trust in machine learning models. In Proceedings of the CHI Conference on Human Factors in Computing Systems
2022
-
[26]
Wang, and T
Rieger, M., M. Wang, and T. Hens (2017). Estimating cumulative prospect theory parameters from an international survey. Theory and Decision\/ 82
2017
-
[27]
Rieger, M. O. and M. Wang (2006). Cumulative prospect theory and the st. petersburg paradox. Economic Theory\/ 28\/ (3), 665--679
2006
-
[28]
Wright, and R
Robertson, T., F. Wright, and R. Dykstra (1988). Order Restricted Statistical Inference . Probability and Statistics Series. John Wiley and Sons
1988
-
[29]
Gerstenberg, and J
Vodrahalli, K., T. Gerstenberg, and J. Zou (2022). Uncalibrated models can improve human- AI collaboration. In Advances in Neural Information Processing Systems , Volume 35, pp.\ 4004--4016
2022
-
[30]
Berkovsky, R
Yu, K., S. Berkovsky, R. Taib, J. Zhou, and F. Chen (2019). Do i trust my machine teammate? an investigation from perception to decision. In Proceedings of the 24th International Conference on Intelligent User Interfaces , pp.\ 460–468. Association for Computing Machinery
2019
-
[31]
Zadrozny, B. and C. Elkan (2001). Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. In Proceedings of the ICML international conference on machine learning , Volume 1, pp.\ 609--616
2001
-
[32]
Zadrozny, B. and C. Elkan (2002). Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of the ACM SIGKDD international conference on Knowledge discovery and data mining , pp.\ 694--699
2002
-
[33]
Ranzato, R
Zeiler, M., M. Ranzato, R. Monga, M. Mao, K. Yang, Q. Le, P. Nguyen, A. Senior, V. Vanhoucke, J. Dean, and G. Hinton (2013). On rectified linear units for speech processing. In Proceedings of the ICASSP International Conference on Acoustics, Speech and Signal Processing , pp.\...
2013
-
[34]
Zhang, Y., Q. V. Liao, and R. K. E. Bellamy (2020). Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision making. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency , pp.\ 295–305
2020
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.