{"id":"e474245a-cae4-4d67-8822-65eb88ffd365","arxiv_id":"2508.18317","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"Calibration alone does not make human decisions follow a model's confidence scores; a prospect-theory based correction does.","lead":"This paper tests whether calibrating an AI model's confidence scores changes how non-expert humans decide and how much they trust the model. It finds calibration alone is not enough; correcting the scores with a behavioral economics idea (prospect theory) is what makes people's decisions line up with the model's predictions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PT correction may have been tuned on the same decision data; without evidence of pre-specification, the claimed correlation gain could be curve fitting.","rationale":"The reader's verdict is UNVERDICTED because the supplied full text is unreadable and only the abstract could be assessed. My stress-test does not change that: the central claim cannot be evaluated without the actual Methods and Results sections. On the abstract alone, the load-bearing assumption is indeed the pre-specification of the PT correction; if that assumption fails, the claim that the PT correction is 'crucial' would be unsupported. Since no clean full text is available to verify this, the honest outcome is to leave the verdict as UNVERDICTED rather than accept, condition, or reject. I agree with the reader's weakest-assumption identification. The proposed split-half test is a concrete way to settle whether the PT advantage generalizes once any possibility of in-sample tuning is removed.","tokens_in":7177,"tokens_out":2888,"duration_ms":35529,"concrete_test":"Obtain a clean copy of arXiv:2508.18317 and inspect the Methods/analysis section for an explicit statement that the Kahneman-Tversky correction, including its functional form and any numerical parameters (e.g., loss-aversion, probability-weighting constants), was fixed before the experiment or pre-registered. If no such statement exists, request the pre-registration or analysis script. Then run a split-half check: fit the PT transform parameters on a random half of participants/trials, and evaluate the human-model correlation on the held-out half for both calibrated-only and calibrated+PT conditions. If the PT condition's advantage over calibrated-only disappears on held-out data (e.g., the confidence interval for the difference includes zero), the headline claim reduces to in-sample curve fitting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that calibration alone is insufficient and that the prospect theory (PT) correction is 'crucial for increasing the correlation between human decisions and the model's predictions.' For this conclusion to hold, the PT correction must have been an a priori intervention, not a transform selected or tuned after looking at the experimental outcomes. The abstract only states that the corrections are 'based on Kahneman and Tversky's prospect theory'; it does not state whether the functional form or its parameters were fixed before data collection. If the PT form or its numerical parameters were chosen using the same decision data that produced the reported correlation, then the increased correlation is an expected in-sample fitting artifact, and the comparison between calibrated-only and calibrated+PT conditions does not establish that PT is causally important. A second, related gap is that the abstract reports no effect on self-reported trust but does not define the correlation metric or report its uncertainty; however, that is secondary. The supplied full text is corrupted mojibake, so no Methods, equations, or statistical details can be checked; this concern is based on the abstract's wording and would be resolved by inspecting the actual manuscript.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an HCI experiment on how calibrating a machine-learning classifier affects non-expert humans' decisions and trust. The abstract states two main findings: (i) calibration alone is not sufficient; applying a prospect-theory (PT) correction to calibrated scores significantly increases the correlation between human decisions and model predictions; and (ii) self-reported answers to \"Do you trust the model more?\" are unaffected by the presentation method. The supplied full text is heavily corrupted mojibake, so no Methods, equations, sample details, or statistical analyses could be inspected. The assessment below is therefore based almost entirely on the abstract and on fragments that survive in the corrupted text.","tokens_in":7427,"tokens_out":2607,"duration_ms":30652,"significance":"If the findings are valid, the paper makes a useful empirical contribution to the HCI/ML calibration literature: it distinguishes the effect of statistical calibration from the effect of a psychologically motivated transformation of calibrated scores, and it reports a dissociation between behavioral alignment and self-reported trust. Anchoring the transformation in an external theory (Kahneman and Tversky's prospect theory) is a strength, provided the functional form and parameters were specified independently of the experimental outcomes. However, the current manuscript as supplied does not provide enough information to verify the central claim, so the significance cannot yet be assessed.","major_comments":[{"comment":"The central claim is unverifiable from the supplied manuscript. The full text is corrupted mojibake, and the abstract reports no sample size, participant population, task design, statistical test, effect size, confidence interval, or exclusion rule. These are load-bearing for the claim that \"the prospect theory correction is crucial for increasing the correlation.\" Without a readable Methods and Results section, I cannot evaluate whether the experimental design supports the causal language in the abstract.","section":"Abstract; full text"},{"comment":"The abstract does not state whether the PT correction's functional form and parameters were fixed before data collection or selected after examining the decision data. If the parameters were tuned on the same data used to compute the reported correlation, the improved correlation is an in-sample fitting artifact and does not establish that PT is causally important. The manuscript needs to report the exact PT equation, the parameter values, their literature source, and a statement of pre-specification (e.g., preregistration or fixed-prior parameters).","section":"Abstract, \"based on Kahneman and Tversky's prospect theory\""},{"comment":"The correlation metric is not defined. The abstract does not say whether this is Pearson or Spearman correlation, whether it is computed per participant and averaged or pooled across participants, or on what scale (raw decisions, decision confidence, binarized predictions). It also reports no uncertainty or confidence interval for the correlation difference across conditions. Without this, the magnitude and reliability of the claimed effect cannot be assessed.","section":"Abstract, \"correlation between decisions and predictions\""},{"comment":"The null finding on self-reported trust is stated without any statistical evidence. A claim of \"no effect\" requires either a test with an equivalence bound or at least a reported effect size and confidence interval. The abstract also does not describe the questionnaire item or its response scale, making it unclear whether the instrument could have detected a difference.","section":"Abstract, \"responses to 'Do you trust the model more?' are unaffected\""}],"minor_comments":[{"comment":"The supplied full text is unreadable mojibake. If this is the submitted version, a clean PDF is required before any further review can be conducted.","section":"Full text"},{"comment":"The phrase \"calibration is not sufficient on its own\" is potentially ambiguous. Since the PT correction is applied to calibrated scores, the contrast is between calibrated scores and calibrated scores plus a psychological transform. Clarify that the conclusion concerns the presentation or scoring transform, not calibration in the statistical sense.","section":"Abstract, wording"},{"comment":"The abstract refers to \"non-expert humans\" but gives no inclusion criteria, recruitment source, or number of participants. This information is needed even for a high-level summary.","section":"Abstract, participant description"}],"recommendation":"uncertain","confidential_remarks":"The supplied full text is corrupted, so I could only review the abstract and scattered fragments. This report should not be read as an endorsement or rejection of the substance. Please obtain a clean, readable version of the manuscript before sending it to referees; if the clean version contains the missing experimental and statistical details, the paper may be suitable for major_revision or further review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I’ll get straight to the point: the research question is worth asking, and the abstract's central contrast — calibrated scores alone don't shift non-expert decisions, but the prospect-theory correction does — is a genuinely new and testable proposition. If that holds up, it tells us something useful about how to present model outputs to people. Credit where due: the intervention is anchored in Kahneman and Tversky's framework rather than invented ad hoc, and the claim that behavioral alignment can diverge from self-reported trust is a nice secondary observation.\n\nNow the soft spots. The supplied full text is corrupted: it’s unreadable mojibake with a header from an unrelated cond-mat paper. So I can only evaluate the abstract. That means the big empirical claims — sample size, statistical tests, the exact definition of “correlation between decisions and predictions,” and how the PT parameters were set — are all uncheckable. Your stress-test note hits the right nerve: if the PT functional form or its parameters were chosen after looking at the decision data, the “correlation gain” could be in-sample curve fitting. But the abstract says only that the corrections are “based on” prospect theory; there’s no way to tell pre-specification from post-hoc fitting without the Methods. That’s not a confirmed flaw — it’s an open question that the full text would answer. The other weaknesses (no effect sizes, no uncertainty) are also invisible because of the corruption; they may be present in the actual manuscript.\n\nBottom line: the idea and the experimental contrast deserve a serious referee, but the version I was given cannot be reviewed. If the real paper is intact, I’d send it to a reviewer with a background in human-AI interaction and behavioral economics, and let them judge the pre-specification issue. The authors should also fix the arXiv upload — submitting a corrupted file is a self-inflicted wound. If I’m asked to cite or build on this tomorrow, I can’t, because I can’t verify the evidence.","headline":"Promising question and novel intervention, but the supplied text is corrupted beyond the abstract, so the empirical claims are impossible to verify.","tokens_in":7893,"tokens_out":2388,"would_cite":false,"duration_ms":26470,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Calibrated probabilities alone do not change human decisions; a prospect-theory correction does.","keywords":["calibration","prospect theory","human decision-making","model trust","human-computer interaction","probability weighting","decision alignment","model confidence"],"falsifier":"Re-run the experiment with the prospect-theory correction's parameters and functional form fixed and pre-registered before any human data are collected. If pre-registered parameters reproduce the correlation gain, the claim stands; if the gain disappears or requires data-driven tuning, the result reduces to post-hoc fitting. A second check: measure human decision accuracy against ground truth; if corrections only raise correlation with model predictions while human accuracy drops, the 'alignment' result is ambiguous.","tokens_in":7088,"feed_emoji":"🎯","tokens_out":3834,"duration_ms":37425,"temperature":0.7,"pith_summary":"The paper tests whether standard probability calibration—making a model's confidence scores match its actual accuracy—changes how non-expert humans use those scores when making decisions. In a human-computer interaction experiment, the authors find that calibration by itself does not increase the correlation between human decisions and the model's predictions. What does increase that correlation is a further, nonlinear correction based on prospect theory, which adjusts reported scores to match how people actually weigh probabilities. Interestingly, self-reported trust in the model does not change with any of the methods, even though the behavioral alignment improves.","feed_headline":"Calibration isn't enough; a prospect-theory fix aligns humans with AI","feed_subtitle":"A behavioral-economics correction to calibrated scores lifts decision-prediction correlation; stated trust doesn't move.","key_machinery":"The prospect theory correction: a nonlinear transformation of calibrated probability scores, grounded in Kahneman and Tversky's probability-weighting function, that maps stated probabilities to the subjective weights humans appear to use when choosing. This transform is the component doing the work: it converts a model's well-calibrated confidence into the form people actually act on.","core_discovery":"The central claim is that calibration, by itself, is not sufficient to make human decisions track a model's predictions more closely. Adding a prospect-theory correction to the calibrated scores—a transformation derived from the behavioral-economics finding that people overweight small probabilities and underweight large ones—is what significantly increases decision-prediction correlation. The paper further shows that this improved behavioral alignment is not accompanied by a change in explicit trust judgments: participants' answers to a direct 'do you trust the model more?' question are unaffected by whether they saw raw, calibrated, or prospect-theory-corrected scores. Thus, the effect ope","pith_inferences":["If the prospect-theory parameters were fixed before the experiment, the result is a free-standing demonstration; if they were tuned on the decision data, the quantitative correlation gain may partly reflect overfitting to the same subjects. Inspecting the design or pre-registration would settle this","The uncoupling of explicit trust from decision alignment suggests a two-channel model of human-AI interaction: stated attitudes and behavioral compliance can move independently; a practical corollary is that trust surveys alone are a poor proxy for actual reliance.","The findings point to a possible design rule: when a model's output is meant to guide a human decision, presentation engineering (e.g., distorting probabilities to compensate for human bias) may be as consequential as statistical calibration. A testable extension would vary the PT weighting parameters across participant groups to find an optimal presentation transform.","One open question the paper leaves implicit is whether the same correction improves decisions against ground truth or only alignment with the model; alignment with a partly wrong model could push human choices in the model's direction without improving real-world outcomes."],"forward_implications":["If the result holds, deploying calibrated models for human-in-the-loop decisions should include a prospect-theory correction, not just calibration, to increase alignment with model predictions.","The null effect on self-reported trust implies that subjective trust questionnaires may not capture the decision-level influence that corrected scores produce; future trust studies should pair attitude measures with behavioral correlation measures.","The claim that increased correlation suggests higher trust is bounded: the paper's direct trust question did not move, so the evidence for trust gain is indirect and should not be oversold.","In safety-critical or high-stakes settings, the finding implies that how scores are presented to human operators matters at least as much as whether the scores are statistically accurate."],"supporting_citations":[],"fun_headline_variants":["Prospect theory, not calibration, aligns human decisions with AI","Calibration alone fails; prospect-theory fix aligns humans with AI","Human-AI decision alignment needs prospect theory, not just calibration","Calibration alone isn't enough; prospect theory aligns humans with AI","Calibration can't do it alone; prospect theory aligns humans with AI"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The prospect-theory correction must have been specified (functional form and parameters) from prior behavioral literature, not chosen by looking at the experimental results; otherwise the reported increase in correlation could be curve fitting.","fun_headline_variants_meta":{"raw":{"variants":["Prospect theory, not calibration, aligns human decisions with AI","Calibration alone fails; prospect-theory fix aligns humans with AI","Human-AI decision alignment needs prospect theory, not just calibration","Calibration alone isn't enough; prospect theory aligns humans with AI","Calibration can't do it alone; prospect theory aligns humans with AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001137,"raw_usage":{"total_tokens":4517,"prompt_tokens":664,"completion_tokens":3853,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":408,"completion_tokens_details":{"reasoning_tokens":3763}},"tokens_in":408,"tokens_out":3853,"duration_ms":29915,"temperature":1.0,"reasoning_tokens":3763,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T16:58:48.403878+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the experiment with the prospect-theory correction's parameters and functional form fixed and pre-registered before any human data are collected. If pre-registered parameters reproduce the correlation gain, the claim stands; if the gain disappears or requires data-driven tuning, the result reduces to post-hoc fitting. A second check: measure human decision accuracy against ground truth; if corrections only raise correlation with model predictions while human accuracy drops, the 'alignment' result is ambiguous.","supporting_citations":[],"review_version":1}