{"id":"f4428426-632c-4d02-b36b-c1449b308a9c","arxiv_id":"2501.16693","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"Low AI confidence reduced clinician trust and agreement and increased decision time, while high confidence corresponded to a small drop in diagnostic accuracy in a 28-participant web experiment.","lead":"This study ran a web experiment with 28 clinicians using an AI breast cancer decision support tool with different explainability levels and confidence scores. Low AI confidence made clinicians trust and agree less and take longer to decide, while high confidence slightly reduced diagnostic accuracy, suggesting overreliance.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The high-confidence overreliance finding is at risk of being an artifact: Section 3.4 codes any agreement above neutral as adoption of the AI label, so the reported 'performance decrease' may be mechanical rather than behavioral.","rationale":"The central claim about overreliance reducing diagnostic accuracy rests on the Table 2 performance result. That result is not interpretable as a measured behavior change because of the Section 3.4 coding rule: trials with agreement above neutral are never accompanied by an actual participant diagnosis, so performance is imputed to be the AI label. High confidence increases agreement (nonsignificantly but positively), mechanically pushing more high-confidence trials onto the AI's accuracy curve, which is around 81%. A small negative performance coefficient under high confidence is therefore a predictable consequence of the outcome construction, not necessarily evidence that clinicians made worse decisions. The same problem does not apply to the trust, agreement, or duration outcomes, but it directly undermines the most novel headline finding. The additional identification problems noted by the reader, including the fixed intervention order and the no-confidence baseline being only the first intervention, reinforce rather than replace this concern. Because existing data do not contain the actual final decisions for the crucial trials, the concern can only be settled by a replication that records a forced-choice diagnosis on every trial. Until then, the high-confidence overreliance claim is unsupported. Since the reader already rejected the preprint on closely related grounds, my assessment does not change the verdict.","tokens_in":16783,"tokens_out":9218,"duration_ms":95205,"concrete_test":"Run a focal replication in which every trial elicits a forced-choice final diagnosis (Healthy/Benign/Malignant) after the agreement rating, and re-estimate the Table 2 performance model using these actual decisions with intervention fixed effects. If the high-confidence coefficient (β = -0.015) is no longer negative and significant, the overreliance effect is an artifact of the agreement-based performance imputation and/or the first-intervention no-confidence baseline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.4 defines performance so that on every trial where the participant rates agreement above neutral, the recorded 'decision' is the AI's suggested label; the participant's own diagnosis is never elicited. The high-confidence performance result in Table 2 (β = -0.015, p = 0.020) is thus not a clean behavioral effect. High confidence tends to raise agreement (β = 0.108, p = 0.121 in the same model), so more high-confidence trials are scored as AI correctness. Since the AI's accuracy is only about 81%, the small negative coefficient relative to the no-confidence reference is exactly what this coding rule would generate if high confidence shifts stated agreement but not actual decisions. The design also cannot fully separate this from the fact that the no-confidence reference is only the first intervention and from the fixed intervention order, but the imputation is the decisive problem: the very trials that define 'overreliance' contain no recorded human decision.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a web-based interrupted time-series experiment with 28 U.S. healthcare professionals who diagnosed breast ultrasound images under an AI-based CDSS with increasing levels of explainability. The authors model how AI confidence scores (none, low, high) affect trust, agreement, diagnostic performance, and diagnosis duration, and they analyze how explainability affects mental demand and stress. They report that low confidence significantly decreased trust and agreement and increased diagnosis duration, while high confidence was associated with a small but significant decrease in performance, which they interpret as overreliance. They also report that the localization intervention increased stress and that demographic factors were associated with perceptions of AI.","tokens_in":16884,"tokens_out":12629,"duration_ms":117801,"significance":"The study addresses a practically important question in clinical decision support, and the low-confidence results—decreased trust and agreement and increased diagnosis duration—are internally consistent, within-participant findings that could inform designs for calibrated reliance. The experiment uses a clinically relevant task and a realistic AI model trained on breast ultrasound data, which strengthens ecological validity relative to abstract prediction tasks. However, the headline claim that high confidence caused overreliance and reduced diagnostic accuracy rests on a performance measure that imputes the AI label on any trial with above-neutral agreement, so the effect is not a clean behavioral outcome. The abstract also overstates the high-confidence trust result, and the discussion contradicts the mental-demand table. These issues are load-bearing for the paper's main conclusions.","major_comments":[{"comment":"The outcome labeled 'diagnostic performance' is not an observed final diagnosis on trials where agreement is above neutral; the participant's decision is 'considered aligned with the AI's suggestion' and is therefore scored as the AI label. Table 2's high-confidence performance coefficient (β = −0.015, p = 0.020) is consequently not a direct measure of decision accuracy: high confidence and agreement are modeled separately but are not independent, and the same agreement ratings are used to construct the performance outcome. Because no independent human diagnosis was elicited on those trials, the interpretation in Section 4.2 and the abstract that high confidence 'led to overreliance, reducing diagnostic accuracy' is not supported by the data as reported.","section":"Section 3.4, Performance"},{"comment":"The abstract states that 'high confidence scores substantially increased trust' and Section 6 repeats that high confidence 'substantially improved trust,' but Table 2 reports β = 0.103, p = 0.169 for high-confidence trust, which Section 4.2 correctly describes as not statistically significant. The abstract and conclusion should be corrected to match the reported result. The conclusion's statement that 'AI confidence score also elevated stress levels' is also unsupported, since stress is modeled by intervention in Table 1 and never by confidence score.","section":"Abstract and Section 6"},{"comment":"The Discussion states that the fourth condition 'resulted in increased mental demand,' but Table 1 reports β = −0.107, p = 0.501 for the 4th Intervention, a nonsignificant decrease, and Section 4.1 describes a slight decrease. This is an internal contradiction that affects the interpretation of RQ1; the text should be aligned with the table.","section":"Section 5.1 versus Table 1"},{"comment":"The experiment uses a fixed intervention order for all participants, and the mixed-effects models in Table 2 include no time or order term. The low- and high-confidence effects are therefore confounded with practice, fatigue, and cumulative interface exposure, so the causal language in Section 5.2 is stronger than the design supports. Please add a sensitivity analysis with session or trial order as a covariate, or explicitly temper the causal claims.","section":"Sections 3.3 and 4.2"},{"comment":"The demographic ANOVAs appear to treat repeated post-intervention observations as independent units; for example, gender on complexity perception is reported as F = 92.97, p = 6.61e-16, which is implausible with 28 participants unless each observation is counted as an independent case. Please specify the unit of analysis and the number of observations per cell, and use participant-level or mixed-effects models before drawing the demographic conclusions in RQ3.","section":"Section 4.3, Table 3"}],"minor_comments":[{"comment":"The agreement and trust scales are described as 5-point Likert scales 'from 0 to 5,' which is six response options; please clarify the actual scale and the neutral point.","section":"Section 3.4"},{"comment":"Section 3.3 says participants give their own diagnosis when agreement is 'below 3,' while Section 3.4 refers to agreement 'above neutral'; specify how the neutral response (e.g., agreement = 3) is handled.","section":"Sections 3.3 and 3.4"},{"comment":"The intervention numbering is unclear: Section 3.3 lists Baseline plus Interventions I–IV, but Tables 1–2 use '1st/2nd/3rd/4th Intervention' without a consistent mapping; label all conditions the same way.","section":"Tables 1 and 2"},{"comment":"The 90% threshold used to dichotomize low versus high AI confidence is not justified; report a sensitivity analysis or motivate the cutoff.","section":"Section 3.4, AI Confidence Score"},{"comment":"There are multiple typographical errors, including 'mental memand' in Section 4.3 and 'V ariable' in the Table 1 and Table 2 captions; please proofread the manuscript.","section":"General"}],"recommendation":"reject","confidential_remarks":"The performance-measure problem is the decisive issue: because final diagnoses were not elicited on above-neutral agreement trials, the high-confidence overreliance claim cannot be repaired by re-analysis of the existing data. If the authors have access to raw per-trial responses and can show that agreement ratings map to actual decisions, the paper might be reconsidered; as it stands, the central claim is not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline claim—that high AI confidence causes overreliance and reduces diagnostic accuracy—does not survive contact with the paper's own outcome coding. The stress-test note is correct: Section 3.4 treats any agreement rating above neutral as adoption of the AI label, and the participant's own diagnosis is never recorded in those trials. So the reported performance decrease for high-confidence trials is partly mechanical: high confidence nudges agreement upward, and the AI is only ~81% accurate. That is a load-bearing flaw, not a minor quibble.\n\nThat said, the paper has real merits. The research question is timely, the breast cancer CDSS with graduated explainability levels is a reasonable testbed, and the low-confidence effects (decreased trust, decreased agreement, increased diagnosis duration) are based on within-participant comparisons and look plausible. That part is a useful contribution, even if the underlying confidence-trust relationship is known from prior work.\n\nNow the soft spots, in proportion. The abstract says high confidence \"substantially increased trust,\" but the model gives β=0.103, p=0.169—not significant. The demographic analysis appears to treat repeated measures as independent; p-values like 10^-22 from 28 participants are not credible. Table 1 reports a negative group variance for mental demand, which indicates estimation problems. And the Discussion directly contradicts Table 1 by claiming the fourth condition increased mental demand when the coefficient is negative and nonsignificant. These are not deep conceptual flaws—they are fixable with reanalysis and corrected claims—but they matter for the paper's credibility.\n\nThe citation pattern looks fine; self-citations point to earlier work on the same AI system, which is expected. The reference list covers the XAI and trust literature adequately.\n\nWho is this for? Researchers studying trust calibration in AI-assisted clinical decision making. The low-confidence findings and the design idea are worth knowing about, but the paper needs a major revision: reanalyze performance using a measure that records actual human decisions, drop or reframe the overreliance claim, correct the demographic analysis, and release the raw data. I would not cite it in its current form, but I would bring it to a reading group as a case study in how outcome construction can create phantom effects. A serious editor should send this to peer review—the flaws are substantial but diagnosable, and the empirical setting is relevant enough to warrant referee time.","headline":"The high-confidence overreliance finding is likely an artifact of the outcome coding, but the low-confidence effects and the study design make this worth a serious referee.","tokens_in":17465,"tokens_out":2140,"would_cite":false,"duration_ms":22237,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Low AI confidence makes clinicians pause; high confidence breeds overreliance.","keywords":["clinical decision support systems","explainable AI","trust calibration","AI confidence scores","cognitive load","breast cancer diagnosis","automation bias","human-AI interaction"],"falsifier":"Re-run the study with a design where participants must always enter their final diagnosis explicitly, without the agreement-rating shortcut. If the high-confidence performance decrement disappears or reverses when the final diagnosis is recorded directly, the paper's overreliance claim is falsified by its own outcome construction.","tokens_in":16520,"feed_emoji":"🩺","tokens_out":3680,"duration_ms":32799,"temperature":0.7,"pith_summary":"The paper tries to show that the way a clinical AI displays its confidence changes how clinicians behave, not just what they say. In an online experiment with 28 healthcare professionals diagnosing breast ultrasound images, low AI confidence significantly lowered trust and agreement and lengthened diagnosis time, while high confidence slightly but significantly reduced diagnostic accuracy. The authors argue this is overreliance: clinicians followed high-confidence suggestions from a system that was only about 80 percent accurate. They also report that one explainability feature, tumor localization with probability estimates, raised stress. The study matters because confidence displays are a design choice that can either support appropriate reliance or nudge clinicians toward automation bias.","feed_headline":"Low AI confidence slows doctors; high confidence dulls accuracy","feed_subtitle":"In a breast-cancer CDSS trial, low confidence made clinicians cautious; high confidence made them overrely.","key_machinery":"The experimental apparatus is a web-based CDSS built on a U-Net segmentation model and a CNN classifier trained on a public breast ultrasound dataset with 780 images and 81 percent accuracy. The mechanism carrying the argument is the staged explainability design: baseline (no AI), classification only, probability distribution, tumor localization, and enhanced localization with high and low confidence regions. The outcome measures are trust and agreement ratings per image, agreement-above-neutral treated as adopting the AI's label, and two NASA-TLX items (mental demand and stress) after each stage. Mixed-effects models estimate the effect of confidence level while controlling for participant-level variation.","core_discovery":"The central claim is that AI confidence scores act as a trust-calibration signal with asymmetric effects: low confidence makes clinicians more cautious and less accepting of AI advice, whereas high confidence increases trust enough to produce measurable overreliance and a small drop in diagnostic performance. Using an interrupted time series design, the authors compared a no-support baseline with four explainability conditions and modeled trust, agreement, performance, and diagnosis duration as functions of confidence level. The effect sizes are modest — low confidence reduced trust by 0.16 points and agreement by 0.19 points on a 5-point scale and increased diagnosis duration, while high confidence reduced performance by a small but statistically significant margin, which the authors attribute to automation bias rather than to the AI's advice being wrong more often.","pith_inferences":["Editorial inference: the asymmetric confidence effect suggests a practical design rule — degrade or flag low-confidence outputs rather than suppress them, since low confidence already triggers useful caution.","Editorial inference: the performance measurement issue could be resolved by a forced-choice final diagnosis per image; such a design would also let the authors separate agreement from adoption.","Editorial inference: testing the same confidence thresholds with a more accurate, calibrated model would show whether overreliance scales with model accuracy or is a fixed human tendency.","Editorial inference: the stress finding for localization suggests that visual complexity — not just confidence — drives cognitive load, and could be tested by isolating localization with and without probability numbers."],"forward_implications":["Displaying low confidence will make clinicians more skeptical and slower, which may be desirable for error-prone AI but costly in time-critical settings.","Displaying high confidence on a system with imperfect accuracy will nudge a small but real fraction of decisions toward incorrect AI labels, so high-confidence displays should be paired with safeguards.","Tumor localization with probability estimates raises stress without improving mental demand, so explainability features are not neutral additions to the interface.","Age, experience, job role, gender, and race are associated with different perceptions of AI usefulness and complexity, so a single explanation design is unlikely to fit all clinician groups."],"supporting_citations":[{"why":"Provides the prior finding that confidence scores help calibrate trust, which the paper extends to a clinical setting.","marker":"[17]"},{"why":"Shows that displaying system confidence influences trust, forming the basis for the confidence-display manipulation.","marker":"[18]"},{"why":"Demonstrates that dynamic confidence updates reduce automation bias, which the paper cites to interpret cautious behavior under low confidence.","marker":"[59]"},{"why":"Introduces the automation misuse framework that the paper uses to explain high-confidence overreliance.","marker":"[11]"},{"why":"Describes the authors' prior architecture for graduated trust levels, which the CDSS in this experiment builds upon.","marker":"[68]"},{"why":"Supplies the public breast ultrasound dataset of 780 images used to train the classifier, setting the 81 percent accuracy context.","marker":"[73]"},{"why":"Provides the NASA-TLX items used to measure mental demand and stress after each intervention.","marker":"[74]"},{"why":"Provides the interrupted time-series methodology that structures the experimental design.","marker":"[67]"}],"fun_headline_variants":["AI overconfidence breeds bias; underconfidence breeds delay","AI confidence: too low slows clinicians, too high impairs accuracy","Asymmetric AI trust: low confidence triggers caution, high confidence triggers bias","High AI trust: overreliance; low AI trust: overcautiousness and delay"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result that high confidence hurts performance rests on treating an agreement rating above neutral as the participant adopting the AI's diagnosis; if that mapping misrepresents what clinicians actually decided, the overreliance finding could be an artifact of how performance was coded rather than a real behavioral change.","fun_headline_variants_meta":{"raw":{"variants":["AI overconfidence breeds bias; underconfidence breeds delay","AI confidence: too low slows clinicians, too high impairs accuracy","Asymmetric AI trust: low confidence triggers caution, high confidence triggers bias","High AI trust: overreliance; low AI trust: overcautiousness and delay"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0013,"raw_usage":{"total_tokens":5281,"prompt_tokens":899,"completion_tokens":4382,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":4304}},"tokens_in":515,"tokens_out":4382,"duration_ms":26455,"temperature":1.0,"reasoning_tokens":4304,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T11:21:39.410491+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the study with a design where participants must always enter their final diagnosis explicitly, without the agreement-rating shortcut. If the high-confidence performance decrement disappears or reverses when the final diagnosis is recorded directly, the paper's overreliance claim is falsified by its own outcome construction.","supporting_citations":[{"cited_title":"Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making,","cited_arxiv_id":null,"evidence_quote":"Provides the prior finding that confidence scores help calibrate trust, which the paper extends to a clinical setting."},{"cited_title":"Towards improving trust in context-aware systems by displaying system con- fidence,","cited_arxiv_id":null,"evidence_quote":"Shows that displaying system confidence influences trust, forming the basis for the confidence-display manipulation."},{"cited_title":"Supporting trust calibration and the effective use of deci- sion aids by presenting dynamic system confi- dence information,","cited_arxiv_id":null,"evidence_quote":"Demonstrates that dynamic confidence updates reduce automation bias, which the paper cites to interpret cautious behavior under low confidence."},{"cited_title":"Humans and au- tomation: Use, misuse, disuse, abuse,","cited_arxiv_id":null,"evidence_quote":"Introduces the automation misuse framework that the paper uses to explain high-confidence overreliance."},{"cited_title":"Cognitive Architecture and Instruc- tional Design,","cited_arxiv_id":null,"evidence_quote":"Describes the authors' prior architecture for graduated trust levels, which the CDSS in this experiment builds upon."},{"cited_title":"An architecture to support graduated levels of trust for cancer diagnosis with ai,","cited_arxiv_id":null,"evidence_quote":"Provides the NASA-TLX items used to measure mental demand and stress after each intervention."},{"cited_title":"In- terrupted time-series analysis and its application to behavioral data,","cited_arxiv_id":null,"evidence_quote":"Provides the interrupted time-series methodology that structures the experimental design."}],"review_version":1}