{"id":"66301165-2ede-4ee8-8b93-065efc95bec6","arxiv_id":"2412.11298","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In a 28-clinician experiment, higher levels of AI explanation did not improve trust or diagnostic accuracy in breast cancer diagnosis and sometimes reduced both.","lead":"This paper reports an experiment with 28 clinicians who diagnosed breast cancer images with and without AI suggestions at four levels of explanation. It finds that adding explanations did not reliably improve trust or accuracy, and the simplest AI suggestion was often the best.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The fixed intervention order (Section 3.1) makes explanation level and session order perfectly collinear, so the significant accuracy drops for the 3rd and 4th interventions may reflect order effects, not explainability.","rationale":"The reader's weakest_assumption correctly identifies the fixed intervention order as a load-bearing threat to the causal interpretation. The mixed-model coefficients comparing later interventions to the first are unidentifiable with respect to order because every participant received the same sequence, so any monotonic learning, fatigue, or motivational trend is entirely aliased with explanation level. This is the most damaging issue because it directly undermines the specific claim that accuracy significantly declines for the 3rd and 4th explanation levels, as well as the recommendation to adopt the simplest intervention. I considered other potential issues—multiple comparisons, the imputed agreement-based performance measure, and small sample size—but these either inflate or affect precision without fully invalidating the causal contrast. The order confound, by contrast, makes the central comparison non-causal regardless of statistical correction. A counterbalanced design is the decisive test, and until that is done, the paper's conclusions must be treated as provisional. The existing CONDITIONAL verdict appropriately requires the authors to address this or soften their claims, so I do not recommend a harsher verdict. The agreement is full because the reader's weakest_assumption is exactly the concern I would emphasize.","tokens_in":13686,"tokens_out":8694,"duration_ms":76211,"concrete_test":"Run a replication with a Latin-square counterbalanced order of the four AI interventions across participants (n ≥ 28), using the same ten images per condition and the same underlying AI suggestions; test whether the 3rd and 4th interventions still yield significantly lower diagnostic accuracy than the 1st. If the pattern does not replicate, the original finding is an order artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All 28 participants completed the baseline and four interventions in the same increasing-explainability sequence, with no counterbalancing or randomization of order (Section 3.1). The mixed-effects model in Table 5 uses the 1st intervention as the reference category, so the coefficients for the 3rd intervention (-0.073, p = 0.005) and 4th intervention (-0.068, p = 0.009) absorb any global temporal trend—fatigue, boredom, growing familiarity with the AI, or changes in attention—because order and intervention are perfectly collinear. Since the 1st intervention always occurs early in the session and the 3rd/4th occur later, the observed declines in diagnostic accuracy and the corresponding recommendation that the 1st intervention is ideal (Section 5.1) cannot be causally attributed to the level of explanation rather than to the passage of time. This confound threatens the paper's central claim that increasing explanations can hurt performance; the current data are equally consistent with a simple time-on-task effect.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an online experiment with 28 clinicians who performed breast cancer diagnosis tasks under five conditions: a no-AI baseline and four CDSS interventions with increasing levels of explainability (classification only, plus probability distribution, plus tumor localization, plus confidence levels). The authors analyze self-reported trust, understandability, perceived accuracy, and behavioral measures (diagnostic accuracy, agreement, decision time) using mixed-effects models, chi-square tests, and ANOVA. They report that higher explainability does not monotonically improve trust or accuracy, with Interventions III and IV showing significantly lower diagnostic accuracy than Intervention I, and they conclude that Intervention I (the simplest design) is ideal. Demographic analyses suggest gender and experience relate to self-reported AI familiarity but not to behavioral measures.","tokens_in":14007,"tokens_out":6188,"duration_ms":56568,"significance":"If the causal claims were supported, the study would make a useful contribution to the human-AI interaction and XAI literature, particularly for clinical decision support. The use of a realistic diagnostic task, multiple trust measures (self-report and behavioral), and a clinician sample are strengths. The paper also explicitly distinguishes self-reported from behavioral outcomes, which is valuable. However, the study is not pre-registered, does not provide code or data, and—as detailed below—its central causal inference is undermined by a design confound that cannot be repaired with the collected data. The significance of the contribution therefore depends on whether the authors can reframe the results as descriptive rather than causal.","major_comments":[{"comment":"The fixed intervention order is a fundamental confound. All participants receive the baseline and then Interventions I through IV in the same increasing-explainability sequence (Section 3.1). The mixed-effects model in Table 5 uses Intervention I as the reference category and includes no term for time or session order. Because intervention level and session order are perfectly collinear within every participant, the coefficients for Interventions III and IV (diagnostic accuracy -0.073, p=0.005; -0.068, p=0.009) absorb any fatigue, practice, boredom, or attention effects. The conclusion in Section 5.1 that 'the 1st intervention with the simplest design would be the ideal choice' is therefore not causally supported; the observed declines may simply reflect time-on-task. Since every participant followed the same order, the design offers no way to separate order from intervention, and adding a time covariate would make the intervention effects unidentifiable. This threatens the paper's central claim that increasing explanations can hurt performance.","section":"Section 3.1 / Section 4.2 / Section 5.1"},{"comment":"Diagnostic accuracy is not measured directly for trials where agreement is above neutral. The text states that if the participant's agreement level is higher than neutral, 'we considered their decision to align with the AI's suggestion'; only when agreement is below neutral are participants asked to provide their own decision. This creates a mechanical dependency: agreement and performance are partly the same variable. In the 4th intervention, agreement is significantly lower (Table 5, -0.149, p=0.048), so a larger fraction of trials enters the performance calculation as actual (and possibly less accurate) human decisions rather than imputed AI decisions. The observed accuracy decline in Interventions III and IV could therefore be an artifact of this measurement rule rather than a true change in decision quality. The authors should either collect explicit decisions on every trial or model performance conditional on the agreement threshold.","section":"Section 3.3.2"},{"comment":"The claim that accuracy when using AI is 'generally higher compared to the baseline intervention without AI, all improvements being statistically significant' is not supported by any statistic shown. Table 5 only contrasts Interventions II, III, and IV with Intervention I; the baseline condition is not included in the mixed-effects model reported. Without reporting the coefficients or tests for the baseline comparison, this claim is unverifiable and should be either documented or removed.","section":"Section 4.2 (Performance paragraph)"}],"minor_comments":[{"comment":"There is an inconsistency between the text and Table 5: the text reports the 3rd Intervention trust coefficient as p = 0.187, while Table 5 gives p = 0.465. One of these is a typo and should be corrected.","section":"Section 4.2 (Trust paragraph)"},{"comment":"The chi-square tests are based on 28 participants with many categories (e.g., five age brackets, three experience brackets). Several expected cell counts are likely below 5, which violates the assumptions of the chi-square test. The authors should report expected counts or use an exact test.","section":"Section 4.1 / Table 3"},{"comment":"No correction for multiple comparisons is applied across the many mixed-effects tests, chi-square tests, and ANOVAs. Some p-values near 0.05 (e.g., perceived accuracy for the 4th intervention, p=0.048; agreement for the 4th intervention, p=0.048) may not survive a Benjamini-Hochberg correction. A pre-specified analysis plan or explicit multiplicity adjustment would strengthen the results.","section":"Section 4.2 / Table 5"},{"comment":"The limitations section does not acknowledge the fixed intervention order or the agreement-based performance imputation, both of which are major threats to causal interpretation. These should be explicitly discussed.","section":"Section 5.2"},{"comment":"There are several language issues: 'such breast cancer' in the Abstract, 'regrad' in the Conclusion, and 'The clinicians were interaction' in Section 5.2. These need copyediting.","section":"Abstract / Conclusion / Section 5.2"},{"comment":"Some reference entries include 'Publisher:' metadata and formatting is inconsistent across entries (e.g., [1], [12], [21]). Ensure a consistent citation style.","section":"References"}],"recommendation":"reject","confidential_remarks":"This paper addresses a timely and important question, but the fixed-order design is a fundamental threat to the central causal claim. Because every participant experiences the same sequence, order and explanation level are perfectly collinear, and the current data cannot separate them. The agreement-based performance imputation adds a second, related problem. I do not see how the stated conclusions can be supported without a new experiment that randomizes or counterbalances intervention order and records explicit decisions on every trial. The authors might consider resubmitting a substantially revised manuscript that reframes the findings as descriptive observations and clearly labels the causal claims as hypotheses for future work, but as it stands, the evidence does not support the paper's title and abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real empirical study with actual clinicians, and the new data are the main contribution. The headline finding—that more explanation doesn't always help and may hurt—isn't new (Ahn et al., Wang & Yin already showed this), but applying it to breast cancer diagnosis with four graduated explanation levels is a useful addition. The paper is well-structured and honest about several limitations.\n\nThe biggest problem is the fixed intervention order. All 28 clinicians go through baseline then the four interventions in the same always-increasing-explainability sequence (Section 3.1). That makes intervention level and session order perfectly collinear. The mixed-effects model uses Intervention 1 as the reference, so the significant negative coefficients for Interventions 3 and 4 (-0.073 and -0.068 for accuracy) could just be fatigue, boredom, or growing familiarity with the interface. The paper's recommendation that Intervention 1 is 'the ideal choice' (Section 5.1) is therefore not well supported. This doesn't kill the broad conclusion that explanations don't monotonically improve things—that pattern is consistent even if order effects are present—but it does mean the causal language should be dialed back or the design acknowledged as a major limitation.\n\nSecond issue: performance is partly imputed. When participants rate agreement above neutral, the paper assumes they followed the AI recommendation. Actual decisions are only recorded when agreement is below neutral. So the performance metric is a mix of recorded and assumed decisions. The paper doesn't test sensitivity to the neutral threshold. That's a moderate concern.\n\nSmaller stuff: no multiple-comparison correction across the many outcomes, and n=28 is small (though typical for clinician studies). The gender/age/experience findings on self-reported measures are interesting but not central. The citation pattern looks honest; they cite Ahn, Wang & Yin, and self-cite for the model's accuracy, which is appropriate.\n\nOn balance, the paper is worth engaging with. The new dataset and the clinical context make it a candidate for peer review, but it needs a careful revision to address the confound—either by reframing the results as descriptive, adding a time covariate (which can't fully fix collinearity), or at least acknowledging the order effect explicitly in limitations. The imputed performance measure also needs a robustness check.\n\nWho's it for: HCI and clinical informatics researchers. I'd send it to a serious referee, with the expectation of major revision.","headline":"New clinical dataset, but the fixed intervention order confounds the main finding; deserves revision rather than rejection.","tokens_in":14381,"tokens_out":2660,"would_cite":true,"duration_ms":25675,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A study of 28 clinicians finds that adding explanation layers to an AI breast-cancer diagnosis tool can reduce diagnostic accuracy, and the simplest interface is the best.","keywords":["explainable AI","clinical decision support","clinician trust","diagnostic accuracy","breast cancer","human-AI interaction","user study","explanation design"],"falsifier":"Run the same experiment with the four explanation conditions presented in randomized or counterbalanced order across clinicians. If the diagnostic accuracy decline at the third and fourth explanation levels disappears or reverses under random ordering, the fixed sequence, not the explanation content, was the cause.","tokens_in":13518,"feed_emoji":"🩺","tokens_out":6156,"duration_ms":53216,"temperature":0.7,"pith_summary":"This paper tries to establish that increasing the amount of explanation an AI gives clinicians does not automatically improve trust or diagnostic accuracy, and can make both worse. In a study of 28 clinicians diagnosing breast tissue images, adding probability distributions and tumor localization to a classification-only recommendation significantly reduced diagnostic accuracy, and the richest explanation condition also lowered understandability, perceived accuracy, and agreement while increasing decision time. The authors conclude that any AI decision support improves accuracy over no support, but that the simplest explanation design is the best choice. The result matters because explanation richness is often treated as a trust-building feature, and the paper gives evidence that it can backfire.","feed_headline":"More AI explanation didn't help clinicians and sometimes hurt accuracy","feed_subtitle":"In 28 clinicians, tumor-localization and confidence explanations reduced accuracy versus classification alone.","key_machinery":"The load-bearing mechanism is a four-rung explanation ladder built on top of the same AI recommender: (1) class prediction only, (2) class prediction with a probability distribution, (3) added tumor localization, and (4) added low/high confidence localization. Each clinician saw all four rungs in that fixed order, with a no-AI baseline first. A mixed-effects model with intervention as a fixed effect and participant as a random grouping is used to compare the later interventions to the first, attempting to isolate the effect of added explanation on trust, accuracy, agreement, decision time, understandability, and perceived accuracy.","core_discovery":"The central claim is that explainability is not a monotonic good in AI-assisted breast cancer diagnosis. Using an interrupted time-series design in which 28 clinicians diagnosed breast tissue images under no support, then under classification-only, then classification plus probabilities, then plus tumor localization, then plus confidence levels, the authors found that the third and fourth explanation levels produced statistically significant drops in diagnostic accuracy relative to the first, with coefficients of -0.073 (p = 0.005) and -0.068 (p = 0.009). The fourth level also significantly reduced understandability, perceived accuracy, and agreement and lengthened decision time, while trust did not change significantly across interventions. Overall, all AI-assisted conditions outperformed the no-AI baseline, and the authors single out the classification-only interface as the ideal design.","pith_inferences":["Because all clinicians saw the explanation levels in the same order, the causal reading depends on an assumption the paper does not test; a counterbalanced replication would separate explanation level from fatigue or practice effects.","The paper does not vary the accuracy of the underlying AI; with a more accurate system, richer explanations might not show the same decline, so the 'simplest is best' recommendation is bounded to this model and task.","A practical extension would replace the fixed ladder with adaptive explanations that give more detail only on request, and measure whether that restores the trust gains without the accuracy loss."],"forward_implications":["Clinicians' diagnostic accuracy can decline when an AI support system adds probability and localization details on top of a plain classification recommendation.","Trust, understandability, and perceived accuracy do not reliably improve as explanation richness increases, so explanation design should be treated as a trade-off rather than a monotonic benefit.","The presence of an AI recommendation, even without explanations, improves diagnostic accuracy over relying on clinical judgment alone in this setting.","Self-reported demographic differences, such as gender differences in AI familiarity, need not translate into differences in actual trust or diagnostic performance."],"supporting_citations":[{"why":"Provides the interrupted time-series design used for the fixed-order multi-intervention experiment.","marker":"[41]"},{"why":"Supplies the AI architecture (U-Net plus CNN) on which the CDSS recommendations are built.","marker":"[42]"},{"why":"Supplies the U-Net segmentation method used to localize tumors in the tissue images.","marker":"[43]"},{"why":"Supplies the 780-image breast ultrasound dataset on which the AI system was trained.","marker":"[44]"},{"why":"Provides the self-reported and behavioral trust measurement categories used in the study.","marker":"[22]"},{"why":"Motivates the study as the closest prior work, which qualitatively assessed xAI tools in pathology without measuring trust and performance.","marker":"[40]"}],"fun_headline_variants":["More AI explanations hurt clinician accuracy in breast cancer diagnosis","Extra AI explanations reduce diagnostic accuracy for clinicians","Explainability can backfire: Extra AI details lower accuracy","Classification-only AI outperforms added explanations","Study: Extra AI explanations lower breast cancer diagnostic accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparisons assume that the fixed order in which all clinicians saw the explanation levels did not itself affect performance, so fatigue, practice, or learning are not responsible for the declines attributed to richer explanations.","fun_headline_variants_meta":{"raw":{"variants":["More AI explanations hurt clinician accuracy in breast cancer diagnosis","Extra AI explanations reduce diagnostic accuracy for clinicians","Explainability can backfire: Extra AI details lower accuracy","Classification-only AI outperforms added explanations","Study: Extra AI explanations lower breast cancer diagnostic accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001215,"raw_usage":{"total_tokens":4948,"prompt_tokens":841,"completion_tokens":4107,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":4035}},"tokens_in":457,"tokens_out":4107,"duration_ms":26856,"temperature":1.0,"reasoning_tokens":4035,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:04:14.582873+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same experiment with the four explanation conditions presented in randomized or counterbalanced order across clinicians. If the diagnostic accuracy decline at the third and fourth explanation levels disappears or reverses under random ordering, the fixed sequence, not the explanation content, was the cause.","supporting_citations":[],"review_version":1}