{"id":"74c398bf-5fa5-4ac8-a5cb-a4a2736720b9","arxiv_id":"2502.03482","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Radiologists using AI for prostate MRI diagnosis beat their unaided performance but still trail the AI alone, and performance feedback does not close that gap.","lead":"In a study with eight radiologists, AI assistance improved prostate cancer diagnoses over humans alone, but doctors still did worse than the AI by itself because they often ignored its advice. The results suggest that feedback about past performance does not fix this, and that pooling several doctor-AI decisions may beat the AI alone.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ensemble claim exceeds what the study's single-split, shared-case design can support","rationale":"The reader's verdict CONDITIONAL with moderate confidence is appropriate. The strongest-claim analysis is largely correct: the paper's two headline findings are (1) human-AI teams outperform humans alone but underperform AI, and (2) the ensemble of human-AI teams can outperform AI alone. The under-reliance claim is well-supported by the behavioral analysis and is consistent with prior work. My stress-test therefore concentrates on the second claim, which is the most load-bearing because it is the paper's main positive contribution and the basis for the \"promising directions\" conclusion. The concern is not that the analysis is dishonest, but that the statistical evidence for the ensemble advantage is weaker than the paper's language suggests: Table 3 shows a significant advantage in AUROC (p=0.034) and accuracy (p=0.046) in Study 1, but Appendix E's common-50 analysis shows the same effect is not significant (p=0.265 for AUROC; p=0.112 for AUROC in Study 2). Since the common-50 is a subset of the Study 1 test set, the effect size is fragile and depends on the case mix. Additional concerns: N=8 radiologists is small, and the ensemble is a post-hoc aggregate that has not been validated as a deployable workflow. The computational test I propose (leave-one-out or split-half) would directly settle whether the ensemble advantage is robust. The reader's weakest assumption about the feedback confound is also valid, but it is less load-bearing for the paper's central claim because the paper's conclusion explicitly says feedback did not lead to significant improvements; the lack of a control arm makes that null result uninterpretable, but the ensemble claim, if wrong, would overstate the paper's most novel finding. Hence agreement: I agree with the reader's identification of the feedback confound as a real limitation, but I weight the ensemble validation as the more load-bearing concern for the central claim.","tokens_in":23459,"tokens_out":1845,"duration_ms":15696,"concrete_test":"Perform a leave-one-radiologist-out or repeated random split-half validation on the Study 1 predictions: for each of 100 random 50/25 case splits, recompute the H+AI majority-vote ensemble AUROC/accuracy and compare against AI alone on the held-out half. If the ensemble advantage over AI is positive in fewer than, say, 80% of splits, or if the mean held-out advantage is within one standard error of zero, the claim \"ensemble of Human-AI teams can outperform AI alone\" is not supported as a generalizable result.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that \"the majority vote of Human-AI teams can outperform AI alone\" rests on a single realization of the test set with substantial case overlap between Study 1 and Study 2. In Table 3, H+AI ensemble beats AI on Study 1 AUROC (0.771 vs 0.730, p=0.034) and accuracy (73.3% vs 69.3%, p=0.046). But no holdout validation or split-half analysis is reported: the same 75 cases (and the 50 common cases, plus their duplicate diagnoses in Study 2) are used both to select the AI model threshold/architecture and to compute the ensemble advantage. With N=8 radiologists and 75 cases, and with bootstrapped p-values computed on the same cases that define the ensemble, the result has high variance and is vulnerable to overfitting; Appendix E already shows the effect weakens (p=0.265 in Study 1; p=0.112 in Study 2) on the common-50 subset, so the claim's strength depends on the particular case mix. The ensemble is also an unvalidated deployment construct: majority vote over 8 radiologists with confidence-based tie-breaking is a post-hoc aggregation rule, and the paper does not test whether this rule would be robust to a different set of radiologists or a different set of cases. A user of this paper should not conclude that \"majority vote of Human-AI teams can outperform AI alone\" is a stable property without out-of-sample or repeated-split evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports two pre-registered human studies with eight practicing radiologists on AI-assisted prostate cancer MRI diagnosis. Study 1 asks radiologists to give an independent diagnosis, then show the AI prediction, then finalize; Study 2 provides individual performance feedback from Study 1 and shows the AI prediction up front without an independent diagnosis. The authors report that human-AI teams outperform humans alone, still underperform the AI alone, that performance feedback did not significantly improve team performance, that upfront AI increases AI adoption, and that a majority vote of human-AI teams can outperform the AI alone.","tokens_in":23678,"tokens_out":2232,"duration_ms":23010,"significance":"The study is valuable because it brings real domain experts into the human-AI decision-making loop, uses biopsy-confirmed labels, and pre-registers the protocol. The behavioral measurements (agreement, follow/overrule rates, confidence, time) are directly observed and bootstrap testing is applied throughout. If the stated effects hold, they strengthen the empirical basis for claims about under-reliance in expert populations and point to ensemble aggregation as a potentially practical deployment strategy. The main weaknesses are the confounded Study 2 design and the lack of out-of-sample validation for the ensemble claim, both of which are load-bearing for central conclusions, so the current manuscript needs revision before the claims are fully supported.","major_comments":[{"comment":"The conclusion that performance feedback did not lead to significant improvements is not supported by the design, because Study 2 changes both the feedback and the AI presentation timing simultaneously. Participants in Study 2 receive performance feedback and also see the AI prediction up front without making an independent diagnosis, whereas Study 1 requires an initial diagnosis and shows the AI afterward. Any observed null effect on accuracy or any change in AI adoption could be due to either manipulation or their interaction, so the paper should not attribute the result to feedback specifically. A control arm with upfront AI but no feedback, or a factorial design, is needed to disentangle these factors.","section":"§3.4, §4.2"},{"comment":"The claim that the majority vote of human-AI teams can outperform AI alone is based on a single test set with shared cases, and the bootstrapped p-values are computed on the same cases used to form the ensemble. Appendix E (Table 13) shows that on the common-50 subset the effect weakens substantially (e.g., Study 1 H+AI ensemble vs AI AUROC p=0.265, accuracy p=0.216; Study 2 AUROC p=0.112, accuracy p=0.229), which indicates the result is highly dependent on the particular case mix. Without a holdout set, repeated split-half evaluation, or an independent set of radiologists/cases, the paper cannot support the general statement that ensemble voting outperforms AI alone.","section":"§4.1, Table 3, Appendix E"},{"comment":"The statement that human-AI teams 'consistently outperform humans alone' overstates the evidence in Study 2, where the human-alone baseline is taken from Study 1 on a different case set and the common-case paired analysis shows non-significant differences in several metrics (e.g., AUROC p=0.074, specificity p=0.450 in Study 2 common-50). The direction is consistent, but the wording should be qualified to reflect which metrics are significant and which are not, especially since multiple metrics are tested without adjustment.","section":"§4.1, Table 1, Table 2"}],"minor_comments":[{"comment":"The interface description mentions 'BWI' as one of the image sequences, while the rest of the paper and the dataset description refer to 'DWI' (diffusion-weighted imaging); this inconsistency should be fixed.","section":"§3.3"},{"comment":"The running header on the first page reads 'Trovato and Tobin, et al.' which appears to be a template artifact and should be replaced with the manuscript's own title and authors.","section":"Global"},{"comment":"The labels in Fig. 6a are somewhat hard to parse; the text describing the four sub-groups would benefit from a clearer mapping between the diagram boxes and the accuracy values in the text.","section":"§4.2"},{"comment":"The statistical testing section does not state whether any correction for multiple comparisons was applied across the many metrics and conditions; even if no correction is intended, it should be stated explicitly so readers can calibrate the reported p-values.","section":"§3.5"},{"comment":"In the demographics description, 'whilte' should be 'while'.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serious empirical study with a difficult-to-recruit expert population, and the raw behavioral data seem carefully collected. The main concerns are methodological rather than conceptual: the Study 2 confound prevents a causal reading of the feedback result, and the ensemble claim needs out-of-sample validation or must be substantially softened. These are fixable in revision, so I do not recommend rejection, but the current claims outrun the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious look, mostly for the right reasons. It gives us one of the few expert-level human-AI studies in a real clinical task, with practicing radiologists, two realistic workflows, pre-registration, and a careful behavioral analysis of when people follow or overrule the AI. The core ordering—human alone < human+AI < AI—is consistent with prior clinical work, but the specific measurements for prostate MRI and the feedback workflow comparison are new, and they add a useful data point. I also give credit for the shared-case subset analysis and the attention checks on the feedback page.\n\nThe soft spots are where you'd expect them. N=8 is small, and Study 2 confounds performance feedback with a changed workflow (no independent diagnosis, AI shown upfront), so the null feedback result cannot be read causally. The many metrics without multiple-comparison correction also inflate the apparent number of significant results. None of these are fatal for the main under-reliance conclusion, which is directionally consistent with prior studies.\n\nThe real problem is the ensemble claim. The abstract and conclusion say the majority vote of human-AI teams can outperform AI alone, but this rests on a single split with the same 75 cases used to set thresholds and evaluate, and no holdout or repeated-split validation. The paper's own Appendix E shows the effect collapses on the common-50 subset: p=0.265 in Study 1 and p=0.112 in Study 2 for the AUROC comparison against AI. That is not a stable property; it is a post-hoc aggregation on the training set. A reader should not walk away believing \"majority vote beats AI\" is a demonstrated result.\n\nMissing data and code also matters here. The authors should release the interface, the de-identified decisions, and the analysis scripts, at least for the ensemble part. Otherwise the key new claim is not independently checkable.\n\nWho is this for? Researchers working on human-AI decision making in medicine, and people designing AI-assisted radiology workflows. They will find the behavioral results and the study design useful, even if the headline claim needs to be read skeptically.\n\nRecommendation: send it to peer review. It is a legitimate empirical study with a rare expert participant pool, and the main under-reliance finding is likely correct. But the reviewers should push hard on the ensemble claim, the Study 2 confound, and the lack of data/code. If the authors weaken the causal language and validate the ensemble on held-out or repeated splits, this could be a solid contribution.","headline":"Solid expert-level study confirming under-reliance in prostate MRI, but the flashy ensemble claim is not supported by the analysis as presented.","tokens_in":24306,"tokens_out":1471,"would_cite":true,"duration_ms":17170,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that radiologists reading prostate MRI with AI help beat their own unaided accuracy but still trail the AI alone because they override it too often, and that majority-voting across AI-assisted radiologists can surpass…","keywords":["human-AI decision making","under-reliance","domain experts","prostate cancer MRI","radiologist performance","AI assistance","majority-vote ensemble","performance feedback"],"falsifier":"Run the same two workflows in a 2x2 design: with and without performance feedback, and with the AI shown before or after the radiologist's own diagnosis. If a feedback-only arm changes reliance or accuracy as much as the up-front-AI arm does, then the paper's conclusion that performance feedback did not significantly improve human-AI teams would be wrong.","tokens_in":23224,"feed_emoji":"🩺","tokens_out":6357,"duration_ms":61756,"temperature":0.7,"pith_summary":"This paper asks whether trained radiologists use AI assistance appropriately when diagnosing prostate cancer from MRI scans, and whether knowing their own performance makes them use it better. It reports that AI-assisted radiologists outperform unaided radiologists, but their team performance still falls short of the AI alone, a gap the authors attribute to under-reliance: doctors override the AI's advice on a sizable share of disagreements, often incorrectly. Giving radiologists performance feedback and showing the AI prediction up front made them follow the AI more often, but did not significantly improve overall accuracy. The most forward-looking result is that a simple majority vote across eight AI-assisted radiologists can outperform the AI alone, suggesting that collective human-AI decision making can be complementary even when individual assisted readers are not.","feed_headline":"Radiologists underuse AI advice, but teams can beat the model","feed_subtitle":"Eight radiologists reading prostate MRIs beat the AI alone only when their AI-assisted votes are combined.","key_machinery":"The experimental engine is a three-condition comparison with practicing radiologists: an independent human read, an AI-assisted read made after seeing the AI's prediction and lesion annotations, and an AI-first read made after performance feedback. The behavioral mechanism measured is reliance: how often a radiologist's final decision agrees with the AI, how often they override it, and whether the override is correct. The ensemble analysis uses a majority vote over the eight radiologists' final predictions, with ties broken by reported confidence. Under-reliance is defined by the gap between the AI-alone decision and the human-AI final decision: when the radiologist and AI disagree, the radiologist keeps an incorrect independent answer often enough to drag team performance below AI-alone.","core_discovery":"The central claim is that expert human-AI teams in prostate MRI diagnosis occupy a middle band: they are more accurate than the same radiologists working alone, but less accurate than the model by itself because of under-reliance. In Study 1, radiologists made an independent diagnosis, then saw the AI prediction, then finalized their decision; in Study 2, after a memory washout and after receiving feedback on their own, the AI's, and their team's previous performance, they saw the AI prediction before diagnosing. Neither workflow produced team accuracy that reached the AI's standalone performance on the paired common-case comparisons, and aggregate performance feedback did not significantly improve the team. Yet when the eight radiologists' AI-assisted final decisions were combined by majority vote, the ensemble outperformed the AI on AUROC, accuracy, and precision. The paper interprets this as evidence that complementary performance is achievable, but not at the level of the individual assisted reader.","pith_inferences":["Beyond the paper's claims, if under-reliance is structural rather than informational, interventions such as requiring radiologists to justify overrides or defaulting to the AI recommendation may be more effective than performance feedback.","Beyond the paper's claims, the ensemble gains could grow if the radiologists were deliberately diverse in experience or in prostate-zonal expertise, because the error patterns of individual readers would be less correlated.","Beyond the paper's claims, the finding that humans are better at catching the AI's false positives than its missed cases suggests a division-of-labor routing strategy: let the AI flag negatives for human confirmation, while escalating AI-positive cases for extra review.","Beyond the paper's claims, a clinical deployment would need a cost-benefit comparison between AI alone and a human-AI ensemble, since the ensemble requires multiple expert reads per case and the AI model itself is inexpensive to run."],"forward_implications":["Deploying AI as a second reader in prostate MRI improves expert accuracy, but the full accuracy of the AI is not realized unless radiologists follow its advice more closely.","Showing AI predictions before the human commits to a diagnosis increases adoption of AI advice and can improve sensitivity, but by itself does not close the gap to AI-alone performance.","Aggregating multiple AI-assisted reads by majority vote is a promising deployment pattern, since the ensemble can outperform both the average radiologist and the AI alone.","Performance feedback about one's own and the AI's accuracy does not, in this setting, significantly improve human-AI team accuracy.","The behavioral pattern of under-reliance seen in crowdworker studies also appears among board-certified domain experts, suggesting the finding is not an artifact of lay participants."],"supporting_citations":[{"why":"Supplies the AI segmentation model used to generate the diagnostic predictions and lesion annotations shown to radiologists.","marker":"[12]"},{"why":"Provides the annotation-efficient labeling approach used to construct training labels for the AI model.","marker":"[5]"},{"why":"Provides the benchmark dataset and international evaluation context for prostate cancer detection on MRI that the study's test sets draw on.","marker":"[33]"},{"why":"Is the prior expert study that achieved complementary human-AI performance, serving as the main positive comparison point.","marker":"[37]"},{"why":"Is the prior multireader mammography study showing human-AI performance can fall short of AI alone, supporting the under-reliance pattern.","marker":"[15]"},{"why":"Is the prior expert study in chest x-ray diagnosis where AI-assisted physicians did not reach AI-alone performance, reinforcing the authors' comparison.","marker":"[29]"},{"why":"Provides experimental evidence on combining human expertise with AI in radiology and frames the real-world deployment question.","marker":"[1]"}],"fun_headline_variants":["AI-assisted radiologists still lag the model, but teams surpass it","Under-reliance keeps expert-AI teams below AI, yet ensembles win","Prostate MRI: expert-AI teams underuse AI, combined votes beat model","Feedback fails to close expert-AI gap; majority vote beats AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The feedback conclusion rests on comparing Study 1 with Study 2, but Study 2 changed both the feedback and the timing of AI (shown up front, with no independent diagnosis) at the same time, with no feedback-only control arm, so the effect of feedback alone is not isolated.","fun_headline_variants_meta":{"raw":{"variants":["AI-assisted radiologists still lag the model, but teams surpass it","Under-reliance keeps expert-AI teams below AI, yet ensembles win","Prostate MRI: expert-AI teams underuse AI, combined votes beat model","Feedback fails to close expert-AI gap; majority vote beats AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":2982,"prompt_tokens":1034,"completion_tokens":1948,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":1869}},"tokens_in":650,"tokens_out":1948,"duration_ms":14293,"temperature":1.0,"reasoning_tokens":1869,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T14:43:17.320048+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same two workflows in a 2x2 design: with and without performance feedback, and with the AI shown before or after the radiologist's own diagnosis. If a feedback-only arm changes reliance or accuracy as much as the up-front-AI arm does, then the paper's conclusion that performance feedback did not significantly improve human-AI teams would be wrong.","supporting_citations":[{"cited_title":"Annotation-efficient cancer detection with report-guided lesion annotation for deep learning-based prostate cancer detection in bpMRI","cited_arxiv_id":"2112.05151","evidence_quote":"Provides the annotation-efficient labeling approach used to construct training labels for the AI model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the prior expert study that achieved complementary human-AI performance, serving as the main positive comparison point."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the prior multireader mammography study showing human-AI performance can fall short of AI alone, supporting the under-reliance pattern."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the prior expert study in chest x-ray diagnosis where AI-assisted physicians did not reach AI-alone performance, reinforcing the authors' comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides experimental evidence on combining human expertise with AI in radiology and frames the real-world deployment question."}],"review_version":1}