{"id":"42c5a473-2826-4be3-955f-e9c60877dc6a","arxiv_id":"2412.00372","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Showing doctors reference X-rays of the AI's predicted diagnosis improved their accuracy in a small online study, compared with saliency maps or AI-only advice.","lead":"When an AI suggests a chest X-ray diagnosis, this new system also shows four example X-rays of that condition so the doctor can compare and confirm. In a study of 69 physicians, this approach gave the highest diagnostic accuracy when the AI was correct, hinting that simple reference images may improve human-AI teamwork in radiology.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline claim that 2FR increases clinician accuracy is not supported by a reported test of the modality effect; the key confidence intervals for 2FR, Saliency, and AI-only overlap (Fig. 4).","rationale":"I read the paper in good faith. The 2FR idea is clearly motivated and the interface is a sensible contribution to the human-AI decision-making literature. The study is small but the within-subject design is a reasonable first step. My stress test focused on the strongest claim, which is that 2FR improves clinician accuracy relative to the other modalities. The problem is not that the result is impossible; it is that the manuscript never reports the statistical test that would establish it. Section 3.5 states that mixed-effects models were constructed for accuracy, but Section 4.1 only reports a p-value for the effect of AI correctness, not for modality. The figure of overlapping confidence intervals makes the missing test particularly salient: 2FR's 95% CI when AI is correct (0.57–0.81) contains the point estimates of both Saliency (0.65) and AI-only (0.64). With just three trials per modality per participant, the observed advantage could easily be chance variation. Another concern is the confound between AI correctness and case difficulty, which the authors themselves note in Section 4.1; this further weakens any causal reading of the 2FR advantage because hard cases are the same cases on which AI fails. I did not find evidence of fraud or fabrication, and the internal logic of the method is sound. I agree partially with the reader: the reader's formal weakest assumption was the labeling and representativeness of retrieved images, but the more decisive issue is the absence of a reported modality-level statistical test. My concrete proposal—a per-trial mixed-effects logistic regression with a planned 2FR contrast—would directly settle whether the claimed accuracy advantage is real or within sampling noise. If the analysis supports the claim, the paper should be accepted; if it does not, the claim should be softened to a hypothesis-generating result. Since the reader already assigned CONDITIONAL, my concern does not change the verdict.","tokens_in":12233,"tokens_out":2464,"duration_ms":25874,"concrete_test":"Obtain the per-trial data (69 participants × 12 trials, with modality, AI correctness, difficulty, clinician diagnosis, and confidence) and fit a mixed-effects logistic regression with fixed effects for modality, AI correctness, difficulty, specialty, and the modality × AI-correctness interaction, plus a random intercept for participant. Report the omnibus test for modality and the planned contrast of 2FR vs each other arm with 95% confidence intervals, both overall and restricted to AI-correct trials. If the 2FR contrast is not statistically significant (p > 0.05) or the confidence interval includes zero after accounting for within-participant clustering, the headline claim of improved accuracy is not established. If raw data cannot be released, the authors should provide the model output table that Section 3.5 says was fitted for accuracy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that 2FR yields higher clinician accuracy than saliency, AI-only, or no AI. The only statistical result reported for accuracy is that AI correctness predicts accuracy (p<0.001); Section 3.5 promises mixed-effects models for accuracy, but Section 4.1 reports no test for modality as a fixed effect and no pairwise comparison of 2FR against the other arms. Fig. 4 gives the key comparison: when AI is correct, 2FR accuracy is 0.69 (95% CI 0.57–0.81), Saliency is 0.65 (0.52–0.77), and AI-only is 0.64 (0.51–0.76). These intervals overlap substantially, and the design has 69 participants with only 3 trials per modality each, so the 5–6 point advantage over Saliency/AI-only is within sampling noise. Because the central claim rests on this difference, and because AI correctness is deliberately set at 66.7% while the paper itself acknowledges (Section 4.1) that cases with incorrect AI predictions may be inherently harder, the study currently demonstrates a plausible interface idea but not a statistically verified accuracy improvement. Even if all 2FR retrieved images were perfectly labeled, the modality-level claim would still be unsupported without a reported statistical test. The retrieval-label issue identified by the reader is secondary to this more fundamental evidentiary gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new human-AI decision support modality called 2-factor retrieval (2FR), in which an AI prediction is presented together with four reference images purportedly showing the same diagnosis, and compares it with saliency-map explanations, AI-only assistance, and no assistance in a chest X-ray diagnosis task. In an online study with 69 physicians, each completing 12 cases (three per modality), the authors report that 2FR yields the highest overall accuracy (about 70% when the AI is correct) and claim particular benefits for radiologists and for low-confidence decisions. The paper argues that 2FR enables a verification-based reasoning step that improves human-AI decision making.","tokens_in":12449,"tokens_out":3377,"duration_ms":31460,"significance":"If the central claim were solidly established, the 2FR idea would be a valuable, low-cost contribution to the human-AI decision-making literature: it requires no new model training and could be paired with any classifier. The study also addresses a relevant gap by evaluating verification-based interfaces against explainability methods in a medical domain. However, the current evidence base is too weak to support the headline accuracy claim: no statistical test of the modality effect is reported, the key confidence intervals overlap broadly, and the subgroup analyses are severely underpowered. The idea is promising, but the paper as written does not meet the evidentiary standard needed to recommend 2FR over existing interfaces.","major_comments":[{"comment":"The central claim that 2FR increases clinician accuracy over the other modalities is not supported by any reported statistical test of the modality effect. Section 3.5 states that mixed-effects models for accuracy were constructed, but Section 4.1 reports only the effect of AI correctness (p<0.001) and no pairwise comparisons among 2FR, Saliency, AI-only, and No AI. In Fig. 4, the 95% confidence intervals for 2FR (0.57–0.81), Saliency (0.52–0.77), and AI-only (0.51–0.76) when the AI is correct overlap substantially, and with N=69 and only three trials per modality per participant, the observed point differences (0.69 vs. 0.65 vs. 0.64) are within sampling noise. Without a reported test, the manuscript's headline claim is not statistically verified.","section":"Section 4.1 and Fig. 4"},{"comment":"The 2FR intervention is underspecified. The text states that four images 'recognized by other physicians to represent that diagnosis' were retrieved, but it does not describe the retrieval algorithm, the source pool from which the images were drawn, or any verification procedure for ensuring that the retrieved images are correctly labeled and representative. Because the mechanism of 2FR depends on the retrieved images being valid examples of the AI-predicted pathology, failing to specify this makes the comparison difficult to interpret and impossible to reproduce. The authors should describe how the reference images were selected and what quality checks were applied to them.","section":"Section 3.4"},{"comment":"The subgroup claims (e.g., '2FR significantly aids expert users,' '2FR is most useful for clinicians with less experience,' and the low-confidence advantage shown in Fig. 10) are not backed by any reported statistical tests for interaction or subgroup effects. With 25 radiologists and 44 non-radiologists, and only three trials per modality per participant, the cell sizes for the experience-by-modality and confidence-by-modality analyses are very small; the differences highlighted in the text (e.g., the 3x/2x low-confidence comparison) could be driven by a small number of responses. The manuscript should report interaction tests or explicitly label these analyses as exploratory.","section":"Section 4.1 and Figs. 5–7"},{"comment":"The interpretation that 'physicians are overly trusting of AI predictions' is confounded by case difficulty. The authors acknowledge that cases with incorrect AI predictions might be inherently harder, but the conclusion that clinicians over-trust AI is drawn from observing that accuracy tracks AI correctness. Because AI correctness is a property of the case, not of the AI assistance, and is deliberately set at 66.7%, this correlation alone does not demonstrate over-reliance. The alignment-based measure described in Section 3.7 (the correlation between physician and AI diagnoses) would be a more direct test, but it is not reported.","section":"Section 4.1"}],"minor_comments":[{"comment":"The 14 diagnosis options given to participants are not listed; the full set of choices should be provided so readers can assess the difficulty and granularity of the task.","section":"Section 3.4"},{"comment":"The 'Easy' versus 'Hard' case labels are not defined; the manuscript does not state who assigned these ratings or what criteria were used, which is important because the paper interprets accuracy differences across these categories.","section":"Section 3.4"},{"comment":"The claim that clinician confidence is 'marginal' across conditions is stated without reporting the mixed-effects model results for confidence that are promised in Section 3.5; the authors should include the relevant statistical output or state that confidence analyses were descriptive only.","section":"Section 4.2"},{"comment":"Fig. 3 would benefit from error bars, and Fig. 4 is presented as a table rather than a figure; the manuscript should be consistent in labeling and should state explicitly that Fig. 4 reports 95% confidence intervals.","section":"Fig. 3 and Fig. 4"}],"recommendation":"major_revision","confidential_remarks":"The paper's central empirical claim outruns its reported statistics. The authors likely have the underlying data to run the mixed-effects models described in Section 3.5, including modality as a fixed effect and pairwise contrasts, so the gap is fixable within the manuscript's scope. I would also encourage the editor to ask for a complete description of the retrieval procedure, as the current description is too vague for replication. The promise of the 2FR idea is real, but the current version does not establish it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The 2FR idea is worth stealing: showing clinicians canonical examples of the AI's predicted diagnosis is a simple, model-agnostic interface twist that could plausibly aid verification. The study is clearly presented, the sample of 69 physicians is decent for a pilot, and they report confidence intervals rather than hiding them. They also honestly note in Section 4.1 that cases where the AI is wrong may be inherently harder, which is a real confound they at least flag.\n\nBut the headline claim—that 2FR increases clinician accuracy—is not backed by the reported statistics. The methods section (3.5) promises linear mixed-effects models for accuracy, but the results section (4.1) only reports a test of AI correctness, not a test of modality as a fixed effect. The key numbers in Fig. 4 tell the story: when the AI is correct, 2FR accuracy is 0.69 (95% CI 0.57–0.81), saliency is 0.65 (0.52–0.77), and AI-only is 0.64 (0.51–0.76). Those intervals overlap heavily. With three trials per modality per physician, the 4–5 point gap is within sampling noise. The subgroup claims about radiologists and low-confidence decisions are even thinner—I would treat them as hypothesis-generating, not evidence.\n\nA second soft spot is the retrieval pipeline. The paper says the four reference images are \"physician-confirmed pathology\" (Fig. 2), but it never specifies the retrieval algorithm or how the confirmation was done. If those images are mislabeled or atypical, the 2FR arm is not actually testing the proposed method. This is a reproducibility issue as much as a validity issue.\n\nWho is this for? People working on human-AI decision support in medical imaging. It's a reasonable pilot that could seed a larger, preregistered study. It deserves serious peer review, but the manuscript needs major revision: a prespecified analysis with pairwise modality comparisons (corrected for multiplicity), a larger sample, and full disclosure of the retrieval and verification process. I'd send it to reviewers, but with the clear expectation that the statistical gap gets fixed before publication.","headline":"The 2FR idea is simple and worth testing, but the paper's central claim outruns its statistics: no modality-level test is reported and the confidence intervals overlap.","tokens_in":13036,"tokens_out":2296,"would_cite":false,"duration_ms":23436,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a 2-factor retrieval interface—an AI diagnosis plus four physician-confirmed reference images—raises clinician accuracy on chest X-rays above saliency maps, AI-only, and no-AI baselines, reaching roughly 70 percent…","keywords":["2-factor retrieval","verification-based AI","human-AI decision making","chest X-ray diagnosis","explainable AI","saliency maps","clinician confidence","human-computer interaction"],"falsifier":"Run the same 12-case study but replace the four physician-confirmed reference images with randomly selected same-label images; if 2FR's accuracy advantage over saliency disappears, the benefit is driven by the specific canonical examples rather than by the verification process itself.","tokens_in":12002,"feed_emoji":"🩻","tokens_out":6202,"duration_ms":51118,"temperature":0.7,"pith_summary":"The paper's central claim is that how an AI prediction is presented to clinicians changes how much it helps. It introduces '2-factor retrieval' (2FR): alongside the AI-predicted diagnosis, the clinician sees four physician-confirmed chest X-ray examples of that same diagnosis. In a study with 69 physicians, 2FR produced the highest diagnostic accuracy of any tested mode, about 70 percent when the AI was correct, outperforming saliency maps, AI-only, and no-AI conditions. The authors interpret this as evidence that a verification-based interface helps clinicians recall and compare the features of a pathology, rather than merely trusting the AI's answer. If true, the result would argue for retrieval-based design in clinical decision support systems.","feed_headline":"Reference images lift AI-assisted chest X-ray accuracy to 70%","feed_subtitle":"Pairing AI predictions with physician-confirmed examples beat saliency, AI-only, and no-AI modes.","key_machinery":"The central object is '2-factor retrieval' (2FR), an interface plus retrieval scheme that, given an AI-predicted diagnosis, displays four physician-confirmed reference images of that diagnosis without processing them through the model. It is called 2-factor because the correct images must be retrieved by the AI and the clinician must associate those images with the pathology in the current X-ray, forming a verification loop. The comparison conditions are saliency maps from the same chest X-ray model, AI-only prediction, and no AI; participants chose one of 14 diagnoses and rated confidence on a 10-point scale. Linear mixed-effects models with a random participant intercept estimate how modality, AI correctness, difficulty, specialty, and experience affect accuracy and confidence.","core_discovery":"On its own terms, this paper discovers that presenting an AI diagnosis alongside four physician-confirmed reference images of that diagnosis yields the highest clinician accuracy among the tested modes of AI assistance. Across 69 physicians reading 12 chest X-rays, 2FR reached about 70 percent accuracy when the AI prediction was correct, compared with roughly 65 percent for saliency maps, 64 percent for AI-only, and 45 percent for no AI. The advantage was strongest for radiologists and for clinicians with less than 11 years of practice, and 2FR was the best-performing modality at low confidence, roughly tripling saliency's accuracy in that subgroup. When the AI was wrong, all modalities fell to about 25 percent, similar to no AI, so the benefit is tied to the AI being right. The authors conclude that a verification-based presentation can improve human-AI diagnostic performance without changing the AI model.","pith_inferences":["A straightforward test left implicit in the paper is to vary the quality of the retrieved set (random same-label images versus canonical examples); this would separate the verification mechanism from simple anchoring on any same-label exemplar.","If the effect is driven by canonical examples, 2FR could be combined with model confidence or uncertainty estimates so that reference images are only shown when the AI prediction is likely to be correct.","The same retrieval-based interface could be tried in other visual diagnostic domains, such as dermatology or ophthalmology, although the paper only reports chest X-rays.","The paper's decoupling of confidence from accuracy suggests that future studies should track decision time and verification behavior, not just final diagnosis, as a predictor of when reference images help."],"forward_implications":["Clinical decision support systems could raise diagnostic accuracy by changing only the interface, not the AI model, if 2FR's effect replicates.","Low-confidence clinicians, who otherwise struggle most, are the clearest beneficiaries of 2FR-based presentation.","2FR does not compensate for an incorrect AI; accuracy under 2FR falls to no-AI levels when the AI errs, so the method inherits the AI's reliability.","Task experts such as radiologists gain more from reference images than from saliency maps, while saliency maps were less helpful for experts than for non-experts.","Because confidence stays roughly flat while accuracy moves, self-assessed confidence is not a reliable proxy for whether AI assistance helped."],"supporting_citations":[{"why":"Supplies the chest X-ray images, labels, and the saliency-map baseline used in the experiment.","marker":"[49]"},{"why":"Motivates the claim that explanations only help when they support verification of the answer.","marker":"[12]"},{"why":"Documents overreliance on AI as the dominant human-AI error and motivates forcing verification.","marker":"[5]"},{"why":"Shows that explanation design affects overreliance, setting up the interface comparison.","marker":"[47]"},{"why":"Provides a prior result that non-task-expert physicians benefit from correct explainable AI advice on X-rays, which this study extends.","marker":"[13]"},{"why":"Documents clinician overreliance in clinical decision-aids, the failure mode 2FR is designed to counter.","marker":"[14]"},{"why":"Provides survey evidence about clinician preferences and AI correctness standards, used to interpret radiologists' results.","marker":"[38]"}],"fun_headline_variants":["X-ray diagnoses get a boost when AI shows doctor-confirmed examples","Why radiologists may trust AI more with 2-factor retrieval","AI plus human-checked reference images bests other explainable tools","New 2-factor method lifts chest X-ray accuracy, especially for radiologists","When AI is right, showing physician-approved examples boosts diagnosis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison's validity depends on the four reference images shown for each AI prediction being correctly labeled and genuinely typical of the predicted pathology, yet the paper does not describe the retrieval algorithm or the verification of these images.","fun_headline_variants_meta":{"raw":{"variants":["X-ray diagnoses get a boost when AI shows doctor-confirmed examples","Why radiologists may trust AI more with 2-factor retrieval","AI plus human-checked reference images bests other explainable tools","New 2-factor method lifts chest X-ray accuracy, especially for radiologists","When AI is right, showing physician-approved examples boosts diagnosis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000403,"raw_usage":{"total_tokens":2087,"prompt_tokens":919,"completion_tokens":1168,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":1078}},"tokens_in":535,"tokens_out":1168,"duration_ms":8724,"temperature":1.0,"reasoning_tokens":1078,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:26:58.598048+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 12-case study but replace the four physician-confirmed reference images with randomly selected same-label images; if 2FR's accuracy advantage over saliency disappears, the benefit is driven by the specific canonical examples rather than by the verification process itself.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that explanation design affects overreliance, setting up the interface comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a prior result that non-task-expert physicians benefit from correct explainable AI advice on X-rays, which this study extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents clinician overreliance in clinical decision-aids, the failure mode 2FR is designed to counter."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides survey evidence about clinician preferences and AI correctness standards, used to interpret radiologists' results."}],"review_version":1}