{"id":"ba6322f0-add2-4e3a-978e-dfba69e27223","arxiv_id":"2607.25933","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new 1,089-case multi-turn multimodal benchmark shows that even top AI medical diagnosticians are fully correct only ~34% of the time and frequently hallucinate reasoning.","lead":"Researchers built ClinMM-Bench, a set of 1,089 real-world clinical cases with 3,760 images, to test how well multimodal AI models diagnose patients when information is revealed step by step. Testing 15 models shows that even the best ones get fully correct diagnoses in only about a third of cases, and often produce unreliable reasoning.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No contamination analysis for the public PMCOA case reports; headline scores and rankings may reflect memorized case outcomes rather than multi-turn diagnostic reasoning.","rationale":"The reader's verdict is CONDITIONAL and the weakest assumption already identifies contamination of the public PMCOA case reports as part of the core threat. My pass isolates contamination as the single most load-bearing concern because the benchmark's central claim depends on the scores reflecting multi-turn diagnostic reasoning rather than parametric memory. The data source is public, the models are trained on public text, and no leakage check is reported anywhere in the manuscript. A temporal/recency test is a concrete, feasible way to settle whether contamination is present: if scores on older, more likely memorized cases are systematically higher than on recent cases, the reported rankings and absolute accuracies are not valid evidence for the paper's interpretation. I do not see a reason to move the verdict: the qualitative finding that completely correct diagnoses are rare would likely remain directionally true even under contamination, but the exact numbers and the proprietary-versus-open-weight ranking would need revision. The paper's expert validation and multi-stage curation are worthwhile but do not address the memorization threat. No formal verification or independent support is present. The code repository is listed, but the benchmark data itself is not released as a versioned artifact, which further limits independent contamination checks; this is consistent with the reader's conditional assessment.","tokens_in":28699,"tokens_out":8716,"duration_ms":92810,"concrete_test":"For each case, record its PMCOA publication date and compute each model's mean diagnostic accuracy and complete-correct rate by publication-date quartile. Then recompute the headline scores and rankings using only the most recent quartile, i.e., the cases least likely to have appeared in any model's pretraining data. If accuracy declines monotonically with recency, or if GPT-5-medium's advantage and the proprietary/open-weight gap shrink or disappear on the recent subset, then contamination is present and the reported numbers should not be interpreted as measuring multi-turn reasoning. To make the test decisive, restrict to cases published after each model's documented training cutoff where available (e.g., Gemma-3-27B, Qwen3-VL-32B).","verdict_should_be":"UNCHANGED","load_bearing_attack":"ClinMM-Bench is built from PMCOA case reports published before September 1, 2025 (§4.1), and the 15 evaluated models (§4.2) were trained on web-scale corpora that almost certainly include PubMed Central. The paper contains no contamination analysis: no membership-inference test, no temporal holdout, no n-gram overlap check, and no canary strings. The Discussion's limitation paragraph only says 'potential biases may still be introduced during data curation and automated evaluation'; it does not address memorization. Because the final diagnosis is deliberately removed from the case presentation during conversion (§4.1.4), a model that recognizes a case from pretraining can output the correct diagnosis without performing the multi-turn multimodal reasoning the benchmark claims to measure. This would inflate the absolute accuracy figures (e.g., GPT-5-medium's 33.88% completely-correct rate) and would confound the proprietary-vs-open-weight comparison, since proprietary models typically train on larger, more recent corpora and may memorize public case reports at higher rates. The qualitative conclusion that completely correct diagnoses are limited might survive, but the benchmark's stated interpretation—as a measure of diagnostic reasoning rather than memorization—would not be supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ClinMM-Bench, a multi-turn multimodal diagnostic benchmark built from 1,089 real-world PMCOA case reports and 3,760 medical images across eight specialties. Each case is converted into a progressive multi-turn dialogue in which clinical information and images are disclosed over time, and the final diagnosis is removed from the case presentation. The authors evaluate 15 MLLMs using a two-level framework: diagnostic accuracy is scored by a dual-LLM consensus mechanism (GPT-5-medium and Claude-4.5-Sonnet), and reasoning quality is scored via atomic-fact decomposition (fact recall, hallucination, fact density). Headline results include GPT-5-medium achieving the highest accuracy score (1.140) but only 33.88% completely correct diagnoses; proprietary models generally outperform open-weight models; and reasoning quality is limited, with fact recall below 0.60 and hallucination scores between 0.085 and 0.237. The paper also compares medical vs. general models and reasoning vs. non-reasoning variants, and identifies five failure modes.","tokens_in":28953,"tokens_out":3735,"duration_ms":36199,"significance":"If the evaluation is valid, this is a substantial contribution: it is, to my knowledge, the largest multi-turn multimodal diagnostic benchmark, with a detailed six-stage curation pipeline, expert validation, two-level scoring that goes beyond correctness, and a clinically meaningful error taxonomy. The code is publicly released, and the use of real-world case reports with human-derived ground-truth diagnoses gives the benchmark external grounding. However, the central empirical claims rest on the validity of the automated evaluation, and the manuscript currently lacks the analyses needed to support that validity: no contamination control for public pretraining data, self-judging evaluation for two of the evaluated models, and LLM-generated reference reasoning without human validation of the fact-decomposition pipeline. These threats are load-bearing for the reported absolute scores and for the proprietary-vs-open-weight comparisons.","major_comments":[{"comment":"No contamination analysis is reported for the PMCOA case reports, which were published before September 1, 2025 and are almost certainly present in the web-scale corpora used to train the evaluated models. The final diagnosis is removed from the case presentation during conversion (§4.1.4), but a model that recognizes a case from pretraining can still output the correct diagnosis without performing the multi-turn reasoning the benchmark claims to measure. This directly threatens the absolute accuracy figures (e.g., GPT-5-medium's 33.88% completely-correct rate) and confounds the proprietary-vs-open-weight comparison. The limitations paragraph in the Discussion mentions only 'potential biases' in curation and automated evaluation, not memorization. Please add a contamination analysis—e.g., n-gram overlap tests, temporal holdouts, canary strings, membership inference, or case-level human i","section":"§4.1, §4.2"},{"comment":"The dual-LLM judge mechanism uses GPT-5-medium and Claude-4.5-Sonnet as independent evaluators, but these are the same models whose outputs are scored (GPT-5-medium evaluates GPT-5-medium; Claude-4.5-Sonnet evaluates Claude-4.5-Sonnet). This creates a self-evaluation loop that can bias both absolute scores and model rankings, particularly for the best-performing proprietary models. The claim that consensus scoring 'reduc[es] single-LLM judgment bias' does not address same-model bias. Please provide external validation: an expert-annotated subset of cases, inter-judge agreement statistics, and/or a re-analysis with held-out independent judges that excludes each model from judging its own outputs.","section":"§4.3.1"},{"comment":"The reference reasoning used for fact recall and hallucination is generated by GPT-4.1 during data conversion, and the atomic-fact extraction and matching are performed by LLMs without reported human validation. Hallucination is defined as the proportion of model-generated atomic facts unsupported by the reference; if the reference omits a true clinical fact, a correct model statement is counted as a hallucination. The manuscript does not report inter-annotator agreement, human spot-checks of the extracted facts, or sensitivity of the metrics to the LLM matcher. Without such validation, the fact recall and hallucination scores are difficult to interpret as measurements of reasoning quality. Please validate the fact decomposition and matching on a random sample against expert annotations and report reliability metrics.","section":"§4.1.4, §4.3.2"}],"minor_comments":[{"comment":"The labels 'Qwen3-VI-4B/8B/32B' use 'VI' while the text and other figures use 'VL'; this inconsistency should be fixed.","section":"Fig. 2a"},{"comment":"The caption says 'circle size and color encode the mean diagnostic accuracy score' but the figure shows fact recall scores; please correct.","section":"Fig. 7 caption"},{"comment":"Several specialties have very small sample sizes (Internal Medicine n=23, Emergency Medicine n=23, Nephrology n=29). Overlapping bootstrap CIs are reported, but formal significance tests or effect sizes with multiple-comparison corrections would strengthen claims about specialty differences and model comparisons.","section":"Supplementary Tables A1–B3"},{"comment":"The caption uses inconsistent subfigure labeling ('a, Data curation' in text but 'b' and 'c' implied in the figure); ensure the panel labels match the text.","section":"Figure 1 caption"},{"comment":"The data-curation flow table would benefit from a column indicating how many cases were excluded at each stage for each reason; the current counts are informative but the drop between collection and inspection is very large (e.g., Radiology 2,582 to 1,387) and not explained.","section":"Supplementary Information D"}],"recommendation":"major_revision","confidential_remarks":"The contamination and self-judging concerns are the two that I would press most strongly. If the authors can add a convincing memorization analysis and re-run the headline results with independent, non-self judges, the benchmark could be a useful contribution. The current limitations paragraph is too weak on both points. I see this as a major revision rather than a rejection because the core benchmark construction and error taxonomy are valuable and the threats are addressable in principle."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"ClinMM-Bench is a real step up from the existing single-turn medical QA and sequential diagnosis benchmarks, and the two-level evaluation idea is worth borrowing. But the headline numbers are not yet trustworthy because the evaluation loop is partly self-referential and the contamination question is wide open.\n\nWhat's new: 1,089 multi-turn multimodal cases from PMCOA with progressive disclosure of history, labs, and images across eight specialties—the largest setup of its kind I know of. The curation pipeline is detailed: dual-LLM suitability checks, automated fidelity scoring, expert review. The evaluation framework also does something genuinely useful—scoring accuracy with a dual-LLM consensus rubric, then separately decomposing model reasoning into atomic facts and measuring fact recall, hallucination, and fact density. The failure-mode taxonomy is well illustrated with concrete examples. The empirical sweep across 15 models is broad, and the findings that reasoning mode doesn't reliably help and scale helps in open-weight models are interesting.\n\nSoft spots, in order of seriousness.\n\nFirst, contamination: all cases come from public PMCOA case reports, and the models were trained on web-scale data that almost certainly includes PubMed Central. The paper runs no contamination check—no membership inference, no n-gram overlap, no temporal holdout. The conversion removes the final diagnosis from the presentation, but a model that recognizes the case from pretraining can still output the correct diagnosis without doing the multi-turn reasoning the benchmark claims to measure. That would inflate absolute accuracy and could shift the proprietary-vs-open-weight comparison. The limitation paragraph mentions 'potential biases' but does not address memorization. This needs to be answered.\n\nSecond, the judges are the same models being scored: GPT-5-medium judges GPT-5-medium, Claude-4.5-Sonnet judges Claude-4.5-Sonnet. Dual-LLM consensus helps, but it doesn't remove the self-judging concern. I'd want a human-rated subset to validate the rubric, or independent judges.\n\nThird, the reference reasoning and atomic-fact matching are both LLM-generated, with no reported human validation of the fact decomposition. That's a moderate concern; the matching instructions are careful, but the metrics inherit any systematic errors.\n\nAlso minor: the data itself doesn't appear to be released, only code. For a benchmark paper, that's a real barrier to reproducibility.\n\nThe central conclusion—that completely correct diagnoses are rare and reasoning quality is unreliable—is probably robust. But the exact scores and rankings in Fig. 2 should be treated cautiously until contamination and judge validity are addressed.\n\nThis paper deserves a serious referee. It's a useful benchmark and the framework is likely to be reused. But I'd send it back for major revision requiring contamination analysis, judge validation, and data release. Worth discussing in our reading group, partly as a case study in benchmark validity.","headline":"Real step up in benchmark scale and design, but self-judging and contamination issues mean the headline numbers need a skeptical read until addressed.","tokens_in":29550,"tokens_out":2678,"would_cite":true,"duration_ms":28114,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On ClinMM-Bench, a new 1,089-case multi-turn benchmark built from real case reports, the strongest multimodal model reaches a completely correct diagnosis in only 33.88% of cases.","keywords":["ClinMM-Bench","multimodal large language models","clinical diagnostic reasoning","multi-turn evaluation","diagnostic accuracy","fact recall","hallucination","medical AI benchmark"],"falsifier":"A human-expert panel re-scoring a random sample of 100–200 model outputs and finding that the LLM judges systematically over-credit models from their own family, or a data-contamination check showing the case reports appear in model training corpora, would invalidate the reported accuracy numbers and rankings.","tokens_in":28577,"feed_emoji":"🩺","tokens_out":5276,"duration_ms":43546,"temperature":0.7,"pith_summary":"The paper introduces ClinMM-Bench, a benchmark of 1,089 real, challenging clinical cases reformatted into multi-turn dialogues in which text and images are disclosed progressively, and uses it to test 15 multimodal large language models. The claim is that current models, including the strongest proprietary ones, do not yet perform reliable diagnostic reasoning under realistic conditions: the best model reaches a completely correct final diagnosis in about a third of cases, and most models mostly produce partially correct diagnoses. The paper also claims that reasoning quality is a separate and limiting dimension—models recall only about half to 60 percent of reference facts, and all produce unsupported statements at non-trivial rates. A sympathetic reader would care because it shifts evaluation from static question-answering to the dynamic, evidence-updating task that clinicians actually face, and it offers a reusable framework for measuring both the answer and the reasoning.","feed_headline":"Best AI fully diagnoses only a third of hard cases","feed_subtitle":"ClinMM-Bench: 1,089 real cases, progressive disclosure; top model fully diagnoses 33.88%.","key_machinery":"The load-bearing mechanism is the two-level evaluation framework. In level one, a dual-LLM consensus mechanism (GPT-5-medium and Claude-4.5-Sonnet act as judges) compares each model's predicted diagnosis against the ground truth and assigns a score of 0 (incorrect), 1 (partially correct), or 2 (completely correct); averaging the two judges produces a consensus score. In level two, both the model's explanation and a reference reasoning text are decomposed into atomic clinical facts—minimal verifiable statements—and three metrics are computed: fact recall (fraction of reference facts the model captured), hallucination (fraction of the model's facts unsupported by the reference), and fact densi","core_discovery":"The central discovery claimed is a measurement: when diagnostics are evaluated in a multi-turn, multi-image setting built from real case reports, every model tested shows a large gap between recognizing the general direction of a diagnosis and naming the exact diagnosis. GPT-5-medium, the best performer, achieves 33.88% completely correct diagnoses; open-weight models exceed 10% completely correct in only one case. Reasoning-quality metrics underline the same gap: the best fact recall is 0.599, hallucination scores range from 0.085 to 0.237, and no model combines high recall, low hallucination, and high fact density. The paper also reports that model scale helps, medical fine-tuning helps sm","pith_inferences":["Editorial inference: because the benchmark is built from published case reports selected for diagnostic challenge, its scores likely underestimate performance on common presentations; the paper's 'limited accuracy' claim should be read as about hard cases, not routine care.","Editorial inference: the failure of reasoning mode to help suggests a testable prediction—adding explicit memory or hypothesis-revision mechanisms to an MLLM should produce larger gains on this benchmark than increasing thinking tokens alone.","Editorial inference: the dual-LLM judges are the same model families being scored, so if judge leniency correlates with model family, the reported ranking could shift; a human-expert re-score on a random subset would provide a calibration check."],"forward_implications":["If ClinMM-Bench reflects real diagnostic difficulty, current MLLMs are not yet safe for independent final diagnosis on challenging cases; their role is closer to triage or differential-diagnosis support.","Accuracy scores alone overstate capability: a model can land a partially correct diagnosis while omitting key evidence and inserting unsupported facts, so deployment monitoring should track reasoning fidelity alongside the final answer.","Scale still matters within open-weight families: larger models consistently improved accuracy and completely-correct rates across the Gemma, MedGemma, and Qwen series.","Medical specialization is not a reliable route to better diagnosis in larger models; its clearest benefit is reduced hallucination, mainly at smaller scales.","Reasoning settings (extended 'thinking' traces) do not reliably improve multi-turn multimodal diagnosis, so progress is more likely to come from better cross-turn memory, visual grounding, and knowledge mapping than from longer chains of thought."],"fun_headline_variants":["AI diagnostic reasoning falls short: top model gets 33.9% fully right","Multi-turn medical exams: best AI still fails most cases","ClinMM-Bench: AI points to right diagnosis, rarely confirms it","GPT-5-medium tops 33.88% on hardest clinical cases","AI sees diagnostic direction but misses exact answer in most cases"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claims rest on the assumption that the dual-LLM consensus judges and the atomic-fact matching pipeline produce unbiased, reliable measurements—and that the evaluated models have not memorized the published case reports used to build the benchmark.","fun_headline_variants_meta":{"raw":{"variants":["AI diagnostic reasoning falls short: top model gets 33.9% fully right","Multi-turn medical exams: best AI still fails most cases","ClinMM-Bench: AI points to right diagnosis, rarely confirms it","GPT-5-medium tops 33.88% on hardest clinical cases","AI sees diagnostic direction but misses exact answer in most cases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000856,"raw_usage":{"total_tokens":3550,"prompt_tokens":737,"completion_tokens":2813,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":2735}},"tokens_in":481,"tokens_out":2813,"duration_ms":17767,"temperature":1.0,"reasoning_tokens":2735,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T01:02:49.357334+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human-expert panel re-scoring a random sample of 100–200 model outputs and finding that the LLM judges systematically over-credit models from their own family, or a data-contamination check showing the case reports appear in model training corpora, would invalidate the reported accuracy numbers and rankings.","supporting_citations":[],"review_version":1}