{"id":"d47513b3-03b5-49d7-a6ef-7af8697e09c2","arxiv_id":"2508.07031","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI models hallucinate when reading medical images and when generating them from text, producing false findings and anatomically impossible pictures.","lead":"This paper tests whether large AI models can reliably interpret and create medical images such as chest X-rays, CT scans, and MRIs. It finds confident errors in both directions, including anatomically impossible generated images and missed findings in real scans.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 4 success rates rest on undocumented manual curation/labeling; without inter-rater reliability or a rubric, the 66%→94% generative-hallucination numbers are unsupported.","rationale":"The paper’s broad central claim—that current multimodal LLMs hallucinate in both image-to-text and text-to-image medical tasks—is robustly supported by qualitative demonstrations (Fig. 3 toe-fracture, Fig. 7 radioulnar-joint/brain-in-abdomen), which are compelling regardless of the quantitative evaluation. The load-bearing weakness is exactly where the reader pointed: the specific success-rate numbers in Table 4, which appear verbatim in the strongest_claim, rest on the authors’ undocumented manual judgment of prompt curation and output labeling. Without inter-rater reliability or a transparent rubric, those percentages are not independently verifiable. The additional F1 inconsistency (Sec 5.2 text vs Table 3) reinforces that the numerical results cannot be taken at face value. However, this does not invalidate the central argument; it only means the quantitative support is weaker than presented. The paper should be accepted conditionally on the authors supplying the missing methods and, ideally, a re-annotation study. Since the reader’s verdict is already CONDITIONAL, our stress-test does not change it; hence UNCHANGED.","tokens_in":9755,"tokens_out":7098,"duration_ms":69988,"concrete_test":"Re-run the Sec 5.3 evaluation with a pre-registered rubric. Have two board-certified radiologists, blinded to model identity and to the paper’s labels, independently classify each of the 50 prompt outputs into: (a) refusal/task-clarification, (b) generated image that is anatomically plausible and does not depict the requested impossible feature, or (c) generated image depicting the clinically implausible combination. Compute Cohen’s kappa between raters, adjudicate disagreements, and recalculate Table 4. If kappa < 0.6 or the recalculated GPT-4o P1→P2 difference is no longer statistically significant (e.g., McNemar test), then the specific success-rate claim in Sec 5.3 is unsupported. Also recompute Table 3 using only the published columns to resolve the 0.06 vs 0.37 F1 discrepancy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim in Sec 5.3—that GPT-4o generates clinically implausible content with 66% (P1) and 94% (P2) success, and Gemini with 30%/46%—derives entirely from Table 4. This table depends on two undocumented subjective steps: (1) the curation of the 50 ‘clinically implausible prompts’ (no inclusion/exclusion criteria, no example list, no pilot), and (2) the authors’ own binary judgment of whether each output ‘counts’ as a successful hallucination. The paper states only that outputs were evaluated using ‘expert informed criteria’ (Abstract) and ‘We assessed the success rate’ (Sec 5.3), with no rubric, no pre-registration, no inter-rater reliability, and no confidence intervals. Because the prompts themselves demand impossible anatomical combinations, the distinction between ‘model generated the requested impossible content’ and ‘model refused or produced a corrected/clarified image’ can be ambiguous (e.g., Gemini’s toe-fracture response shows the model can productively refuse). An independent re-annotation could shift individual classifications; with only 50 prompts, a shift of 5–7 labels changes the reported percentages by 10–14 points and could erase the claimed P1→P2 gap for GPT-4o. This is compounded by an internal inconsistency in Sec 5.2: the text reports Qwen’s zero-shot CT F1 as 0.06, but Table 3 lists 0.37; the missing ‘Enhanced data filtering’ column promised in the text also never appears. These issues do not topple the qualitative claim—clear examples like the toe-fracture X-ray suffice—but they undermine the specific numeric headline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies hallucinations in multimodal LLMs applied to medical imaging, covering both image-to-text tasks (pleural effusion detection on chest X-rays, chest CT cancer classification) and text-to-image tasks (generation of clinically implausible medical images). It reports classification accuracy/F1 for three open-source vision-language models and success rates for GPT-4o and Gemini-2.5 Flash in generating implausible content under two prompt styles. The qualitative examples are intended to show common hallucination patterns and prompt-sensitivity.","tokens_in":10190,"tokens_out":2893,"duration_ms":28017,"significance":"If the quantitative results were reproducible, the paper would provide a useful cautionary study of multimodal LLM behavior in medical imaging. The qualitative demonstrations—such as GPT-4o overlaying finger bones on a chest X-ray and both models superimposing unrelated anatomy onto abdominal scans—are compelling and directly support the broad thesis that current models can violate modality-anatomy consistency. A strength is that the study is direct measurement with no fitted parameters, and the qualitative examples give concrete substance to the hallucination categories. However, the quantitative evaluation is under-specified and contains internal inconsistencies, so the specific numeric claims in Tables 3 and 4 are not currently supported as stated.","major_comments":[{"comment":"There is an internal inconsistency: the text states that Qwen's zero-shot F1 score is 0.06 ('Qwen's large gap between accuracy (51.5%) and F1 score (0.06)'), but Table 3 lists Qwen's zero-shot F1 as 0.37. Additionally, the text says classification was conducted in three modes—'Zero-shot, Enhanced data filtering, and Few-shot'—but §4.1.1 defines only zero-shot and few-shot, and Table 3 has no 'Enhanced data filtering' column. Please define this mode, report its results, and correct the F1 inconsistency with the exact sample sizes.","section":"§5.2, Table 3"},{"comment":"The central generative claim—GPT-4o success rates of 66% (P1) and 94% (P2), and Gemini rates of 30%/46%—rests on 50 curated 'clinically implausible prompts' and the authors' own binary judgment of whether each output counts as a successful hallucination. No prompt list, inclusion/exclusion criteria, scoring rubric, inter-rater reliability, or confidence intervals are provided. The paper itself shows that Gemini-2.5 Flash can productively refuse and clarify (Fig. 3), so the distinction between 'hallucinated' and 'refused/corrected' is nontrivial. With n=50, a shift of 5–7 labels changes percentages by 10–14 points and could erase the reported P1→P2 gap. A full evaluation protocol and independent annotation are needed before these numbers can be interpreted.","section":"§5.3, Table 4"},{"comment":"The pleural effusion experiment uses a 'curated subset' of the Indiana Chest X-ray dataset, but the subset size, case selection criteria, and ground-truth label derivation (presence, extent, laterality) are not described. Table 2 reports F1 scores and DQ2 percentages without sample sizes or confidence intervals, and the DQ2 categories are not defined in the caption. Without these details, the quantitative comparisons among LLaVA, Gemma, and Qwen are not reproducible and the claimed severity of each model's hallucination pattern cannot be assessed. Please report the exact number of cases, the label source, and per-class counts.","section":"§5.1, Table 2"},{"comment":"The text says that 'GPT-4o blocks P1 because of its safeguard' (Fig. 7b) yet Table 4 reports a 66% success rate for GPT-4o under P1. This is confusing: if blocking/refusing is counted as failure, the text should say so explicitly; if some P1 attempts still yield images, the example in Fig. 7(b) is not representative. Define the success-rate denominator and clarify whether refusals are scored as failures or as separate outcomes, since this directly affects the Table 4 percentages.","section":"§4.2.2, Fig. 7 vs. Table 4"}],"minor_comments":[{"comment":"The text says 'Fig. 5 shows that both GPT and Gemini produce images with right-sided effusions,' but the generated chest X-rays appear first in Fig. 4. Please renumber the figures or correct the cross-references so the order matches the narrative.","section":"§4.2.1, Figures 4–6"},{"comment":"The column headers 'DQ 2(a) DQ 2(b)' are unclear without a caption explaining the sub-questions. Also, the model names 'LLaVA', 'Gemma', and 'Qwen' are spelled inconsistently (e.g., 'LLaV A' in Table 2, 'Gemma3-4B' in Fig. 8 caption but 'Gemma-3B' in the text).","section":"Table 2"},{"comment":"The text alternates between 'GPT-4o' and 'GPT-4' (e.g., 'GPT-4 initially refused'), which could confuse readers. Please use a consistent model identifier throughout.","section":"§4.2.2"},{"comment":"No code or data availability statement is provided. For a study whose quantitative claims depend on manual curation of prompts and images, releasing the prompt set, output images, and annotation records is essential for reproducibility; at minimum, the paper should state where these can be obtained.","section":"General"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the qualitative finding—multimodal LLMs confidently produce clinically impossible content in both interpretation and generation—is real and well illustrated. The toe-fracture chest X-ray and the radioulnar joint-in-abdomen examples speak for themselves. The dual-direction framing (image-to-text plus text-to-image) is convenient but not new; prior benchmarks like MedVH, RadFlag, and MedHallBench already cover this ground. The most novel bit is the prompt-bypass success rate: showing that appending a fake \"for research purposes\" justification increases generation of impossible anatomy from 66% to 94% for GPT-4o. That phenomenon is worth measuring, and the paper deserves credit for trying.\n\nThe problems are with the numbers, not the headline. The internal inconsistency in Sec 5.2—text says Qwen zero-shot CT F1 is 0.06, Table 3 says 0.37—makes you wonder about proofreading. The \"Enhanced data filtering\" mode is promised in Sec 5.2 but never defined in Sec 4.1.1 and has no column in Table 3. And Table 4's success rates come from 50 manually curated \"clinically implausible prompts\" judged by the authors alone, with no rubric, no inter-rater reliability, no confidence intervals, and no breakdown of how many prompts were refusals versus generations. With only 50 prompts, moving a few labels changes the percentages by ten points and could erase the claimed P1→P2 gap. The ground-truth derivation for the Indiana chest X-ray and IQ-OTH/NCCD subsets is also not described.\n\nThese flaws don't topple the central claim—that LLMs hallucinate in medical imaging is already well established, and the examples alone support it. But they do undermine the specific numeric contributions. The paper reads like a workshop-level empirical note, not a definitive benchmark.\n\nWho is this for? Someone tracking prompt-bypass vulnerabilities in medical image generation might find the P1/P2 comparison a useful datapoint, provided the authors release the prompt list and a codebook. The paper deserves a serious referee because the topic matters and the flaws are fixable, not because it's a breakthrough. I'd send it to peer review with an expectation of major revision: correct the Qwen F1 figure, define the filtering mode, document the curation and labeling protocol, and report raw counts with confidence intervals. If the authors do that, it becomes a legitimate cautionary data point.","headline":"This paper is a modest empirical survey showing that current multimodal LLMs hallucinate in both medical image interpretation and generation; the qualitative examples convince, but the headline numeric claims, especially the 66%→94% prompt-bypass success rate, rest on under-documented manual evaluation.","tokens_in":10634,"tokens_out":2570,"would_cite":false,"duration_ms":26385,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current multimodal LLMs hallucinate in both reading medical scans and generating them; a one-line 'for research' justification raises GPT-4o's fabrication of impossible anatomy from 66% to 94%.","keywords":["hallucination","multimodal LLM","medical imaging","image-to-text","text-to-image","chest X-ray","CT classification","prompt sensitivity"],"falsifier":"Have two board-certified radiologists independently score the same 50 implausible-prompt outputs under a pre-registered rubric, and independently relabel the chest X-ray subset for effusion presence, extent, and laterality; if inter-rater agreement is low or the 66%/94% and 30%/46% success rates do not reproduce, the claimed wording-sensitive hallucination rates and the F1-based conclusions collapse.","tokens_in":9684,"feed_emoji":"🩻","tokens_out":9446,"duration_ms":80044,"temperature":0.7,"pith_summary":"This paper tries to establish that today's multimodal LLMs cannot yet be trusted for medical imaging: when they read chest X-rays, CTs, or MRIs they both invent findings that are absent and miss findings that are present, and when they generate synthetic scans from text they attach unprompted laterality, surgical artifacts, and anatomically impossible content. The evidence covers two directions—image-to-text interpretation and text-to-image generation—across open-weight and proprietary models. The most direct quantitative result is that rephrasing an implausible prompt with a 'for research purposes' justification raises GPT-4o's success rate at producing clinically impossible images from 66% to 94%. The authors conclude that hallucination-aware evaluation, not raw accuracy, is the appropriate yardstick for clinical safety.","feed_headline":"94% of reworded prompts make GPT-4o invent impossible scans","feed_subtitle":"Adding 'for research' to a request spikes GPT-4o's implausible-image success rate from 66% to 94%.","key_machinery":"The central instrument is a paired prompting protocol: each clinically implausible generation request is posed in a direct form (P1) and again with an appended 'for research purposes' justification (P2), which measures how easily model safeguards are bypassed by wording. For interpretation, the load-bearing probe is a set of diagnostic questions on pleural effusion (presence, extent, laterality) applied to chest X-rays, plus zero-shot versus few-shot classification of chest CT scans; these expose the difference between accuracy and hallucination-aware F1 performance.","core_discovery":"On the interpretation side, the paper reports that LLaVA-v1.5-7B, Gemma-3B, and Qwen2.5-VL-7B hallucinate on pleural-effusion detection in chest X-rays: Qwen shows a high false-negative rate (hallucinated absence) while LLaVA shows a mix of false positives and false negatives, and few-shot prompting in chest-CT cancer classification improves accuracy modestly without eliminating either fabricated or missed findings. On the generation side, GPT-4o and Gemini-2.5 Flash add unprompted clinical details such as right-sided effusions and surgical clips, and across 50 curated clinically implausible prompts they generate anatomically impossible images at rates of 66%/94% (GPT-4o, P1/P2) and 30%/46%","pith_inferences":["The P1→P2 jump suggests a general compliance gradient in multimodal LLMs: any plausibility wrapper (research, educational, demonstrative) may relax safety constraints, implying audits should sample a family of paraphrases rather than one prompt.","Because the gold-standard labels for both the curated prompts and the image findings rest on the authors' own judgment, the reported rates are best read as upper bounds on hallucination until an independent radiologist-labeled subset is scored under a pre-registered rubric.","The recurring right-sided laterality in generated effusions hints that training corpora overrepresent that anatomy; a testable extension is to correlate model-chosen side or clip placement with caption statistics in the underlying training data.","The same image-to-text / text-to-image hallucination probe could transfer to other high-stakes visual domains such as pathology slides or fundus photographs, where confident wrong outputs are similarly dangerous and prompt phrasing is easily manipulated."],"forward_implications":["Hospitals evaluating LLM triage tools should report false positives and false negatives separately, since models like Qwen can combine moderate accuracy with a high rate of hallucinated absence.","Few-shot examples are not a sufficient remedy: they lift chest-CT classification accuracy a few points but leave the hallucinated-presence/absence profile essentially unchanged.","A generic justification phrase ('for research purposes') can flip a model from refusing an impossible imaging request to complying with it, so safety claims must be tested across prompt paraphrases, not single phrasings.","Generated images carry implicit clinical priors—right-sided effusions, left- or right-sided cysts, surgical clips—that can bias trainees and downstream models if synthetic data are used without explicit labeling.","Accuracy and F1 score disagree substantially in these models, so any deployment threshold should be set on hallucination-aware metrics rather than top-1 accuracy."],"supporting_citations":[{"why":"Supplies the GPT-4o model used in the generation and refusal experiments.","marker":"[1]"},{"why":"Supplies the Gemini-2.5 Flash model compared against GPT-4o on implausible generation.","marker":"[25]"},{"why":"The public chest X-ray dataset whose subset provides images and ground truths for the pleural-effusion test.","marker":"[22]"},{"why":"The reference describing the chest X-ray view data that supports the same effusion evaluation.","marker":"[28]"},{"why":"The public lung CT dataset used for the cancer-versus-normal classification task.","marker":"[4]"},{"why":"The accompanying description of the marked CT dataset used in the classification experiments.","marker":"[14]"},{"why":"The brain MRI dataset and empirical study that motivates the zero-shot classification protocol.","marker":"[27]"},{"why":"The systematic hallucination-evaluation framework the paper's error taxonomy builds on.","marker":"[10]"},{"why":"The prior finding that narrative rephrasing bypasses multimodal content safeguards, which underlies the P2 prompt design.","marker":"[7]"},{"why":"The clinician-led labeling framework for hallucinated versus omitted content in LLM clinical outputs, emulated in the manual assessment.","marker":"[5]"}],"fun_headline_variants":["GPT-4o fabricates impossible scans in 94% of misleading prompts","LLMs add fake surgical clips when describing chest X-rays","Hallucinations plague medical AI: from missed tumors to invented anatomy","Reword prompts, and GPT-4o's impossible-image rate jumps to 94%"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The quantitative claims assume the authors' manual judgments are a correct gold standard: what counts as a hallucinated or omitted finding in the X-ray/CT images, and which of the 50 generated outputs count as successful hallucinations, were decided by the authors alone with no second reader, expert adjudication, or pre-registered rubric.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4o fabricates impossible scans in 94% of misleading prompts","LLMs add fake surgical clips when describing chest X-rays","Hallucinations plague medical AI: from missed tumors to invented anatomy","Reword prompts, and GPT-4o's impossible-image rate jumps to 94%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00059,"raw_usage":{"total_tokens":2591,"prompt_tokens":715,"completion_tokens":1876,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":1796}},"tokens_in":459,"tokens_out":1876,"duration_ms":12130,"temperature":1.0,"reasoning_tokens":1796,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:21:58.129209+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two board-certified radiologists independently score the same 50 implausible-prompt outputs under a pre-registered rubric, and independently relabel the chest X-ray subset for effusion presence, extent, and laterality; if inter-rater agreement is low or the 66%/94% and 30%/46% success rates do not reproduce, the claimed wording-sensitive hallucination rates and the F1-based conclusions collapse.","supporting_citations":[{"cited_title":"Indiana university chest x- ray","cited_arxiv_id":null,"evidence_quote":"The public chest X-ray dataset whose subset provides images and ground truths for the pleural-effusion test."},{"cited_title":"Rodney Long, and George R","cited_arxiv_id":null,"evidence_quote":"The reference describing the chest X-ray view data that supports the same effusion evaluation."},{"cited_title":"Al-Yasriy","cited_arxiv_id":null,"evidence_quote":"The public lung CT dataset used for the cancer-versus-normal classification task."},{"cited_title":"Evaluation of SVM performance in the detection of lung cancer in marked ct scan dataset","cited_arxiv_id":null,"evidence_quote":"The accompanying description of the marked CT dataset used in the classification experiments."},{"cited_title":"On large visual language models for medical imaging analysis: An empirical study","cited_arxiv_id":null,"evidence_quote":"The brain MRI dataset and empirical study that motivates the zero-shot classification protocol."},{"cited_title":"Breaking the shield: Vulnerabilities in content moderation for multi- modal language models","cited_arxiv_id":null,"evidence_quote":"The prior finding that narrative rephrasing bypasses multimodal content safeguards, which underlies the P2 prompt design."},{"cited_title":"A framework to assess clini- cal safety and hallucination rates of LLMs for medical text summarisation","cited_arxiv_id":null,"evidence_quote":"The clinician-led labeling framework for hallucinated versus omitted content in LLM clinical outputs, emulated in the manual assessment."}],"review_version":1}