{"id":"12c9d673-7642-4db1-82f7-376bdf6eb976","arxiv_id":"2505.19031","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Fine-tuning medical vision-language models on the Med-MIM multi-image instruction dataset improves their scores on the authors' multi-image benchmarks, but the held-in benchmark is drawn from the same data used for training.","lead":"This paper introduces Med-MIM, a dataset of 83,200 medical question-answer pairs built from multi-image clinical data, and fine-tunes two vision-language models on it. The dataset and its companion benchmark aim to give medical AI the ability to compare, track over time, and cross-reference multiple images, which is how doctors actually work.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Held-in benchmark is derived from the training data with no described split; Table 1 gains may be memorization, so the central generalization claim is not yet evidenced.","rationale":"The reader's weakest assumption is the same as the critical load-bearing concern: no evidence that the evaluation is decontaminated. The paper's own text confirms a circular setup: the held-in benchmark is 'derived from the Med-MIM instruction dataset,' and fine-tuning uses that dataset with no split described. Because the abstract and conclusion generalize from benchmark success to enhanced multi-image visual ability, the possibility that the fine-tuned models have memorized the evaluation examples voids the core claim. The held-out benchmarks are only partial relief because their construction reuses the GPT-4o synthesis protocol and, as Table 1 shows, GPT-4o itself outperforms both fine-tuned models on several held-out settings, so the 'superior performance on held-out' phrasing in the abstract is also too strong. However, this is a correctable evaluation issue rather than a fatal flaw: the dataset and models could still be valuable, and a clean split would settle it. Thus I agree with the reader and leave the CONDITIONAL verdict unchanged.","tokens_in":8532,"tokens_out":5115,"duration_ms":47827,"concrete_test":"Construct a strict disjoint split at the study/patient level: for each held-in benchmark example, remove from the fine-tuning set every Med-MIM instruction sample sharing the same patient/study identifier or, for the composed subset, the same source image triple; retrain Med-Mantis and MIM-LLaVA-Med on the remaining samples and recompute Table 1. If the fine-tuning deltas on the disjoint held-in set, such as temporal close +35.44, collapse to near zero, the held-in claim is refuted as memorization; if the gains persist, the generalization claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 says the held-in Med-MIM benchmark is 'derived from the Med-MIM instruction dataset,' containing 2,968 closed-type and 256 open-type examples, while the two models are fine-tuned on 'the Med-MIM instruction dataset' for 3 epochs. The paper never states that the benchmark examples are disjoint from the fine-tuning examples, nor does it describe a patient-level split or de-duplication. Consequently the large held-in improvements in Table 1, e.g., MIM-LLaVA-Med temporal close +35.44 and Med-Mantis co-reference close +44.69, could reflect memorization of exact or near-duplicate QA pairs. This is not a minor statistical issue: the central conclusion that 'fine-tuning with the Med-MIM instruction dataset significantly enhances multi-image visual abilities' rests on these numbers. The held-out benchmarks (MIM-RAD, MIM-ODIR) do not clean this up because they are produced by the same GPT-4o 'composed/inherent' generation procedures as the training data; improved scores there can come from matching response format and synthetic answer style rather than transferable medical multi-image reasoning. The paper would need to demonstrate a completely disjoint evaluation before the abstract's generalization claim is supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Med-MIM, a medical multi-image instruction dataset of 83.2K QA pairs spanning four visual abilities (temporal understanding, reasoning, comparison, co-reference), and a corresponding Med-MIM benchmark with held-in and held-out subsets. The authors fine-tune LLaVA-Med and Mantis on Med-MIM to produce MIM-LLaVA-Med and Med-Mantis, and report that both achieve superior performance on held-in and held-out benchmark subsets, concluding that fine-tuning with Med-MIM significantly enhances multi-image visual abilities in the medical domain.","tokens_in":8803,"tokens_out":2016,"duration_ms":19347,"significance":"If the claims are substantiated, the Med-MIM dataset and benchmark would be a useful resource for training and evaluating medical LVLMs on multi-image reasoning, providing a taxonomy of abilities and two tuned model checkpoints. The paper is honest in reporting open-ended results and includes ablations on four visual abilities and dataset size. The held-out benchmarks (MIM-RAD, MIM-ODIR) are a constructive attempt to test generalization, and the positive deltas over base models in those benchmarks are suggestive. However, the evaluation design has a load-bearing weakness: the held-in benchmark is derived from the same instruction data used for training, and the held-out benchmarks are produced by the same GPT-4o synthesis pipeline as the training data, so the current evidence does not clearly separate learned generalization from memorization or style transfer.","major_comments":[{"comment":"The held-in Med-MIM benchmark is described as 'derived from the Med-MIM instruction dataset,' yet the paper does not state that the benchmark examples are disjoint from the 83.2K instruction samples used for fine-tuning. No patient-level split, image-level de-duplication, or question-level filtering is reported. Consequently, the large held-in gains in Table 1 (e.g., MIM-LLaVA-Med temporal close +35.44, Med-Mantis co-reference close +44.69) could reflect memorization of exact or near-duplicate QA pairs rather than improved multi-image visual ability. The central conclusion that fine-tuning 'significantly enhances multi-image visual abilities' rests on these numbers, so a disjoint, de-duplicated evaluation split is necessary.","section":"Section 2, 'Multi-image Visual Abilities Evaluation via Med-MIM Benchmark'"},{"comment":"The held-out benchmarks MIM-RAD and MIM-ODIR are constructed using the same generation procedures as the Med-MIM instruction dataset: MIM-RAD uses the composed-multi-image procedure with location suffixes, and MIM-ODIR applies the same GPT-4o prompt/refinement pipeline as the inherent Med-MIM data. This makes the held-out evaluation distribution closely resemble the training distribution in format and synthetic answer style. The reported improvements over base models on MIM-RAD and MIM-ODIR may thus partly reflect adaptation to the GPT-4o answer format rather than transferable medical multi-image reasoning. The paper should either use existing independently constructed medical multi-image benchmarks or explicitly demonstrate that performance holds on benchmarks not generated by the Med-MIM pipeline.","section":"Section 2, 'Held-out part' and Section 3, Table 1"},{"comment":"All reported numbers are from single runs without error bars or significance tests. The open-ended subsets are very small (e.g., 30 examples each for temporal and reasoning on the held-in benchmark), so observed deltas of a few points can easily be within random variation. For example, the MIM-LLaVA-Med vs. LLaVA-Med held-out MIM-RAD open improvement is only +0.22 percentage points, and several other open deltas are in the 2–7 point range. The paper should report variance across multiple seeds (or bootstrap confidence intervals) and should note the small sample size caveat in the discussion.","section":"Section 3, Table 1 and ablation figures"}],"minor_comments":[{"comment":"The paper states that Med-MIM is 'the first focused effort on medical multi-image analysis'; this is a strong claim that would benefit from a more careful comparison with prior multi-image medical VLM works beyond the references cited, such as recent multi-image medical LVLM papers that may have appeared by the time of submission.","section":"Abstract and Introduction"},{"comment":"Equation (1) is a standard next-token NLL loss; the notation p(r_k|I, q_{1:S}, r_{1:k-1}) is clear, but the sentence immediately after the equation repeats the definition. Consider tightening this paragraph.","section":"Section 2, 'Instruction Tuning'"},{"comment":"Fig. 4(c) reports results for '0%, 25%, 50%, 75%, and 100%' of the instruction dataset. The 0% point is presumably the base model without fine-tuning; this should be stated explicitly in the caption or text to avoid ambiguity.","section":"Section 3, 'Ablation on Four Visual Abilities'"},{"comment":"The co-reference category is applied to both the inherent and composed subsets, but the inherent co-reference examples ask about image location (e.g., 'which image... represents the view that captures the breast from above?'), which is closer to spatial/view understanding than co-reference as defined in the natural-image MANTIS taxonomy. The paper should clarify the relationship or justify the category naming.","section":"Section 2, 'Co-reference'"},{"comment":"The manuscript does not include a limitations section. Given the evaluation overlap concerns above, a short paragraph stating the potential distribution shift between synthetic GPT-4o-generated QA pairs and real clinical multi-image questions, and the lack of error bars, would improve scientific transparency.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper makes a useful resource contribution (dataset, benchmark, tuned models) but the evaluation is not yet convincing as evidence of generalization. The held-in benchmark appears to be drawn from the same pool as the training data, and the held-out benchmarks are generated by the same pipeline. The central claims in the abstract and conclusion are therefore stronger than the evidence supports. This is fixable with a properly disjoint evaluation split, additional independent held-out benchmarks, and variance-aware reporting. I recommend major revision rather than rejection because the resource itself is potentially valuable and the core idea is reasonable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi Jian,\n\nThe thing to know: the Med-MIM dataset is a real asset, and if you work on medical LVLMs you'll probably want to cite it. The paper's central claim—that fine-tuning on this dataset 'significantly enhances multi-image visual abilities'—is not yet supported by the evaluation, because the held-in benchmark is carved from the same instruction data used for training and the held-out benchmarks are generated with the same GPT-4o procedure.\n\nWhat's new: this is the first focused multi-image instruction dataset for medicine, 83.2K QA pairs spanning temporal, reasoning, comparison, and co-reference abilities, built from MS-CXR-T, EMBED, LUMIERE plus a composed subset from LLaVA-Med. That's a useful resource. The fine-tuning recipes (interleaved image-text, NLL) are standard, but the models do show consistent positive deltas over their bases on held-out sets: MIM-LLaVA-Med gains ~2-7 points on MIM-RAD/ODIR close-ended, Med-Mantis larger on MIM-ODIR. The ablations are sensible and show each ability subset matters.\n\nThe soft spots are real and load-bearing. Section 2 says the held-in benchmark is 'derived from the Med-MIM instruction dataset' with no described train/test split or de-duplication. The big held-in jumps—co-reference close +44.69, temporal close +35.44—are exactly what you'd see from memorization. The held-out sets don't fix this because they're synthesized by the same 'composed/inherent' pipeline, so the models may just be matching answer style. Also: single runs, no error bars, and some open-ended subsets have only 30 examples. And the abstract's 'superior performance on held-out' is too strong—GPT-4o still beats both fine-tuned models on MIM-RAD close (69.0 vs 65.3) and MIM-ODIR close (45.7 vs 32.3).\n\nWho should read it: anyone building medical multi-image LVLMs or evaluating them. The dataset itself is the contribution. The evaluation needs a proper disjoint split (patient-level), ideally a human-validated held-out set, error bars, and a more careful claim. I'd send it to peer review, but I'd ask for those changes. The central idea holds up; the evidence does not yet.\n\nCheers,\n[Name]","headline":"Med-MIM is a genuinely new dataset for medical multi-image QA, but the paper's headline claims outrun its evaluation: the held-in benchmark is drawn from the training data and the held-out benchmarks share the same synthesis pipeline.","tokens_in":9314,"tokens_out":3257,"would_cite":true,"duration_ms":26468,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that an 83,200-question instruction dataset, Med-MIM, can teach medical vision-language models to reason across multiple images.","keywords":["medical large vision-language models","multi-image understanding","instruction tuning","visual question answering","temporal understanding","co-reference","medical benchmark","GPT-4o generation"],"falsifier":"Assemble a new multi-image medical QA test from a source never touched by the Med-MIM generation pipeline, such as fresh human-written questions about images from a different institution; if Med-Mantis and MIM-LLaVA-Med then show no advantage over their base models, the dataset's effect is distribution-specific rather than a general multi-image ability.","tokens_in":8312,"feed_emoji":"🩻","tokens_out":6717,"duration_ms":57022,"temperature":0.7,"pith_summary":"Medical vision-language models answer questions about single images well, but clinical work often needs two or three images at once: a patient's previous and current chest X-ray, two mammographic views, several MRI sequences. This paper claims that a new instruction dataset called Med-MIM, with 83,200 question-answer pairs spanning four multi-image abilities, can teach medical LVLMs to handle such cases. The authors fine-tune two open-source models, Med-Mantis and MIM-LLaVA-Med, on this dataset and report that both beat six strong baselines on their held-in benchmark and improve on their base versions on two held-out benchmarks. If the claim holds, Med-MIM is a reusable resource for upgrading medical AI from single-image perception to multi-image clinical reasoning.","feed_headline":"83.2K training pairs give medical AI multi-image skills","feed_subtitle":"Two fine-tuned models beat six baselines on temporal, reasoning, comparison, and co-reference medical questions.","key_machinery":"The Med-MIM instruction dataset is the engine: 83.2K QA pairs in five imaging domains, organized into inherent and composed subsets, with each sample tagged to one of four visual abilities. Inherent samples are generated by GPT-4o from medical reports attached to naturally multi-image collections; composed samples are made by taking single-image LLaVA-Med QA pairs and adding location prefixes like 'in the first image.' The four ability labels drive both training and the held-in benchmark, while held-out benchmarks MIM-RAD and MIM-ODIR are built with the same generation recipe from VQA-RAD and ODIR. Interleaved image-text formatting and standard negative log-likelihood fine-tuning let the two base models ingest up to three images at a time.","core_discovery":"The central claim is that instruction tuning with the Med-MIM dataset produces the gain: fine-tuning on 83.2K multi-image QA pairs converts a general multi-image model (Mantis) and a single-image medical model (LLaVA-Med) into Med-Mantis and MIM-LLaVA-Med, both of which outperform the original models and all open-source baselines on the Med-MIM benchmark. On the held-in subset, Med-Mantis reaches 74.86% close-ended accuracy on temporal questions and 80.40% on co-reference, while MIM-LLaVA-Med improves over LLaVA-Med by 35.44 points on temporal close-ended accuracy. On the held-out MIM-RAD and MIM-ODIR benchmarks, both fine-tuned models show consistent gains over their base versions, and MIM-LLaVA-Med attains the best open-source scores on MIM-ODIR. The paper interprets these results as evidence that the Med-MIM instruction dataset effectively enhances multi-image understanding in the medical domain.","pith_inferences":["A natural next experiment is to fine-tune a larger or stronger general-purpose multi-image backbone on Med-MIM to see whether the medical gains compound with model scale.","Because the composed subset is built by adding location prefixes to single-image QA pairs, it plausibly teaches co-reference mostly through surface phrasing; whether the skill survives rephrased location questions is an open testable question.","The same inherent/composed recipe could be applied to other paired medical sources, such as longitudinal CT or pathology slide series, to see whether the multi-image abilities generalize beyond chest X-ray and mammography.","The ablations leave the inherent-versus-composed mix unexplored factorially; a controlled experiment varying only that mix would show which data source carries each ability."],"forward_implications":["The Med-MIM instruction dataset can be reused as a training resource by any open-source medical LVLM, so future models can acquire multi-image capability without collecting new clinical image pairs.","The four-ability split gives researchers a standard way to diagnose where a medical LVLM fails: temporal change, multi-view reasoning, image comparison, or locating content across images.","The Med-MIM benchmark provides a quantitative target for medical multi-image QA, with close-ended accuracy and open-ended lexical scores, so model progress on this skill can be tracked over time.","The reported gains imply that supervised fine-tuning on multi-image instruction data is a viable path for open-source medical models to approach or exceed closed-source performance on held-in clinical multi-image questions."],"supporting_citations":[{"why":"It supplies the longitudinal chest X-ray data with paired reports that Med-MIM uses to build temporal and comparison instruction samples.","marker":"[3]"},{"why":"It supplies the multi-view mammography data used to construct reasoning and co-reference samples in the Med-MIM instruction dataset.","marker":"[7]"},{"why":"It supplies the longitudinal glioblastoma MRI data that Med-MIM draws on for temporal, comparison, and co-reference samples.","marker":"[23]"},{"why":"It provides the single-image medical VQA data used to create the composed multi-image subset, and its model is one of the two bases fine-tuned into MIM-LLaVA-Med.","marker":"[12]"},{"why":"It provides the MANTIS base model and the four-category multi-image ability taxonomy that organizes Med-MIM into temporal, reasoning, comparison, and co-reference subsets.","marker":"[8]"},{"why":"It is the language-only generator used to produce the inherent Med-MIM QA pairs and the MIM-ODIR held-out benchmark, and it is also one of the evaluated baselines.","marker":"[6]"},{"why":"It supplies the single-image medical QA pairs that are regrouped and location-annotated to form the MIM-RAD held-out benchmark.","marker":"[11]"},{"why":"It supplies the paired fundus images without QA pairs that are turned into the MIM-ODIR held-out benchmark through the same generation procedure.","marker":"[13]"}],"fun_headline_variants":["83.2K training pairs give medical AI multi-image skills","Instruction tuning boosts medical VLMs on multi-image tasks","New dataset sharpens medical AI on comparison and co-reference","Med-MIM dataset teaches medical AI to handle multiple images","Fine-tuned medical VLMs beat baselines on multi-image benchmarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim depends on the benchmarks being genuinely different from the training data, so the reported gains measure transferable multi-image skill rather than memorization of near-duplicate examples.","fun_headline_variants_meta":{"raw":{"variants":["83.2K training pairs give medical AI multi-image skills","Instruction tuning boosts medical VLMs on multi-image tasks","New dataset sharpens medical AI on comparison and co-reference","Med-MIM dataset teaches medical AI to handle multiple images","Fine-tuned medical VLMs beat baselines on multi-image benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1293,"prompt_tokens":994,"completion_tokens":299,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":216}},"tokens_in":610,"tokens_out":299,"duration_ms":3310,"temperature":1.0,"reasoning_tokens":216,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:20:28.599990+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Assemble a new multi-image medical QA test from a source never touched by the Med-MIM generation pipeline, such as fresh human-written questions about images from a different institution; if Med-Mantis and MIM-LLaVA-Med then show no advantage over their base models, the dataset's effect is distribution-specific rather than a general multi-image ability.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the longitudinal chest X-ray data with paired reports that Med-MIM uses to build temporal and comparison instruction samples."},{"cited_title":"Radiology: Artificial Intelligence5(1), e220047 (2023)","cited_arxiv_id":null,"evidence_quote":"It supplies the multi-view mammography data used to construct reasoning and co-reference samples in the Med-MIM instruction dataset."},{"cited_title":"Scientific data9(1), 768 (2022)","cited_arxiv_id":null,"evidence_quote":"It supplies the longitudinal glioblastoma MRI data that Med-MIM draws on for temporal, comparison, and co-reference samples."},{"cited_title":"Transactions on Machine Learning Research 2024 (2024), https://openreview.net/forum?id=skLtdUVaJa","cited_arxiv_id":null,"evidence_quote":"It provides the MANTIS base model and the four-category multi-image ability taxonomy that organizes Med-MIM into temporal, reasoning, comparison, and co-reference subsets."},{"cited_title":"In: Benchmarking, Measur- ing, and Optimizing: Third BenchCouncil International Symposium, Bench 2020, 10 X","cited_arxiv_id":null,"evidence_quote":"It supplies the paired fundus images without QA pairs that are turned into the MIM-ODIR held-out benchmark through the same generation procedure."}],"review_version":1}