{"id":"e4364668-a2d3-4c81-9cfc-6871ceaa7df2","arxiv_id":"2505.14726","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Fine-tuning BLIP on ROCO improves standard captioning metrics, but the model's captions remain clinically unreliable and the paper's qualitative claim is contradicted by its own data.","lead":"This preprint fine-tunes the BLIP vision-language model on radiology images from the ROCO dataset and reports improved captioning scores over zero-shot baselines. The paper's own clinical evaluation shows the fine-tuned model still hallucinates findings and misses key details, so the claimed qualitative improvement is not supported.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's qualitative-improvement claim is contradicted by Table 2: all four fine-tuned captions are graded Incorrect, Misleading, Misleading, or Incomplete, and none is clinically correct.","rationale":"I read the paper in good faith: it compares several vision-language baselines, fine-tunes BLIP on ROCO, reports standard captioning metrics, visualizes attention, and includes an ablation. The central claim, as stated in the abstract, is that domain-specific fine-tuning improves performance on both quantitative and qualitative evaluation metrics. What would have to be true is that the fine-tuned model's outputs are not only closer to reference captions in lexical/semantic terms but also better in some clinical or qualitative sense. The paper's own Table 2 is the decisive evidence against this. Every fine-tuned caption is graded as clinically incorrect, misleading, or incomplete, and the paper explicitly says improved metrics do not guarantee clinical correctness. That is not a manufactured weakness; it is the authors' own evaluation table contradicting the abstract. The most load-bearing consequence is that the headline qualitative claim cannot stand as written. Secondary issues, such as the unspecified held-out subset and the minor inconsistency where encoder-only BERTScore F1 (0.7289) slightly exceeds full fine-tuning (0.7284) despite the claim that full fine-tuning yields the best results overall, reinforce the need for revision but are less central. I found no basis to change the reader's REJECT verdict: the contradiction between the abstract's qualitative claim and Table 2 is real, concrete, and directly testable with a blinded clinical scoring check. I therefore mark the verdict as unchanged rather than moving it to a different category.","tokens_in":6573,"tokens_out":4759,"duration_ms":46571,"concrete_test":"Re-score the four Table 2 cases, or better a random 50-image held-out sample, with a blinded clinician using a fixed rubric covering modality, laterality, findings, anatomy, and overall clinical correctness. Compare BLIP base versus fine-tuned using McNemar's test on 'fully clinically correct' and on 'contains a hallucinated finding.' If the fine-tuned model shows no significant improvement, or if zero fine-tuned captions are fully correct, the abstract's qualitative claim is unsupported and the paper must be revised to claim only quantitative similarity gains plus documented hallucination risks.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim says fine-tuning 'significantly improves performance across both quantitative and qualitative evaluation metrics.' The quantitative part is at least plausible as n-gram/semantic similarity to ROCO captions. The qualitative part is the load-bearing weakness. Table 2, the paper's own clinical scoring, grades the fine-tuned BLIP as Incorrect (chest X-ray), Misleading (brain MRI), Misleading (knee X-ray), and Incomplete (abdominal US). None of the four fine-tuned captions is clinically correct; at least three introduce a hallucinated or wrong finding (pleural effusion, left-sided hyperintensity, osteolytic lesion) or miss the key diagnosis (hydronephrosis). Moving from 'Poor' or generic to 'Misleading' or 'Incomplete' is not an unambiguous qualitative improvement, and for the chest X-ray both base and fine-tuned are Incorrect. The paper itself concedes that 'improved metrics do not guarantee clinical correctness.' Because the abstract explicitly claims qualitative improvement, Table 2 refutes the claim as written. The claim would need to be narrowed to lexical and semantic similarity gains, with clinical correctness explicitly reported as still failing, before the abstract is accurate.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents MedBLIP, a project that fine-tunes the BLIP image-captioning model on the ROCO radiology dataset and compares it against zero-shot BLIP, BLIP-2, BLIP-2 Instruct, Gemini 1.5 Flash, and ViT-GPT2. The authors report quantitative gains for fine-tuned BLIP on CIDEr, SPICE, BERTScore, and cosine similarity, show cross-attention visualizations, and conduct an ablation of full, encoder-only, and decoder-only fine-tuning. The paper claims that domain-specific fine-tuning improves both quantitative and qualitative performance, while also acknowledging in the text that fine-tuned captions can hallucinate findings or miss key diagnoses.","tokens_in":6784,"tokens_out":6247,"duration_ms":55339,"significance":"If the quantitative results were accompanied by proper statistical support, the paper would provide a modest confirmation that fine-tuning a general vision-language model on radiology captions improves lexical and semantic similarity to reference captions. The attention visualizations and the explicit clinical failure analysis are useful as cautionary evidence that standard captioning metrics do not imply clinical correctness. The paper does not ship code, specify the held-out evaluation size, or report any variance across runs, and its central positive claim is contradicted by its own clinical table: all four fine-tuned examples are graded clinically incorrect, misleading, or incomplete. The negative result is valuable, but the paper as written overstates the benefits of fine-tuning.","major_comments":[{"comment":"The abstract's claim that fine-tuning significantly improves performance across both quantitative and qualitative evaluation metrics is contradicted by the paper's own clinical evaluation. Table 2 grades the four fine-tuned captions as Incorrect (chest X-ray), Misleading (brain MRI), Misleading (knee X-ray), and Incomplete (abdominal US), and none is clinically correct. The paper itself states that \"improved metrics do not guarantee clinical correctness.\" The qualitative claim should be narrowed to lexical and semantic similarity to ROCO references, with the persistent clinical failures explicitly acknowledged in the abstract and conclusion.","section":"Abstract and §4.3, Table 2"},{"comment":"The efficiency claim is misattributed. The abstract says \"decoder-only fine-tuning (encoder-frozen) offers a strong performance baseline with 5% lower training time than full fine-tuning,\" but §4.5 reports encoder-only training at 4h59m versus 5h15m for full fine-tuning, which is the 5% saving. No decoder-only training time is reported. Moreover, Table 4 shows decoder-only fine-tuning has the lowest CIDEr (0.0664), SPICE (0.0275), and BERTScore F1 (0.7228) among the three configurations, so calling it a strong performance baseline is unsupported. The abstract should refer to encoder-only fine-tuning, or the authors should report decoder-only timing and justify the claim.","section":"Abstract and §4.5, Table 4"},{"comment":"The quantitative evaluation lacks essential experimental details and statistical rigor. The size and composition of the \"held-out validation subset\" are not specified, metrics are reported from a single run without standard deviations, and no significance tests are provided. Given that the absolute scores are low (CIDEr 0.0917, SPICE 0.0409) and the gaps to some baselines are small, the claim that fine-tuned BLIP \"consistently outperformed\" all baselines cannot be assessed as robust. The authors should report the number of test images, the distribution of modalities, and at least mean±std over multiple seeds.","section":"§4.1 and Table 1"},{"comment":"The clinical evaluation is based on only four hand-picked examples with no selection criteria, no inter-rater reliability, and no systematic scoring rubric. This is anecdotal evidence, not a qualitative evaluation metric. It cannot support the abstract's claim of improved qualitative performance. A larger, randomly sampled or independently reviewed clinical assessment would be needed to make any statement about clinical correctness; in its absence, the authors should explicitly state that clinical correctness remains an open failure.","section":"§3.5 and Table 2"},{"comment":"The use of ROCO figure captions as ground truth is a strong assumption that is not discussed. These captions are extracted from open-access publications and may be descriptive figure captions rather than verified clinical reports. While this does not create circularity because all models are scored against the same references, it limits the clinical validity of the reported quantitative improvements and should be stated as a caveat in the evaluation section.","section":"§3.1"}],"minor_comments":[{"comment":"The GitHub link \"github.com/Med Img Captioning\" contains a space and is not a usable URL; it should be either corrected or removed.","section":"Abstract and §5"},{"comment":"The training setup says \"1–3 epochs\" with early stopping, but the actual epoch count is never reported; please specify the epoch at which training stopped for each configuration.","section":"§3.3"},{"comment":"There are typographical issues in the metric names, including \"BER TScore\" and \"Hugging Face.For\"; these should be fixed.","section":"§3.5"},{"comment":"The decoding parameters (beam size = 4, max length = 128) are given but not justified, and it is unclear how many reference captions are used per image in the ROCO subset; please clarify.","section":"§4.1"},{"comment":"Several references are incomplete or suspicious: [1] is listed as \"Journal of Important Results,\" [9] uses implausibly round page numbers, and [7] lacks venue information; these need correction.","section":"References"},{"comment":"The figure caption says \"Each method generates different interpretations\" but does not specify which images are shown or how the attention maps are aggregated; please add this information.","section":"Figure 1"},{"comment":"The model name \"MedBLIP\" is identical to the title of reference [2], an existing medical vision-language model; this naming collision is confusing and should be clarified or renamed.","section":"Title and §2.3"},{"comment":"Gemini 1.5 Flash is reported with CIDEr 0.0 and SPICE 0.0; this is suspicious and likely due to output format mismatch, and it should be explained or excluded from the comparison.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reads as a preliminary project report rather than a polished journal submission. The central positive claim is overstated, and the paper's own clinical table undermines it. However, the errors are fixable by rewriting the abstract and conclusion, correcting the efficiency claim, and adding the missing experimental details. If the authors are unwilling to make these substantive revisions, rejection would be appropriate. The naming collision with the existing MedBLIP paper (reference [2]) also needs attention."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read for you: this is a routine fine-tuning study that is saved, barely, by an unusually honest clinical evaluation section. The authors fine-tune BLIP on ROCO, compare against BLIP-2, BLIP-2 Instruct, Gemini 1.5 Flash, and ViT-GPT2, and report CIDEr, SPICE, BERTScore, and cosine similarity. Fine-tuned BLIP wins on all four metrics. That part is expected transfer learning and not novel; prior cited work already fine-tunes BLIP-family models for medical captioning.\n\nWhat is actually good: Table 2. It grades four base vs fine-tuned captions against ground truth on modality, laterality, finding, and anatomy. Every fine-tuned caption is clinically wrong, misleading, or incomplete. The paper then says, in plain words, that improved metrics do not guarantee clinical correctness. That is the right message, and it is more than most captioning papers do.\n\nThe soft spot is the abstract. It claims fine-tuning significantly improves qualitative evaluation metrics. Table 2 refutes that: none of the four fine-tuned captions is clinically correct, and three introduce a hallucinated or wrong finding. The body contradicts the abstract; the abstract needs to be narrowed to lexical/semantic similarity, with clinical correctness explicitly reported as still failing. There are also smaller inconsistencies: the abstract says decoder-only fine-tuning gives 5% lower training time, but Section 4.5 reports encoder-only as the efficient variant (4h59 vs 5h15), and the conclusion credits encoder-only. The text mentions BLEU-4, METEOR, and CLIP similarity, but Table 1 reports different metrics. Evaluation is single-run, no split sizes, no significance tests, and no code or data despite a github placeholder.\n\nNone of that makes the quantitative result suspect in a loaded way; the numbers are at least consistent across ablation variants, and the paper is not circular. But as a contribution it is thin: no new method, dataset, or derivation, and the one interesting message (metrics don't equal clinical correctness) is buried under an overstated abstract. Who is this for? Someone writing about evaluation methodology for medical captioning might use the clinical table as anecdotal evidence; anyone else can skip. If I were the editor, I would not spend full referee cycles on it as is; after a careful rewrite that makes the negative clinical result the point, a workshop or short-paper venue could be right. My recommendation: desk-reject the current version, not because the experiment is worthless, but because the paper's own data contradicts its headline and the loose ends are too many.","headline":"A routine fine-tuning study whose own clinical evaluation table undercuts the abstract's qualitative-improvement claim; the honest negative result is the only part worth keeping.","tokens_in":7300,"tokens_out":3110,"would_cite":false,"duration_ms":30769,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning the BLIP model on radiology captions lifts measured scores and attention, yet every example still fails clinical review.","keywords":["medical image captioning","BLIP","fine-tuning","ROCO dataset","radiology","vision-language models","attention visualization","clinical correctness evaluation"],"falsifier":"A reader could take the four cases in the paper's clinical-evaluation table, expand them to 100 randomly selected held-out images, and have two blinded radiologists score each fine-tuned caption for whether it correctly states modality, laterality, primary finding, and anatomical location. If fine-tuned captions do not receive significantly more 'correct' verdicts than the zero-shot baseline, the qualitative-improvement claim is falsified; the table already shows this is a live possibility.","tokens_in":6386,"feed_emoji":"🩻","tokens_out":8877,"duration_ms":74980,"temperature":0.7,"pith_summary":"This paper asks whether fine-tuning a general-purpose image-captioning model on radiology figure captions makes it usable for medical image captioning. The authors report that full fine-tuning of the BLIP model on a subset of the ROCO radiology dataset raises lexical and semantic alignment scores—CIDEr from 0.0294 to 0.0917, SPICE from 0.0171 to 0.0409, BERTScore (Bio ClinicalBERT) F1 from 0.7078 to 0.7284—and produces more localized decoder attention over anatomical regions. They also find, in their own clinical-correctness evaluation, that every fine-tuned example still contains an incorrect, misleading, poor, or incomplete clinical statement. The pith is that domain adaptation measurably aligns a general model with radiology vocabulary and visual grounding, while exposing that standard captioning metrics do not certify clinical safety.","feed_headline":"Fine-tuning BLIP lifts radiology caption scores, not clinical accuracy","feed_subtitle":"Domain adaptation raises CIDEr from 0.029 to 0.092 and sharpens attention, yet every sample still fails clinical review.","key_machinery":"The load-bearing component is BLIP itself, a vision-language architecture with a ViT-B/16 image encoder and a BERT-style text decoder. The paper fine-tunes this pair on radiology image-caption pairs with cross-entropy loss and teacher forcing, then reads decoder cross-attention from the final layer—averaged over heads and overlaid on the image—as a window into visual grounding. The ablation (encoder-only, decoder-only, full) uses the same machinery to attribute the gains to language-side versus vision-side adaptation.","core_discovery":"The central claim, stated the way a sympathetic reader would take it, is that domain-specific full fine-tuning of BLIP on ROCO is the most effective of the tested routes to medical captioning: it achieves the best measured CIDEr, SPICE, and BERTScore while making decoder attention more localized, and decoder-only fine-tuning offers a competitive economy. The paper's own clinical-correctness table, however, marks every fine-tuned output as Incorrect, Misleading, Poor, or Incomplete, so the authors conclude that improved metrics and attention do not guarantee clinically accurate captions. Thus the discovery is double: adaptation helps alignment, and alignment is not enough.","pith_inferences":["A blinded clinician preference test on a larger held-out set would directly test whether the metric gains correspond to usable captions; the paper's clinical table suggests fine-tuned captions would often still be rejected.","The paper's 5% training-time claim is attributed to decoder-only fine-tuning in the abstract but to encoder-only training in the results; reconciling this would clarify which configuration actually saves time.","Attention localization could be repurposed as a human-in-the-loop verification aid: the paper shows focused attention on the correct region can accompany a hallucinated finding, so attention maps should flag, not certify, clinical statements.","One could build a composite failure score from the paper's four clinical dimensions—modality, laterality, finding, anatomical specificity—to penalize hallucinations and missed findings in a way CIDEr and BERTScore do not."],"forward_implications":["If the paper's central claim holds, full fine-tuning of BLIP on ROCO gives the best measured captioning scores in this comparison, with CIDEr rising from 0.0294 to 0.0917 and SPICE from 0.0171 to 0.0409.","Decoder-only fine-tuning is a competitive cheaper baseline, offering most of the measured benefit while updating fewer parameters and, as the paper reports, a roughly 5% training-time saving in at least one configuration.","Encoder-only fine-tuning scores below the other strategies, indicating that adapting the decoder's language model matters more than adapting the visual encoder alone.","Standard overlap and embedding metrics can improve while every generated example remains clinically unacceptable, so medical deployment needs clinical verification rather than metric-based acceptance."],"supporting_citations":[{"why":"Supplies the BLIP architecture and pretrained weights that the paper fine-tunes on radiology data.","marker":"[6]"},{"why":"Provides the ROCO radiology image-caption sample used for fine-tuning and evaluation.","marker":"[4]"},{"why":"Cited alongside the original as the updated ROCO dataset release.","marker":"[3]"},{"why":"Defines the BLIP-2 zero-shot baseline that the fine-tuned model must beat.","marker":"[7]"},{"why":"Defines the ViT-GPT2 encoder-decoder baseline used in the comparison.","marker":"[9]"},{"why":"Documents generalization and interpretability challenges in radiology report generation, motivating the paper's clinical-correctness review.","marker":"[8]"}],"fun_headline_variants":["MedBLIP: Better scores, still fails clinical review","Fine-tuned BLIP boosts metrics, not medical accuracy","Radiology captions improve, but all still clinically wrong","Decoder-only fine-tuning: 5% faster, nearly as good","BLIP fine-tuned for radiology: metrics up, trust down"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that ROCO's figure captions and the four reported scores measure what matters for medical captions; the paper's own clinical review shows that all fine-tuned outputs still carry incorrect, misleading, poor, or incomplete medical statements.","fun_headline_variants_meta":{"raw":{"variants":["MedBLIP: Better scores, still fails clinical review","Fine-tuned BLIP boosts metrics, not medical accuracy","Radiology captions improve, but all still clinically wrong","Decoder-only fine-tuning: 5% faster, nearly as good","BLIP fine-tuned for radiology: metrics up, trust down"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1259,"prompt_tokens":893,"completion_tokens":366,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":279}},"tokens_in":509,"tokens_out":366,"duration_ms":3616,"temperature":1.0,"reasoning_tokens":279,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:09:44.953212+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could take the four cases in the paper's clinical-evaluation table, expand them to 100 randomly selected held-out images, and have two blinded radiologists score each fine-tuned caption for whether it correctly states modality, laterality, primary finding, and anatomical location. If fine-tuned captions do not receive significantly more 'correct' verdicts than the zero-shot baseline, the qualitative-improvement claim is falsified; the table already shows this is a live possibility.","supporting_citations":[{"cited_title":"Blip: Bootstrapping language-image pre-training for unified vision-language understand- ing and generation","cited_arxiv_id":null,"evidence_quote":"Supplies the BLIP architecture and pretrained weights that the paper fine-tunes on radiology data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ROCO radiology image-caption sample used for fine-tuning and evaluation."},{"cited_title":"Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models, 2023","cited_arxiv_id":null,"evidence_quote":"Defines the BLIP-2 zero-shot baseline that the fine-tuned model must beat."},{"cited_title":"Image caption generation using vision transformer and gpt architecture","cited_arxiv_id":null,"evidence_quote":"Defines the ViT-GPT2 encoder-decoder baseline used in the comparison."}],"review_version":1}