{"id":"47142bc6-9353-4920-89d3-9d73adda1ddf","arxiv_id":"2505.14064","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"NOVA is a new evaluation-only benchmark of rare brain MRI pathologies where GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B all exhibit large performance drops across anomaly localization, image captioning, and diagnostic reasoning.","lead":"A new benchmark, NOVA, tests AI models on 906 brain MRI scans covering 281 rare diseases, with expert-annotated boxes and clinical histories. Top vision-language models score far below clinical usefulness on localization, captioning, and diagnosis, showing they do not generalize to rare conditions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Task 1 is not double-blinded: annotators saw full clinical history and diagnosis (Sec. 3.2), so localization scores may reward clinical priors rather than visual anomaly detection.","rationale":"The reader's weakest assumption is exactly the load-bearing concern I identify: the abstract promises 'double-blinded expert bounding-box annotations,' but Section 3.2 states annotators reviewed the full Eurorad clinical description and metadata. This is an internal inconsistency that directly affects the validity of Task 1, the anchor task of the benchmark. If radiologists used clinical context to place boxes, the model-vs-human comparison is unfair in a specific, quantifiable way: models are evaluated image-only, while ground truth may encode information not present in the image. This would not merely be a mild limitation; it would mean the localization scores do not cleanly measure visual anomaly detection, undermining one of the paper's three headline claims. The suggested re-annotation study would settle this empirically. I do not think this concern alone justifies rejection, because Tasks 2 and 3 remain informative and the dataset is a useful resource, but it does justify the conditional verdict: the authors should either re-annotate image-only or demonstrate that image-only annotations agree with the released ones. I agree with the reader rather than proposing a different concern; the GPT-4o-as-judge issue and the missing UAD baselines are real but secondary, since they affect only one task or the strength of the comparison, not the core validity of Task 1.","tokens_in":11254,"tokens_out":3394,"duration_ms":44193,"concrete_test":"Sample 150 cases stratified across the six diagnostic categories. Have two or more independent radiologists who are not authors draw bounding boxes from the MRI slice alone, with no clinical history, diagnosis, or Eurorad metadata. Compare these image-only boxes against the released boxes using the same greedy IoU matching protocol (IoU > 0.3 consensus). Report the detection rate of released boxes that are matched by image-only boxes and the median IoU. If a large fraction of released boxes are unmatched (e.g., >30–50%), Task 1 is substantially contaminated by clinical priors; if matched, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The benchmark's central claim of rigorous, 'double-blinded' expert bounding-box annotations is contradicted by the annotation protocol in Section 3.2: 'Each case was independently labeled by two readers, who reviewed the full original Eurorad clinical description and associated metadata to inform their annotations.' Because annotators saw the clinical history and diagnosis, their boxes may reflect knowledge of where pathology is expected (e.g., 'right frontal lesion' in the history) rather than what is visually detectable on the MRI slice. Models in Task 1 receive only the image, so low localization scores could measure an annotation advantage rather than pure visual anomaly detection failure. This is an internal inconsistency, not a matter of consensus: the abstract's 'double-blinded' claim is false under the stated protocol. The senior adjudication step does not remove the bias, since the adjudicator also had clinical context. If the released boxes contain regions that are not visually identifiable without clinical priors, then Task 1 is not a valid visual localization benchmark, and the headline conclusion about VLM unreliability is overstated for that task.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"NOVA is an evaluation-only benchmark compiled from 906 Eurorad brain MRI cases covering 281 rare diagnoses. The paper defines three tasks: anomaly localization from a single released MRI slice via bounding boxes, image captioning, and diagnostic reasoning from clinical history plus image caption. The authors benchmark GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B and report substantial performance drops on all three tasks relative to natural-image benchmarks, arguing that current vision-language models are unreliable for rare brain MRI analysis in zero-shot settings. The dataset and annotations are released under a CC BY-NC-SA license.","tokens_in":11478,"tokens_out":4358,"duration_ms":46336,"significance":"If the validity concerns are resolved, NOVA addresses a genuine gap: no existing benchmark jointly evaluates anomaly localization, clinical captioning, and diagnostic reasoning on rare brain MRI. The breadth of diagnoses (281), the real-world acquisition heterogeneity, the expert multi-reader annotation protocol, and the inference-only design are strengths, and the authors explicitly acknowledge possible pretraining contamination. The paper also avoids fitted parameters, so the benchmark is not self-fulfilling in the usual sense. However, the two central validity issues—unblinded ground-truth boxes for Task 1 and GPT-4o serving as both judge and contestant in Task 3—must be addressed before the benchmark can support the headline claim that VLMs are unreliable in open-world clinical brain MRI.","major_comments":[{"comment":"The abstract states that NOVA provides 'double-blinded expert bounding-box annotations,' but Section 3.2 reports that 'Each case was independently labeled by two readers, who reviewed the full original Eurorad clinical description and associated metadata to inform their annotations.' The readers therefore knew the clinical history and likely the diagnosis before drawing boxes, and the senior adjudicator also had this context. Since Task 1 (Section 4.1) gives models only the image, low localization scores conflate visual anomaly detection with the information asymmetry between annotators and models. This is an internal inconsistency: the stated protocol contradicts the 'double-blinded' claim. The authors should either re-annotate a representative subset with image-only readers to quantify the effect, or rescope Task 1 as localization with clinical priors and soften the corresponding claims about VLM visual unreliability.","section":"§3.2, Abstract"},{"comment":"Task 3 uses GPT-4o 'to perform semantic matching between predictions and ground truth labels,' while GPT-4o is itself one of the models evaluated on the same task. This introduces a self-scoring circularity: GPT-4o's own predictions may be judged more leniently than those of other models if the judge favors its own output distribution. The authors should use a different judge model (e.g., a strong open-source model) or human raters, and should report inter-annotator agreement on a sample of diagnostic predictions to confirm that the semantic matching is not biased.","section":"§4.3, Table 3"},{"comment":"The release format is described as 'uniformly sized 480×480 grayscale PNG slices,' and Section 6 acknowledges that only 2D slices are provided, yet the manuscript repeatedly refers to 'scans' and claims to evaluate anomaly localization in brain MRI. A single axial, sagittal, or coronal slice may not contain the full lesion burden or the most representative plane, and the paper does not state how the slice was selected for each case. This limits the external validity of the localization and captioning results. The authors should document the slice-selection rule and report the proportion of cases in which the selected slice actually contains the annotated abnormality.","section":"§3.3, §6"}],"minor_comments":[{"comment":"The spelling 'NOVA' and 'NOV A' is used inconsistently; please unify the dataset name.","section":"Throughout"},{"comment":"The sentence 'over 600 false-positive boxes were recorded' is ambiguous because Table 1 reports FP30 values of 899, 1163, and 672 per model; the text should explicitly state whether this is per-model or total.","section":"§5.1, Table 1"},{"comment":"Please clarify whether each 'case' is represented by exactly one slice or by multiple slices, and explain how multiple ground-truth bounding boxes per case map onto the released slice(s).","section":"§3.3"},{"comment":"The column header 'Cov.' is not defined in the caption; please state explicitly that it denotes the fraction of ground-truth labels covered by the model's prediction vocabulary.","section":"Table 3"},{"comment":"The binary normal-versus-abnormal F1 values in Table 2 are extremely low (2.4–11.3%); since all Eurorad cases are pathological, the paper should explain how the binary label is derived and whether the metric measures any meaningful signal for these models.","section":"§4.2, Table 2"}],"recommendation":"major_revision","confidential_remarks":"The unblinded annotation issue is the most consequential risk: it directly undermines Task 1's validity and the abstract's 'double-blinded' claim. This is likely fixable within the manuscript's scope by re-annotating a subset in an image-only condition and reporting the difference, or by clearly reframing Task 1 as a clinically informed localization benchmark. The GPT-4o judge issue is also serious but easier to repair. I would not reject, because the resource itself is valuable and the baseline results are informative even if task framing must change."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this paper. First, NOVA is a genuinely useful new resource: roughly 900 brain MRIs, 281 rare diagnoses, expert bounding boxes, captions, and clinical histories, all in a single evaluation-only benchmark. Second, the paper's central 'double-blinded' claim is false under its own protocol, and that weakens Task 1's validity more than any other single issue.\n\nWhat's new: the dataset itself and the joint evaluation of localization, captioning, and diagnostic reasoning. There's no comparable brain MRI benchmark with this pathology diversity. The authors also handle the baseline experiments honestly—they flag possible training-data contamination and frame the results as upper bounds. The baseline numbers are low across the board, which is credible and clinically relevant.\n\nThe soft spots, in order of severity. The annotation protocol (Section 3.2) says each reader 'reviewed the full original Eurorad clinical description and associated metadata to inform their annotations.' That is not double-blinded. The annotators knew the diagnosis and history, so the boxes may encode clinical priors. In Task 1, models get only the image, so their low scores could partly reflect an information asymmetry rather than pure visual failure. The paper should acknowledge this and release a version of boxes drawn without clinical context, or at least test sensitivity. The adjudicator also had the history, so the bias isn't removed.\n\nNext, Task 3 uses GPT-4o as the semantic judge while GPT-4o is also one of the benchmarked models. That's a mild circularity. It's not catastrophic—the matching is simple—but a fixed external judge or a human-checked sample would be cleaner.\n\nMinor: no error bars, only three models, no UAD baselines even though the paper positions itself in that literature. The 2D slice limitation is acknowledged and defensible for broad accessibility.\n\nOverall, the benchmark itself is a meaningful contribution and the core observation—strong VLMs degrade sharply on rare pathology—is probably true. The annotation-protocol issue is serious but fixable; the circularity is minor. This paper deserves referee time.\n\nRecommendation: send it to review, and push for an honest revision of the 'double-blinded' language and a sensitivity analysis for Task 1.","headline":"NOVA is a genuinely useful new benchmark that is undermined by a false 'double-blinded' claim in its annotation protocol, yet the dataset and baseline results are worth serious consideration.","tokens_in":12029,"tokens_out":2937,"would_cite":true,"duration_ms":28519,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-07T15:40:48.261044+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}