{"id":"26c5bf79-c876-42c8-8456-d85b3f429fe4","arxiv_id":"2607.15241","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across nine Medico 2025 systems, answer accuracy on GI endoscopy VQA does not predict explanation faithfulness; structured reasoning and grounding correlate with better trustworthiness.","lead":"This paper looks back at a 2025 medical imaging challenge where nine teams built AI systems that answer questions about endoscopy images and explain their answers. It finds that good leaderboard scores don't guarantee trustworthy clinical explanations, and recommends ways to change how such systems are evaluated.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Subtask 2 trustworthiness scores come entirely from an unvalidated Qwen3-30B-A3B judge; the central 'explainability > fluency' claim collapses if that judge is biased.","rationale":"The reader's weakest assumption is exactly the load-bearing concern I identify: the Subtask 2 adjudicator is unvalidated and potentially biased, and all trustworthiness conclusions flow through it. The paper's own threat-to-validity statement (Section V.D) admits this. I agree with the reader that this makes the paper CONDITIONAL rather than ACCEPT, but I do not see grounds to move the verdict further: the paper is careful in hedging its claims as correlational, and the proposed clinician-validation check is a feasible path to strengthen or refute the central claim. The attack is not an internal inconsistency; it is an external-validity risk, which is precisely why the paper's own caveat is the most important evidence in the manuscript.","tokens_in":10451,"tokens_out":2816,"duration_ms":26165,"concrete_test":"Re-score a stratified sample of ~200 Subtask 2 responses per team (minimum 500 total) with 2–3 clinicians blinded to team identity, using the same five rubric dimensions. Compute rank correlation (e.g., Kendall's tau) between Qwen3-30B-A3B scores and clinician scores, and test for a judge-family interaction: do teams using Qwen-family backbones receive systematically higher Qwen3 scores than non-Qwen teams at matched clinician scores? If the correlation is low or the interaction is significant, the Subtask 2 rankings and the 'fluency vs. faithfulness' gap are judge artifacts and the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that answer-level gains do not translate into faithful/complete reasoning and that structured reasoning/grounding are more reliable—rests on Subtask 2 rubric scores from Qwen3-30B-A3B across all five dimensions (Correctness, Faithfulness, Clinical Relevance, Clarity, Completeness). The paper itself flags in Section V.D that the judge 'lacked clinician validation at scale here and may introduce alignment bias if submitted systems use related model families.' Yet Sections VII.D and VIII use these same scores to conclude that fluency is easier to optimize than faithfulness and to rank Team Nepal ahead of others. If Qwen3-30B-A3B systematically rewards fluent text or outputs from Qwen-family backbones, the observed 'clarity high, faithfulness low' pattern could be a judge artifact, and the design-axis lessons (e.g., self-probing > fluency) would not be supported. Because no clinician-rated or otherwise independent validation of the rubric scores appears anywhere in the paper, the strongest claim is currently a single-model proxy with no external anchor.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a retrospective analysis of the MediaEval Medico 2025 challenge, comparing nine documented systems for GI endoscopy VQA and explainable reasoning. It combines official leaderboard scores, organizer-released semantic adjudication, and rubric-based LLM scores to examine how answer accuracy, explanation faithfulness/completeness, robustness to private-set shift, and design choices (PEFT, self-probing, grounding) relate. The main claim is that lexical answer gains do not necessarily translate into faithful and complete clinical reasoning, and that methods enforcing structured reasoning and explicit grounding show more reliable behavior, although the evidence is explicitly correlational rather than ablation-based. The paper concludes with five practical recommendations for medical VQA evaluation, including semantic adjudication, image-level splits, explanation schemas, and calibration metadata.","tokens_in":10621,"tokens_out":4756,"duration_ms":41009,"significance":"The paper's value lies in using a real shared task to produce concrete, transferable observations about evaluation practice. It is unusually candid about limitations: it explicitly labels evidence as correlational, judge-dependent, and non-ablation-based, and it proposes specific, actionable evaluation upgrades. If the central finding holds, it strengthens the case for moving beyond BLEU-style metrics in medical VQA and for standardizing explanation artifacts. However, the central finding currently rests on a single unvalidated LLM judge, and the design-axis comparisons have very small per-cell sample sizes. These issues do not invalidate the contribution, but they limit the strength of the conclusions that can be drawn without additional validation.","major_comments":[{"comment":"The central RQ4 finding—that fluent language is easier to optimize than faithful, clinically grounded justification—depends entirely on Qwen3-30B-A3B rubric scores. The paper itself admits this judge 'lacked clinician validation at scale here and may introduce alignment bias if submitted systems use related model families' (Section V.D). No sensitivity analysis, second judge, inter-judge agreement, or human spot-check is reported. Because the same judge also produces the Subtask 1 semantic adjudication and the official Subtask 2 ranking, a systematic preference for fluent text could generate the observed clarity-faithfulness gap and the team ordering without reflecting clinically meaningful explanation quality. The manuscript's caveat that these are 'structured proxies rather than substitutes for clinician-led assessment' is appropriate but does not resolve the load-bearing nature of the","section":"Section V.D and Section VII.D"},{"comment":"The design-axis comparison is based on one or two teams per family. Within the self-probing family, Team Nepal has the highest faithfulness (0.74) while IReL@IIT(BHU) has the lowest (0.27), a spread larger than most between-family differences. The paper acknowledges this is exploratory, but Section VIII still concludes that 'explicit reasoning structure, grounding constraints, and calibrated confidence reporting are consistently associated with stronger trust signals.' With this sample, 'consistently' is unsupported. Please either present per-family aggregates with variance and a formal descriptive comparison, or limit the conclusion to the specific teams observed rather than to design families.","section":"Section VII.D and Table I"},{"comment":"The leakage risk is noted but not resolved. The paper states that no complete per-team overlap audit was available, and that Lama4Vision reported substantial image-level overlap. Despite this, Test→Private BLEU drift is used as a robustness measure and teams are compared by it (e.g., CVG-IBA −0.048 vs EndoVision −0.005). If overlap differs across teams, the private-set comparison conflates generalization with contamination. The caveat that drift alone is insufficient is helpful, but the analysis in Section VII.C should either incorporate overlap information as a covariate or be explicitly demoted to a hypothesis-generating observation rather than a team-level robustness ranking.","section":"Section IV and Section VII.C"}],"minor_comments":[{"comment":"Table II(a) reports only the top four teams, but the text claims behavior 'across all nine teams with both splits.' Please provide the full table (or an appendix) to support the reported mean drift and top-quartile statistics.","section":"Table II(a) and Section VII.A"},{"comment":"The non-monotonic complexity profile (BLEU 0.356 at L1, 0.323 at L2, 0.416 at L3) is reported only for the test split. Clarify whether the private set shows the same pattern, and if not, discuss the difference.","section":"Section VII.B"},{"comment":"finding_presence is listed as a strongest class in Subtask 1 (0.929, semantic adjudication) but as the hardest in Subtask 2 (correctness 0.099, rubric score). Label the subtask and metric in both places to avoid an apparent contradiction.","section":"Section VII.B vs Section VII.D"},{"comment":"The caption says the plot is 'from Team Nepal.' Clarify whether these are Team Nepal's outputs or aggregate scores across teams, and what the y-axis scale represents.","section":"Fig. 2(b)"},{"comment":"The replacement of the originally proposed expert evaluation with LLM adjudication is a significant protocol change. Consider discussing its implications more prominently, since it directly affects how readers interpret the absence of clinician validation.","section":"Section V.B"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-structured challenge retrospective with valuable recommendations, and the authors' honesty about limitations is commendable. However, the strongest claims about explanation trustworthiness hinge on a single unvalidated judge that the authors themselves flag as potentially biased. For a serious journal, the manuscript needs either a human-validated sample (even a small one) or a clear downgrade of the central claim. The design-axis conclusions also need to be tempered or supported with per-family error bars. The paper might be better suited to an evaluation-focused venue unless these points are strengthened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is not a new architecture or a new benchmark. It is a post-hoc, cross-system analysis of the Medico 2025 challenge on GI endoscopy VQA, written by the challenge organizers. The interesting part is the quantified evidence that leaderboard accuracy (BLEU/ROUGE) does not align with explanation quality: they show BLEU drift between test and private splits, class-wise semantic adjudication profiles, non-monotonic complexity trajectories, and a consistent clarity-high / faithfulness-low pattern. The five evaluation recommendations in Section VIII are practical and worth adopting. The paper is honestly written: it repeatedly says the evidence is correlational, not ablation-based, and flags split-hygiene concerns from one team. That degree of candor is welcome.\n\nThe main soft spot is exactly what the stress-test says. All Subtask 2 trustworthiness scores come from Qwen3-30B-A3B as a rubric-based judge, with no clinician validation at scale. The paper itself admits in V.D that this judge may introduce alignment bias. The central claim — that fluency is easier to optimize than faithful, complete reasoning — is derived from the gaps between Clarity, Faithfulness, and Completeness in that judge's outputs. If that model systematically rewards fluent prose or Qwen-family outputs, the whole interpretation shifts. The authors do label this as a proxy, but they could have done more: for example, report a small human-annotated subset, or show that score patterns hold across two or three different judge models. As it stands, the central comparison is a single-model proxy with no external anchor.\n\nOther soft spots: family-level comparisons have one or two teams per cell, so the design-axis conclusions are exploratory. No code or analysis data are released, so the numerical diagnostics are not independently checkable. And the authors are the challenge organizers and dataset creators, yet there is no explicit conflict-of-interest statement. None of these are disqualifying, but they matter for a paper whose thesis is about evaluation methodology. If the authors expect us to adopt their recommendations, they should follow them: release the analysis pipeline, report all nine teams in the main tables, validate or more strongly caveat the judge, and add a COI statement.\n\nBottom line: the paper deserves a serious referee. It is a useful, honest, incremental contribution to evaluation practice for a safety-relevant niche. It should be sent to review, with the expectation of moderate revision toward more transparent artifacts and stronger caveats on the judge. I would bring it to reading group and cite it if I were working on medical VQA evaluation.","headline":"A well-hedged retrospective of a MedVQA challenge; the central descriptive claims are plausible, but the trustworthiness rankings rest entirely on an unvalidated LLM judge and no analysis artifacts are shipped.","tokens_in":11161,"tokens_out":3829,"would_cite":true,"duration_ms":29703,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In medical visual question answering, top lexical answer scores do not reliably translate into faithful and complete clinical explanations, so evaluation must look beyond leaderboards.","keywords":["multimodal VQA","medical visual question answering","explainability","faithfulness","robust evaluation","parameter-efficient fine-tuning","gastrointestinal endoscopy","semantic adjudication"],"falsifier":"Collect answers and explanations from the nine systems and have both a panel of gastrointestinal clinicians and the rubric-based LLM judge score them. If the clinicians rank systems differently from the judge, or if an answer-optimized system with no grounding is judged as faithful as a structured-grounding system, then the central claim that answer accuracy decouples from reasoning trustworthiness loses its evidentiary base.","tokens_in":10302,"feed_emoji":"🩺","tokens_out":5471,"duration_ms":46560,"temperature":0.7,"pith_summary":"This paper retrospectively analyzes nine systems that participated in a gastrointestinal endoscopy visual question answering (VQA) challenge, comparing how well they answer questions and how well they explain those answers. It argues that parameter-efficiently adapted models can top lexical leaderboards without producing clinically trustworthy reasoning: explanation quality varies widely on faithfulness, clinical relevance, and completeness, while fluency stays high. Systems that enforce structured reasoning and explicit visual grounding show more reliable explanation behavior, but the evidence is correlational, not causal. The authors conclude that medical VQA evaluation needs semantic correctness checks, evidence-linked explanation standards, leakage-aware data splits, and lightweight robustness and calibration tests.","feed_headline":"Medical VQA accuracy does not guarantee trustworthy reasoning","feed_subtitle":"Top-scoring answers often come with weak clinical explanations; evaluation needs more than word-overlap metrics.","key_machinery":"The central machinery is a cross-system comparison: nine independently built systems, most adapting pretrained vision-language backbones with parameter-efficient fine-tuning, were scored on two connected tasks (answer generation and explanation generation). Explanation quality was measured by a rubric-based large-language-model adjudicator on five dimensions (correctness, faithfulness, clinical relevance, clarity, completeness). The analysis then maps design choices—self-probing pipelines that generate auxiliary clinical sub-questions before the final answer, multi-task grounded learning with vision-language grounding supervision, unified answer-plus-explanation heads, and answer-focused bas","core_discovery":"The paper's central discovery is a decoupling: a system can rank at or near the top on standard lexical answer metrics while producing explanations that score notably lower on faithfulness and completeness, and vice versa. In the nine-system comparison, the largest score ranges across systems were in faithfulness, clinical relevance, and completeness, not in clarity, indicating that fluent language is easier to achieve than grounded justification. The paper also documents that difficulty is non-monotonic with question complexity and that fine-grained spatial/color classes are the persistent weak points. Its main claim follows from these patterns: because answer-level gains do not reliably tr","pith_inferences":["The correlational link between structured reasoning or grounding and better explanation scores would be much stronger if tested by ablation on a single model; a controlled comparison turning these design elements on and off is the natural next experiment the paper does not run.","Because the explanation judge is a large language model that may favor fluent text, the observed clarity-versus-faithfulness gap could partly be judge bias; having clinicians score a sample of the same explanations would test whether the ranking survives human review.","The answer-reasoning decoupling likely generalizes to other medical imaging VQA domains, such as radiology or pathology, where spatial and color language is similarly central and lexical benchmarks are known to be weak.","The paper's proposed standardized explanation schema could be turned into a clinical safety filter: any explanation lacking an evidence region or a confidence estimate would be automatically flagged, a testable product implication the paper leaves implicit."],"forward_implications":["Benchmark rankings based on lexical overlap should be supplemented with semantic adjudication, exact-match counts, and normalized yes/no checks before any deployment decision is made.","Explanation outputs should be standardized with evidence links, such as visual regions and confidence values, so faithfulness and completeness can be audited.","Medical VQA data splits should be image-disjoint and accompanied by explicit overlap audits, since QA-level splitting creates a structural leakage risk.","Systems should be required to report lightweight robustness checks, such as corruption and transformation tests, along with calibration metadata.","Question classes requiring fine spatial or color discrimination should be reported separately, since they drive persistent failure even in otherwise strong systems."],"fun_headline_variants":["Top VQA scores don't guarantee faithful clinical reasoning","Medical VQA: accuracy ≠ trustworthy explanations","Leaderboard winners may fail at clinical reasoning","VQA accuracy masks weak clinical explanations","Beyond accuracy: the real test for medical VQA"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole comparison of explanation quality rests on the assumption that a large language model's rubric scores are a valid proxy for clinically meaningful explanation quality; the paper acknowledges the judge was not validated by clinicians and may be biased if submitted systems share its model family.","fun_headline_variants_meta":{"raw":{"variants":["Top VQA scores don't guarantee faithful clinical reasoning","Medical VQA: accuracy ≠ trustworthy explanations","Leaderboard winners may fail at clinical reasoning","VQA accuracy masks weak clinical explanations","Beyond accuracy: the real test for medical VQA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000151,"raw_usage":{"total_tokens":979,"prompt_tokens":625,"completion_tokens":354,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":369,"completion_tokens_details":{"reasoning_tokens":285}},"tokens_in":369,"tokens_out":354,"duration_ms":3687,"temperature":1.0,"reasoning_tokens":285,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T23:44:51.574891+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect answers and explanations from the nine systems and have both a panel of gastrointestinal clinicians and the rubric-based LLM judge score them. If the clinicians rank systems differently from the judge, or if an answer-optimized system with no grounding is judged as faithful as a structured-grounding system, then the central claim that answer accuracy decouples from reasoning trustworthiness loses its evidentiary base.","supporting_citations":[],"review_version":1}