{"id":"5cdbdc08-a74d-40c7-8b9b-5c630b693b4c","arxiv_id":"2505.18319","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MatVQA is a new materials-science visual QA benchmark with automated shortcut removal, and current multimodal LLMs score at most about 52% on it.","lead":"The paper introduces MatVQA, a benchmark of 1,325 multiple-choice questions that test multimodal AI models on visual reasoning from materials science figures. It also presents MArxivAgent, an automated pipeline that removes text-only shortcuts so that models must inspect the image itself.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stage 2 caption-shortcut removal verifies only consistency with an LLM reasoning path, not that the final gold answer is determinable from the image; with a small unreported audit, MatVQA's labels and thus its headline accuracy may be ungrounded.","rationale":"The reader's weakest assumption—that the two-stage refinement preserves the correctness and image-groundedness of the gold labels—is exactly the load-bearing point, and the manuscript provides no direct evidence for it. Algorithm 1's Stage 2 uses an LLM-based consistency checker that only enforces agreement with the extracted reasoning path, not with the original Q&A pair or with the figure. This is a circularity risk because the same style of LLM is used to generate, refine, and check the items; the only non-circular check is the 20% human audit, and Section 4.3 reports no audit scores, no inter-annotator agreement, and no procedure for determining whether experts could answer from the image alone. The ablation in Table 3 shows that rewriting reduces model accuracy, but that is compatible with the rewritten items becoming ambiguous or relying on priors rather than becoming genuinely visual. The small Quantitative split (7 items) is a secondary concern that does not affect the main causal-heavy conclusion, so the central issue remains label validity. If the proposed expert grounding study shows high expert–gold agreement and high inter-expert agreement, the concern would be resolved and the benchmark's central claim would be supported. If not, the accuracy numbers in Table 2 cannot be interpreted as measuring visual-scientific reasoning. The reader's CONDITIONAL verdict is therefore appropriate; this stress-test does not change it.","tokens_in":16017,"tokens_out":3840,"duration_ms":32771,"concrete_test":"Conduct a blind expert grounding study on a random sample (e.g., 100) of final MatVQA items. Give two independent materials-science experts (not authors) the image, the question, and the options, with the caption and source paper withheld. Ask each expert to (i) select the correct option and (ii) rate whether the gold answer is determinable from the image alone. Compute expert–gold agreement, inter-expert agreement, and the proportion of items flagged as ambiguous or unanswerable. If expert–gold agreement is not significantly above chance (25% for 4 options), or inter-expert agreement is low, or more than 10% of items are flagged as not image-answerable, then Stage 2 has not preserved answerability and the benchmark scores are not a clean measure of visual grounding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that MatVQA measures fine-grained visual-scientific reasoning and that 51.9% accuracy reflects a large capability gap—requires that every final MCQ's gold answer is correct and uniquely determined by the figure. Algorithm 1's Stage 2 does not establish this. After the rewriter modifies the question and options (line 26), Checker_caption only verifies that the revised pair follows the same extracted reasoning chain (lines 27–28); it never checks that the new correct answer is entailed by the image, nor that the options are mutually exclusive under the image. The paper's own description (Section 4.2) says Stage 2 'results in larger modification on generated questions' than Stage 1, so semantic drift is possible. The human audit (Section 4.3) is limited to a random 20% sample, has no reported scores, no inter-annotator agreement, and no check that experts solved items from the image alone without the paper's caption or context. Since Table 3's large accuracy drops only show that the rewriting made LLMs worse, they do not show that the rewritten questions became more visually grounded; they could have become ambiguous or reliant on priors. If even a fraction of gold labels are wrong or underdetermined, the headline accuracy numbers and the claim that current MLLMs perform poorly at visual-scientific reasoning are not a valid measure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MatVQA, a 1,325-item multiple-choice visual question answering benchmark for materials science, automatically generated from 44 recent arXiv papers via the MArxivAgent pipeline. The questions are organized into four structure-property-performance (SPP) reasoning tasks (causal, comparative, quantitative, hypothetical), derived from extracted reasoning chains, and refined in two stages that iteratively remove language shortcuts and caption shortcuts using LLM evaluators, rewriters, and consistency checkers. The authors benchmark 17 open- and closed-source MLLMs and report that the best model, Claude-3.7-Sonnet, reaches 51.9% overall accuracy, concluding that current MLLMs have a substantial gap in fine-grained visual-scientific reasoning. A random 20% of items is said to have been audited by two materials-science experts. The abstract also claims a comparison with human experts, but no human accuracy numbers appear in the main text.","tokens_in":16311,"tokens_out":8274,"duration_ms":71307,"significance":"If the gold labels are trustworthy, MatVQA would be a useful and scalable evaluation resource for multimodal scientific reasoning in materials science. The paper has clear strengths: the dataset and evaluation code are publicly released; the benchmark spans a wide range of materials subfields; the two-stage shortcut-removal procedure is a sensible response to known textual-bias problems; and the systematic evaluation of 17 models is a useful contribution. However, the central validity claim rests on the correctness and visual groundedness of the gold answers, and that grounding is not established by the evidence currently reported. The human audit is described only qualitatively, the image-only answerability of final items is not checked, the seven-item quantitative split is too small to support per-task accuracy claims, and at least one table contains internally inconsistent numbers. These issues are addressable in revision, but they are load-bearing for the paper's headline claims.","major_comments":[{"comment":"The human quality audit is described as a random 20% sample reviewed by two materials-science experts for scientific accuracy, logical consistency, and contextual relevance, with 'uniformly high marks,' but no scores, no per-criterion distributions, no inter-annotator agreement, and no resolution procedure for disagreements are reported. Because this audit is the only evidence presented that the machine-generated gold labels are correct, the absence of quantitative audit results is a load-bearing gap. Please report the full audit outcome, including how many items were accepted, revised, or rejected, the actual score ranges, and agreement statistics.","section":"§4.3"},{"comment":"Stage 2 of the refinement loop verifies only that (i) the evaluator can no longer answer correctly from the stem, options, and caption, and (ii) the rewritten question-answer pair follows the same extracted reasoning chain (Algorithm 1, lines 25-30). It does not check that the gold answer is entailed by the image alone, that the image determines a unique option, or that the rewritten options are mutually exclusive under the image. Since Section 4.2 states that Stage 2 'results in larger modification on generated questions,' the risk of semantic drift or ambiguity is real. The reported accuracy drops therefore show that rewriting made the LLMs fail, not that the rewritten questions are genuinely visually grounded. Add an image-only validation: have experts or a carefully controlled protocol answer a sample from the figure alone, without captions or surrounding text, and report the proportion of items with a uniquely determinable correct answer.","section":"§4.2 / Algorithm 1"},{"comment":"Figure 4 shows the correct answer changing from A in the raw sample to B after language-shortcut removal, with no explanation. If Stage 1 is claimed to preserve semantic equivalence, a changed gold label suggests either the raw label was incorrect or the rewriting altered the scientific content. Please clarify this example and describe how label changes during refinement are handled; otherwise the stability of the gold standard is not established.","section":"Figure 4"},{"comment":"The quantitative split contains only 7 items, yet Table 2 reports per-split accuracies for it (e.g., 57.1%, 28.5%) and Section 5.1 draws conclusions such as quantitative items appearing 'easy' and large models outperforming small models by +9.5 pp on this split. With 7 items, a single answer shifts accuracy by about 14 percentage points, so these task-level comparisons are statistically unstable. Either expand the quantitative split or remove per-split quantitative conclusions and report confidence intervals instead.","section":"§5.1 / Table 2"},{"comment":"The abstract and introduction claim that a subset of models was compared against human experts, but the main text contains no human accuracy results, no expert protocol, and no human-baseline table. For a benchmark whose purpose is to show that MLLMs perform poorly relative to research-level human reasoning, the human baseline is essential evidence. Add expert performance numbers with a clear protocol (e.g., image-only, no caption, fixed time budget) or remove the claim.","section":"Abstract / §5.1"},{"comment":"Several rows in Table 2 are internally inconsistent with the split sizes in Table 1. For example, Claude-3.5-Haiku is reported with overall 44.7%, but its split accuracies (Caus 32.9%, Hypo 38.3%, Quan 57.1%, Comp 37.5%) weight to approximately 34.5% given the 950/112/256/7 split; MOL-VL-7B's reported overall of 23.6% is also inconsistent with its split accuracies, which weight to about 28.4%. Please recompute and report corrected overall accuracies, or explain the discrepancy if a different evaluation subset was used.","section":"Table 2"}],"minor_comments":[{"comment":"There are several typographical issues: 'aviod' in Section 4.2, 'consistancy' and 'consistencey' in Section 4.1, 'deatails' and 'splited' in Section 5.1, and 'in varies domain' in the contributions list in Section 1. Please copy-edit the manuscript.","section":"§4.2, §4.1, §5.1"},{"comment":"The inline text in Figures 2 and 3 is partially garbled, with fragments such as 'BCCX' and '21.79° rotated TBLG' in the diagram boxes. A redrawn, larger version with legible example questions would make the pipeline much easier to follow.","section":"Figures 2 and 3"},{"comment":"The error-analysis examples in Figures 6 and 7 show blank or misordered options (e.g., options (2), (3), (4) are empty in the rendered excerpt). Please ensure the full question text and all options are visible in the final version.","section":"Appendix C"},{"comment":"Table 3 reports large accuracy drops across refinement stages but no confidence intervals or significance tests. Because the same items are evaluated before and after refinement, a paired test (e.g., McNemar's test) and confidence intervals would strengthen the claim that the drops are systematic rather than noise.","section":"Table 3"},{"comment":"The comparison table lists MacBench without a complete citation; reference [5] gives only a title and no venue, arXiv identifier, or year. Please provide full bibliographic details for all benchmark citations.","section":"Table 1(a)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the dataset release is a potentially useful community resource. However, the validity of the gold labels is the central issue: the human audit is not quantified, the image-only answerability of final items is not established, and at least one table contains arithmetic inconsistencies. I would recommend asking for a revised version with a full audit report, an image-only human validation study, corrected tables, and removal or qualification of the seven-item quantitative conclusions. These are substantial but fixable changes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on the MatVQA paper. The benchmark is a real contribution—first materials-science visual QA at research level, 1325 items from recent arXiv papers, and the caption-shortcut removal stage is genuinely new and well-motivated. The pipeline (MArxivAgent) is clever and the two-stage refinement produces large, consistent accuracy drops (roughly 12% then 19%), which shows the rewriting is doing something nontrivial. The released data and code are a plus. That's the good part.\n\nThe soft spot is right where the stress-test says: Stage 2 verifies only that the rewritten question-answer pair is consistent with the original reasoning chain, not that the correct answer is actually entailed by the image. The paper even says Stage 2 produces 'larger modification' than Stage 1, so semantic drift is plausible. The accuracy drops after refinement show models get worse, but not why—they could be failing on ambiguous or prior-solvable questions just as easily as on genuinely visual ones. Without an image-entailment check or a proper human baseline with experts solving from the figure alone, the headline number (51.9% top model) doesn't yet support the claim that current MLLMs are bad at visual-scientific reasoning.\n\nThe human audit is too thin to fix this: 20% of items, two experts, no reported scores, no inter-annotator agreement, no indication that experts were isolated from captions. The Quantitative split is 7 items, so the per-task numbers there are noise. And the abstract and intro promise a comparison against human experts that never appears in the experiments. That's a missing result, not a minor omission.\n\nSo: the paper is worth engaging with, and I'd send it to peer review—the benchmark and pipeline deserve scrutiny and the authors have done real work. But I'd expect heavy revision: report the audit fully, run a proper human baseline with image-only access, add confidence intervals, and either add an image-entailment check to the pipeline or at least analyze whether rewritten questions are answerable without the figure. My own verdict is conditional, not accept. It's a solid resource in the making, not yet a solid evaluation standard.","headline":"A genuinely useful new benchmark and pipeline, but the gold labels aren't shown to require the image, so the headline accuracy numbers don't yet support the central claim.","tokens_in":16855,"tokens_out":2527,"would_cite":false,"duration_ms":20661,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MatVQA argues that current multimodal AI models can answer only about half of research-level materials-science questions that require reading figures, not just text.","keywords":["MatVQA","MArxivAgent","multimodal large language models","visual question answering","materials science","shortcut elimination","structure-property-performance reasoning","scientific reasoning benchmark"],"falsifier":"Take the released MatVQA questions, replace every image with a blank placeholder while keeping the stem and options, and run the same models under the benchmark's prompt; if any model scores far above random (25% for four-option items), textual shortcuts remain and the central claim fails.","tokens_in":15793,"feed_emoji":"🔬","tokens_out":6746,"duration_ms":50471,"temperature":0.7,"pith_summary":"The paper introduces MatVQA, a benchmark of 1,325 multiple-choice questions that test whether multimodal AI models can reason about materials-science figures at research level, not just read text. The authors argue that existing materials-science QA datasets are mostly text-based and can be solved with shortcuts, so they build an automated pipeline that rewrites questions until the image is genuinely required. On this benchmark the best of 17 models reaches 51.9% accuracy, which the paper reads as a large open gap in visual-scientific reasoning. If correct, MatVQA gives the field an evaluation standard that separates genuine figure-reading reasoning from language and caption guessing, and it scales automatically from new literature.","feed_headline":"Top AI models score only 52% on new materials-science visual quiz","feed_subtitle":"A 1,325-question benchmark forces models to read microscopy and diffraction figures instead of guessing from captions.","key_machinery":"MArxivAgent, an automated three-stage pipeline. It parses papers into text and figures, extracts a verifiable reasoning path structured around structure, property, performance, processing, and environment, and generates multiple-choice questions; then it runs an evaluator agent that answers without the image to detect language shortcuts, rewrites the item, and repeats with a caption-only evaluator to detect caption shortcuts. A consistency checker ensures each rewrite preserves the original scientific claim or reasoning path. This two-stage loop is what carries the benchmark's claim to measure visual-scientific reasoning rather than text guessing.","core_discovery":"MatVQA claims to be the first benchmark for research-level multimodal reasoning in materials science. It contains 1,325 questions organized into four structure-property-performance tasks: quantitative, comparative, causal, and hypothetical variation. Generated by the MArxivAgent pipeline from recent materials-science papers and 378 unique figures, each item is passed through iterative refinement that removes language shortcuts (answerable from wording) and caption shortcuts (answerable from the figure caption), forcing models to inspect low-level visual features such as diffraction peaks and lattice fringes. Benchmarking 17 open- and closed-source multimodal models under chain-of-thought prompting, the best model scores 51.9%, with ablations showing raw questions are answered with roughly 76–83% accuracy before refinement and roughly 39–52% after, evidence that the shortcut removal meaningfully raises difficulty.","pith_inferences":["Beyond the paper, the same shortcut-elimination recipe could be ported to other image-heavy sciences, such as biology microscopy or medical imaging, wherever captions and wording can leak answers; each field would need its own reasoning-path ontology.","The heavy reliance on LLM-generated reasoning paths means part of MatVQA's validity depends on the generator's domain knowledge; an independent human re-judgement of all 1,325 gold answers, not just a random 20%, would directly test for hidden label noise.","The tiny quantitative split, only 7 items, makes the 57% quantitative score statistically fragile; expanding that split should be a priority before drawing conclusions about numerical structure-property-performance reasoning.","A testable extension: measure whether a model pre-finetuned on MatVQA-style visual question pairs transfers to unseen materials papers; if transfer is strong, the benchmark could become a training resource as well as an evaluation set."],"forward_implications":["If MatVQA measures what it claims, research-level multimodal reasoning in materials science is currently unsolved: the best model is barely above half accuracy, and materials-finetuned models fall below 24%.","The automated pipeline makes benchmark construction scalable: the same procedure can be re-run on new papers, which the paper plans to use to expand MatVQA to roughly 12,000 questions.","The ablation quantifies that caption shortcuts are the larger leak: removing language shortcuts drops accuracy by about 10–15 percentage points, and additionally removing caption shortcuts drops it by another 18–20 points.","Comparing task types shows comparative reasoning is the weakest area and causal reasoning dominates the benchmark, so future model improvement should target joint perception of multiple structures and multi-hop causal chains."],"supporting_citations":[{"why":"Supplies the forensic study of multimodal biases and the language-shortcut removal method that MatVQA's two-stage refinement extends, along with the microscopy benchmark it positions against.","marker":"[13]"},{"why":"Represents the text-based scientific QA benchmark that MatVQA says fails to capture visual research-level reasoning.","marker":"[8]"},{"why":"MaScQA is the text-based materials science QA dataset that MatVQA aims to move beyond by requiring visual evidence.","marker":"[49]"},{"why":"ScienceQA is the existing multimodal science benchmark whose high-school exam level MatVQA contrasts with research-level materials tasks.","marker":"[30]"},{"why":"Provides the chain-of-thought prompt template used to evaluate all 17 models in the benchmark.","marker":"[48]"},{"why":"LLM4Mat-Bench is the large materials property-prediction benchmark that MatVQA differentiates from by targeting visual structure-property-performance reasoning.","marker":"[41]"}],"fun_headline_variants":["AI models flunk new materials-science visual reasoning test","New benchmark forces AI to read microscopy figures, not captions","MatVQA: 1,325 questions that trip up multimodal AI","Best AI scores 52% on new visual materials-science benchmark","Why AI struggles with materials images: new MatVQA benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the iterative rewrites preserve scientifically correct, figure-grounded answers; because only a random 20% of items are human-audited, any ambiguity or leakage that survives would make the model scores misread as visual-scientific reasoning.","fun_headline_variants_meta":{"raw":{"variants":["AI models flunk new materials-science visual reasoning test","New benchmark forces AI to read microscopy figures, not captions","MatVQA: 1,325 questions that trip up multimodal AI","Best AI scores 52% on new visual materials-science benchmark","Why AI struggles with materials images: new MatVQA benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1379,"prompt_tokens":967,"completion_tokens":412,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":324}},"tokens_in":583,"tokens_out":412,"duration_ms":3666,"temperature":1.0,"reasoning_tokens":324,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:32:50.926257+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the released MatVQA questions, replace every image with a blank placeholder while keeping the stem and options, and run the same models under the benchmark's prompt; if any model scores far above random (25% for four-option items), textual shortcuts remain and the central claim fails.","supporting_citations":[{"cited_title":"The sciqa scientific question answering benchmark for scholarly knowledge","cited_arxiv_id":null,"evidence_quote":"Represents the text-based scientific QA benchmark that MatVQA says fails to capture visual research-level reasoning."},{"cited_title":"Mascqa: investigating materials science knowledge of large language models","cited_arxiv_id":null,"evidence_quote":"MaScQA is the text-based materials science QA dataset that MatVQA aims to move beyond by requiring visual evidence."}],"review_version":1}