{"id":"764ba69b-7c7d-44b5-804f-992225078a46","arxiv_id":"2607.10240","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Official short-answer VQA scores undercount semantic success by several points on text-rich benchmarks because automatic evaluators reject acceptable surface-form variants, with sensitivity structured by answer-contract type.","lead":"Short-answer VQA benchmarks often mark correct model answers wrong when wording differs from the expected string. This paper shows that on text-rich datasets up to half of official errors are semantically fine, so leaderboard numbers mix real understanding with surface-form compliance.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged text-only judge limit; that limit is mitigated and does not overturn the central undercount claim.","rationale":"The strongest claim is an empirical measurement claim, not a theoretical derivation. The manuscript supplies multi-model, multi-benchmark quantitative evidence (Tables 3–7), dual-judge robustness, human validation, contract-mix tables, perturbation flips, and a released CPU-only repair pipeline that recovers part of the undercount without relying on the judge. The reader's weakest_assumption is precisely the softest remaining point, but it is already quantified, cross-checked, and does not invert the direction or structure of the findings. Closed-source sample sizes and the heuristic taxonomy are secondary and do not carry the central claim. Therefore no verdict adjustment is warranted: CONDITIONAL with high confidence remains appropriate.","tokens_in":17544,"tokens_out":520,"duration_ms":5727,"concrete_test":"On a stratified 200-item subsample of judge-accepted official errors drawn from ST-VQA/TextVQA readout and multi-span buckets, re-label with humans who also see the image (question+image+gold+output). If human-with-image acceptance falls more than ~5–8 pp below the text-only judge rate, the undercount magnitudes need downward revision; otherwise the proxy remains adequate for the headline claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (up to ~half of official errors on text-rich short-answer VQA are surface-form false negatives, structured by answer contract) rests on treating a text-only LLM judge as a proxy for semantic acceptability. The reader correctly flags residual risk: the judge never sees the image, residual disagreements concentrate in other-open/scalar, and visual grounding is not re-verified. However, this is already the paper's own stated limitation (§6), is bounded by a stratified 570-item human audit (97.6% precision, κ=0.917 across all buckets/models/benchmarks), and is corroborated by a second independent text-only judge with near-perfect per-dataset FN-rate correlation (r=0.999). Heuristic RA-Eval and deterministic CPU-only bidirectional repair further recover a non-trivial fraction of the same pool without any LLM judge. No stronger internal inconsistency or unmitigated load-bearing flaw is present that would reverse the measured undercount pattern or the contract-structure finding.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This paper argues that short-answer VQA leaderboard scores conflate semantic correctness with surface-form survival under automatic evaluators. Across six VLMs and six benchmarks, the authors audit over 37k official errors with a text-only LLM semantic judge (human-validated at 97.6% precision, 95.5% recall, κ=0.917) and a second independent judge that reproduces the same benchmark-level false-negative pattern. On text-rich sets such as ST-VQA, up to ~48% of official errors are judged semantically acceptable; the undercount is structured by an operational four-bucket answer-contract taxonomy (scalar, readout, identifier, multi-span), with extractive and multi-span answers far more evaluator-sensitive than scalar ones. Benign prompt/context rewrites flip item-level official correctness at substantial rates, and a deterministic CPU-only bidirectional contract repair recovers about +1.34 pp overall without re-running models. The authors conclude that official scores should be reported with semantic audits and contract-mix diagnostics.","tokens_in":17857,"tokens_out":1302,"duration_ms":20133,"significance":"If the measured undercount and contract structure hold, the paper is a useful measurement contribution for multimodal evaluation: it shows that adjacent score gaps and official error pools on text-rich short-answer VQA can mix genuine failures with surface-form rejections, and that this mixture is not uniform across answer types. Strengths include scale (37k+ audited errors), dual-judge robustness with near-perfect per-dataset FN-rate correlation (r=0.999), stratified human validation, an operational contract taxonomy anchored partly in benchmark-native metadata, and a released deterministic repair pipeline that recovers a non-trivial fraction of the undercount without LLM judges. The practical reporting template and the separation of official, true-overlay, and judge layers are actionable for both benchmark users and dataset designers.","major_comments":[{"comment":"§3.3 and §5.1: The central undercount claim treats text-only semantic equivalence (question + gold + output, no image) as the operational definition of “semantically acceptable.” Human validation and the second judge also operate on this text-side interface, so the audit correctly measures official false negatives relative to gold references, not re-verified visual grounding. That is a defensible diagnostic quantity, but the main results still occasionally read as if they quantify pure task success. Please state more explicitly in §5.1 (not only §6) that judge-accepted errors are gold-equivalent surface variants, and that residual risk of gold-reference incompleteness or ungrounded but gold-matching paraphrases is not eliminated by the audit.","section":"§3.3, §5.1"},{"comment":"Appendix B vs. §5.2: For raw judge-FN separation, dataset-only grouping has higher η² (0.050) than the contract taxonomy (0.024), while contracts are stronger for perturbation flips. The paper already notes this, but the main-text claim that instability “follows answer contract rather than benchmark name alone” (RQ2 / abstract) should be tightened to match the evidence: contracts explain a useful cross-benchmark structure and perturbation sensitivity, but do not dominate dataset identity for baseline FN rates. A short quantitative caveat in §5.2 would prevent over-reading Table 4–5 as a full replacement for benchmark-level analysis.","section":"§5.2, Appendix B"}],"minor_comments":[{"comment":"Table 3: “Acc. Err.” is easy to misread as an accuracy-style error rate; the footnote helps, but renaming the column (e.g., “Judge-accepted share of official errors”) would reduce confusion.","section":"Table 3"},{"comment":"Table 7 and surrounding text: “ST-VAQ” / “TextVAQ” appear as typos for ST-VQA / TextVQA in the repair table and Figure 3 captions.","section":"Table 7, Figure 3"},{"comment":"§4 Perturbation protocol: B3 is described as an adversarial control excluded from main flip-rate analysis; a one-line statement of what B3 does (empty/minimal instruction) would make the exclusion criterion clearer without appendix diving.","section":"§4"},{"comment":"Eq. (2) is conceptual (S_off = f(S_sem, κ)) and never operationalized beyond the diagnostic split; either keep it clearly labeled as notation-only or briefly say that f is the benchmark’s string/ANLS/relaxed rule rather than a fitted function.","section":"§3.1"},{"comment":"Closed-source rows in Table 3 use stratified samples; pointing readers more visibly from the table caption to Appendix G CIs would help avoid treating those Acc. Err. figures (e.g., DocVQA GPT-5.4 77.8%) as full-pool estimates.","section":"Table 3, Appendix G"},{"comment":"A few reference entries appear only loosely related to VQA evaluation methodology (e.g., multi-sensor tracking, image fusion, federated GNN). Trimming or relocating peripheral citations would tighten the related-work signal.","section":"References"}],"recommendation":"minor_revision","confidential_remarks":"Solid evaluation-methodology paper with careful multi-layer validation; the text-only judge limitation is real but already largely owned by the authors and bounded by human + second-judge checks plus non-LLM repair. I do not see a load-bearing flaw that would reverse the undercount pattern. Fit is good for a CV/ML venue that publishes benchmark and evaluation analyses. Citation list is a bit padded with loosely related prior work from the author group; not a reason to reject, but worth a quiet editorial note if space is tight."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing to know is that this is a careful empirical audit, not a theory paper. On ST-VQA and TextVQA, 21–48% of official automatic errors are judge-accepted as semantically fine; the undercount is large enough to swamp adjacent model gaps, and it tracks answer contract (readout and multi-span far worse than scalar) more than dataset name. That is the result that matters for anyone who still mines official error pools or compares leaderboard deltas.\n\nWhat is new is the scale and the packaging: 37k+ audited errors across six VLMs and six benchmarks, a human-validated text-only judge (97.6% precision, κ=0.917 on 570 stratified items), a second independent judge that reproduces the same per-dataset FN pattern (r=0.999), contract-mix tables that explain why ChartQA-M barely moves while ST-VQA jumps ~9 pp, benign prompt/context flips of 6–23%, and a deterministic CPU-only bidirectional repair that recovers ~1.3 pp overall without re-running models. Prior work already said exact match is brittle; this paper quantifies how much, where, and how much is recoverable. The S_off = f(S_sem, κ) framing is just bookkeeping, but it keeps the claims clean.\n\nSoft spots are real but already flagged by the authors and do not reverse the pattern. The judge never sees the image, so visual grounding is not re-checked; residual disagreements sit in other-open and scalar. Closed-source numbers are on samples. The four-bucket taxonomy is heuristic. None of that invents the undercount: RA-Eval substring recovery and the break-free reference-side repair recover a non-trivial slice without any LLM judge. Citation pattern is normal; the math is simple counting and agreement stats, not load-bearing derivation.\n\nThis is for people who build or report short-answer VQA and adjacent extractive multimodal eval. It will not rewrite the field, but it should change how error analyses and score gaps are written up. I would send it to peer review; a serious referee will tighten the taxonomy language and the closed-source CIs, not reject the measurement. Worth engaging.","headline":"Solid measurement paper: up to half of official errors on text-rich short-answer VQA are surface-form false negatives, structured by answer contract, with dual-judge audit and partial CPU repair.","tokens_in":18458,"tokens_out":552,"would_cite":true,"duration_ms":4923,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Short-answer VQA scores mix semantic correctness with surface-form compliance, and on text-rich benchmarks up to half of official errors are acceptable answers rejected only for wording.","keywords":["visual question answering","evaluation methodology","vision-language models","benchmark analysis","answer contracts","short-answer VQA","evaluator instability"],"falsifier":"A large human re-audit of the same official-error pool that found most judge-accepted false negatives were actually wrong or visually unsupported, and that extractive and multi-span items were not more often falsely rejected than scalar items, would overturn both the undercount size and the contract-structure claim.","tokens_in":18444,"feed_emoji":"⚖️","tokens_out":688,"duration_ms":16437,"temperature":0.7,"pith_summary":"Short-answer visual question answering benchmarks still score models mainly by string matching, so a leaderboard number confuses two different things: whether the answer is right in meaning, and whether it matches the form the automatic scorer expects. Auditing more than 37,000 official errors from six vision–language models on six benchmarks with a human-validated semantic judge shows that this gap is large on text-rich datasets: as much as about half of the marked errors are answers humans and the judge accept as correct. The problem is not uniform. It tracks answer type—extractive readouts and multi-span answers are far more sensitive than simple numbers or yes/no—and even mild prompt rewrites flip official item-level outcomes at high rates without changing the task. A cheap deterministic repair of answer and reference forms recovers part of the undercount. The practical claim is that official short-answer VQA scores stay useful only if they are reported with semantic audits and answer-type breakdowns.","feed_headline":"Up to half of VQA 'errors' are correct but wrong form","feed_subtitle":"Text-rich short-answer scores mix meaning with wording; extractive answers take the hit hardest.","key_machinery":"The answer contract: the implicit agreement about what form an acceptable answer must take, split operationally into scalar, extractive-readout, identifier-like, and multi-span buckets so that official score can be read as semantic acceptability filtered by contract compatibility.","core_discovery":"Across six vision–language models and six short-answer VQA benchmarks, a human-validated text-only semantic judge finds that 21–48% of official errors on text-rich datasets are semantically acceptable answers rejected purely for surface-form mismatch, producing multi-point score undercounts. The undercount follows answer contract more than benchmark name: extractive readout and multi-span answers are much more evaluator-sensitive than scalar answers. Benign prompt and context rewrites further flip official item correctness at substantial rates, and CPU-only contract repair recovers a measurable fraction of the false negatives.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Up to half of short-answer VQA errors are form not meaning","VQA scores mix correctness with surface-form match","Text-rich VQA: up to 48% of errors are semantically right","Extractive answers drive most evaluator-sensitive VQA fails","Semantic audits show multi-point undercount in short VQA"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The claim depends on treating a text-only language-model judge—given only the question, gold answer, and model output, never the image—as a faithful proxy for whether an answer is semantically acceptable.","fun_headline_variants_meta":{"raw":{"variants":["Up to half of short-answer VQA errors are form not meaning","VQA scores mix correctness with surface-form match","Text-rich VQA: up to 48% of errors are semantically right","Extractive answers drive most evaluator-sensitive VQA fails","Semantic audits show multi-point undercount in short VQA"]},"model":"grok-4.5","effort":"low","cost_usd":0.003706,"raw_usage":{"total_tokens":1217,"prompt_tokens":803,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":37060000,"prompt_tokens_details":{"text_tokens":803,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":342,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":803,"tokens_out":72,"duration_ms":5578,"temperature":1.0,"reasoning_tokens":342,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T13:16:55.363894+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A large human re-audit of the same official-error pool that found most judge-accepted false negatives were actually wrong or visually unsupported, and that extractive and multi-span items were not more often falsely rejected than scalar items, would overturn both the undercount size and the contract-structure claim.","supporting_citations":[],"review_version":1}