{"id":"ab357b20-144d-4324-8639-fb4709a0d96d","arxiv_id":"2508.02645","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A VQA variance-analysis abstract is attached to a different paper about projectile motion reconstruction, so the claimed study is not present in the submitted text.","lead":"The abstract promises an analysis of performance variance across 14 visual question answering benchmarks, but the supplied full text is an unrelated 3D Gaussian splatting paper. This mismatch prevents any evaluation of the claimed VQA study.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's VQA-variance claims have no supporting body: the supplied full text is an unrelated 3D reconstruction paper, so the central claim is unevaluable from this submission.","rationale":"The reader's verdict is UNVERDICTED, and I agree that no scientific verdict on the VQA claims is possible from the supplied text. The reader's weakest_assumption targets a downstream risk: that Cloze-style rewording could change task validity and make variance reductions artifacts. That is a reasonable secondary concern, but it is not the load-bearing issue. The load-bearing issue is prior to any such assumption: the submitted manuscript body contains no VQA experiments, no 14-benchmark variance analysis, and no Cloze evaluation at all. The supplied full text is a coherent 3D reconstruction paper, but it does not address the abstract's claims. Under the review rule that all supplied text is evidence, this mismatch is decisive. I would not move the verdict to REJECT because the possibility remains that the official arXiv listing contains the VQA study and the supplied full text is a packaging error; the correct scientific posture is to withhold judgment pending the actual source. Hence the reader's UNVERDICTED verdict stands unchanged, and my load-bearing concern is closest to the reader's rationale even though it differs from the specific weakest_assumption listed.","tokens_in":7621,"tokens_out":3980,"duration_ms":42606,"concrete_test":"Retrieve the official source of arXiv:2508.02645 (TeX/HTML/PDF) and search the body for the terms 'VQA', 'benchmark', 'stochastic', 'seed', and 'Cloze'. If none of the 14 named VQA benchmarks or any variance/Cloze experiment appears in the methods or results, the mismatch is confirmed and the abstract's claims have no supporting evidence in the submission. If such content does appear in the official source but is missing from the supplied full text, then the supplied text is corrupt and the verdict should be based on the official version.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that MLLM evaluation on VQA benchmarks shows significant variance from stochastic outputs, seed sensitivity, and hyperparameters, and that Cloze-style evaluation can reduce that variance. For this claim to hold, the paper needs variance measurements across the announced 14 VQA benchmarks, seed/framework experiments, and a Cloze-style evaluation comparison. The full text supplied for arXiv:2508.02645 contains none of this: its abstract, method, experiments, and ablations all concern PMGS, a 3D Gaussian Splatting system for projectile-motion reconstruction. No VQA benchmark, multimodal LLM, stochasticity analysis, seed study, or Cloze evaluation appears outside the title/abstract. This is not a subtle hidden assumption; it is a complete internal mismatch between the abstract's promised study and the manuscript's content. The central claim therefore cannot be checked from the submitted document, and any scientific verdict on the VQA claims is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript as submitted is internally inconsistent: the title and abstract announce an empirical study of variance in multimodal large language model evaluation across 14 VQA benchmarks, with seed sensitivity, hyperparameter configuration, instruction finetuning, and Cloze-style evaluation as the central objects of study. The full text, however, is a completely different paper on PMGS, a 3D Gaussian Splatting system for reconstructing projectile motion. The abstract's claims about stochastic model outputs, benchmark ranking instability, and variance-aware reporting are not supported by any methods, experiments, tables, or results in the body of the submission. As a result, the promised VQA study is absent from the submitted artifact, and the central claims cannot be evaluated.","tokens_in":7795,"tokens_out":1566,"duration_ms":19916,"significance":"If the claims in the abstract were properly established, the work would address a genuinely important issue: current VQA benchmark evaluations often rely on point estimates, and documenting variance from seeds, frameworks, and hyperparameters would be a useful corrective for the community. The proposed Cloze-style evaluation as a variance-reduction strategy is also a plausible and testable idea. However, the submitted manuscript provides none of the required evidence. There are no variance measurements, no benchmark results, no seed or framework comparisons, and no Cloze evaluation. The PMGS paper that constitutes the full text is unrelated to the abstract, so the significance of the claimed contribution cannot be assessed from this submission.","major_comments":[{"comment":"The abstract's central claim is that MLLM evaluation on 14 VQA benchmarks exhibits significant variance from stochastic outputs, training seed sensitivity, and hyperparameter configurations, and that Cloze-style evaluation can reduce this variance. The full text contains no VQA benchmarks, no MLLM evaluations, no stochasticity analysis, no seed experiments, and no Cloze-style comparisons; instead it describes PMGS, a 3D Gaussian Splatting system for projectile-motion reconstruction. This is not a missing detail but a complete mismatch between the promised study and the supplied evidence, so the abstract's claims are unsupported by the manuscript.","section":"Abstract vs. Full Text"},{"comment":"The experimental sections report reconstruction metrics (PSNR, SSIM, LPIPS, IoU, ATE, RMSE) on projectile-motion datasets and compare methods such as 4DGS, DynamicGS, MotionGS, and CFGS. None of these experiments address variance in VQA, seed sensitivity, framework non-determinism, model scale, instruction finetuning, or Cloze-style evaluation. The load-bearing empirical claims of the abstract therefore have no supporting results in the paper.","section":"Experiments"},{"comment":"The conclusion recommends variance-aware evaluation methodologies, but this recommendation is not grounded in any data presented in the manuscript. The paper does not demonstrate ranking instability, quantify stochastic output variance, or show that Cloze-style evaluation reduces variance, so the central recommendation is an unsupported assertion rather than a finding.","section":"Conclusion"}],"minor_comments":[{"comment":"The manuscript title, abstract, and full text should describe the same research; in this submission they describe two unrelated papers, which makes the artifact unsuitable for review in its current form.","section":"General"}],"recommendation":"reject","confidential_remarks":"This appears to be a submission error: the full text is a different paper (PMGS on 3D reconstruction) and contains none of the announced VQA variance study. If the authors intended to submit a VQA paper, the correct manuscript was not uploaded. As it stands, the submission cannot be reviewed for the claims in its abstract. I would recommend rejecting this artifact and inviting a corrected resubmission if the proper VQA manuscript is available."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nPunchline: what is under arXiv:2508.02645 is not the VQA variance study promised by the title and abstract. The full text is a different paper, PMGS, about 3D Gaussian Splatting reconstruction of projectile motion. The abstract's claims about variance across 14 VQA benchmarks, seed sensitivity, and Cloze-style evaluation have no methods, data, or results anywhere in the submission. The stress-test note is right: this is a complete internal mismatch, not a subtle gap.\n\nWhat the supplied material does contain is a reasonably coherent 3D reconstruction paper. PMGS proposes a two-stage pipeline for target modeling and motion recovery, adds an acceleration consistency constraint, a dynamic simulated annealing schedule, and a Kalman fusion scheme, with ablations and a code link. On its own terms that portion looks like a plausible engineering contribution, though I have not verified the code or dataset, and there are details I would want checked, such as the acceleration sign conventions and the fixed noise covariances in the Kalman update. But it does not advance the VQA argument at all.\n\nSoft spots: the central VQA claims are unevaluable. No benchmark list, no variance estimates, no seed/framework experiments, no Cloze comparison. The weakest assumption, that rewording questions for Cloze-style evaluation changes the task, is only a passing concern in the abstract, not something the paper investigates. There is also no citation pattern to audit for the VQA part because no VQA literature appears beyond the abstract.\n\nBottom line: this submission cannot go to peer review as the VQA paper. It should be returned as a desk reject, and the authors should resubmit the actual VQA manuscript when it exists. The PMGS content, if genuine, belongs in its own submission under its own title. For anyone interested in VQA evaluation methodology, this arXiv entry currently provides nothing to engage with.","headline":"The abstract promises VQA variance analysis; the actual paper is an unrelated 3D reconstruction manuscript, so the advertised study is unevaluable.","tokens_in":8247,"tokens_out":2846,"would_cite":false,"duration_ms":34144,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that VQA benchmark scores, reported as point estimates, hide variance from stochastic outputs, training seeds, and hyperparameters, and that Cloze-style evaluation should be part of the fix.","keywords":["visual question answering","multimodal large language models","benchmark evaluation variance","point estimates","training seed sensitivity","hyperparameter sensitivity","Cloze-style evaluation","stochastic model outputs"],"falsifier":"Take a fixed set of VQA models and run each on the same benchmark across many seeds and decoding settings; if the score spread is small relative to the gaps between models, the premise of hidden variance collapses. A second check is to compare answer distributions between open and Cloze forms: if rewording shifts which questions are answered correctly, then any variance reduction may be a task change rather than a reliability improvement.","tokens_in":7433,"feed_emoji":"🎲","tokens_out":6925,"duration_ms":70645,"temperature":0.7,"pith_summary":"The paper's target claim is that the standard way of evaluating multimodal language models on visual question answering (VQA) benchmarks—reporting a single accuracy number per model—is misleading, because model outputs, training seeds, and hyperparameter settings introduce variance large enough to shuffle rankings. It proposes to measure that variance explicitly across a set of fourteen VQA benchmarks and to test Cloze-style evaluation, where a question is reformulated as a fill-in-the-blank task, as a way to reduce stochasticity. If the claim is right, published benchmark tables are unstable and evaluation practices should report error bars, multiple seeds, and configuration sweeps. The supplied full text of this paper, however, contains a different manuscript on projectile-motion reconstruction, so the stated analysis and experiments are not present in the document reviewed here.","feed_headline":"VQA benchmark rankings shuffle across seeds and settings","feed_subtitle":"A paper argues for variance-aware reporting and Cloze-style questions so leaderboards stop depending on luck.","key_machinery":"The central object is a variance-analysis protocol for benchmark evaluation: instead of a single accuracy point, the paper tracks the spread of scores across stochastic inference runs, training seeds, framework implementations, model scale, and instruction-finetuning regimes. The proposed stabilizer is Cloze-style evaluation, in which open-ended questions are turned into masked-blank or single-answer completions that constrain the model's response space. The argument is that constraining the output distribution reduces the measurement noise that makes point estimates uninformative.","core_discovery":"On the paper's own terms, the central discovery is that point-estimate benchmarks for multimodal visual question answering are unreliable: a model's measured accuracy depends on random generation, training seed, and framework nondeterminism, and the resulting spread can be as large as the gaps that decide leaderboard ordering. The paper therefore argues that evaluation should be variance-aware, and it explores Cloze-style reformulation of questions as an assessment strategy that might lower variance while preserving the skill being tested. Since the supplied full text contains no VQA experiments, this discovery is presented as the paper's program rather than as something demonstrated in the reviewed document.","pith_inferences":["The paper's premise could be tested meta-analytically by re-running publicly available checkpoints under different seeds and decoding temperatures, without retraining.","A risk the paper leaves implicit is that Cloze-style questions may favour certain answer formats, so the variance reduction could be bought at the cost of evaluating a narrower skill.","The supplied full text is a different paper; until the actual experiments appear, the claim should be read as a proposal rather than a measured finding.","If variance is as large as claimed, ensembling across seeds at inference time could be a cheaper reliability fix than changing annotation format."],"forward_implications":["Published VQA leaderboards would need to report confidence intervals rather than single accuracy numbers.","Model comparisons would require multiple training seeds or repeated decoding runs before a ranking can be trusted.","If Cloze-style evaluation works, benchmarks could adopt fill-in-the-blank questions as a low-cost way to stabilise scores.","Variance-aware reporting would make it easier to tell whether a new model improves on substance or just on noise.","Hyperparameter and framework sensitivity would become a standard reporting axis, not an omitted detail."],"supporting_citations":[],"fun_headline_variants":["VQA scores swing with seed and sampling, study finds","Leaderboards unreliable: VQA variance laid bare","Cloze tests could steady wobbly VQA benchmarks","Why your VQA rank might be a fluke","Randomness rules VQA benchmarks, urges variance-aware reporting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recommendation rests on the assumption that Cloze-style reformulation reduces stochasticity without changing what the benchmark measures.","fun_headline_variants_meta":{"raw":{"variants":["VQA scores swing with seed and sampling, study finds","Leaderboards unreliable: VQA variance laid bare","Cloze tests could steady wobbly VQA benchmarks","Why your VQA rank might be a fluke","Randomness rules VQA benchmarks, urges variance-aware reporting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000203,"raw_usage":{"total_tokens":1325,"prompt_tokens":825,"completion_tokens":500,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":422}},"tokens_in":441,"tokens_out":500,"duration_ms":6618,"temperature":1.0,"reasoning_tokens":422,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:52:51.198210+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fixed set of VQA models and run each on the same benchmark across many seeds and decoding settings; if the score spread is small relative to the gaps between models, the premise of hidden variance collapses. A second check is to compare answer distributions between open and Cloze forms: if rewording shifts which questions are answered correctly, then any variance reduction may be a task change rather than a reliability improvement.","supporting_citations":[],"review_version":1}