{"id":"cbd5d99d-e43c-4354-853c-0f80ce75ee97","arxiv_id":"2411.17558","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that organizes VQA methods from feature extraction through MLLM reasoning, datasets, and metrics, without introducing new experimental results.","lead":"This paper surveys visual question answering (VQA), covering feature extraction, multimodal fusion, knowledge reasoning, datasets, and evaluation, with emphasis on multimodal large language models. It is a reference map for researchers entering VQA or comparing MLLM-based methods, though its value is limited by weak editing and unreliable references.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's reference list contains many unrelated template entries and Table 6 mixes incompatible evaluation protocols, so the 'up-to-date synthesis' claim is not currently supported.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the survey's usefulness as a current reference depends on accurate citations and comparable performance tables, and both are compromised. My stress-test confirms this concern with concrete examples: unrelated template references, duplicate entries, and Table 6 rows that mix different evaluation splits and shot settings. This is a correctness risk for the central claim, not merely a style issue, because a survey reader cannot tell which numbers are directly comparable. The reader's CONDITIONAL verdict is therefore appropriate: the manuscript's taxonomy and coverage may be useful, but the scholarly apparatus must be fixed before the survey can serve as a reliable reference. My review does not identify a basis for moving the verdict to ACCEPT or REJECT; the problems are mechanical and fixable, but they are currently load-bearing.","tokens_in":52724,"tokens_out":3023,"duration_ms":32380,"concrete_test":"Run an automated citation audit: parse every in-text bracketed citation and every numbered bibliography entry, compute the set of entries never cited in the body, and manually classify those entries as VQA-relevant, unrelated, or template/duplicate. If more than a small fraction (e.g., 5%) of the bibliography is uncited, unrelated, or duplicated, the manuscript must be revised to remove or properly integrate these entries. Separately, for Table 6, reproduce each reported accuracy from the cited original paper under the exact split and shot setting; flag any row that cannot be reproduced or that conflates test-dev with test-std, or zero-shot with few-shot.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central value as an up-to-date reference depends on the correctness and relevance of its bibliography and comparative performance tables. That condition is not met. The reference list contains numerous entries unrelated to VQA and never cited in the body, including Clifford algebra [5], amsthm [16], sensor networks [10, 76], and ACM template artifacts such as [98], [99], [103], and [420]. It also contains duplicates ([45]/[46], [68]/[69], [72]/[73], [436]/[437]) and several near-identical entries. This is not a one-off typo; it indicates that the bibliography was not systematically checked against the manuscript. Table 6 compounds the problem by mixing incompatible settings: VQAv2 rows use test-std but some entries are zero-shot, GQA rows separately list test-dev, Test2019, Test2020, and Test2021 without noting that these are different evaluation splits, and VizWiz rows mix zero-shot and few-shot results across different dataset versions. A reader cannot verify the survey's comparative claims without re-deriving each number from the original papers. The prose taxonomy may be salvageable, but the central claim of a reliable, up-to-date synthesis is currently unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of Visual Question Answering (VQA), organized around a taxonomy of natural-language understanding of images and text and of natural-language inference/knowledge reasoning. It reviews feature extraction, fusion mechanisms, vision-language pretraining, multimodal large language models, knowledge sources and reasoning, datasets, evaluation metrics, and open challenges. The abstract claims to provide an up-to-date synthesis of VQA with particular attention to MLLMs. The presentation includes several large tables summarizing models, datasets, and comparative performance, along with figures and equations taken from or inspired by prior work.","tokens_in":52935,"tokens_out":3631,"duration_ms":33061,"significance":"If accurate, the survey could serve as a broad reference for VQA, especially for readers seeking a single entry point to the transition from conventional models to MLLMs and to knowledge-based reasoning. The proposed taxonomy in Fig. 2 and the coverage of recent MLLM techniques (Sections 3 and 4.4) are potentially useful organizing contributions. The paper also explicitly claims timeliness and exhaustiveness, which raises the bar for factual reliability. On the current submission, however, the central value of the survey is undermined by a contaminated bibliography, missing appendices referenced in the body, and comparative tables that mix incompatible evaluation protocols. These issues must be addressed before the survey can be used as a reference.","major_comments":[{"comment":"The comparative performance table mixes fundamentally different evaluation settings without adequate qualification. VQAv2 rows include test-std results for models evaluated in zero-shot, few-shot, and fine-tuned settings; GQA rows separately list test-dev, Test2019, Test2020, and Test2021 without noting that these are different test splits; and VizWiz rows mix zero-shot and few-shot results across different dataset versions. The surrounding text only states that 'variation in experimental configurations' can lead to 'substantial diminution in performance' but does not tell the reader which numbers are directly comparable. As it stands, a reader cannot verify or use the comparative claims without returning to each original paper. The table should either be split by protocol and split, with the shot setting and dataset version stated per row, or removed.","section":"Table 6 (Sec. 5.4)"},{"comment":"The reference list contains many entries that are unrelated to VQA and never cited in the body, including the Clifford algebra package [5], the amsthm package [16], wireless sensor network surveys [10, 76], and ACM template artifacts such as [98], [99], [103], and [420]. There are also duplicate entries ([45]/[46], [68]/[69], [72]/[73], [436]/[437]) and near-identical repeated entries for the same work ([205]-[211]). This demonstrates that the bibliography was not systematically checked against the manuscript. Because a survey's value rests on its sourcing, the entire reference list needs to be reconstructed from the in-text citations, with each entry verified against the original source, and all irrelevant, duplicate, and template entries removed.","section":"References (whole bibliography)"},{"comment":"The body text repeatedly refers to appendices that are not present in the submission: Appendix A for embedding improvements, Appendix B for vision-language pretraining variants, Appendix C for knowledge sources, Appendix D for knowledge extraction, Appendix E for additional datasets, and Appendix F for additional metrics. For example, Sec. 2.1.3 says 'We give the detailed improvement methods in Appendix. A,' but no appendix follows. A survey that promises these details in the body is incomplete without them. Either include the appendices in the submission or remove all references to them and fold the necessary content into the main text.","section":"Appendices (Secs. 2.1.3, 2.2.3, 4.1, 4.2.2, 5.1, 5.3)"},{"comment":"The table and surrounding text contain factual errors that undermine the paper's claim of providing reliable information about latest models. The model name 'mOLUG-owl2' in Table 3 should be 'mPLUG-OWL2'; the text in Sec. 3.3.3 refers to 'Genimi' instead of 'Gemini'; the LLaVA-1.5 row reports 78.5 on VQAv2 as 'few-shot,' which is misleading because LLaVA-1.5 is a fine-tuned model, not a few-shot method; and the InternVL2 row in Table 3 has blank performance entries, so the row provides no information. Every entry in Table 3 should be checked against the cited papers, and the text should be corrected accordingly.","section":"Table 3 and Sec. 3.3.3"}],"minor_comments":[{"comment":"There are many typographical and spelling errors that should be corrected in copyediting: 'Accuruacy' in the Table 6 header, 'Knowldege' in Sec. 4.1, 'extrctor' in Sec. 2.1.2, 'applys' in Sec. 2.1.1, 'avarage' and 'Imgae' in Table 5, and 'breif' in Sec. 5.3.2.","section":"Throughout"},{"comment":"The description of VGG-Net is imprecise: 'VGG-Net increases the convolutional layers to 19' is only one configuration; VGG-16 is equally common. The phrase about ResNet 'weakening strong connections' is unclear and should be rewritten.","section":"Sec. 2.1.1"},{"comment":"The variables in Eq. (5) do not match the prose: the text says that textual tokens {x'_Q_i} are aligned with visual objects {x'_S_i}, but the same index i appears in both sequences. The equation should use different indices (e.g., i for the question node and j for the scene-graph node) so that the alignment is well defined.","section":"Sec. 2.2.2, Eq. (5)"},{"comment":"The sentence 'which has proven especially effective in sophisticated tasks like math reasoning' is a sentence fragment and the connection to InternVL is unclear; it should be rewritten as a complete sentence or merged with the preceding claim.","section":"Sec. 3.3.3"},{"comment":"The table lists 'AMN 2020 VGG-Net Word2Vec Attention MovieQA' but the year and architecture for AMN are not supported by the citation in the text; please verify the entry against the cited paper and correct the row or remove it.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be an unfinished draft: the bibliography contains template and unrelated entries, appendices referenced in the text are missing, and tables contain typos and inconsistent comparisons. These are not subtle issues; they suggest that the submission was not thoroughly checked before posting. The prose taxonomy and the survey structure are defensible, and the central claim of providing an up-to-date synthesis can be salvaged with substantial revision. Given the pervasiveness of the reference-list contamination and the potential for factual errors in the model summaries, I recommend that the revision be verified by at least one independent check of all citations and all numeric entries in the tables."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The survey has a solid organizational core. The taxonomy contrasting natural language understanding and inference, with a focus on MLLMs and knowledge reasoning, is timely and could serve as a useful entry point for newcomers. The paper does a good job comparing prior surveys and covering recent models, datasets, and evaluation metrics. The prose sections on knowledge reasoning and MLLM techniques are readable and well-structured.\n\nThe soft spots are significant. The reference list is contaminated: it includes unrelated template entries (Clifford algebra, amsthm, sensor networks, ACM placeholder books), duplicates, and many references never cited in the body. That is not a small typo; it means the bibliography was not checked against the manuscript. Table 6 compounds the problem by mixing zero-shot and few-shot results and different evaluation splits (test-dev, Test2019, Test2020, Test2021), so the comparative claims are not verifiable. There are also numerous typos ('mOLUG-owl2', 'Genimi', 'Accuruacy', 'Knowldege'). These are fixable, but they currently undermine the central claim of an up-to-date, reliable synthesis.\n\nAs a survey, no new empirical result is expected, and the central argument that VQA is shifting toward MLLM-based reasoning holds up. But a survey is only as good as its sources. The reference list and performance table need major attention.\n\nRecommendation: This deserves serious peer review with the expectation of major revision. The authors should verify every reference, remove or correct duplicates and template artifacts, and either standardize Table 6 or clearly flag the experimental protocols. I'd bring it to a reading group as a cautionary example of why reference hygiene matters, but I wouldn't cite it in its current form.","headline":"Useful taxonomy, but contaminated references and a mixed-protocol performance table make it unreliable as a reference.","tokens_in":53465,"tokens_out":2995,"would_cite":false,"duration_ms":27240,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey organizes visual question answering into understanding and inference, with multimodal large language models as the field's newest stage.","keywords":["visual question answering","multimodal large language models","vision-language pretraining","multimodal fusion","knowledge reasoning","VQA datasets","evaluation metrics","survey"],"falsifier":"Check each bibliographic entry for a corresponding in-text citation: entries such as a Clifford algebra software package, a LaTeX style manual, or sensor-network surveys that appear in the references but never in the body would show the bibliography is not load-bearing. Separately, inspect Table 6 to see whether each row states its dataset split and shot setting (zero-shot, few-shot, or fine-tuned); rows that mix settings without labels would invalidate the reported comparisons.","tokens_in":52544,"feed_emoji":"🧠","tokens_out":6231,"duration_ms":112484,"temperature":0.7,"pith_summary":"This survey sets out to establish a two-part synthesis of Visual Question Answering (VQA), the task of producing an answer $A$ to a question $Q$ about an image $V$. The two parts are natural language understanding of images and text, and natural language inference over image-question information. It argues that the field has moved from separate feature extractors and fusion modules, through vision-language pretraining, to multimodal large language models (MLLMs)—large language models extended to accept visual input—that add instruction tuning, in-context learning, and chain-of-thought reasoning. The survey also organizes knowledge-based VQA into knowledge sources, extraction methods, and one-hop versus multi-hop reasoning, and it catalogues datasets, MLLM benchmarks, and evaluation metrics. A sympathetic reader would use this as a current reference for choosing models, datasets, and metrics and for locating open problems such as data bias, explainability, and answer generation.","feed_headline":"Survey maps VQA: from feature fusion to MLLM reasoning","feed_subtitle":"A taxonomy splits visual question answering into understanding image and text, then reasoning over knowledge.","key_machinery":"The carrying device is the paper's taxonomy of the VQA task (its Fig. 2), which splits the field along a perception-to-cognition axis: natural language understanding of image and text versus natural language inference, with knowledge reasoning and MLLM reasoning as sub-branches. This taxonomy does the work of the survey's argument: it gives every model, fusion module, knowledge source, and dataset a slot, so that the claimed up-to-date synthesis becomes a single map rather than a chronological list. The paper's secondary machinery is the distinction between internal and external knowledge, which organizes the knowledge-reasoning section.","core_discovery":"The paper's central claim is that VQA is best understood through one taxonomy: natural language understanding of images and text on the perception side, and natural language inference on the cognition side, with multimodal large language models as the newest stage of both. On the understanding side, the survey traces visual and textual feature extraction, embedding improvements, fusion by vector operations, attention and graph neural networks, and dual-stream versus single-stream pretraining. On the inference side, it distinguishes internal from external knowledge, entity-based from feature-based extraction, and conventional, one-hop, and multi-hop reasoning, including memory-based, graph-based, and implicit methods. The same taxonomy then places current MLLM techniques—instruction tuning, in-context learning, multimodal chain-of-thought, and tool-aided reasoning—as the latest answer to both perceptual and cognitive demands, with datasets and benchmarks as the evaluation layer.","pith_inferences":["Editorial inference: if the perception/inference split is right, a testable prediction is that gains from new MLLMs on external-knowledge benchmarks will come more from better retrieval and tool use than from larger parameter counts, since the survey groups knowledge access separately from raw comprehension.","Editorial inference: the taxonomy implies that classic VQA datasets will increasingly be treated as sub-benchmarks inside general MLLM evaluations, which may make benchmark-specific leaderboards less informative over time.","Editorial inference: the survey's treatment of chain-of-thought and tool-aided reasoning suggests that reasoning in VQA may fragment into distinct skills such as symbolic, temporal, spatial, and commonsense reasoning, each needing its own evaluation rather than a single accuracy number."],"forward_implications":["If the taxonomy is right, the historical VQA pipeline of separate visual and textual feature extractors is being absorbed into image-to-text alignment modules inside MLLMs, so fusion research has largely shifted to alignment architecture.","On this account, knowledge-based VQA has two live routes—internal knowledge already stored in model parameters and external knowledge retrieved from knowledge bases or passages—and multi-hop reasoning is where most of the remaining difficulty lies.","The survey's reading implies that MLLM benchmarks such as MME, SEED-Bench, and MathVista are becoming the de facto evaluation layer for VQA, testing perception, reasoning, hallucination, and specialized domains at once.","If the stated open problems are taken seriously, the next round of VQA progress will need indirect visual information, dataset-debiasing, explainable answers, and generative rather than selection-based answer production."],"supporting_citations":[{"why":"Introduces the VQA task and VQA v1 dataset that anchors the survey's definition of VQA and its earliest models.","marker":"[23]"},{"why":"Presents VQAv2, which the survey treats as the leading benchmark and the dataset designed to reduce language priors.","marker":"[123]"},{"why":"Defines OK-VQA, the main benchmark for knowledge-based VQA that organizes the external-knowledge discussion.","marker":"[292]"},{"why":"Supplies FVQA and its supporting-fact knowledge base, the reference point for knowledge extraction and reasoning methods.","marker":"[424]"},{"why":"Provides LXMERT, the cross-modal transformer backbone used by several surveyed understanding and knowledge methods.","marker":"[402]"},{"why":"Proposes MuKEA, which the survey cites as the representative implicit multi-hop reasoning method with multimodal knowledge graph completion.","marker":"[87]"},{"why":"Introduces Flamingo, the few-shot visual-language model the survey uses to anchor MLLM-based VQA.","marker":"[14]"},{"why":"Presents BLIP-2 and the Q-Former alignment approach, which the survey treats as a key image-to-text alignment mechanism in MLLMs.","marker":"[242]"}],"fun_headline_variants":["VQA survey: perception and inference, tied by MLLMs","From feature fusion to MLLM: VQA's full taxonomy","Understanding vs inference: VQA mapped for MLLMs","Survey decodes VQA: perception then cognition with MLLMs","MLLMs unify VQA's two pillars: understanding and inference"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's usefulness depends on the assumption that its bibliography and performance tables accurately represent the state of the field; if many listed references are irrelevant or its score comparisons mix incompatible test conditions, the claimed up-to-date synthesis is not established.","fun_headline_variants_meta":{"raw":{"variants":["VQA survey: perception and inference, tied by MLLMs","From feature fusion to MLLM: VQA's full taxonomy","Understanding vs inference: VQA mapped for MLLMs","Survey decodes VQA: perception then cognition with MLLMs","MLLMs unify VQA's two pillars: understanding and inference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000933,"raw_usage":{"total_tokens":3955,"prompt_tokens":871,"completion_tokens":3084,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":487,"completion_tokens_details":{"reasoning_tokens":2994}},"tokens_in":487,"tokens_out":3084,"duration_ms":19599,"temperature":1.0,"reasoning_tokens":2994,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:58:28.970599+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check each bibliographic entry for a corresponding in-text citation: entries such as a Clifford algebra software package, a LaTeX style manual, or sensor-network surveys that appear in the references but never in the body would show the bibliography is not load-bearing. Separately, inspect Table 6 to see whether each row states its dataset split and shot setting (zero-shot, few-shot, or fine-tuned); rows that mix settings without labels would invalidate the reported comparisons.","supporting_citations":[{"cited_title":"Ok-vqa: A visual question answering benchmark requiring external knowledge","cited_arxiv_id":null,"evidence_quote":"Defines OK-VQA, the main benchmark for knowledge-based VQA that organizes the external-knowledge discussion."}],"review_version":1}