{"id":"a82b2a89-dff6-400c-b9db-f8b61181402a","arxiv_id":"2501.07109","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey tracing the evolution of visual question answering from 2015 CNN-LSTM models through attention mechanisms, modular networks, vision-language pretraining, and large multimodal models.","lead":"This preprint is a survey paper that reviews how visual question answering (VQA) systems evolved from early deep learning models to modern transformer-based vision-language models. It is a broad literature overview aimed at newcomers, not a paper that introduces a new method or result.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Citation and terminology errors undermine the survey's reliability; verify all references before acceptance.","rationale":"The reader's verdict of REJECT is well-supported. The paper is a survey, so its central claim is to provide a reliable, comprehensive overview. That claim requires accurate citations and consistent terminology. The identified errors are frequent and load-bearing: a reader using this survey to navigate the VQA literature would be misled about foundational works (DAQUAR), modern models (ViLT), and even the scope of the field (Video vs. Visual QA). These are not stylistic preferences or debatable interpretations; they are objective misattributions and inconsistent subject matter in the later sections. The terminology drift in Sections 7, 8, and 9 is particularly damaging because it systematically changes the topic from image-based VQA to video VQA without explanation. The reader's weakest assumption also identified the same core issue: the survey's reliability depends on accurate references, and that assumption fails. Therefore, the verdict should remain REJECT, with revision required to verify all citations and restore consistent 'Visual Question Answering' terminology throughout.","tokens_in":20592,"tokens_out":1199,"duration_ms":10139,"concrete_test":"Independently verify every citation in the reference list against its in-text claim, focusing on Section 5 (ViLT), Section 2.4 (DAQUAR), Section 3.3 (SMN), and Sections 7–8 (Video vs. Visual). If more than a few citations cannot be matched to correct papers, the survey fails as a reliable reference.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's value as a survey rests entirely on accurate attribution and consistent terminology. Several concrete errors undermine this: (1) ViLT is cited as (Radford et al., 2021) in Section 5, matching the CLIP reference, but ViLT is by Kim et al., 2021; (2) DAQUAR is attributed to (Malinowski et al., 2014), but the cited paper is 'Multimodal Learning with Deep CNNs,' not the DAQUAR dataset paper (Malinowski and Fritz, 2014); (3) Section 3.3 credits 'Spatial Memory Network' to (Chen et al., 2017), but the reference is 'Spatial Memory for Context Reasoning in Object Detection,' not a VQA model; (4) Sections 7 and 8 and the Conclusion systematically replace 'Visual' with 'Video' Question Answering, a terminological drift that changes the subject matter. Because these errors are frequent, they invalidate the paper's claim to be a comprehensive, trustworthy overview. This is not a question of consensus; it is a correctness failure in the survey's core content.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a survey of Visual Question Answering (VQA), organized by a claimed dichotomy between extractive and abstractive approaches. It spans early CNN-LSTM models, attention mechanisms, compositional reasoning, vision-language pre-training, domain-specific applications, and future directions, and it includes a large summary table. The abstract and opening sections describe VQA as starting in 2015 and claim the survey is comprehensive.","tokens_in":91,"tokens_out":6719,"duration_ms":118264,"significance":"If the survey were reliable, it would serve as a useful entry point for newcomers to VQA. The chronological organization and the table of models are potentially valuable. The paper also attempts to cover a wide range of subareas, including medical VQA and recent large multimodal models. However, the manuscript currently contains numerous citation errors, internal inconsistencies, and a systematic conflation of image-based VQA with video question answering, which invalidate its central claim of being a trustworthy, comprehensive overview. The strengths—a broad structure and an extensive table—are outweighed by the fact that the details are frequently incorrect.","major_comments":[{"comment":"The paper attributes ViLT to (Radford et al., 2021), the same citation used for CLIP; ViLT is by Kim et al. (2021) and does not appear in the reference list at all. Section 5 states 'Models such as CLIP (Radford et al., 2021) and ViLT (Radford et al., 2021)', which is a direct misattribution of a different architecture. Similarly, Section 2.4 attributes the DAQUAR dataset to (Malinowski et al., 2014), but the reference given is 'Multimodal Learning with Deep Convolutional Neural Networks', not the DAQUAR dataset paper (Malinowski and Fritz, 2014). These are not isolated typos; they affect foundational works and make the survey unreliable as a literature map.","section":"Section 5 and Table 1"},{"comment":"The survey repeatedly redefines VQA as 'Video Question Answering'. The opening sentences of Sections 7, 8, and 9 all use the definition 'Video Question Answering (VQA)', whereas Sections 1–6 are about image-based VQA. This is not a harmless abbreviation: the challenges and future directions discuss temporal reasoning, video datasets, and video-specific models (VideoBERT, Frozen in Time, VideoDistill) as if they were part of the image-VQA story, without acknowledging that the scope has shifted. The paper's own conclusion then frames the entire survey as an introduction to video question answering, contradicting the abstract and the earlier sections.","section":"Sections 7, 8, and 9"},{"comment":"The central extractive/abstractive dichotomy is applied inconsistently. For example, Stacked Attention Networks are described in Section 3.3 as an extractive method, but Table 1 classifies 'Stacked Attention Networks for Image Question Answering' as abstractive. Multimodal Compact Bilinear pooling is described in Section 3.2 under the extractive paradigm, yet Table 1 lists 'Multimodal Compact Bilinear Pooling' as abstractive. The paper never defines the criteria for labeling a model extractive or abstractive, so the framework cannot be applied in a principled way and its use throughout the survey is not reliable.","section":"Section 3.3 and Table 1"},{"comment":"The paper attributes GPT-4V to 'Radford and Narasimhan (2018)', which is the reference for the original GPT paper, despite also citing (Li et al., 2024c) for GPT-4V in the same paragraph. Section 3.3 also misattributes the 'Knowing When to Look' image-captioning paper (Lu et al., 2017) to a VQA mixed-attention mechanism. Such errors are frequent enough that the survey cannot be used as a reliable secondary source, and they undercut the paper's claim of a comprehensive overview.","section":"Section 6.1"},{"comment":"The comprehensiveness claim is undercut by both omissions and irrelevant entries. The survey omits several influential VQA systems, such as LLaVA, InstructBLIP, and other recent large multimodal models, despite discussing such models in Section 8. Conversely, Table 1 includes entries that are not VQA works, such as 'Finding Structure in Time' (Elman, 1990) and 'ImageNet Classification' (2012), and it contains a duplicate VisualBERT row. The table therefore does not support the abstract's assertion that this is a comprehensive overview of VQA's evolution.","section":"Table 1 and Sections 1–2"}],"minor_comments":[{"comment":"The paper writes 'introduced VisualQA'; the dataset is 'VQA' or 'Visual Question Answering', not 'VisualQA'.","section":"Section 1.2"},{"comment":"The sentence 'Each section is divided into two main paradigms: extractive and abstractive.the Section 2 reviews...' has a missing space, a lowercase 't', and an extraneous 'the' before 'Section 2'.","section":"Section 1.4"},{"comment":"The opening line of Section 5 says 'Vision-Question Answering (VQA)' instead of 'Visual Question Answering'.","section":"Section 5"},{"comment":"The paper inconsistently spells 'VilBERT' and 'ViLBERT', and the reference list contains two entries for the same ViLBERT paper (Lu et al., 2019a and 2019b).","section":"Sections 4.4 and 5"},{"comment":"The VisualBERT row appears twice with identical wording, and several rows list 'NaN' in the 'Datasets Used' column (e.g., 'Learning Transferable Visual Models'), indicating the table data was not carefully curated.","section":"Table 1"},{"comment":"Several works cited in the text are missing from the reference list, including the actual ViLT paper (Kim et al., 2021), the DAQUAR paper (Malinowski and Fritz, 2014), and a correct UNITER citation; the existing UNITER entry (Chen et al., 2020) has an author list that does not match the published paper.","section":"References"},{"comment":"The year for Visual Genome is listed as 2017 in the table but as 2016 in the references; please make these entries consistent.","section":"Table 1"}],"recommendation":"reject","confidential_remarks":"The errors are so pervasive that I do not think this manuscript is ready for formal review. The survey mixes image-based VQA with video question answering, misassigns references, and contains internal contradictions about the field's timeline (e.g., the abstract says VQA began in 2015, while DAQUAR is discussed as earlier in the same paper). Even a major revision would require re-checking essentially every citation and redefining the scope. A separate, carefully verified survey on image VQA could be viable, but this manuscript in its current form does not meet the reliability bar for a survey article."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The survey is organized around a clean extractive/abstractive dichotomy, and the chronological structure from CNN-LSTM fusion through attention and pretraining to large multimodal models gives a newcomer a reasonable skeleton of the field. The table of models is broad and the coverage of domain-specific VQA (medical, movie, fashion, sports) is more inclusive than many older surveys. If the citations were accurate, this would be a passable orientation document.\n\nThe problems are in the details, and they are not minor. I checked several of the references flagged in the stress test and they land: ViLT is attributed to Radford et al. 2021 (the CLIP paper) in Section 5, when ViLT is Kim et al. 2021; the DAQUAR discussion cites Malinowski et al. 2014, but that reference is the multimodal CNN paper, not the DAQUAR dataset paper; the Spatial Memory Network is credited to a paper about object detection, not VQA. Section 7, Section 8, and the Conclusion systematically substitute “Video Question Answering” for “Visual Question Answering,” which is not a typo but a shift in subject matter that would mislead a reader trying to understand the state of image-based VQA. There are also internal inconsistencies, like UNITER appearing as both extractive and abstractive in different places, and the table listing the same VisualBERT paper twice.\n\nThe reader’s verdict of REJECT is fair. A survey’s only product is trust, and these errors break that trust at the load-bearing level: a newcomer cannot use the reference list to find the original papers. The abstractive/extractive framing is fine as an organizing device, but it is not a contribution that rescues the paper. I do not see evidence of bad faith; the authors appear to have compiled the survey carelessly, not maliciously. But carelessness of this frequency is enough to desk-reject.\n\nWho should read it? Someone who wants a quick, uncritical map of VQA history and is willing to verify every citation themselves. That is a small audience. A serious referee could fix the problems, but the revision would require rechecking hundreds of references and rewriting two sections and the conclusion. That is a major overhaul, not a light touch.\n\nRecommendation: reject, but the authors could resubmit after a thorough citation audit and after fixing the Visual/Video terminology. The raw material for a usable survey exists, but it is not in a publishable state.","headline":"Frequent citation errors and a systematic Visual/Video drift make this VQA survey unreliable as a reference, despite a sensible structure and a useful extractive/abstractive framing.","tokens_in":21315,"tokens_out":618,"would_cite":false,"duration_ms":8296,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey maps a decade of visual question answering through an extractive-versus-abstractive lens.","keywords":["Visual Question Answering","survey","vision-language pretraining","attention mechanisms","multimodal learning","extractive vs abstractive","transformer models","VQA datasets"],"falsifier":"Look up any sample of the paper's in-text citations and check whether the cited document actually makes the claim attributed to it; for example, the text describes (Chen et al., 2017) as a Spatial Memory Network for VQA, while the reference list entry points to a paper on spatial memory for object detection, so that single check can directly test the survey's central reliability claim.","tokens_in":20300,"feed_emoji":"🖼️","tokens_out":4944,"duration_ms":46668,"temperature":0.7,"pith_summary":"This survey sets out to give a structured overview of how Visual Question Answering has evolved from its 2015 formalization to today's large multimodal models. It organizes the field's history around a contrast between extractive systems, which pick answers from a fixed set, and abstractive systems, which generate open-ended responses. A careful reader would care because the paper supplies a map of the field's key models, datasets, and techniques, and frames current debates about bias, reasoning, and multimodal pretraining against that history. The paper's contribution, if its literature account holds up, is a usable roadmap for newcomers and a coherent vocabulary for describing where VQA has been.","feed_headline":"A survey maps 10 years of visual question answering","feed_subtitle":"Attention, transformers, and pretraining arranged along an extractive-versus-abstractive arc.","key_machinery":"The central organizing device is the extractive-versus-abstractive paradigm distinction, applied section by section: extractive VQA treats answering as selecting or grounding a predefined answer, while abstractive VQA treats it as generating a free-form, context-aware response. Carrying the argument alongside this axis are the technical mechanisms the survey identifies as milestone drivers: attention in its stacked, co-attention, and bottom-up/top-down forms; compositional reasoning modules such as neural module networks, scene graphs, and MAC; and transformer-based vision-language pretraining represented by LXMERT, ViLBERT, UNITER, OSCAR, CLIP, BLIP-2, and Flamingo. The survey uses these mechanisms as the stages of its chronological narrative and as the basis for its tabulated summaries of models and methods.","core_discovery":"The paper claims that the entire trajectory of VQA can be understood as movement along two axes: a chronological sequence of technical leaps, from deep CNN-LSTM fusion and bilinear pooling through attention and compositional reasoning to transformer-based vision-language pretraining, and a persistent paradigm split between extractive answer retrieval and abstractive free-form generation. It argues that transformers and large-scale multimodal pretraining have been the decisive drivers of recent progress, and that the same extractive/abstractive split reappears in domain-specific VQA for medicine and entertainment. On the paper's own terms, the central discovery is not a new model but an organizing narrative: one framework under which the field's major benchmarks, architectures, and open problems can be laid out coherently.","pith_inferences":["The extractive/abstractive axis could be sharpened into a testable design spectrum: a model's position on it predicts whether its failures show up as wrong label choices or as ungrounded fluent text, which would give evaluators a cheap diagnostic.","The framework suggests a concrete next benchmark: hold the image constant while varying only the extractive-versus-abstractive demand of the question, isolating what each paradigm contributes.","If the survey's historical narrative is right, future VQA progress will come less from new fusion tricks and more from pretraining data scale and external-knowledge integration, since those are the levers the surveyed history shows moving performance.","Readers should anchor on the abstract and early sections when using the framework, because some later passages speak of 'Video Question Answering' where 'Visual Question Answering' is meant."],"forward_implications":["A newcomer to VQA can use the survey as a staged reading list: CNN-LSTM fusion, bilinear pooling, attention, compositional reasoning, vision-language pretraining, then large multimodal models.","The extractive/abstractive split gives a vocabulary for comparing models across eras, so that a 2016 attention model and a 2023 multimodal LLM can be discussed in the same terms.","Domain-specific VQA in medical imaging, movie understanding, fashion, and scientific figures inherits the same paradigm split and the same dependence on benchmark datasets.","The paper's stated open challenges, including dataset bias, interpretability, and the need for common-sense and external knowledge, define the next targets for VQA research.","If the narrative is right, transformer-based vision-language pretraining, rather than task-specific fusion architectures, is what drove the largest accuracy gains in VQA's recent history."],"supporting_citations":[{"why":"Supplies the VQA v1.0 dataset and the formal task definition the survey uses as its origin point.","marker":"Antol et al., 2015b"},{"why":"Supplies the transformer architecture the survey credits for recent progress.","marker":"Vaswani et al., 2023"},{"why":"Supplies ViLBERT, an early vision-language pretraining model the survey treats as a turning point.","marker":"Lu et al., 2019b"},{"why":"Supplies the bottom-up/top-down attention framework the survey calls a major breakthrough.","marker":"Anderson et al., 2018"},{"why":"Supplies the MAC network, the survey's exemplar of compositional attention reasoning.","marker":"Hudson and Manning, 2018"},{"why":"Supplies VQA v2.0, the benchmark used to show pretrained transformers surpassing earlier models.","marker":"Goyal et al., 2017"},{"why":"Supplies Flamingo, the survey's example of few-shot multimodal pretraining in the abstractive paradigm.","marker":"Alayrac et al., 2022"},{"why":"Supplies BLIP-2, the survey's example of efficient bootstrapped vision-language pretraining.","marker":"Li et al., 2023"}],"fun_headline_variants":["VQA's decade mapped: transformers, attention, and the extractive-abstractive divide","Survey: 10 years of visual question answering, one framework","The extractive-abstractive split: a new lens on VQA history","VQA's evolution: from CNN-LSTM to multimodal pretraining"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's value depends on its literature summaries being accurate and on its terminology staying stable; the manuscript contains multiple citation mismatches and at least one passage that says 'Video Question Answering' where 'Visual Question Answering' is meant, so if these errors are representative, the map it draws is unreliable.","fun_headline_variants_meta":{"raw":{"variants":["VQA's decade mapped: transformers, attention, and the extractive-abstractive divide","Survey: 10 years of visual question answering, one framework","The extractive-abstractive split: a new lens on VQA history","VQA's evolution: from CNN-LSTM to multimodal pretraining"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000713,"raw_usage":{"total_tokens":3177,"prompt_tokens":887,"completion_tokens":2290,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":2209}},"tokens_in":503,"tokens_out":2290,"duration_ms":15577,"temperature":1.0,"reasoning_tokens":2209,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:48:19.194859+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Look up any sample of the paper's in-text citations and check whether the cited document actually makes the claim attributed to it; for example, the text describes (Chen et al., 2017) as a Spatial Memory Network for VQA, while the reference list entry points to a paper on spatial memory for object detection, so that single check can directly test the survey's central reliability claim.","supporting_citations":[],"review_version":1}