Pith. sign in

REVIEW 5 major objections 7 minor 98 references

The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering

T0 review · 5 major / 7 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A survey maps a decade of visual question answering through an extractive-versus-abstractive lens.

desk verdict Frequent citation errors and a systematic Visual/Video drift make this VQA survey unreliable as a reference, despite a sensible structure and a useful extractive/abstractive framing. read the letter →

arxiv 2501.07109 v1 pith:W6VQWM6M submitted 2025-01-13 cs.CV

classification cs.CV
keywords VisualQuestionAnsweringsurveyvision-languagepretrainingattentionmechanismsmultimodallearningextractivevsabstractivetransformermodelsVQAdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey sets out to give a structured overview of how Visual Question Answering has evolved from its 2015 formalization to today's large multimodal models. It organizes the field's history around a contrast between extractive systems, which pick answers from a fixed set, and abstractive systems, which generate open-ended responses. A careful reader would care because the paper supplies a map of the field's key models, datasets, and techniques, and frames current debates about bias, reasoning, and multimodal pretraining against that history. The paper's contribution, if its literature account holds up, is a usable roadmap for newcomers and a coherent vocabulary for describing where VQA has been.

What carries the argument

The central organizing device is the extractive-versus-abstractive paradigm distinction, applied section by section: extractive VQA treats answering as selecting or grounding a predefined answer, while abstractive VQA treats it as generating a free-form, context-aware response. Carrying the argument alongside this axis are the technical mechanisms the survey identifies as milestone drivers: attention in its stacked, co-attention, and bottom-up/top-down forms; compositional reasoning modules such as neural module networks, scene graphs, and MAC; and transformer-based vision-language pretraining represented by LXMERT, ViLBERT, UNITER, OSCAR, CLIP, BLIP-2, and Flamingo. The survey uses these mechanisms as the stages of its chronological narrative and as the basis for its tabulated summaries of models and methods.

What would settle it

Look up any sample of the paper's in-text citations and check whether the cited document actually makes the claim attributed to it; for example, the text describes (Chen et al., 2017) as a Spatial Memory Network for VQA, while the reference list entry points to a paper on spatial memory for object detection, so that single check can directly test the survey's central reliability claim.

Watch

Extended reading notes

Core claim

The paper claims that the entire trajectory of VQA can be understood as movement along two axes: a chronological sequence of technical leaps, from deep CNN-LSTM fusion and bilinear pooling through attention and compositional reasoning to transformer-based vision-language pretraining, and a persistent paradigm split between extractive answer retrieval and abstractive free-form generation. It argues that transformers and large-scale multimodal pretraining have been the decisive drivers of recent progress, and that the same extractive/abstractive split reappears in domain-specific VQA for medicine and entertainment. On the paper's own terms, the central discovery is not a new model but an organizing narrative: one framework under which the field's major benchmarks, architectures, and open problems can be laid out coherently.

Load-bearing premise

The survey's value depends on its literature summaries being accurate and on its terminology staying stable; the manuscript contains multiple citation mismatches and at least one passage that says 'Video Question Answering' where 'Visual Question Answering' is meant, so if these errors are representative, the map it draws is unreliable.

Editorial extensions

If this is right

  • A newcomer to VQA can use the survey as a staged reading list: CNN-LSTM fusion, bilinear pooling, attention, compositional reasoning, vision-language pretraining, then large multimodal models.
  • The extractive/abstractive split gives a vocabulary for comparing models across eras, so that a 2016 attention model and a 2023 multimodal LLM can be discussed in the same terms.
  • Domain-specific VQA in medical imaging, movie understanding, fashion, and scientific figures inherits the same paradigm split and the same dependence on benchmark datasets.
  • The paper's stated open challenges, including dataset bias, interpretability, and the need for common-sense and external knowledge, define the next targets for VQA research.
  • If the narrative is right, transformer-based vision-language pretraining, rather than task-specific fusion architectures, is what drove the largest accuracy gains in VQA's recent history.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The extractive/abstractive axis could be sharpened into a testable design spectrum: a model's position on it predicts whether its failures show up as wrong label choices or as ungrounded fluent text, which would give evaluators a cheap diagnostic.
  • The framework suggests a concrete next benchmark: hold the image constant while varying only the extractive-versus-abstractive demand of the question, isolating what each paradigm contributes.
  • If the survey's historical narrative is right, future VQA progress will come less from new fusion tricks and more from pretraining data scale and external-knowledge integration, since those are the levers the surveyed history shows moving performance.
  • Readers should anchor on the abstract and early sections when using the framework, because some later passages speak of 'Video Question Answering' where 'Visual Question Answering' is meant.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. This paper is a survey of Visual Question Answering (VQA), organized by a claimed dichotomy between extractive and abstractive approaches. It spans early CNN-LSTM models, attention mechanisms, compositional reasoning, vision-language pre-training, domain-specific applications, and future directions, and it includes a large summary table. The abstract and opening sections describe VQA as starting in 2015 and claim the survey is comprehensive.

Significance. If the survey were reliable, it would serve as a useful entry point for newcomers to VQA. The chronological organization and the table of models are potentially valuable. The paper also attempts to cover a wide range of subareas, including medical VQA and recent large multimodal models. However, the manuscript currently contains numerous citation errors, internal inconsistencies, and a systematic conflation of image-based VQA with video question answering, which invalidate its central claim of being a trustworthy, comprehensive overview. The strengths—a broad structure and an extensive table—are outweighed by the fact that the details are frequently incorrect.

major comments (5)
  1. [Section 5 and Table 1] The paper attributes ViLT to (Radford et al., 2021), the same citation used for CLIP; ViLT is by Kim et al. (2021) and does not appear in the reference list at all. Section 5 states 'Models such as CLIP (Radford et al., 2021) and ViLT (Radford et al., 2021)', which is a direct misattribution of a different architecture. Similarly, Section 2.4 attributes the DAQUAR dataset to (Malinowski et al., 2014), but the reference given is 'Multimodal Learning with Deep Convolutional Neural Networks', not the DAQUAR dataset paper (Malinowski and Fritz, 2014). These are not isolated typos; they affect foundational works and make the survey unreliable as a literature map.
  2. [Sections 7, 8, and 9] The survey repeatedly redefines VQA as 'Video Question Answering'. The opening sentences of Sections 7, 8, and 9 all use the definition 'Video Question Answering (VQA)', whereas Sections 1–6 are about image-based VQA. This is not a harmless abbreviation: the challenges and future directions discuss temporal reasoning, video datasets, and video-specific models (VideoBERT, Frozen in Time, VideoDistill) as if they were part of the image-VQA story, without acknowledging that the scope has shifted. The paper's own conclusion then frames the entire survey as an introduction to video question answering, contradicting the abstract and the earlier sections.
  3. [Section 3.3 and Table 1] The central extractive/abstractive dichotomy is applied inconsistently. For example, Stacked Attention Networks are described in Section 3.3 as an extractive method, but Table 1 classifies 'Stacked Attention Networks for Image Question Answering' as abstractive. Multimodal Compact Bilinear pooling is described in Section 3.2 under the extractive paradigm, yet Table 1 lists 'Multimodal Compact Bilinear Pooling' as abstractive. The paper never defines the criteria for labeling a model extractive or abstractive, so the framework cannot be applied in a principled way and its use throughout the survey is not reliable.
  4. [Section 6.1] The paper attributes GPT-4V to 'Radford and Narasimhan (2018)', which is the reference for the original GPT paper, despite also citing (Li et al., 2024c) for GPT-4V in the same paragraph. Section 3.3 also misattributes the 'Knowing When to Look' image-captioning paper (Lu et al., 2017) to a VQA mixed-attention mechanism. Such errors are frequent enough that the survey cannot be used as a reliable secondary source, and they undercut the paper's claim of a comprehensive overview.
  5. [Table 1 and Sections 1–2] The comprehensiveness claim is undercut by both omissions and irrelevant entries. The survey omits several influential VQA systems, such as LLaVA, InstructBLIP, and other recent large multimodal models, despite discussing such models in Section 8. Conversely, Table 1 includes entries that are not VQA works, such as 'Finding Structure in Time' (Elman, 1990) and 'ImageNet Classification' (2012), and it contains a duplicate VisualBERT row. The table therefore does not support the abstract's assertion that this is a comprehensive overview of VQA's evolution.
minor comments (7)
  1. [Section 1.2] The paper writes 'introduced VisualQA'; the dataset is 'VQA' or 'Visual Question Answering', not 'VisualQA'.
  2. [Section 1.4] The sentence 'Each section is divided into two main paradigms: extractive and abstractive.the Section 2 reviews...' has a missing space, a lowercase 't', and an extraneous 'the' before 'Section 2'.
  3. [Section 5] The opening line of Section 5 says 'Vision-Question Answering (VQA)' instead of 'Visual Question Answering'.
  4. [Sections 4.4 and 5] The paper inconsistently spells 'VilBERT' and 'ViLBERT', and the reference list contains two entries for the same ViLBERT paper (Lu et al., 2019a and 2019b).
  5. [Table 1] The VisualBERT row appears twice with identical wording, and several rows list 'NaN' in the 'Datasets Used' column (e.g., 'Learning Transferable Visual Models'), indicating the table data was not carefully curated.
  6. [References] Several works cited in the text are missing from the reference list, including the actual ViLT paper (Kim et al., 2021), the DAQUAR paper (Malinowski and Fritz, 2014), and a correct UNITER citation; the existing UNITER entry (Chen et al., 2020) has an author list that does not match the published paper.
  7. [Table 1] The year for Visual Genome is listed as 2017 in the table but as 2016 in the references; please make these entries consistent.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey makes no derivation-based claims, and its self-citations are not load-bearing.

full rationale

This is a survey paper rather than a derivation or empirical study, so the standard circularity patterns do not apply. The central claim is that the paper offers a structured overview of the evolution of Visual Question Answering, organizing prior work under extractive and abstractive paradigms. There is no fitted parameter that is later relabeled as a prediction, no equation that reduces to its own input, and no uniqueness theorem imported from the authors' prior work to force a conclusion. The paper's value depends on accurate citation and attribution, and the skeptic's findings identify genuine correctness problems, such as ViLT being attributed to Radford et al. (2021), DAQUAR being tied to a paper that is not the DAQUAR dataset paper, the Spatial Memory Network being credited to a non-VQA object-detection paper, and the systematic replacement of 'Visual' with 'Video' in later sections. These are reliability and scholarship concerns, not circularity: they do not make the survey's descriptive claims equivalent to their inputs by construction. One author, Asif Ekbal, is a coauthor on some cited works, including the code-mixed VQA system by Khan et al. (2021), but those citations are used as ordinary literature references in the future-directions discussion and are not load-bearing for the survey's organization or any derived conclusion. Machine-checked or externally verifiable support is not needed for a narrative survey, and no self-citation chain is invoked to forbid alternative viewpoints. Accordingly, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

As a survey, the paper introduces no free parameters and no invented entities. The main ledger items are the factual assumptions about the cited literature and the organizing taxonomy. The high number of citation mismatches directly undermines the first assumption.

assumptions (3)
  • domain assumption The cited references accurately represent the models and datasets they are claimed to describe.
    The entire survey rests on the correctness of its literature summaries. This assumption fails in several places, for example the ViLT citation is attributed to (Radford et al., 2021) and the Spatial Memory Network is cited with an object detection paper.
  • ad hoc to paper The extractive versus abstractive paradigm dichotomy can meaningfully classify every VQA model discussed.
    The paper applies this dichotomy to all models, including image captioning models, CLIP, and large multimodal models, without justifying that it is a useful or field-standard taxonomy.
  • domain assumption The paper's narrative accurately reflects the historical sequence and significance of VQA milestones.
    The survey claims to trace the evolution of VQA, but several descriptions are generic and may not match the primary sources, as evidenced by the citation mismatches.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering." pith.science (2026). https://pith.science/paper/W6VQWM6M

@misc{pith2026250107109,
  author       = {Pith},
  title        = {Pith review of: The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/W6VQWM6M}},
  note         = {Machine review of arXiv:2501.07109}
}
read the original abstract

Visual Question Answering (VQA) is an interdisciplinary field that bridges the gap between computer vision (CV) and natural language processing(NLP), enabling Artificial Intelligence(AI) systems to answer questions about images. Since its inception in 2015, VQA has rapidly evolved, driven by advances in deep learning, attention mechanisms, and transformer-based models. This survey traces the journey of VQA from its early days, through major breakthroughs, such as attention mechanisms, compositional reasoning, and the rise of vision-language pre-training methods. We highlight key models, datasets, and techniques that shaped the development of VQA systems, emphasizing the pivotal role of transformer architectures and multimodal pre-training in driving recent progress. Additionally, we explore specialized applications of VQA in domains like healthcare and discuss ongoing challenges, such as dataset bias, model interpretability, and the need for common-sense reasoning. Lastly, we discuss the emerging trends in large multimodal language models and the integration of external knowledge, offering insights into the future directions of VQA. This paper aims to provide a comprehensive overview of the evolution of VQA, highlighting both its current state and potential advancements.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

98 extracted references · 40 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda, Timothy Chung, Bala Krishna S Vegesna, Abhipsha Das, Anthony Susevski, Ryan Sze-Yin Chan, SM Iftekhar Uddin, Shayekh Bin Islam, et al. 2024. Maya: An instruction finetuned multilingual multimodal model. arXiv e-prints, pages arXiv--2412

  4. [4]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski,...

  5. [5]

    Anderson, X

    P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR

  6. [6]

    Andreas, M

    J. Andreas, M. Rohrbach, T. Darrell, and D. Klein. 2016. Neural module networks. In ICML

  7. [7]

    Antol, A

    S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, and C. L. Lawrence Zitnick. 2015 a . Vqa: Visual question answering. In ICCV

  8. [8]

    Lawrence Zitnick, and Devi Parikh

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 b . http://arxiv.org/abs/1505.00468 VQA: visual question answering . CoRR, abs/1505.00468

Show all 98 references
  1. [9]

    Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2016. http://arxiv.org/abs/1409.0473 Neural machine translation by jointly learning to align and translate

  2. [10]

    Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2022. http://arxiv.org/abs/2104.00650 Frozen in time: A joint video and image encoder for end-to-end retrieval

  3. [11]

    Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Laila Bashmal, and Mansour Zuair. 2023. https://doi.org/10.3390/bioengineering10030380 Vision–language model for visual question answering in medical imagery . Bioengineering, 10(3)

  4. [12]

    Hedi Ben-younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome. 2017. http://arxiv.org/abs/1705.06676 Mutan: Multimodal tucker fusion for visual question answering

  5. [13]

    Sarath Chandar, Sungjin Ahn, Hugo Larochelle, Pascal Vincent, Gerald Tesauro, and Yoshua Bengio. 2016. http://arxiv.org/abs/1605.07427 Hierarchical memory networks

  6. [14]

    L. Chen, H. Kornblith, M. Swersky, and M. Norouzi. 2020. Uniter: Universal image-text representation learning. In ECCV

  7. [15]

    Long Chen, Hanwang Jiang, Jin-Hwa Xiao, Shih-Fu Shi, and Shuicheng Chen. 2017. Spatial memory for context reasoning in object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 4086--4096

  8. [16]

    Marks, and Jonathan Le Roux

    Anoop Cherian, Chiori Hori, Tim K. Marks, and Jonathan Le Roux. 2022. http://arxiv.org/abs/2202.09277 (2.5+1)d spatio-temporal scene graphs for video question answering

  9. [17]

    Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. http://arxiv.org/abs/1406.1078 Learning phrase representations using rnn encoder-decoder for statistical machine translation

  10. [18]

    Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, and Dhruv Batra. 2017. http://arxiv.org/abs/1611.08669 Visual dialog

  11. [19]

    J. Deng, W. Dong, R. Socher, L. Li, and F. Li. 2009. Imagenet: A large-scale hierarchical image database. In CVPR 2009

  12. [20]

    J. L. Elman. 1990. Finding structure in time. Cognitive Science

  13. [21]

    Everingham, L

    M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. 2010. The pascal visual object classes (voc) challenge. In IJCV 2010

  14. [22]

    Fukui, D

    A. Fukui, D. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach. 2016. Multimodal compact bilinear pooling for visual question answering and visual grounding. In CVPR

  15. [23]

    Kunihiko Fukushima. 1980. https://doi.org/10.1007/BF00344251 Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position . Biological Cybernetics, 36(4):193--202

  16. [24]

    H. Gao, J. Mao, J. Zhou, T. Huang, L. Xu, and Y. Wang. 2015. Are you talking to a machine? dataset and methods for multilingual image question answering. In CVPR

  17. [25]

    Donald Geman, Stuart Geman, Neil Hallonquist, and Laurent Younes. 2015. https://doi.org/10.1073/pnas.1422953112 Visual turing test for computer vision systems . Proceedings of the National Academy of Sciences, 112(12):3618--3623

  18. [26]

    Girshick, J

    R. Girshick, J. Donahue, T. Darrell, and J. Malik. 2014 a . Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR 2014

  19. [27]

    Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014 b . http://arxiv.org/abs/1311.2524 Rich feature hierarchies for accurate object detection and semantic segmentation

  20. [28]

    Goyal, T

    Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh. 2017. Visual question answering in the wild. In CVPR

  21. [29]

    Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P

    Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. 2018. http://arxiv.org/abs/1802.08218 Vizwiz grand challenge: Answering visual questions from blind people

  22. [30]

    Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Matthew Lungren, and Henning M\"uller

    Sadid A. Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Matthew Lungren, and Henning M\"uller. 2018. Overview of the ImageCLEF 2018 medical domain visual question answering task. In CLEF2018 Working Notes, CEUR Workshop Proceedings, Avignon, France. CEUR-WS.org < http://ceur-ws.org >

  23. [31]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. http://arxiv.org/abs/1512.03385 Deep residual learning for image recognition

  24. [33]

    Sepp Hochreiter and Jürgen Schmidhuber. 1997 b . https://doi.org/10.1162/neco.1997.9.8.1735 Long short-term memory . Neural computation, 9:1735--80

  25. [34]

    D. A. Hudson and C. D. Manning. 2018. Compositional attention networks for machine reasoning. In NeurIPS

  26. [35]

    Hudson and Christopher D

    Drew A. Hudson and Christopher D. Manning. 2019. http://arxiv.org/abs/1902.09506 Gqa: A new dataset for real-world visual reasoning and compositional question answering

  27. [36]

    Johnson, B

    J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick. 2017 a . Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR 2017

  28. [37]

    Johnson, B

    J. Johnson, B. Hariharan, L. van der Maaten, J. Hoffman, L. Fei-Fei, and C. L. Zitnick. 2017 b . Inferring scene structure and generating stories from images. In CVPR

  29. [38]

    Kushal Kafle and Christopher Kanan. 2017. http://arxiv.org/abs/1703.09684 An analysis of visual question answering algorithms

  30. [39]

    Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Akos Kadar, Adam Trischler, and Yoshua Bengio. 2018. http://arxiv.org/abs/1710.07300 Figureqa: An annotated figure dataset for visual reasoning

  31. [40]

    Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. Computer Vision and Pattern Recognition (CVPR)

  32. [41]

    Aisha Urooj Khan, Amir Mazaheri, Niels da Vitoria Lobo, and Mubarak Shah. 2020. http://arxiv.org/abs/2010.14095 Mmft-bert: Multimodal fusion transformer with bert encodings for visual question answering

  33. [42]

    Humair Raj Khan, Deepak Gupta, and Asif Ekbal. 2021. http://arxiv.org/abs/2109.04653 Towards developing a multilingual and code-mixed visual question answering system by knowledge distillation

  34. [43]

    Deva Priyakumar, and C

    Yash Khare, Viraj Bagal, Minesh Mathew, Adithi Devi, U. Deva Priyakumar, and C. V. Jawahar. 2021. https://doi.org/10.48550/arXiv.2104.01394 Mmbert: Multimodal bert pretraining for improved medical vqa

  35. [44]

    J. Kim, Y. Jun, H. Zhang, and J. Kim. 2017. Bilinear attention networks for multimodal learning. In CVPR

  36. [45]

    J. Kim, Y. Jun, H. Zhang, and J. Kim. 2018. Learning to answer visual questions with attention on attention. In CVPR

  37. [46]

    Krishna, Y

    R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, ..., and L. Fei-Fei. 2017. Visual genome: Connecting language and vision using crowdsourced dense annotations. In IJCV, volume 123, pages 32--73

  38. [47]

    Shamma, Michael S

    Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li. 2016. http://arxiv.org/abs/1602.07332 Visual genome: Connecting language and vision using cr...

  39. [48]

    Krizhevsky, I

    A. Krizhevsky, I. Sutskever, and G. E. Hinton. 2012 a . Imagenet classification with deep convolutional neural networks. In NeurIPS 2012

  40. [49]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 b . https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf Imagenet classification with deep convolutional neural networks . In Advances in Neural Information Processing...

  41. [50]

    Jason Joseph Lau, Soumya Gayen, Dina Demner, and Asma Ben Abacha. 2018. https://doi.org/10.17605/OSF.IO/89KPS Visual question answering in radiology (vqa-rad) . Open Science Framework. A dataset of clinically generated visual questions and answers about radiology images

  42. [51]

    Haopeng Li, Andong Deng, Qiuhong Ke, Jun Liu, Hossein Rahmani, Yulan Guo, Bernt Schiele, and Chen Chen. 2024 a . http://arxiv.org/abs/2401.01505 Sports-qa: A large-scale video question answering benchmark for complex and professional sports

  43. [52]

    J. Li, H. Tan, L. Wang, M. Yang, and S. C. H. Hoi. 2019. Visualbert: A simple and performant visual language model. In NeurIPS

  44. [53]

    Jian Li, Weiheng Lu, Hao Fei, Meng Luo, Ming Dai, Min Xia, Yizhang Jin, Zhenye Gan, Ding Qi, Chaoyou Fu, Ying Tai, Wankou Yang, Yabiao Wang, and Chengjie Wang. 2024 b . http://arxiv.org/abs/2408.08632 A survey on benchmarks of multimodal large language models

  45. [54]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. http://arxiv.org/abs/2301.12597 Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

  46. [55]

    X. Li, H. Tan, L. Wang, M. Yang, and S. C. H. Hoi. 2020. Oscar: Object-semantics aware pre-training for vision-language tasks. In CVPR

  47. [56]

    Yingshu Li, Yunyi Liu, Zhanyu Wang, Xinyu Liang, Lei Wang, Lingqiao Liu, Leyang Cui, Zhaopeng Tu, Longyue Wang, and Luping Zhou. 2024 c . http://arxiv.org/abs/2310.20381 A systematic evaluation of gpt-4v's multimodal capability for medical image analysis

  48. [57]

    T. Y. Lin, M. Maire, S. Belongie, et al. 2014. Microsoft coco: Common objects in context. In ECCV 2014

  49. [58]

    Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015. https://doi.org/10.1109/CVPR.2015.7298965 Fully convolutional networks for semantic segmentation . In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3431--3440

  50. [59]

    J. Lu, Z. Yang, H. Mobahi, and D. Parikh. 2019 a . Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Proceedings of NeurIPS 2019

  51. [60]

    Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 b . http://arxiv.org/abs/1908.02265 Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

  52. [61]

    Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. 2017. http://arxiv.org/abs/1612.01887 Knowing when to look: Adaptive attention via a visual sentinel for image captioning

  53. [62]

    Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016. Hierarchical question-image co-attention for visual question answering. In Advances in neural information processing systems (NeurIPS), pages 289--297

  54. [63]

    Sandeep Maddu and Viziananda Row Sanapala. 2024. https://doi.org/10.1145/3695766 A survey on nlp tasks, resources and techniques for low-resource telugu-english code-mixed text . ACM Trans. Asian Low-Resour. Lang. Inf. Process. Just Accepted

  55. [64]

    Malinowski, M

    M. Malinowski, M. Rohrbach, and A. Vedaldi. 2014. Multimodal learning with deep convolutional neural networks. In ICCV

  56. [65]

    M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz. 1993. Building a large annotated corpus of english: The penn treebank. In ACL 1993

  57. [66]

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. http://arxiv.org/abs/1906.00067 Ok-vqa: A visual question answering benchmark requiring external knowledge

  58. [67]

    Khapra, and Pratyush Kumar

    Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra, and Pratyush Kumar. 2020. http://arxiv.org/abs/1909.00997 Plotqa: Reasoning over scientific plots

  59. [68]

    Mikolov, K

    T. Mikolov, K. Chen, G. Corrado, and J. Dean. 2013 a . Efficient estimation of word representations in vector space. In NeurIPS 2013

  60. [69]

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 b . http://arxiv.org/abs/1301.3781 Efficient estimation of word representations in vector space

  61. [70]

    Jong Hak Moon, Hyungyung Lee, Woncheol Shin, Young-Hak Kim, and Edward Choi. 2022. https://doi.org/10.1109/jbhi.2022.3207502 Multi-modal understanding and generation for medical images and text via vision-language pre-training . IEEE Journal of Biomedical and Health Informatic...

  62. [71]

    Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. 2017. Dual attention networks for multimodal reasoning and matching. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), pages 299--307

  63. [72]

    Narasimhan, P

    K. Narasimhan, P. Dixit, and A. Gupta. 2022. Knowledge-augmented neural networks for visual question answering. arXiv preprint arXiv:2203.13843

  64. [73]

    Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021. https://doi.org/10.18653/v1/2021.mrl-1.11 Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages . In Proceedings of the 1st Workshop on Multilingual Representation ...

  65. [74]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, O. Agrawal, and I. Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of CVPR 2021

  66. [75]

    Alec Radford and Karthik Narasimhan. 2018. https://api.semanticscholar.org/CorpusID:49313245 Improving language understanding by generative pre-training

  67. [76]

    Dai, Nissan Hajaj, Michaela Hardt, Peter J

    Alvin Rajkomar, Eyal Oren, Kai Chen, Andrew M. Dai, Nissan Hajaj, Michaela Hardt, Peter J. Liu, Xiaobing Liu, Jake Marcus, Mimi Sun, Patrik Sundberg, Hector Yee, Kun Zhang, Yi Zhang, Gerardo Flores, Gavin E Duggan, Jamie Irvine, Quoc V. Le, Kurt Litsch, Alexander Mossin, Justi...

  68. [77]

    Karen Simonyan and Andrew Zisserman. 2015. http://arxiv.org/abs/1409.1556 Very deep convolutional networks for large-scale image recognition

  69. [78]

    Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. http://arxiv.org/abs/1904.08920 Towards vqa models that can read

  70. [79]

    Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019. http://arxiv.org/abs/1811.00491 A corpus for reasoning about natural language grounded in photographs

  71. [80]

    Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019. http://arxiv.org/abs/1904.01766 Videobert: A joint model for video and language representation learning

  72. [81]

    Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. http://arxiv.org/abs/1409.3215 Sequence to sequence learning with neural networks

  73. [82]

    Tan and M

    H. Tan and M. Bansal. 2019. Lxmert: Learning cross-modality encoder representations from transformers. In Proceedings of IJCNLP 2019

  74. [83]

    Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016. http://arxiv.org/abs/1512.02902 Movieqa: Understanding stories in movies through question-answering

  75. [84]

    Teney, L

    D. Teney, L. Shen, L. Demszky, and K. Saenko. 2017. Graph neural networks for visual question answering. In CVPR

  76. [85]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2023. http://arxiv.org/abs/1706.03762 Attention is all you need

  77. [86]

    Vinyals, A

    O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. 2015. Show and tell: A neural image caption generator. In CVPR 2015

  78. [87]

    Viola and M

    P. Viola and M. Jones. 2001. https://doi.org/10.1109/CVPR.2001.990517 Rapid object detection using a boosted cascade of simple features . In Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. CVPR 2001, volume 1, pages I--I

  79. [88]

    Min Wang, Ata Mahjoubfar, and Anupama Joshi. 2022. http://arxiv.org/abs/2208.11253 Fashionvqa: A domain-specific visual question answering system

  80. [89]

    Yizhou Wang, Ruiyi Zhang, Haoliang Wang, Uttaran Bhattacharya, Yun Fu, and Gang Wu. 2023. http://arxiv.org/abs/2312.02310 Vaquita: Enhancing alignment in llm-assisted video understanding

  81. [90]

    J. Xiao, J. Hays, K. Ehinger, A. Oliva, and A. Torralba. 2010. Sun database: Large-scale scene recognition from abbey to zoo. In CVPR 2010

  82. [91]

    Zemel, and Yoshua Bengio

    Kelvin Xu, Jimmy Lei Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio. 2015. Show, attend and tell: neural image caption generation with visual attention. In Proceedings of the 32nd International Conference on Internatio...

  83. [92]

    Z. Yang, X. He, J. Gao, L. Deng, and A. Smola. 2016. Stacked attention networks for image question answering. In CVPR

  84. [93]

    Z. Yang, J. Mao, T. Huang, L. Xu, and Y. Wang. 2018. Dynamic scene graph for visual question answering. In CVPR

  85. [94]

    K. Yi, J. Shen, Z. Zhang, H. Zhang, and A. Hengel. 2018. Neural-symbolic visual reasoning and generation. In ECCV

  86. [95]

    Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. 2019. http://arxiv.org/abs/1906.10770 Deep modular co-attention networks for visual question answering

  87. [96]

    Amir Zadeh, Michael Chan, Paul Pu Liang, Edmund Tong, and Louis-Philippe Morency. 2019. Social-iq: A question answering benchmark for artificial social intelligence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  88. [97]

    Zhong et al

    Z. Zhong et al. 2022. https://arxiv.org/abs/2203.01225 Videoqa: A comprehensive survey of datasets and methods . arXiv preprint arXiv:2203.01225

  89. [98]

    Corso, and Jianfeng Gao

    Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2019. http://arxiv.org/abs/1909.11059 Unified vision-language pre-training for image captioning and vqa

  90. [99]

    Bo Zou, Chao Yang, Yu Qiao, Chengbin Quan, and Youjian Zhao. 2024. http://arxiv.org/abs/2404.00973 Videodistill: Language-aware vision distillation for video question answering

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.