Pith. sign in

REVIEW 3 major objections 5 minor 229 references

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This survey claims to be the first to organize vision-language methods that use pre-trained models according to the classic challenge each method tackles, and it backs the map with comparative tables across images and videos.

desk verdict A useful survey with a sensible challenge-based taxonomy, despite a few citation slip-ups and an asserted rather than argued choice of four challenges. read the letter →

arxiv 2412.08158 v1 pith:GYLNSXOC submitted 2024-12-11 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords vision-languagetaskspre-trainedmodelslargelanguagedatascarcitychain-of-thoughtreasoningopen-vocabularygeneralizationtaskdiversity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the right way to understand the recent wave of vision-language methods is to look at which classic bottleneck each method attacks, and that large pre-trained models are the common ingredient that lets those attacks work. It identifies four bottlenecks—scarce annotated data, increasingly complex reasoning, poor generalization to novel samples, and task diversity—and files recent methods under one of them, covering both language models and vision-language models, images and videos, and classic and recent work. If the map is right, a researcher entering the field can pick a challenge and immediately see which pre-training paradigm has been used against it and how well it works. The paper also catalogs risks the pre-trained models bring, such as hallucination, stale knowledge, and concept-association bias, and points to mitigation directions.

What carries the argument

The machinery is the survey's challenge-based taxonomy itself, laid out in Table II. Each of the four challenges anchors a family of paradigms: data scarcity is attacked by direct inference, uni-modal training, or pseudo-pair generation; escalating reasoning complexity is attacked by divide-and-conquer or chain-of-thought decomposition; generalization to novel samples is attacked by extracting semantic context from a language model or distilling teacher knowledge from a vision-language model; and task diversity is attacked by continual learning or by planning with natural language or code statements. Illustrated pipeline diagrams make each paradigm concrete, and the benchmark tables anchor each family to measured performance.

What would settle it

Find a vision-language method whose stated motivation is purely efficiency or safety—cutting inference cost or preventing harmful outputs—with no dependence on the four challenges; its absence from the survey's taxonomy would show that the challenge set is not exhaustive.

Watch

Extended reading notes

Core claim

The paper's central claim is that vision-language research today is best understood as a set of responses to four classic challenges, and that pre-trained models supply the capabilities—language priors, a shared image-text space, world knowledge, and in-context flexibility—that let those responses work where earlier methods failed. It argues that before pre-training, each challenge was only partially addressed: semi- and weakly supervised methods overfitted, fixed multi-step reasoners could not scale their reasoning hops, knowledge-base lookups were rigid and thin, and multi-task models suffered catastrophic forgetting. Pre-trained models change the options by enabling direct inference on test samples, training from unlabeled uni-modal data through the CLIP common space, generating pseudo-paired data, divide-and-conquer and chain-of-thought reasoning, extracting semantic context from language models, distilling teacher knowledge from vision-language models, continual learning, and language-model-planned tool use. The paper supports the map with comparative tables showing that methods in each paradigm beat their pre-training-era baselines.

Load-bearing premise

The taxonomy holds only if the four challenges—data scarcity, escalating reasoning complexity, generalization to novel samples, and task diversity—are the main bottlenecks and are distinct enough that each method can be filed under one of them.

Editorial extensions

If this is right

  • If the survey's map is right, a practitioner short on annotated data has three proven routes: direct inference with an LLM plus CLIP, text-only training through the CLIP common space, or generating pseudo-paired data, with the pseudo-pair route approaching fully supervised captioning scores.
  • Decomposing questions or reasoning paths—divide-and-conquer or chain-of-thought—consistently beats one-step reasoning on OK-VQA, A-OKVQA, VCR, SNLI-VE, and ScienceQA, so complexity should be attacked by decomposition rather than by bigger single-step models.
  • For novel samples, querying an LLM for class descriptions improves CLIP's open-vocabulary classification, and distilling a VLM into a close-set detector raises novel-class average precision while keeping base-class performance.
  • A single general system with an LLM planner calling tools can cover many tasks zero-shot, making task-specific training unnecessary for the covered tasks.
  • The same pre-trained capabilities carry risks—hallucination, outdated knowledge, concept-association bias, and compositional confusion—so methods that use them should include verification, retrieval, or code execution to counter those risks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the four-challenge frame doubles as a design checklist; a new vision-language method can be positioned by naming the bottleneck it attacks, which suggests the taxonomy will seed future method papers even if its boundaries blur.
  • Editorial inference: the survey's risk section implies that the next performance gains will come from pairing methods—for example, LLM semantic context to compensate what VLM distillation loses, or retrieval to fix outdated knowledge—though the paper only hints at such combinations.
  • Editorial inference: one could test the taxonomy's completeness by checking whether methods motivated purely by efficiency or safety (e.g., inference-cost reduction or harm avoidance) can be classified; the survey does not cover such motivations.
  • Editorial inference: since the four challenges overlap in practice—novel samples often require reasoning, and task diversity is a generalization problem across tasks—a quantitative study measuring how single methods score on more than one challenge could refine or merge the categories.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This survey organizes recent methods that integrate large pre-trained models (both LLMs and VLMs) into vision-language tasks around four challenges: data scarcity, escalating reasoning complexity, generalization to novel samples, and task diversity. For each challenge, it reviews the corresponding paradigms (e.g., direct inference, chain-of-thought, knowledge distillation, LLM-as-planner), illustrates them with pipeline figures, and collects performance tables for image captioning, complex reasoning benchmarks, open-vocabulary classification and detection, and general modular systems. The paper closes with a discussion of four risks introduced by pre-trained models (hallucination, outdated knowledge, concept association bias, and compositional concept confusion) and suggests mitigation strategies. The central claim, stated in Section I, is that this is the first survey focused specifically on how vision-language tasks benefit from large pre-trained models and the first to categorize methods according to the challenges they tackle.

Significance. If the taxonomy is accepted, the survey provides a useful organizational map of a fast-moving area and a convenient entry point for researchers: the pipeline diagrams in Figures 2 through 5 are pedagogical, and the performance tables make cross-paradigm comparisons accessible. The paper covers both images and videos, both discriminative and generative models, and includes a thoughtful discussion of risks that is often absent from method-oriented surveys. The comparisons in Tables III-VI are concrete and falsifiable in that they name specific methods and benchmarks. The main value is synthetic: it brings together method families that are usually scattered across separate papers and frames them under a small set of recurring challenges. No machine-checked artifacts are shipped, but the claims are checkable against the cited literature.

major comments (3)
  1. [Section III and Table II] The paper's central contribution is the challenge-based taxonomy, but the taxonomy is asserted rather than validated, and the assignment of methods to exactly one challenge is often ambiguous. For instance, ZeroCap [61] is filed under 'Data Scarcity – Direct inference on test samples,' yet its goal is to caption images never seen during training, which overlaps with the definition of 'Generalization to Novel Samples' given in Section III-C; likewise, the pseudo-paired-data methods [84]-[86] create synthetic samples that address data scarcity and simultaneously improve robustness to novel samples. The survey should either provide a concrete decision rule for assigning methods, allow a method to appear under multiple challenges, or explicitly frame the four challenges as one useful partition rather than the 'main challenges' faced by all models. Without this, the claimed novelty of categorizing methods by challenge is not fully supported.
  2. [Section III and Section VI] The paper claims that the four challenges are the 'main challenges' and describes the survey as comprehensive, but the taxonomy omits challenges that are prominent in current vision-language research, such as alignment and safety, computational efficiency, and robustness beyond novel-sample generalization. The four risks discussed in Section VI (hallucination, outdated knowledge, concept association bias, and compositional concept confusion) are not mapped back to the challenge taxonomy, so it is unclear how these risks interact with the four categories. This omission is not fatal if the paper narrows its scope, but as written the 'comprehensive' claim in the abstract and Section I is stronger than the taxonomy supports.
  3. [Section V-D] The grouping of continual learning and LLM-as-planner under 'Task Diversity' is not self-evident. Continual learning addresses catastrophic forgetting and parameter efficiency across a sequence of tasks, not primarily the diversity of input-output workflows; LLM-as-planner systems address compositional reasoning, modularity, and tool use. The survey should explain why these paradigms are classified under task diversity rather than under reasoning complexity or generalization, and should state how 'task diversity' is distinguished from the other three challenges.
minor comments (5)
  1. [Section V-C and Table V] The sentence 'Table V reports the comparison results of methods [9], [10], [12], [203]' is inconsistent with the actual content of Table V, which lists only VCD [9], LCDAtt [10], and CPHC [12]; reference [203] (OVR-CNN) appears in Table VI as an object-detection baseline and is not an image-classification method. Please correct the text or add the missing row to Table V.
  2. [Section II] The sentence 'Table I shows the differences between our survey and the existing related surveys in terms of content and coverage' is immediately repeated with slightly different wording; the duplicate should be removed.
  3. [Table VII] In Table VII, the planning-format entry for MM-VID reads 'natrual lang', which appears to be a typo for 'natural lang'.
  4. [Section VI-A] The benchmark name 'HALLUSIONBENCH' should be written as 'HallusionBench' to match the cited paper and standard usage.
  5. [Header and metadata] The page header 'JOURNAL OF IEEE, VOL. 14, NO. 8, AUGUST 2021' appears to be a leftover from a template, since the arXiv submission is dated December 2024 and the actual journal/volume information is not provided; this should be corrected or removed.

Circularity Check

0 steps flagged · score 1.0 of 10

No material circularity; the survey contains only peripheral self-citations used as method examples, and its central taxonomy is asserted rather than derived, which is a validity concern, not a circularity concern.

full rationale

This is a survey paper, not a derivation or prediction pipeline. There is no fitted parameter that is later relabeled as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in through citation. The paper's central claim is that it is pioneering in focusing on how vision-language tasks benefit from pre-trained models and in categorizing methods by the challenges they tackle (Section I: 'To the best of our knowledge, this survey is pioneering in its focus on this topic and in categorizing methods according to the challenges they tackle.'). That claim is a novelty assertion, not a derived result, so it cannot be circular in the sense of reducing to its own input. The four challenges in Section III (data scarcity, escalating reasoning complexity, generalization to novel samples, and task diversity) are organizing categories introduced by the survey; the paper does not pretend to derive them from first principles. Whether the taxonomy is exhaustive or whether individual methods are assigned to the right bucket is a correctness and completeness concern, not a circularity concern. The only self-citations by the survey authors are reference [195] (Y. Qi, W. Zhao, and X. Wu, 'Relational distant supervision for image captioning without image-text pairs') and reference [204] (Y. Shi, X. Wu, H. Lin, and J. Luo, 'Commonsense knowledge prompting for few-shot action recognition in videos'). Both are cited as examples of existing methods: [195] appears in the summary of pre-pretraining-era approaches as 'constructing pseudo-paired data based on object relationships [195]', and [204] appears in Section V-C as 'Shi et al. [204] enhance the semantics of action categories by using BERT [52] to collect text proposals containing language descriptions of actions.' Neither citation is used to justify the survey's novelty claim, its taxonomy, or any evaluative conclusion; both are ordinary literature pointers within a survey. Consequently, these self-citations are not load-bearing and do not raise the circularity score beyond 1. The paper is also self-contained against external benchmarks: Tables III-VI compare independently published methods on standard datasets, and the survey's discussion of risks is drawn from external studies. Therefore, the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey introduces no free parameters and no new postulated entities. Its central claim rests on framing assumptions about the field, the reliability of cited benchmark numbers, and the transferable capabilities of pre-trained models. These are reasonable assumptions for a review, but they are not demonstrated within the paper.

assumptions (3)
  • domain assumption Vision-language tasks face exactly four main challenges: data scarcity, escalating reasoning complexity, generalization to novel samples, and task diversity.
    Section III presents these four challenges as the organizing principle of the survey without arguing that they are exhaustive or non-overlapping. The entire taxonomy in Table II depends on this partition.
  • domain assumption The benchmark numbers reported in the cited papers are accurate and comparable across the tables in this survey.
    The survey aggregates results from different papers, backbones, and experimental protocols in Tables III through VI. The conclusions about which methods outperform others assume these numbers can be meaningfully compared.
  • domain assumption Large pre-trained models possess transferable capabilities, including zero-shot inference, stored world knowledge, and a shared vision-language embedding space.
    The survey treats these capabilities as established background from the cited literature, not as something it proves. This assumption underlies every paradigm described in Section V.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey." pith.science (2026). https://pith.science/paper/GYLNSXOC

@misc{pith2026241208158,
  author       = {Pith},
  title        = {Pith review of: How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GYLNSXOC}},
  note         = {Machine review of arXiv:2412.08158}
}
read the original abstract

The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelligence and continuously attracts the research community's attention. Despite the improvements in overall performance, classic challenges still exist in vision-language tasks and hinder the development of this area. In recent years, the rise of pre-trained models is driving the research on vision-language tasks. Thanks to the massive scale of training data and model parameters, pre-trained models have exhibited excellent performance in numerous downstream tasks. Inspired by the powerful capabilities of pre-trained models, new paradigms have emerged to solve the classic challenges. Such methods have become mainstream in current research with increasing attention and rapid advances. In this paper, we present a comprehensive overview of how vision-language tasks benefit from pre-trained models. First, we review several main challenges in vision-language tasks and discuss the limitations of previous solutions before the era of pre-training. Next, we summarize the recent advances in incorporating pre-trained models to address the challenges in vision-language tasks. Finally, we analyze the potential risks associated with the inherent limitations of pre-trained models and discuss possible solutions, attempting to provide future research directions.

Figures

Figures reproduced from arXiv: 2412.08158 by the authors.

Figure 1
Figure 1. An illustration of four classic challenges in vision-language tasks. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Paradigms for addressing the data scarcity challenge in vision-language tasks with the help of pre-trained models. (a) shows the paradigm of integrating [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. 1) Divide-and-conquer: Thanks to the synergy between LLMs and VLMs, the main question in a complex visual￾language reasoning task can be decomposed into a series of sub-questions about visual details and sub-questions about factual knowledge related to the main question. Since the sub￾questions are easier for models to answer than the original question and the information answered helps reasoning, the divide-and-con… view at source ↗
Figures from the paper (3 more)
Figure 3
Figure 3. Figure 3: Paradigms of using pre-trained models to conquer the challenge of escalating reasoning complexity in vision-language tasks. (a) shows the basic idea [PITH_FULL_IMAGE:figures/full_fig_p010_3.png]
Figure 4
Figure 4. Figure 4: Paradigms of using pre-trained models to conquer the generalization challenge to novel samples in vision-language tasks. (a) shows the basic idea of [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Paradigm of using pre-trained models to conquer the task diversity challenge in vision-language tasks. (a) shows the basic idea of applying continual [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

229 extracted references · 15 canonical work pages

  1. [9]

    Visual classification via description from large language models,

    S. Menon and C. V ondrick, “Visual classification via description from large language models,” in The Eleventh International Conference on Learning Representations, 2022

  2. [10]

    Learning concise and descriptive attributes for visual recognition,

    A. Yan, Y . Wang, Y . Zhong, C. Dong, Z. He, Y . Lu, W. Y . Wang, J. Shang, and J. McAuley, “Learning concise and descriptive attributes for visual recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3090–3100

  3. [12]

    Chatgpt-powered hierarchical comparisons for image classification,

    Z. Ren, Y . Su, and X. Liu, “Chatgpt-powered hierarchical comparisons for image classification,” Advances in Neural Information Processing Systems, vol. 36, 2024

  4. [203]

    Open-vocabulary object detection using captions,

    A. Zareian, K. D. Rosa, D. H. Hu, and S.-F. Chang, “Open-vocabulary object detection using captions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 393–14 402

  5. [110]

    Open-vocabulary object detec- tion via vision and language knowledge distillation,

    X. Gu, T.-Y . Lin, W. Kuo, and Y . Cui, “Open-vocabulary object detec- tion via vision and language knowledge distillation,” in International Conference on Learning Representations , 2022

  6. [114]

    Open- vocabulary object detection upon frozen vision and language models,

    W. Kuo, Y . Cui, X. Gu, A. Piergiovanni, and A. Angelova, “Open- vocabulary object detection upon frozen vision and language models,” in The Eleventh International Conference on Learning Representations, 2023

  7. [115]

    Cohoz: Contrastive multimodal prompt tuning for hierarchical open- set zero-shot recognition,

    N. Liao, Y . Liu, L. Xiaobo, C. Lei, G. Wang, X.-S. Hua, and J. Yan, “Cohoz: Contrastive multimodal prompt tuning for hierarchical open- set zero-shot recognition,” in Proceedings of the 30th ACM Interna- tional Conference on Multimedia , 2022, pp. 3262–3271

  8. [117]

    Lmc: Large model collaboration with cross-assessment for training-free open-set object recognition,

    H. Qu, X. Hui, Y . Cai, and J. Liu, “Lmc: Large model collaboration with cross-assessment for training-free open-set object recognition,” Advances in Neural Information Processing Systems , vol. 36, 2024

  9. [61]

    Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,

    Y . Tewel, Y . Shalev, I. Schwartz, and L. Wolf, “Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,” in Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 17 918–17 928

  10. [84]

    Image captioning with multi-context synthetic data,

    F. Ma, Y . Zhou, F. Rao, Y . Zhang, and X. Sun, “Image captioning with multi-context synthetic data,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 5, 2024, pp. 4089–4097

  11. [86]

    Freemask: Synthetic images with dense annotations make stronger segmentation models,

    L. Yang, X. Xu, B. Kang, Y . Shi, and H. Zhao, “Freemask: Synthetic images with dense annotations make stronger segmentation models,” arXiv preprint arXiv:2310.15160 , 2023

Show all 229 references
  1. [1]

    Multimodal research in vision and language: A review of current and emerging trends,

    S. Uppal, S. Bhagat, D. Hazarika, N. Majumder, S. Poria, R. Zimmer- mann, and A. Zadeh, “Multimodal research in vision and language: A review of current and emerging trends,” Information Fusion , vol. 77, pp. 149–171, 2022

  2. [2]

    Vision+ language applications: A survey,

    Y . Zhou and N. Shimada, “Vision+ language applications: A survey,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 826–842

  3. [3]

    Bottom-up and top-down attention for image captioning and visual question answering,

    P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 6077–6086

  4. [4]

    Swinbert: End-to-end transformers with sparse attention for video captioning,

    K. Lin, L. Li, C.-C. Lin, F. Ahmed, Z. Gan, Z. Liu, Y . Lu, and L. Wang, “Swinbert: End-to-end transformers with sparse attention for video captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 17 949–17 958

  5. [5]

    Vqa: Visual question answering,

    S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 2425– 2433

  6. [6]

    Movieqa: Understanding stories in movies through question- answering,

    M. Tapaswi, Y . Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler, “Movieqa: Understanding stories in movies through question- answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4631–4640

  7. [7]

    Context-aware attention network for image-text retrieval,

    Q. Zhang, Z. Lei, Z. Zhang, and S. Z. Li, “Context-aware attention network for image-text retrieval,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 3536–3545

  8. [8]

    Fine-grained video-text retrieval with hierarchical graph reasoning,

    S. Chen, Y . Zhao, Q. Jin, and Q. Wu, “Fine-grained video-text retrieval with hierarchical graph reasoning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 10 638–10 647

  9. [11]

    Exploring large language models for multi-modal out-of-distribution detection,

    Y . Dai, H. Lang, K. Zeng, F. Huang, and Y . Li, “Exploring large language models for multi-modal out-of-distribution detection,” in The 2023 Conference on Empirical Methods in Natural Language Processing, 2023

  10. [13]

    Vision-language pre-training: Basics, recent advances, and future trends,

    Z. Gan, L. Li, C. Li, L. Wang, Z. Liu, J. Gao et al., “Vision-language pre-training: Basics, recent advances, and future trends,” Foundations and Trends® in Computer Graphics and Vision , vol. 14, no. 3–4, pp. 163–352, 2022. JOURNAL OF IEEE, VOL. 14, NO. 8, AUGUST 2021 19

  11. [14]

    Show and tell: A neural image caption generator,

    O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” inProceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3156–3164

  12. [15]

    Deep correlation for matching images and text,

    F. Yan and K. Mikolajczyk, “Deep correlation for matching images and text,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3441–3450

  13. [16]

    Draw: A recurrent neural network for image generation,

    K. Gregor, I. Danihelka, A. Graves, D. Rezende, and D. Wierstra, “Draw: A recurrent neural network for image generation,” in Interna- tional conference on machine learning. PMLR, 2015, pp. 1462–1471

  14. [17]

    Memory- attended recurrent network for video captioning,

    W. Pei, J. Zhang, X. Wang, L. Ke, X. Shen, and Y .-W. Tai, “Memory- attended recurrent network for video captioning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 8347–8356

  15. [18]

    Heterogeneous memory enhanced multimodal attention model for video question answering,

    C. Fan, X. Zhang, S. Zhang, W. Wang, C. Zhang, and H. Huang, “Heterogeneous memory enhanced multimodal attention model for video question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 1999–2007

  16. [19]

    Mocogan: Decompos- ing motion and content for video generation,

    S. Tulyakov, M.-Y . Liu, X. Yang, and J. Kautz, “Mocogan: Decompos- ing motion and content for video generation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 1526–1535

  17. [20]

    Rethinking the bottom-up framework for query-based video localiza- tion,

    L. Chen, C. Lu, S. Tang, J. Xiao, D. Zhang, C. Tan, and X. Li, “Rethinking the bottom-up framework for query-based video localiza- tion,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 10 551–10 558

  18. [21]

    Object detection in 20 years: A survey,

    Z. Zou, K. Chen, Z. Shi, Y . Guo, and J. Ye, “Object detection in 20 years: A survey,” Proceedings of the IEEE , vol. 111, no. 3, pp. 257– 276, 2023

  19. [22]

    Neural motifs: Scene graph parsing with global context,

    R. Zellers, M. Yatskar, S. Thomson, and Y . Choi, “Neural motifs: Scene graph parsing with global context,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 5831–5840

  20. [23]

    From recognition to cognition: Visual commonsense reasoning,

    R. Zellers, Y . Bisk, A. Farhadi, and Y . Choi, “From recognition to cognition: Visual commonsense reasoning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6720–6731

  21. [24]

    Visual entailment: A novel task for fine-grained image understanding,

    N. Xie, F. Lai, D. Doran, and A. Kadav, “Visual entailment: A novel task for fine-grained image understanding,” arXiv preprint arXiv:1901.06706, 2019

  22. [25]

    The abduction of sherlock holmes: A dataset for visual abductive reasoning,

    J. Hessel, J. D. Hwang, J. S. Park, R. Zellers, C. Bhagavatula, A. Rohrbach, K. Saenko, and Y . Choi, “The abduction of sherlock holmes: A dataset for visual abductive reasoning,” in European Con- ference on Computer Vision . Springer, 2022, pp. 558–575

  23. [26]

    Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional im- ages,

    N. Bitton-Guetta, Y . Bitton, J. Hessel, L. Schmidt, Y . Elovici, G. Stanovsky, and R. Schwartz, “Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional im- ages,” in Proceedings of the IEEE/CVF International Conference on Computer Vision...

  24. [27]

    Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,

    S. Zhong, Z. Huang, S. Gao, W. Wen, L. Lin, M. Zitnik, and P. Zhou, “Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024

  25. [28]

    Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,

    P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys, vol. 55, no. 9, pp. 1–35, 2023

  26. [29]

    Llama: Open and efficient foundation language models,

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023

  27. [30]

    Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

    W.-L. Chiang, Z. Li, Z. Lin, Y . Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y . Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/

  28. [31]

    Qwen technical report,

    J. Bai, S. Bai, Y . Chu, Z. Cui, K. Dang, X. Deng, Y . Fan, W. Ge, Y . Han, F. Huang et al. , “Qwen technical report,” arXiv preprint arXiv:2309.16609, 2023

  29. [32]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763

  30. [33]

    Scaling up visual and vision-language representation learning with noisy text supervision,

    C. Jia, Y . Yang, Y . Xia, Y .-T. Chen, Z. Parekh, H. Pham, Q. Le, Y .-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in International conference on machine learning . PMLR, 2021, pp. 4904–4916

  31. [34]

    Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,

    H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, S. Piao, and F. Wei, “Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,” Advances in Neural Infor- mation Processing Systems , vol. 35, pp. 32 897–32 912, 2022

  32. [35]

    Filip: Fine-grained interactive language-image pre-training,

    L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu, “Filip: Fine-grained interactive language-image pre-training,” arXiv preprint arXiv:2111.07783 , 2021

  33. [36]

    Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

    J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in International conference on machine learning . PMLR, 2022, pp. 12 888–12 900

  34. [37]

    Minigpt-4: En- hancing vision-language understanding with advanced large language models,

    D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: En- hancing vision-language understanding with advanced large language models,” arXiv preprint arXiv:2304.10592 , 2023

  35. [38]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” in NeurIPS, 2023

  36. [39]

    Video-llama: An instruction-tuned audio-visual language model for video understanding,

    H. Zhang, X. Li, and L. Bing, “Video-llama: An instruction-tuned audio-visual language model for video understanding,” arXiv preprint arXiv:2306.02858, 2023

  37. [40]

    A survey of large language models,

    W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Dong et al. , “A survey of large language models,” arXiv preprint arXiv:2303.18223 , 2023

  38. [41]

    A comprehensive overview of large language models,

    H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Barnes, and A. Mian, “A comprehensive overview of large language models,” arXiv preprint arXiv:2307.06435 , 2023

  39. [42]

    A survey on evaluation of large language models,

    Y . Chang, X. Wang, J. Wang, Y . Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wang et al. , “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 3, pp. 1–45, 2024

  40. [43]

    Large-scale multi-modal pre-trained models: A com- prehensive survey,

    X. Wang, G. Chen, G. Qian, P. Gao, X.-Y . Wei, Y . Wang, Y . Tian, and W. Gao, “Large-scale multi-modal pre-trained models: A com- prehensive survey,” Machine Intelligence Research, vol. 20, no. 4, pp. 447–482, 2023

  41. [44]

    Foundational models defining a new era in vision: A survey and outlook,

    M. Awais, M. Naseer, S. Khan, R. M. Anwer, H. Cholakkal, M. Shah, M.-H. Yang, and F. S. Khan, “Foundational models defining a new era in vision: A survey and outlook,” arXiv preprint arXiv:2307.13721 , 2023

  42. [45]

    A survey on multimodal large language models,

    S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen, “A survey on multimodal large language models,” arXiv preprint arXiv:2306.13549, 2023

  43. [46]

    Mm-llms: Recent advances in multimodal large language models,

    D. Zhang, Y . Yu, C. Li, J. Dong, D. Su, C. Chu, and D. Yu, “Mm-llms: Recent advances in multimodal large language models,” arXiv preprint arXiv:2401.13601, 2024

  44. [47]

    Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning,

    Y . Wang, W. Chen, X. Han, X. Lin, H. Zhao, Y . Liu, B. Zhai, J. Yuan, Q. You, and H. Yang, “Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning,” arXiv preprint arXiv:2401.06805 , 2024

  45. [48]

    Video understanding with large language models: A survey,

    Y . Tang, J. Bi, S. Xu, L. Song, S. Liang, T. Wang, D. Zhang, J. An, J. Lin, R. Zhu et al., “Video understanding with large language models: A survey,” arXiv preprint arXiv:2312.17432 , 2023

  46. [49]

    Vision-language models for vision tasks: A survey,

    J. Zhang, J. Huang, S. Jin, and S. Lu, “Vision-language models for vision tasks: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

  47. [50]

    Llms meet multimodal generation and editing: A survey,

    Y . He, Z. Liu, J. Chen, Z. Tian, H. Liu, X. Chi, R. Liu, R. Yuan, Y . Xing, W. Wang et al. , “Llms meet multimodal generation and editing: A survey,” arXiv preprint arXiv:2405.19334 , 2024

  48. [51]

    Generalized out-of-distribution detection and beyond in vision language model era: A survey,

    A. Miyai, J. Yang, J. Zhang, Y . Ming, Y . Lin, Q. Yu, G. Irie, S. Joty, Y . Li, H. Li et al. , “Generalized out-of-distribution detection and beyond in vision language model era: A survey,” arXiv preprint arXiv:2407.21794, 2024

  49. [52]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, 2018

  50. [53]

    Ernie: Enhanced representation through knowledge integration,

    Y . Sun, S. Wang, Y . Li, S. Feng, X. Chen, H. Zhang, X. Tian, D. Zhu, H. Tian, and H. Wu, “Ernie: Enhanced representation through knowledge integration,” arXiv preprint arXiv:1904.09223 , 2019

  51. [54]

    Language mod- els are few-shot learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language mod- els are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020

  52. [55]

    Image captioning with very scarce supervised data: Adversarial semi-supervised learning approach,

    D.-J. Kim, J. Choi, T.-H. Oh, and I. S. Kweon, “Image captioning with very scarce supervised data: Adversarial semi-supervised learning approach,” arXiv preprint arXiv:1909.02201 , 2019

  53. [56]

    Semi-supervised cross-modal retrieval with label prediction,

    D. Mandal, P. Rao, and S. Biswas, “Semi-supervised cross-modal retrieval with label prediction,” IEEE Transactions on Multimedia , vol. 22, no. 9, pp. 2345–2353, 2019. JOURNAL OF IEEE, VOL. 14, NO. 8, AUGUST 2021 20

  54. [57]

    Weakly supervised dense event captioning in videos,

    X. Duan, W. Huang, C. Gan, J. Wang, W. Zhu, and J. Huang, “Weakly supervised dense event captioning in videos,” Advances in Neural Information Processing Systems , vol. 31, 2018

  55. [58]

    Weakly-supervised visual-retriever-reader for knowledge-based question answering,

    M. Luo, Y . Zeng, P. Banerjee, and C. Baral, “Weakly-supervised visual-retriever-reader for knowledge-based question answering,” arXiv preprint arXiv:2109.04014, 2021

  56. [59]

    Unsupervised image captioning,

    Y . Feng, L. Ma, W. Liu, and J. Luo, “Unsupervised image captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4125–4134

  57. [60]

    Towards unsupervised image captioning with shared multimodal embeddings,

    I. Laina, C. Rupprecht, and N. Navab, “Towards unsupervised image captioning with shared multimodal embeddings,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 7414–7424

  58. [62]

    Language models can see: Plugging visual controls in text generation,

    Y . Su, T. Lan, Y . Liu, F. Liu, D. Yogatama, Y . Wang, L. Kong, and N. Collier, “Language models can see: Plugging visual controls in text generation,” arXiv preprint arXiv:2205.02655 , 2022

  59. [63]

    Conzic: Controllable zero-shot image captioning by sampling-based polishing,

    Z. Zeng, H. Zhang, R. Lu, D. Wang, B. Chen, and Z. Wang, “Conzic: Controllable zero-shot image captioning by sampling-based polishing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 23 465–23 476

  60. [64]

    Mea- cap: Memory-augmented zero-shot image captioning,

    Z. Zeng, Y . Xie, H. Zhang, C. Chen, Z. Wang, and B. Chen, “Mea- cap: Memory-augmented zero-shot image captioning,” arXiv preprint arXiv:2403.03715, 2024

  61. [65]

    Text-only training for visual storytelling,

    Y . Wang, W. Zhou, Z. Lu, and H. Li, “Text-only training for visual storytelling,” in Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 3686–3695

  62. [66]

    Zero- shot video captioning with evolving pseudo-tokens,

    Y . Tewel, Y . Shalev, R. Nadler, I. Schwartz, and L. Wolf, “Zero- shot video captioning with evolving pseudo-tokens,” arXiv preprint arXiv:2207.11100, 2022

  63. [67]

    Zero-shot dense video captioning by jointly optimizing text and moment,

    Y . Jo, S. Lee, A. S. Lee, H. Lee, H. Oh, and M. Seo, “Zero-shot dense video captioning by jointly optimizing text and moment,” arXiv preprint arXiv:2307.02682, 2023

  64. [68]

    Clip models are few-shot learners: Empirical studies on vqa and visual entailment,

    H. Song, L. Dong, W.-N. Zhang, T. Liu, and F. Wei, “Clip models are few-shot learners: Empirical studies on vqa and visual entailment,” arXiv preprint arXiv:2203.07190 , 2022

  65. [69]

    Towards counterfactual image manipulation via clip,

    Y . Yu, F. Zhan, R. Wu, J. Zhang, S. Lu, M. Cui, X. Xie, X.-S. Hua, and C. Miao, “Towards counterfactual image manipulation via clip,” in Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 3637–3645

  66. [70]

    An empirical study of gpt-3 for few-shot knowledge-based vqa,

    Z. Yang, Z. Gan, J. Wang, X. Hu, Y . Lu, Z. Liu, and L. Wang, “An empirical study of gpt-3 for few-shot knowledge-based vqa,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 3, 2022, pp. 3081–3089

  67. [71]

    Plug-and- play vqa: Zero-shot vqa by conjoining large pretrained models with zero training,

    A. M. H. Tiong, J. Li, B. Li, S. Savarese, and S. C. Hoi, “Plug-and- play vqa: Zero-shot vqa by conjoining large pretrained models with zero training,” arXiv preprint arXiv:2210.08773 , 2022

  68. [72]

    Language models with image descriptors are strong few-shot video-language learners,

    Z. Wang, M. Li, R. Xu, L. Zhou, J. Lei, X. Lin, S. Wang, Z. Yang, C. Zhu, D. Hoiem et al. , “Language models with image descriptors are strong few-shot video-language learners,” Advances in Neural Information Processing Systems , vol. 35, pp. 8483–8497, 2022

  69. [73]

    Language as the medium: Multimodal video classification through text only,

    L. Hanu, A. L. Ver ˝o, and J. Thewlis, “Language as the medium: Multimodal video classification through text only,” arXiv preprint arXiv:2309.10783, 2023

  70. [74]

    A video is worth 4096 tokens: Verbalize story videos to understand them in zero shot,

    A. Bhattacharya, Y . K. Singla, B. Krishnamurthy, R. R. Shah, and C. Chen, “A video is worth 4096 tokens: Verbalize story videos to understand them in zero shot,” arXiv preprint arXiv:2305.09758, 2023

  71. [75]

    Retrieving-to-answer: Zero-shot video question answering with frozen large language models,

    J. Pan, Z. Lin, Y . Ge, X. Zhu, R. Zhang, Y . Wang, Y . Qiao, and H. Li, “Retrieving-to-answer: Zero-shot video question answering with frozen large language models,” arXiv preprint arXiv:2306.11732 , 2023

  72. [76]

    Text-only training for image captioning using noise-injected clip,

    D. Nukrai, R. Mokady, and A. Globerson, “Text-only training for image captioning using noise-injected clip,” arXiv preprint arXiv:2211.00575, 2022

  73. [77]

    I can’t believe there’s no images! learning visual tasks using only language data,

    S. Gu, C. Clark, and A. Kembhavi, “I can’t believe there’s no images! learning visual tasks using only language data,” arXiv preprint arXiv:2211.09778, 2022

  74. [78]

    Decap: Decoding clip la- tents for zero-shot captioning via text-only training,

    W. Li, L. Zhu, L. Wen, and Y . Yang, “Decap: Decoding clip la- tents for zero-shot captioning via text-only training,” arXiv preprint arXiv:2303.03032, 2023

  75. [79]

    From association to generation: Text- only captioning by unsupervised cross-modal mapping,

    J. Wang, M. Yan, and Y . Zhang, “From association to generation: Text- only captioning by unsupervised cross-modal mapping,” arXiv preprint arXiv:2304.13273, 2023

  76. [80]

    Zero-shot image captioning by anchor-augmented vision-language space alignment,

    J. Wang, Y . Zhang, M. Yan, J. Zhang, and J. Sang, “Zero-shot image captioning by anchor-augmented vision-language space alignment,” arXiv preprint arXiv:2211.07275 , 2022

  77. [81]

    Transferable decoding with visual entities for zero-shot image captioning,

    J. Fei, T. Wang, J. Zhang, Z. He, C. Wang, and F. Zheng, “Transferable decoding with visual entities for zero-shot image captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3136–3146

  78. [82]

    Language-free training for zero-shot video grounding,

    D. Kim, J. Park, J. Lee, S. Park, and K. Sohn, “Language-free training for zero-shot video grounding,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2539–2548

  79. [83]

    Clip-gen: Language- free training of a text-to-image generator with clip,

    Z. Wang, W. Liu, Q. He, X. Wu, and Z. Yi, “Clip-gen: Language- free training of a text-to-image generator with clip,” arXiv preprint arXiv:2203.00386, 2022

  80. [85]

    Improving cross-modal alignment with synthetic pairs for text-only image captioning,

    Z. Liu, J. Liu, and F. Ma, “Improving cross-modal alignment with synthetic pairs for text-only image captioning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 4, 2024, pp. 3864–3872

  81. [87]

    Towards language-free training for text-to-image generation,

    Y . Zhou, R. Zhang, C. Chen, C. Li, C. Tensmeyer, T. Yu, J. Gu, J. Xu, and T. Sun, “Towards language-free training for text-to-image generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 17 907–17 917

  82. [88]

    See, think, confirm: Interactive prompting between vision and language models for knowledge-based visual reasoning,

    Z. Chen, Q. Zhou, Y . Shen, Y . Hong, H. Zhang, and C. Gan, “See, think, confirm: Interactive prompting between vision and language models for knowledge-based visual reasoning,” arXiv preprint arXiv:2301.05226 , 2023

  83. [89]

    Vicor: Bridging visual un- derstanding and commonsense reasoning with large language models,

    K. Zhou, K. Lee, T. Misu, and X. E. Wang, “Vicor: Bridging visual un- derstanding and commonsense reasoning with large language models,” arXiv preprint arXiv:2310.05872 , 2023

  84. [90]

    Domino: A dual-system for multi-step visual language reasoning,

    P. Wang, O. Golovneva, A. Aghajanyan, X. Ren, M. Chen, A. Celiky- ilmaz, and M. Fazel-Zarandi, “Domino: A dual-system for multi-step visual language reasoning,” arXiv preprint arXiv:2310.02804 , 2023

  85. [91]

    IdealGPT: Iteratively decomposing vision and language reasoning via large language models,

    H. You, R. Sun, Z. Wang, L. Chen, G. Wang, H. Ayyubi, K.-W. Chang, and S.-F. Chang, “IdealGPT: Iteratively decomposing vision and language reasoning via large language models,” in The 2023 Conference on Empirical Methods in Natural Language Processing, 2023. [Online]. Availabl...

  86. [92]

    Good questions help zero-shot image reasoning,

    K. Yang, T. Shen, X. Tian, X. Geng, C. Tao, D. Tao, and T. Zhou, “Good questions help zero-shot image reasoning,” arXiv preprint arXiv:2312.01598, 2023

  87. [93]

    The art of SOCRATIC QUESTIONING: Recursive thinking with large language models,

    J. Qi, Z. Xu, Y . Shen, M. Liu, D. Jin, Q. Wang, and L. Huang, “The art of SOCRATIC QUESTIONING: Recursive thinking with large language models,” in The 2023 Conference on Empirical Methods in Natural Language Processing , 2023. [Online]. Available: https://openreview.net/forum...

  88. [94]

    Filling the image information gap for VQA: Prompting large language models to proactively ask questions,

    Z. Wang, C. Chen, P. Li, and Y . Liu, “Filling the image information gap for VQA: Prompting large language models to proactively ask questions,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Associa...

  89. [95]

    Multimodal multi-hop question answering through a conversation between tools and efficiently finetuned large language models,

    H. Rajabzadeh, S. Wang, H. J. Kwon, and B. Liu, “Multimodal multi-hop question answering through a conversation between tools and efficiently finetuned large language models,” arXiv preprint arXiv:2309.08922, 2023

  90. [96]

    Learn to explain: Multimodal reasoning via thought chains for science question answering,

    P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” Advances in Neural Information Processing Systems , vol. 35, pp. 2507–2521, 2022

  91. [97]

    Multimodal chain-of-thought reasoning in language models,

    Z. Zhang, A. Zhang, M. Li, H. Zhao, G. Karypis, and A. Smola, “Multimodal chain-of-thought reasoning in language models,” arXiv preprint arXiv:2302.00923, 2023

  92. [98]

    T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering,

    L. Wang, Y . Hu, J. He, X. Xu, N. Liu, H. Liu, and H. T. Shen, “T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, 2024, pp...

  93. [99]

    Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning,

    D. Mondal, S. Modi, S. Panda, R. Singh, and G. S. Rao, “Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning,” arXiv preprint arXiv:2401.12863, 2024. JOURNAL OF IEEE, VOL. 14, NO. 8, AUGUST 2021 21

  94. [100]

    Measuring and improving chain-of-thought reasoning in vision-language models,

    Y . Chen, K. Sikka, M. Cogswell, H. Ji, and A. Divakaran, “Measuring and improving chain-of-thought reasoning in vision-language models,” arXiv preprint arXiv:2309.04461 , 2023

  95. [101]

    Efficient end-to-end visual document understanding with rationale distillation,

    W. Zhu, A. Agarwal, M. Joshi, R. Jia, J. Thomason, and K. Toutanova, “Efficient end-to-end visual document understanding with rationale distillation,” arXiv preprint arXiv:2311.09612 , 2023

  96. [102]

    Compositional chain- of-thought prompting for large multimodal models,

    C. Mitra, B. Huang, T. Darrell, and R. Herzig, “Compositional chain- of-thought prompting for large multimodal models,” arXiv preprint arXiv:2311.17076, 2023

  97. [103]

    Cocot: Contrastive chain-of-thought prompting for large multimodal models with multiple image inputs,

    D. Zhang, J. Yang, H. Lyu, Z. Jin, Y . Yao, M. Chen, and J. Luo, “Cocot: Contrastive chain-of-thought prompting for large multimodal models with multiple image inputs,” arXiv preprint arXiv:2401.02582, 2024

  98. [104]

    Chain of images for intuitively reasoning,

    F. Meng, H. Yang, Y . Wang, and M. Zhang, “Chain of images for intuitively reasoning,” arXiv preprint arXiv:2311.09241 , 2023

  99. [105]

    Visual chain of thought: Bridging logical gaps with multimodal infillings,

    D. Rose, V . Himakunthala, A. Ouyang, R. He, A. Mei, Y . Lu, M. Saxon, C. Sonar, D. Mirza, and W. Y . Wang, “Visual chain of thought: Bridging logical gaps with multimodal infillings,” arXiv preprint arXiv:2305.02317, 2023

  100. [106]

    Let’s think frame by frame: Evaluating video chain of thought with video infilling and prediction,

    V . Himakunthala, A. Ouyang, D. Rose, R. He, A. Mei, Y . Lu, C. Sonar, M. Saxon, and W. Y . Wang, “Let’s think frame by frame: Evaluating video chain of thought with video infilling and prediction,” arXiv preprint arXiv:2305.13903, 2023

  101. [107]

    Zero-shot visual relation detection via composite visual cues from large language models,

    L. Li, J. Xiao, G. Chen, J. Shao, Y . Zhuang, and L. Chen, “Zero-shot visual relation detection via composite visual cues from large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024

  102. [108]

    Generating action-conditioned prompts for open- vocabulary video action recognition,

    C. Jia, M. Luo, X. Chang, Z. Dang, M. Han, M. Wang, G. Dai, S. Dang, and J. Wang, “Generating action-conditioned prompts for open- vocabulary video action recognition,”arXiv preprint arXiv:2312.02226, 2023

  103. [109]

    Video- prompter: an ensemble of foundational models for zero-shot video understanding,

    A. Yousaf, M. Naseer, S. Khan, F. S. Khan, and M. Shah, “Video- prompter: an ensemble of foundational models for zero-shot video understanding,” arXiv preprint arXiv:2310.15324 , 2023

  104. [111]

    Aligning bag of regions for open-vocabulary object detection,

    S. Wu, W. Zhang, S. Jin, W. Liu, and C. C. Loy, “Aligning bag of regions for open-vocabulary object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 15 254–15 264

  105. [112]

    Object-aware distillation pyramid for open-vocabulary object detection,

    L. Wang, Y . Liu, P. Du, Z. Ding, Y . Liao, Q. Qi, B. Chen, and S. Liu, “Object-aware distillation pyramid for open-vocabulary object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 11 186–11 196

  106. [113]

    Open vocabulary object detection with pseudo bounding-box labels,

    M. Gao, C. Xing, J. C. Niebles, J. Li, R. Xu, W. Liu, and C. Xiong, “Open vocabulary object detection with pseudo bounding-box labels,” in European Conference on Computer Vision . Springer, 2022, pp. 266–282

  107. [116]

    M-tuning: Prompt tuning with mitigated label bias in open-set scenarios,

    N. Liao, X. Zhang, M. Cao, J. Yan, and Q. Tian, “M-tuning: Prompt tuning with mitigated label bias in open-set scenarios,” arXiv preprint arXiv:2303.05122, 2023

  108. [118]

    Decouple before interact: Multi-modal prompt learning for continual visual question answering,

    Z. Qian, X. Wang, X. Duan, P. Qin, Y . Li, and W. Zhu, “Decouple before interact: Multi-modal prompt learning for continual visual question answering,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2953–2962

  109. [119]

    Coin: A benchmark of continual instruction tuning for multimodel large language model,

    C. Chen, J. Zhu, X. Luo, H. Shen, L. Gao, and J. Song, “Coin: A benchmark of continual instruction tuning for multimodel large language model,” arXiv preprint arXiv:2403.08350 , 2024

  110. [120]

    Beyond anti-forgetting: Multimodal continual instruction tuning with positive forward transfer,

    J. Zheng, Q. Ma, Z. Liu, B. Wu, and H. Feng, “Beyond anti-forgetting: Multimodal continual instruction tuning with positive forward transfer,” arXiv preprint arXiv:2401.09181 , 2024

  111. [121]

    Mm-react: Prompting chatgpt for multimodal reasoning and action,

    Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang, “Mm-react: Prompting chatgpt for multimodal reasoning and action,” arXiv preprint arXiv:2303.11381 , 2023

  112. [122]

    Visual chatgpt: Talking, drawing and editing with visual foundation models,

    C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan, “Visual chatgpt: Talking, drawing and editing with visual foundation models,” arXiv preprint arXiv:2303.04671, 2023

  113. [123]

    Chameleon: Plug-and-play compositional reasoning with large language models,

    P. Lu, B. Peng, H. Cheng, M. Galley, K.-W. Chang, Y . N. Wu, S.-C. Zhu, and J. Gao, “Chameleon: Plug-and-play compositional reasoning with large language models,” ArXiv, vol. abs/2304.09842,

  114. [124]

    Mm-vid: Advancing video understanding with gpt-4v (ision),

    K. Lin, F. Ahmed, L. Li, C.-C. Lin, E. Azarnasab, Z. Yang, J. Wang, L. Liang, Z. Liu, Y . Luet al., “Mm-vid: Advancing video understanding with gpt-4v (ision),” arXiv preprint arXiv:2310.19773 , 2023

  115. [125]

    Visual programming: Compositional visual reasoning without training,

    T. Gupta and A. Kembhavi, “Visual programming: Compositional visual reasoning without training,” 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022

  116. [126]

    Vipergpt: Visual inference via python execution for reasoning,

    D. Sur’is, S. Menon, and C. V ondrick, “Vipergpt: Visual inference via python execution for reasoning,” 2023 IEEE/CVF International Conference on Computer Vision (ICCV) , 2023

  117. [127]

    Zero-shot video question answering with procedural programs,

    R. Choudhury, K. Niinuma, K. M. Kitani, and L. A. Jeni, “Zero-shot video question answering with procedural programs,” ArXiv, vol. abs/2312.00937, 2023. [Online]. Available: https://api.semanticscholar. org/CorpusID:265608945

  118. [128]

    Towards truly zero-shot compositional visual reasoning with llms as programmers,

    A. Stani’c, S. Caelles, and M. Tschannen, “Towards truly zero-shot compositional visual reasoning with llms as programmers,” ArXiv, vol. abs/2401.01974, 2024. [Online]. Available: https://api.semanticscholar. org/CorpusID:266755924

  119. [129]

    Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,

    J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” in Pro- ceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp...

  120. [130]

    Microsoft coco: Common objects in context,

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 . Springer,...

  121. [131]

    Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,

    B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hocken- maier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. ...

  122. [132]

    Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,

    X. Wang, J. Wu, J. Chen, L. Li, Y .-F. Wang, and W. Y . Wang, “Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 4581–4591

  123. [133]

    Msr-vtt: A large video description dataset for bridging video and language,

    J. Xu, T. Mei, T. Yao, and Y . Rui, “Msr-vtt: A large video description dataset for bridging video and language,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 5288–5296

  124. [134]

    A dataset for movie description,

    A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele, “A dataset for movie description,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3202–3212

  125. [135]

    Funqa: Towards surprising video comprehension,

    B. Xie, S. Zhang, Z. Zhou, B. Li, Y . Zhang, J. Hessel, J. Yang, and Z. Liu, “Funqa: Towards surprising video comprehension,” arXiv preprint arXiv:2306.14899, 2023

  126. [136]

    Visual abductive reason- ing,

    C. Liang, W. Wang, T. Zhou, and Y . Yang, “Visual abductive reason- ing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 15 565–15 575

  127. [137]

    Ask, attend and answer: Exploring question- guided spatial attention for visual question answering,

    H. Xu and K. Saenko, “Ask, attend and answer: Exploring question- guided spatial attention for visual question answering,” in Com- puter Vision–ECCV 2016: 14th European Conference, Amsterdam, the Netherlands, October 11–14, 2016, Proceedings, Part VII 14. Springer, 2016, pp. 451–466

  128. [138]

    Dynamic key-value memory enhanced multi- step graph reasoning for knowledge-based visual question answering,

    M. Li and M.-F. Moens, “Dynamic key-value memory enhanced multi- step graph reasoning for knowledge-based visual question answering,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 10, 2022, pp. 10 983–10 992

  129. [139]

    Explore multi-step reason- ing in video question answering,

    X. Song, Y . Shi, X. Chen, and Y . Han, “Explore multi-step reason- ing in video question answering,” in Proceedings of the 26th ACM international conference on Multimedia , 2018, pp. 239–247

  130. [140]

    Knowledge editing for large language models: A survey,

    S. Wang, Y . Zhu, H. Liu, Z. Zheng, C. Chen et al., “Knowledge editing for large language models: A survey,”arXiv preprint arXiv:2310.16218, 2023

  131. [141]

    Noc-rek: novel object captioning with retrieved vocabulary from external knowledge,

    D. M. V o, H. Chen, A. Sugimoto, and H. Nakayama, “Noc-rek: novel object captioning with retrieved vocabulary from external knowledge,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022, pp. 17 979–17 987. JOURNAL OF IEEE, VOL. 14, NO. 8...

  132. [142]

    Straight to the facts: Learning knowledge base retrieval for factual visual question answering,

    M. Narasimhan and A. G. Schwing, “Straight to the facts: Learning knowledge base retrieval for factual visual question answering,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 451–468

  133. [143]

    Ask me anything: Free-form visual question answering based on knowledge from external sources,

    Q. Wu, P. Wang, C. Shen, A. Dick, and A. Van Den Hengel, “Ask me anything: Free-form visual question answering based on knowledge from external sources,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4622–4630

  134. [144]

    Incorporating external knowledge to answer open-domain visual questions with dynamic memory networks,

    G. Li, H. Su, and W. Zhu, “Incorporating external knowledge to answer open-domain visual questions with dynamic memory networks,” arXiv preprint arXiv:1712.00733, 2017

  135. [145]

    Fvqa: Fact-based visual question answering,

    P. Wang, Q. Wu, C. Shen, A. Dick, and A. Van Den Hengel, “Fvqa: Fact-based visual question answering,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 10, pp. 2413–2427, 2017

  136. [146]

    Zero-shot visual question answering using knowledge graph,

    Z. Chen, J. Chen, Y . Geng, J. Z. Pan, Z. Yuan, and H. Chen, “Zero-shot visual question answering using knowledge graph,” in The Semantic Web–ISWC 2021: 20th International Semantic Web Conference, ISWC 2021, Virtual Event, October 24–28, 2021, Proceedings 20 . Springer, 2021, ...

  137. [147]

    Wordnet: a lexical database for english,

    G. A. Miller, “Wordnet: a lexical database for english,” Communica- tions of the ACM , vol. 38, no. 11, pp. 39–41, 1995

  138. [148]

    Wikidata: a free collaborative knowl- edgebase,

    D. Vrande ˇci´c and M. Kr ¨otzsch, “Wikidata: a free collaborative knowl- edgebase,” Communications of the ACM , vol. 57, no. 10, pp. 78–85, 2014

  139. [149]

    Unit: Multimodal multitask learning with a unified transformer,

    R. Hu and A. Singh, “Unit: Multimodal multitask learning with a unified transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2021, pp. 1439–1449

  140. [150]

    Multi-task learning of hierarchical vision-language representation,

    D.-K. Nguyen and T. Okatani, “Multi-task learning of hierarchical vision-language representation,” in Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2019, pp. 10 492– 10 501

  141. [151]

    12-in-1: Multi-task vision and language representation learning,

    J. Lu, V . Goswami, M. Rohrbach, D. Parikh, and S. Lee, “12-in-1: Multi-task vision and language representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, 2020, pp. 10 437–10 446

  142. [152]

    Finetuned language models are zero-shot learners,

    J. Wei, M. Bosma, V . Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V . Le, “Finetuned language models are zero-shot learners,” in International Conference on Learning Representations , 2021

  143. [153]

    Instruction tuning for large language models: A survey,

    S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu et al., “Instruction tuning for large language models: A survey,” arXiv preprint arXiv:2308.10792 , 2023

  144. [154]

    Training language models to follow instructions with human feedback,

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022

  145. [155]

    Deep reinforcement learning from human preferences,

    P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” Advances in neural information processing systems , vol. 30, 2017

  146. [156]

    Align before fuse: Vision and language representation learning with momentum distillation,

    J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” Advances in neural information processing systems, vol. 34, pp. 9694–9705, 2021

  147. [157]

    Modeling caption diversity in contrastive vision-language pretraining,

    S. Lavoie, P. Kirichenko, M. Ibrahim, M. Assran, A. G. Wildon, A. Courville, and N. Ballas, “Modeling caption diversity in contrastive vision-language pretraining,” arXiv preprint arXiv:2405.00740 , 2024

  148. [158]

    Videoclip: Contrastive pre-training for zero-shot video-text understanding,

    H. Xu, G. Ghosh, P.-Y . Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer, “Videoclip: Contrastive pre-training for zero-shot video-text understanding,” arXiv preprint arXiv:2109.14084, 2021

  149. [159]

    Actionclip: A new paradigm for video action recognition,

    M. Wang, J. Xing, and Y . Liu, “Actionclip: A new paradigm for video action recognition,” arXiv preprint arXiv:2109.08472 , 2021

  150. [160]

    Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

    J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,” arXiv preprint arXiv:2301.12597 , 2023

  151. [161]

    High-resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 684–10 695

  152. [162]

    Aligning large multimodal models with factually augmented rlhf,

    Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y . Shen, C. Gan, L.-Y . Gui, Y .-X. Wang, Y . Yanget al. , “Aligning large multimodal models with factually augmented rlhf,” arXiv preprint arXiv:2309.14525 , 2023

  153. [163]

    Using human feedback to fine-tune diffusion models without any reward model,

    K. Yang, J. Tao, J. Lyu, C. Ge, J. Chen, W. Shen, X. Zhu, and X. Li, “Using human feedback to fine-tune diffusion models without any reward model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 8941–8951

  154. [164]

    Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,

    R. Krishna, Y . Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y . Kalantidis, L.-J. Li, D. A. Shammaet al., “Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,” International journal of computer vision , vol. 123, pp. 32–73, 2017

  155. [165]

    Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,

    P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of the 56th Annual Meeting of the Asso- ciation for Computational Linguistics (Volume 1: Long Papers) , 2018, pp....

  156. [166]

    Frozen in time: A joint video and image encoder for end-to-end retrieval,

    M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 1728–1738

  157. [167]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024

  158. [168]

    Videochat: Chat-centric video understanding,

    K. Li, Y . He, Y . Wang, Y . Li, W. Wang, P. Luo, Y . Wang, L. Wang, and Y . Qiao, “Videochat: Chat-centric video understanding,”arXiv preprint arXiv:2305.06355, 2023

  159. [169]

    Mimic-it: Multi-modal in-context instruction tuning,

    B. Li, Y . Zhang, L. Chen, J. Wang, F. Pu, J. Yang, C. Li, and Z. Liu, “Mimic-it: Multi-modal in-context instruction tuning,” arXiv preprint arXiv:2306.05425, 2023

  160. [170]

    Otter: A multi-modal model with in-context instruction tuning,

    B. Li, Y . Zhang, L. Chen, J. Wang, J. Yang, and Z. Liu, “Otter: A multi-modal model with in-context instruction tuning,” arXiv preprint arXiv:2305.03726, 2023

  161. [171]

    Nocaps: Novel object captioning at scale,

    H. Agrawal, K. Desai, Y . Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson, “Nocaps: Novel object captioning at scale,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 8948–8957

  162. [172]

    Ok-vqa: A visual question answering benchmark requiring external knowledge,

    K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, 2019, pp. 3195–3204

  163. [173]

    Vizwiz grand challenge: Answering visual questions from blind people,

    D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham, “Vizwiz grand challenge: Answering visual questions from blind people,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3608–3617

  164. [174]

    Modeling context in referring expressions,

    L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11- 14, 2016, Proceedings, Part II 14 . Springer, 2016, pp. 69–85

  165. [175]

    Seed-bench: Benchmarking multimodal llms with generative comprehension,

    B. Li, R. Wang, G. Wang, Y . Ge, Y . Ge, and Y . Shan, “Seed-bench: Benchmarking multimodal llms with generative comprehension,” arXiv preprint arXiv:2307.16125, 2023

  166. [176]

    Mmbench: Is your multi-modal model an all-around player?

    Y . Liu, H. Duan, Y . Zhang, B. Li, S. Zhang, W. Zhao, Y . Yuan, J. Wang, C. He, Z. Liu et al., “Mmbench: Is your multi-modal model an all-around player?” in European Conference on Computer Vision . Springer, 2025, pp. 216–233

  167. [177]

    Teaching clip to count to ten,

    R. Paiss, A. Ephrat, O. Tov, S. Zada, I. Mosseri, M. Irani, and T. Dekel, “Teaching clip to count to ten,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3170–3180

  168. [178]

    Eyes wide shut? exploring the visual shortcomings of multimodal llms,

    S. Tong, Z. Liu, Y . Zhai, Y . Ma, Y . LeCun, and S. Xie, “Eyes wide shut? exploring the visual shortcomings of multimodal llms,” arXiv preprint arXiv:2401.06209, 2024

  169. [179]

    Do vision-language models understand compound nouns?

    S. Kumar, S. Ghosh, S. Sakshi, U. Tyagi, and D. Manocha, “Do vision-language models understand compound nouns?” arXiv preprint arXiv:2404.00419, 2024

  170. [180]

    cola: A benchmark for compositional text-to-image re- trieval,

    A. Ray, F. Radenovic, A. Dubey, B. Plummer, R. Krishna, and K. Saenko, “cola: A benchmark for compositional text-to-image re- trieval,” Advances in Neural Information Processing Systems , vol. 36, 2024

  171. [181]

    Crepe: Can vision-language foundation models reason compositionally?

    Z. Ma, J. Hong, M. O. Gul, M. Gandhi, I. Gao, and R. Krishna, “Crepe: Can vision-language foundation models reason compositionally?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 10 910–10 921

  172. [182]

    When and why vision-language models behave like bags-of-words, and what to do about it?

    M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou, “When and why vision-language models behave like bags-of-words, and what to do about it?” in The Eleventh International Conference on Learning Representations, 2022

  173. [183]

    Sugar- crepe: Fixing hackable benchmarks for vision-language compositional- ity,

    C.-Y . Hsieh, J. Zhang, Z. Ma, A. Kembhavi, and R. Krishna, “Sugar- crepe: Fixing hackable benchmarks for vision-language compositional- ity,” Advances in neural information processing systems, vol. 36, 2024

  174. [184]

    Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models,

    T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y . Yacoob et al. , “Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models,” in Proceedings of the IEEE/CVF Conference on...

  175. [185]

    Muffin or chihuahua? challenging multimodal large language models with multipanel vqa,

    Y . Fan, J. Gu, K. Zhou, Q. Yan, S. Jiang, C.-C. Kuo, Y . Zhao, X. Guan, and X. Wang, “Muffin or chihuahua? challenging multimodal large language models with multipanel vqa,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: ...

  176. [186]

    Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge,

    A. Wang, B. Wu, S. Chen, Z. Chen, H. Guan, W.-N. Lee, L. E. Li, and C. Gan, “Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 384–13 394

  177. [187]

    What if the tv was off? examining counterfactual reasoning abilities of multi- modal language models,

    L. Zhang, X. Zhai, Z. Zhao, Y . Zong, X. Wen, and B. Zhao, “What if the tv was off? examining counterfactual reasoning abilities of multi- modal language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 21 853–21 862

  178. [188]

    Language models are unsupervised multitask learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019

  179. [189]

    Parallel refinements for lexically constrained text generation with bart,

    X. He, “Parallel refinements for lexically constrained text generation with bart,” arXiv preprint arXiv:2109.12487 , 2021

  180. [190]

    Exploring the limits of transfer learn- ing with a unified text-to-text transformer,

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learn- ing with a unified text-to-text transformer,” The Journal of Machine Learning Research, vol. 21, no. 1, pp. 5485–5551, 2020

  181. [191]

    A style-based generator architecture for generative adversarial networks,

    T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 4401–4410

  182. [192]

    Clipcap: Clip prefix for image captioning,

    R. Mokady, A. Hertz, and A. H. Bermano, “Clipcap: Clip prefix for image captioning,” arXiv preprint arXiv:2111.09734 , 2021

  183. [193]

    Taming transformers for high-resolution image synthesis,

    P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 12 873–12 883

  184. [194]

    Recurrent relational memory network for unsupervised image captioning,

    D. Guo, Y . Wang, P. Song, and M. Wang, “Recurrent relational memory network for unsupervised image captioning,” arXiv preprint arXiv:2006.13611, 2020

  185. [195]

    Relational distant supervision for image captioning without image-text pairs,

    Y . Qi, W. Zhao, and X. Wu, “Relational distant supervision for image captioning without image-text pairs,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 5, 2024, pp. 4524– 4532

  186. [196]

    Bleu: a method for automatic evaluation of machine translation,

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics , 2002, pp. 311–318

  187. [197]

    Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,

    S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,” in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , 2005, pp. 65–72

  188. [198]

    Cider: Consensus- based image description evaluation,

    R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus- based image description evaluation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 4566–4575

  189. [199]

    Spice: Semantic propositional image caption evaluation,

    P. Anderson, B. Fernando, M. Johnson, and S. Gould, “Spice: Semantic propositional image caption evaluation,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, Octo- ber 11-14, 2016, Proceedings, Part V 14. Springer, 2016, pp. 382–398

  190. [200]

    Chain-of-thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhou et al., “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022

  191. [201]

    A-okvqa: A benchmark for visual question answering using world knowledge,

    D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-okvqa: A benchmark for visual question answering using world knowledge,” in European Conference on Computer Vision . Springer, 2022, pp. 146–162

  192. [202]

    Lvis: A dataset for large vocabulary instance segmentation,

    A. Gupta, P. Dollar, and R. Girshick, “Lvis: A dataset for large vocabulary instance segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 5356–5364

  193. [204]

    Commonsense knowledge prompt- ing for few-shot action recognition in videos,

    Y . Shi, X. Wu, H. Lin, and J. Luo, “Commonsense knowledge prompt- ing for few-shot action recognition in videos,” IEEE Transactions on Multimedia, 2024

  194. [205]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255

  195. [206]

    The caltech-ucsd birds-200-2011 dataset,

    C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie, “The caltech-ucsd birds-200-2011 dataset,” 2011

  196. [207]

    Food-101–mining discriminative components with random forests,

    L. Bossard, M. Guillaumin, and L. Van Gool, “Food-101–mining discriminative components with random forests,” in Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, Septem- ber 6-12, 2014, Proceedings, Part VI 13. Springer, 2014, pp. 446–461

  197. [208]

    Places: A 10 million image database for scene recognition,

    B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba, “Places: A 10 million image database for scene recognition,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 6, pp. 1452– 1464, 2017

  198. [209]

    Cats and dogs,

    O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar, “Cats and dogs,” in 2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012, pp. 3498–3505

  199. [210]

    De- scribing textures in the wild,

    M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi, “De- scribing textures in the wild,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2014, pp. 3606–3613

  200. [211]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020

  201. [212]

    Align before fuse: Vision and language representation learning with momentum distillation,

    J. Li, R. R. Selvaraju, A. D. Gotmare, S. Joty, C. Xiong, and S. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in NeurIPS, 2021

  202. [213]

    Zero-shot object detection,

    A. Bansal, K. Sikka, G. Sharma, R. Chellappa, and A. Divakaran, “Zero-shot object detection,” in Proceedings of the European confer- ence on computer vision (ECCV) , 2018, pp. 384–400

  203. [214]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778

  204. [215]

    Evaluating large language models trained on code,

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockmanet al., “Evaluating large language models trained on code,” arXiv preprint arXiv:2107.03374 , 2021

  205. [216]

    Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models,

    F. Liu, T. Guan, Z. Li, L. Chen, Y . Yacoob, D. Manocha, and T. Zhou, “Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models,” arXiv preprint arXiv:2310...

  206. [217]

    Plausible may not be faithful: Probing object hallucination in vision-language pre-training,

    W. Dai, Z. Liu, Z. Ji, D. Su, and P. Fung, “Plausible may not be faithful: Probing object hallucination in vision-language pre-training,” arXiv preprint arXiv:2210.07688 , 2022

  207. [218]

    Siren’s song in the ai ocean: a survey on hal- lucination in large language models,

    Y . Zhang, Y . Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y . Zhang, Y . Chenet al., “Siren’s song in the ai ocean: a survey on hal- lucination in large language models,” arXiv preprint arXiv:2309.01219, 2023

  208. [219]

    How language model hallucinations can snowball,

    M. Zhang, O. Press, W. Merrill, A. Liu, and N. A. Smith, “How language model hallucinations can snowball,” arXiv preprint arXiv:2305.13534, 2023

  209. [220]

    Con- tinual learning for large language models: A survey,

    T. Wu, L. Luo, Y .-F. Li, S. Pan, T.-T. Vu, and G. Haffari, “Con- tinual learning for large language models: A survey,” arXiv preprint arXiv:2402.01364, 2024

  210. [221]

    Visrag: Vision-based retrieval-augmented generation on multi-modality documents,

    S. Yu, C. Tang, B. Xu, J. Cui, J. Ran, Y . Yan, Z. Liu, S. Wang, X. Han, Z. Liu et al., “Visrag: Vision-based retrieval-augmented generation on multi-modality documents,” arXiv preprint arXiv:2410.10594 , 2024

  211. [222]

    Unirag: Universal retrieval augmentation for multi-modal large language mod- els,

    S. Sharifymoghaddam, S. Upadhyay, W. Chen, and J. Lin, “Unirag: Universal retrieval augmentation for multi-modal large language mod- els,” arXiv preprint arXiv:2405.10311 , 2024

  212. [223]

    Wiki-llava: Hierarchical retrieval-augmented gen- eration for multimodal llms,

    D. Caffagni, F. Cocchi, N. Moratelli, S. Sarto, M. Cornia, L. Baraldi, and R. Cucchiara, “Wiki-llava: Hierarchical retrieval-augmented gen- eration for multimodal llms,” in Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition , 2024, pp. 1818– 1826

  213. [224]

    Snapntell: Enhanc- ing entity-centric visual question answering with retrieval augmented multimodal llm,

    J. Qiu, A. Madotto, Z. Lin, P. A. Crook, Y . E. Xu, X. L. Dong, C. Faloutsos, L. Li, B. Damavandi, and S. Moon, “Snapntell: Enhanc- ing entity-centric visual question answering with retrieval augmented multimodal llm,” arXiv preprint arXiv:2403.04735 , 2024

  214. [225]

    Winoground: Probing vision and language models for visio-linguistic compositionality,

    T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross, “Winoground: Probing vision and language models for visio-linguistic compositionality,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 5238–5248

  215. [226]

    When are lemons purple? the concept association bias of vision-language models,

    Y . Tang, Y . Yamada, Y . Zhang, and I. Yildirim, “When are lemons purple? the concept association bias of vision-language models,” in JOURNAL OF IEEE, VOL. 14, NO. 8, AUGUST 2021 24 Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, ...

  216. [227]

    No token left behind: Explainability- aided image classification and generation,

    R. Paiss, H. Chefer, and L. Wolf, “No token left behind: Explainability- aided image classification and generation,” in European Conference on Computer Vision. Springer, 2022, pp. 334–350

  217. [228]

    Valse: A task-independent benchmark for vision and language models centered on linguistic phenomena,

    L. Parcalabescu, M. Cafagna, L. Muradjan, A. Frank, I. Calixto, and A. Gatt, “Valse: A task-independent benchmark for vision and language models centered on linguistic phenomena,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volum...

  218. [2023]

    Available: https://api.semanticscholar.org/CorpusID: 258212542

    [Online]. Available: https://api.semanticscholar.org/CorpusID: 258212542

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.