Pith. sign in

REVIEW 2 major objections 3 minor 158 cited by

A Survey on Multimodal Large Language Models

T0 review · 2 major / 3 minor · reviewed 2026-05-16 · grok-4.3

Pith's one-line read Multimodal large language models use LLMs as a central brain to handle images and other inputs with new emergent reasoning skills.

desk verdict This is a useful organizing survey on multimodal LLMs that maps the literature clearly but introduces no new methods or results. read the letter →

arxiv 2306.13549 v4 pith:ADBOX4AN submitted 2023-06-23 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords multimodallargelanguagemodelsMLLMGPT-4Vemergentcapabilitiesreasoningvision-languagehallucinationartificialgeneralintelligence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey reviews the fast rise of multimodal large language models that combine large language models with visual and other data sources. It covers their basic structure, training on mixed datasets, and evaluation on tasks such as image description and reasoning. The work examines extensions to finer details, more data types, languages, and real-world uses, plus problems like false outputs from images. It closes by listing current limits and open research directions in a field that may lead toward broader artificial intelligence systems.

What carries the argument

The central object is the large language model used as a unifying brain to process and reason over combined multimodal inputs through shared architectures and joint training.

What would settle it

A new review identifying many important recent MLLM papers or key developments absent from this survey and its linked repository would show the summary is incomplete.

Watch

Extended reading notes

Core claim

The paper claims that multimodal large language models, represented by GPT-4V, use powerful large language models as a brain to perform multimodal tasks and display surprising emergent capabilities such as writing stories based on images and OCR-free math reasoning that are rare in traditional multimodal methods, while summarizing their formulation, architecture, training strategy, data, evaluation, extensions to more granularity modalities languages and scenarios, multimodal hallucination, extended techniques including M-ICL M-CoT and LAVR, challenges, and promising directions.

Load-bearing premise

The survey assumes that the cited literature and the associated GitHub repository together provide a sufficiently complete and up-to-date picture of the rapidly evolving MLLM field.

Editorial extensions

If this is right

  • MLLMs can be extended to support finer granularity, additional modalities, more languages, and complex scenarios.
  • Techniques such as multimodal in-context learning, multimodal chain-of-thought reasoning, and LLM-aided visual reasoning improve performance on multimodal tasks.
  • Tackling multimodal hallucination is required for dependable real-world applications.
  • Continued progress in this area may open a route toward artificial general intelligence.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Unified LLM-centered models may replace earlier separate-modality approaches in many vision-language settings.
  • Adding real-time video or audio streams could test whether current emergent skills scale to continuous inputs.
  • The linked repository underscores the value of living resources for tracking fast-changing research areas.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper is a survey tracing recent progress on Multimodal Large Language Models (MLLMs). It begins with the basic formulation and related concepts of architecture, training strategy, data, and evaluation. It then covers extensions supporting greater granularity, additional modalities, languages, and scenarios, followed by multimodal hallucination and techniques including Multimodal In-Context Learning (M-ICL), Multimodal Chain-of-Thought (M-CoT), and LLM-Aided Visual Reasoning (LAVR). The survey concludes with challenges, promising directions, and an associated GitHub repository for updates.

Significance. If the coverage proves comprehensive, the survey supplies a useful organizational framework for the fast-moving MLLM field, explicitly crediting emergent capabilities such as image-based story writing and OCR-free math reasoning while pointing to an open GitHub repository that collects latest papers. This combination of structured delineation and a living resource strengthens its value as a reference for researchers working on vision-language integration.

major comments (2)
  1. [Evaluation] The evaluation section does not quantify how well current benchmarks capture the emergent capabilities highlighted in the abstract (e.g., story writing from images); without such analysis the contrast with traditional multimodal methods remains qualitative and weakens the motivation for the survey's scope.
  2. [Training and Data] In the training and data section, the discussion of data curation omits explicit comparison of scale, filtering, and alignment procedures across representative models (LLaVA, MiniGPT-4, etc.), which is load-bearing for readers seeking to reproduce or extend the reported performance trends.
minor comments (3)
  1. [Abstract] The abstract repeats motivational phrasing about AGI that could be shortened without loss of clarity.
  2. [Architecture] Figure captions for architecture diagrams should explicitly label each component (vision encoder, projector, LLM backbone) to match the textual description.
  3. [Introduction] The GitHub repository is mentioned only in the abstract; a short dedicated paragraph in the introduction describing its maintenance policy and coverage criteria would improve usability.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the encouraging assessment and the specific comments, which help clarify areas where the survey can be strengthened. We address each major comment below and outline the corresponding revisions.

read point-by-point responses
  1. Referee: [Evaluation] The evaluation section does not quantify how well current benchmarks capture the emergent capabilities highlighted in the abstract (e.g., story writing from images); without such analysis the contrast with traditional multimodal methods remains qualitative and weakens the motivation for the survey's scope.

    Authors: We acknowledge that the evaluation section primarily summarizes existing benchmarks and notes emergent capabilities without providing quantitative metrics on benchmark coverage. As this is a survey, we do not introduce new empirical evaluations; however, we will expand the section with a dedicated paragraph discussing the limitations of current benchmarks in capturing capabilities such as image-based story writing and OCR-free reasoning, referencing any available meta-analyses or studies that quantify these gaps. This addition will make the contrast with traditional methods more explicit while remaining within the survey's scope. revision: partial

  2. Referee: [Training and Data] In the training and data section, the discussion of data curation omits explicit comparison of scale, filtering, and alignment procedures across representative models (LLaVA, MiniGPT-4, etc.), which is load-bearing for readers seeking to reproduce or extend the reported performance trends.

    Authors: We agree that a side-by-side comparison would improve utility for readers. We will insert a new table in the training and data section that explicitly compares data scale, filtering strategies, and alignment procedures for representative models including LLaVA, MiniGPT-4, and others, based on details reported in their original papers. This table will directly address reproducibility needs. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; descriptive survey of external literature

full rationale

This paper is a literature survey with no original derivations, equations, quantitative predictions, or first-principles results. Its contribution is organizational: delineating architectures, training strategies, data, evaluations, extensions, hallucination, and techniques like M-ICL and M-CoT drawn from cited external works. The abstract's reference to emergent capabilities is presented as motivation from prior examples rather than a derived claim. No self-citations function as load-bearing justifications for novel results, and no steps reduce to fitted inputs or self-definitions by construction.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

As a survey the paper introduces no free parameters, axioms, or invented entities; all technical content is drawn from the referenced prior literature.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Multimodal Large Language Models." pith.science (2026). https://pith.science/paper/ADBOX4AN

@misc{pith2026230613549,
  author       = {Pith},
  title        = {Pith review of: A Survey on Multimodal Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ADBOX4AN}},
  note         = {Machine review of arXiv:2306.13549}
}
read the original abstract

Recently, Multimodal Large Language Model (MLLM) represented by GPT-4V has been a new rising research hotspot, which uses powerful Large Language Models (LLMs) as a brain to perform multimodal tasks. The surprising emergent capabilities of MLLM, such as writing stories based on images and OCR-free math reasoning, are rare in traditional multimodal methods, suggesting a potential path to artificial general intelligence. To this end, both academia and industry have endeavored to develop MLLMs that can compete with or even better than GPT-4V, pushing the limit of research at a surprising speed. In this paper, we aim to trace and summarize the recent progress of MLLMs. First of all, we present the basic formulation of MLLM and delineate its related concepts, including architecture, training strategy and data, as well as evaluation. Then, we introduce research topics about how MLLMs can be extended to support more granularity, modalities, languages, and scenarios. We continue with multimodal hallucination and extended techniques, including Multimodal ICL (M-ICL), Multimodal CoT (M-CoT), and LLM-Aided Visual Reasoning (LAVR). To conclude the paper, we discuss existing challenges and point out promising research directions. In light of the fact that the era of MLLM has only just begun, we will keep updating this survey and hope it can inspire more research. An associated GitHub link collecting the latest papers is available at https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 158 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 158 Pith citations

  1. MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

    cs.CV 2024-08 conditional novelty 8.0 of 10

    MME-RealWorld is the largest manually annotated high-resolution benchmark for MLLMs, where even the best models achieve less than 60% accuracy on challenging real-world tasks.

  2. Seeing is Coding: On the Effectiveness of Vision Language Models in Code Understanding

    cs.CL 2026-02 unverdicted novelty 7.0 of 10

    Multimodal LLMs process code as images to achieve up to 8x token compression, with visual cues like syntax highlighting aiding tasks and clone detection remaining resilient or even improving under compression.

  3. Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection

    cs.AI 2025-12 unverdicted novelty 7.0 of 10

    ForenAgent lets MLLMs create and iteratively improve low-level Python tools for image forgery detection via a two-stage training pipeline and a new 100k-image benchmark dataset.

  4. Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration

    cs.CV 2025-09 conditional novelty 7.0 of 10

    Even the strongest tested vision-language model, GPT-4o, answers only about a third of the new AVI-Math aerial-imagery math questions correctly.

  5. Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions

    cs.CL 2024-05 unverdicted novelty 7.0 of 10

    Introduces YesBut benchmark showing state-of-the-art multimodal models lag humans on interpreting humorous contradictions in comics.

  6. HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

    cs.CV 2023-10 unverdicted novelty 7.0 of 10

    HallusionBench shows GPT-4V reaches only 31.42% accuracy on paired questions testing language hallucination and visual illusion in LVLMs, with other models below 16%.

  7. Time-LLM: Time Series Forecasting by Reprogramming Large Language Models

    cs.LG 2023-10 conditional novelty 7.0 of 10

    Time-LLM reprograms frozen LLMs for time series forecasting via text prototypes and Prompt-as-Prefix, outperforming specialized models in standard, few-shot, and zero-shot settings.

  8. DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

    cs.CV 2026-02 conditional novelty 6.5 of 10

    The paper reframes image retrieval as agentic exploration over personal visual histories and shows the best tested multimodal agent scores only 28.7 exact match on its new DISBench benchmark.

  9. Uncertainty-Aware Decision Making in Multimodal Large Language Models

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A survey organizing multimodal LLM uncertainty research around a source-signal-calibration-action framework, arguing that uncertainty is only useful when it changes what the system does.

  10. When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A null-image diagnostic shows VLM attribute hallucinations track the visual logit margin, not the language prior, and a routed Calib/Abstain/Adapt framework reduces them.

  11. Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.

  12. Agentic Re-Casting using Agentic Re-Simulations

    hep-ph 2026-07 conditional novelty 6.0 of 10

    An agentic AI system with a physicist in the loop re-casts an ATLAS ttZ measurement into a global top-quark SMEFT fit and recovers injected coloron Wilson coefficients in a repeatable benchmark.

  13. BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

    cs.CV 2026-07 conditional novelty 6.0 of 10

    BVS combines early-stop attention rollout priors with a scale-aware non-stationary kernel and GP-UCB to locate tiny objects in UHR images more accurately and with fewer MLLM queries than prior visual-search methods.

  14. Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    OPPO is an evidence-aware preference optimization that contrasts faithful responses under varying visual evidence strengths to reduce hallucinations in MLLMs.

  15. StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    StoryVideoQA provides the largest auto-generated deep video understanding dataset to date with 363K QAs across TV and movies, paired with the PlotTree agent for hierarchical plot-based reasoning that existing VideoQA ...

  16. Investigating Adversarial Robustness of Multi-modal Large Language Models

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Robust vision encoders from multimodal adversarial pretraining transfer to MLLMs and deliver large gains in adversarial captioning and VQA performance, while test-time stochastic transformations provide an effective b...

  17. Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    LTS-FS locates hallucination-relevant layers in LVLMs via causal attribution on a constructed dataset and applies sparse layerwise feature steering to mitigate hallucinations while preserving general task performance.

  18. Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems

    cs.IR 2026-01 unverdicted novelty 6.0 of 10

    W-RAC decouples extraction from semantic planning via structured units and LLM grouping to match traditional retrieval performance at roughly 10x lower LLM token cost.

  19. Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning

    cs.AI 2025-11 conditional novelty 6.0 of 10

    An MLLM unlearning method and benchmark that aim to erase targeted private facts while preserving image understanding.

  20. MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    MESH, a three-layer video hallucination benchmark, shows LVMs ace basic objects and coarse traits but slip badly on fine character details and multi-subject actions in longer clips.

  21. MoPEQ: Mixture of Mixed Precision Quantized Experts

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Assigning 2, 3, or 4 bits to MoE experts by Hessian trace sensitivity keeps VLM accuracy close to uniform 4-bit while reducing model size.

  22. Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new benchmark shows VLM projection layers retain most of their alignment accuracy on object classes never seen during alignment training, with mechanistic evidence pointing to FFN key-value memory.

  23. Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring

    cs.NI 2025-08 unverdicted novelty 6.0 of 10

    DUAL-Health is an uncertainty-aware multimodal fusion framework that quantifies sensor noise, customizes fusion weights accordingly, and aligns modality distributions to improve outdoor health monitoring.

  24. Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Multi-TW is the first Traditional Chinese benchmark to evaluate multimodal models on both image-text and audio-text questions while also measuring inference latency.

  25. Region-based Cluster Discrimination for Visual Representation Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RICE improves vision encoders by applying cluster discrimination at the region level and unifying object and OCR classification targets in one pretraining framework.

  26. The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents

    cs.CR 2025-07 unverdicted novelty 6.0 of 10

    A multi-agent audio-language model framework can automatically profile private attributes, such as age, health, and income, directly from general audio recordings.

  27. Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction

    eess.AS 2025-07 conditional novelty 6.0 of 10

    A causal audio language model with continuous-valued tokens and masked next-token prediction matches diffusion-based text-to-audio quality with smaller, streamable models.

  28. How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

    cs.CV 2025-07 unverdicted novelty 6.0 of 10

    Multimodal foundation models achieve respectable but sub-specialist performance on semantic vision tasks and weaker results on geometric tasks when evaluated through prompt chaining on established benchmarks.

  29. UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A fine-tuned small multimodal LLM outperforms much larger general models on urban tasks in a new benchmark, with caveats about benchmark overlap with training data.

  30. Object-aware Sound Source Localization via Audio-Visual Scene Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A sound localization model trained with MLLM-generated foreground and background captions and two new losses better distinguishes sounding objects from visually similar silent ones.

  31. Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    CAMO hides harmful instructions across text and image using masked keywords and math-puzzle clues, making several LVLMs answer banned queries while evading common safety filters.

  32. CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 9,356-pair Chinese multimodal financial benchmark reveals that state-of-the-art multimodal LLMs, including GPT-4V, still score below 53% on objective and 39% on subjective financial chart tasks.

  33. MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference

    cs.LG 2025-06 conditional novelty 6.0 of 10

    MadaKV adaptively splits the KV cache budget by attention-head modality preference and compensates across layers, cutting cache memory by 80-95% and speeding decoding by 1.3-1.5x with small accuracy loss.

  34. CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An inference-time attention-shift intervention aligns non-English queries' cross-modal attention with English, cutting multilingual object hallucination in LVLMs on POPE and MME.

  35. Generative Next POI Recommendation with Semantic ID

    cs.IR 2025-06 conditional novelty 6.0 of 10

    GNPR-SID assigns points of interest hierarchical semantic codes via a residual quantized VAE and fine-tunes an LLM to predict the next code, improving next-POI accuracy on three benchmarks.

  36. Vision-Based Assistive Technologies for People with Cerebral Visual Impairment: A Review and Focus Study

    cs.HC 2025-05 accept novelty 6.0 of 10

    A scoping review and focus groups show that vision-based assistive technology has largely ignored cerebral visual impairment, and identify seven challenges and device opportunities for this group.

  37. ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ID-Align improves high-resolution VLM performance by reusing thumbnail position IDs for high-resolution image tokens, yielding small but positive gains on several benchmarks.

  38. Backdoor Cleaning without External Guidance in MLLM Fine-tuning

    cs.CR 2025-05 conditional novelty 6.0 of 10

    BYE filters backdoored training images from MLLM fine-tuning by clustering low attention entropy across selected layers.

  39. Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

    cs.CV 2025-05 unverdicted novelty 6.0 of 10

    Chain-of-Focus enables VLMs to adaptively search and zoom on important image areas via a two-stage SFT and RL pipeline on a custom 3K-sample dataset, yielding 5% gains on the V* benchmark across resolutions from 224 to 4K.

  40. FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    FEALLM is a multimodal LLM fine-tuned on a new, aligned facial expression and action unit reasoning dataset, reporting improved facial emotion analysis on its benchmark and zero-shot gains on RAF-DB, AffectNet, BP4D, ...

  41. Assessing the Performance of Analog Training for Transfer Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    The c-TTv2 analog algorithm fine-tuned a Swin-ViT transformer on CIFAR-100 subsets in simulation, landed within about 2% of digital transfer learning, and tolerated up to 10-15% weight-transfer noise.

  42. SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data

    cs.CV 2025-04 conditional novelty 6.0 of 10

    SpaRE builds a 3.4M-question synthetic spatial QA dataset from captioning sources and reports large gains on VSR and What's Up after fine-tuning Qwen2-VL models.

  43. Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

    cs.CL 2025-04 conditional novelty 6.0 of 10

    MMLA combines 61K multimodal utterances across six semantic dimensions; even the best fine-tuned multimodal LLM reaches only about 69% accuracy, exposing current limits in cognitive-level language understanding.

  44. Media Content Atlas: A Pipeline to Explore and Investigate Multidimensional Media Space using Multimodal LLMs

    cs.HC 2025-04 conditional novelty 6.0 of 10

    A pipeline combining CLIP, LLaVA, and BERTopic clusters and retrieves content from 1.12 million smartphone screenshots, with expert-rated topic relevance of 96%.

  45. DMM: Building a Versatile Image Generation Model via Distillation-Based Model Merging

    cs.CV 2025-04 conditional novelty 6.0 of 10

    DMM trains a single style-promptable diffusion model to reproduce the outputs of multiple teacher models, achieving a merged-model FIDt of 77.51 versus a reference of 74.91.

  46. When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?

    cs.CV 2025-03 unverdicted novelty 6.0 of 10

    Presents YesBut (V2) benchmark and shows state-of-the-art VLMs significantly underperform humans on tasks requiring comparative reasoning for contradictory humor in comics.

  47. FixDrive: Automatically Repairing Autonomous Vehicle Driving Behaviour for $0.08 per Violation

    cs.SE 2025-02 conditional novelty 6.0 of 10

    FixDrive automatically synthesizes µDrive driving-strategy repairs from near-miss and violation moments in AV logs, using GPT-4 Turbo and STL-based localization, with per-violation costs around $0.08.

  48. UniCoRN: Unified Commented Retrieval Network with LMMs

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A frozen multimodal LLM is extended with a retrieval adapter and an entity adapter to retrieve a relevant image and generate a supportive textual comment.

  49. E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection

    cs.LG 2025-02 conditional novelty 6.0 of 10

    E2LVLM improves out-of-context misinformation detection by having a vision-language model rerank and rewrite retrieved evidence, then fine-tune on LVLM-generated explanations.

  50. HCMRM: A High-Consistency Multimodal Relevance Model for Search Ads

    cs.IR 2025-02 conditional novelty 6.0 of 10

    Adding pseudo-query pre-training and a hierarchical softmax loss to an ALBEF-style multimodal model improves query-video relevance for short video search ads.

  51. Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Using large-scale adversarially pretrained vision encoders in LLaVA yields 2x and 1.5x robustness gains on captioning and VQA, and cuts jailbreak success rates by over 10% relative to CLIP fine-tuning baselines.

  52. MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models

    cs.AI 2025-02 conditional novelty 6.0 of 10

    The MM-IQ benchmark shows state-of-the-art multimodal models score 33% on abstract visual reasoning puzzles versus 25% chance and 51% for humans.

  53. AiGet: Transforming Everyday Moments into Hidden Knowledge Discovery with AI Assistance on Smart Glasses

    cs.HC 2025-01 conditional novelty 6.0 of 10

    A proactive smart-glasses AI that infers learning desires from gaze and context can deliver personalized, low-disruption knowledge during everyday activities.

  54. AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    AKVQ-VL quantizes VLM KV caches to mostly 2 bits with attention-aware token protection and Walsh-Hadamard outlier removal, staying near FP16 accuracy on MileBench.

  55. Multi-aspect Knowledge Distillation with Large Language Model

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A multi-aspect knowledge distillation method that appends binary question-answer logits from an MLLM to a classifier's output improves fine-grained image classification accuracy by up to about 6 points.

  56. T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    T2ISafety is a large annotated benchmark plus a fine-tuned MLLM evaluator (ImageGuard) for measuring toxicity, privacy, and fairness in text-to-image models.

  57. Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Eagle2-9B matches or outperforms much larger vision-language models on many benchmarks through a carefully constructed post-training data strategy.

  58. Verifying Cross-modal Entity Consistency in News using Vision-language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Vision-language models with web-sourced evidence images can verify cross-modal entity consistency in news, outperforming a CNN baseline for locations and events in a zero-shot setting.

  59. AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A 600K-question audio-visual trustworthiness benchmark reveals that current AVLLMs are brittle under mismatches and missing modalities, and CAVPref, a calibrated preference-optimization method, improves their accuracy.

  60. IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    IllusionBench+ is a 1,051-image benchmark with 5,548 QA pairs and golden descriptions for testing vision-language models on classic, real-scene, Ishihara, and trap visual illusions, and it finds top models such as GPT...

See all 158 Pith citations

Reference graph

Works this paper leans on

209 extracted references · 209 canonical work pages · cited by 158 Pith papers (see all)

  1. [1]

    Keyphrase Extraction Using Deep Recurrent Neural Networks on Twitter

    W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong et al. , “A survey of large language models,” arXiv:2303.18223, 2023. 1

  2. [2]

    Chatgpt: A language model for conversational ai,

    OpenAI, “Chatgpt: A language model for conversational ai,” OpenAI, Tech. Rep., 2023. [Online]. Available: https: //www.openai.com/research/chatgpt 1, 6

  3. [3]
  4. [4]

    Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality,

    W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez et al. , “Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality,”

  5. [5]

    Available: https://vicuna.lmsys.org 1, 3, 4

    [Online]. Available: https://vicuna.lmsys.org 1, 3, 4

  6. [6]

    LLaMA: Open and Efficient Foundation Language Models

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al. , “Llama: Open and efficient foundation language models,” arXiv:2302.13971, 2023. 1, 3, 4

  7. [7]

    Instruction Tuning with GPT-4

    B. Peng, C. Li, P . He, M. Galley, and J. Gao, “Instruction tuning with gpt-4,” arXiv:2304.03277, 2023. 1

  8. [8]

    Lan- guage models are few-shot learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P . Dhari- wal, A. Neelakantan, P . Shyam, G. Sastry, A. Askell et al. , “Lan- guage models are few-shot learners,” NeurIPS, 2020. 1, 3, 6

Show all 209 references
  1. [9]

    Chain of thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou, “Chain of thought prompting elicits reasoning in large language models,” arXiv:2201.11903, 2022. 1, 12

  2. [10]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” arXiv:2304.02643, 2023. 1, 9

  3. [11]

    Aligning and prompting everything all at once for universal visual perception,

    Y. Shen, C. Fu, P . Chen, M. Zhang, K. Li, X. Sun, Y. Wu, S. Lin, and R. Ji, “Aligning and prompting everything all at once for universal visual perception,” in CVPR, 2024. 1

  4. [12]

    Dino: Detr with improved denoising anchor boxes for end-to-end object detection,

    H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. M. Ni, and H.-Y. Shum, “Dino: Detr with improved denoising anchor boxes for end-to-end object detection,” arXiv:2203.03605, 2022. 1

  5. [13]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V . Khalidov, P . Fernandez, D. Haziza, F. Massa, A. El-Noubyet al., “Dinov2: Learning robust visual features without supervision,” arXiv:2304.07193, 2023. 1

  6. [14]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agar- wal, G. Sastry, A. Askell, P . Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML, 2021. 1, 2, 3, 5

  7. [15]

    Align before fuse: Vision and language representation learning with momentum distillation,

    J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” NeurIPS, 2021. 1

  8. [16]

    Uniter: Universal image-text representation learn- ing,

    Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “Uniter: Universal image-text representation learn- ing,” in ECCV, 2020. 1

  9. [17]

    Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,

    P . Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang, “Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” in ICML, 2022. 1

  10. [18]

    Unifying vision-and- language tasks via text generation,

    J. Cho, J. Lei, H. Tan, and M. Bansal, “Unifying vision-and- language tasks via text generation,” in ICML, 2021. 1

  11. [19]

    Simvlm: Simple visual language model pretraining with weak supervision,

    Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao, “Simvlm: Simple visual language model pretraining with weak supervision,” arXiv:2108.10904, 2021. 1

  12. [20]

    Finetuned language models are zero- shot learners,

    J. Wei, M. Bosma, V . Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V . Le, “Finetuned language models are zero- shot learners,” arXiv:2109.01652, 2021. 1, 6, 11

  13. [21]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” arXiv:2304.08485, 2023. 1, 4, 6, 7, 8, 9, 10

  14. [22]

    Minigpt-4: Enhancing vision-language understanding with advanced large language models,

    D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with advanced large language models,” arXiv:2304.10592, 2023. 1, 2, 6, 7

  15. [23]

    Mm-react: Prompting chatgpt for multimodal reasoning and action,

    Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang, “Mm-react: Prompting chatgpt for multimodal reasoning and action,” arXiv:2303.11381, 2023. 1, 11, 12, 13

  16. [24]

    Palm-e: An embodied multimodal language model,

    D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu et al. , “Palm-e: An embodied multimodal language model,” arXiv:2303.03378, 2023. 1

  17. [25]

    Openflamingo: An open-source framework for training large autoregressive vision-language models,

    A. Awadalla, I. Gao, J. Gardner, J. Hessel, Y. Hanafy, W. Zhu, K. Marathe, Y. Bitton, S. Gadre, S. Sagawa et al., “Openflamingo: An open-source framework for training large autoregressive vision-language models,” arXiv:2308.01390, 2023. 1

  18. [26]

    Videochat: Chat-centric video understanding,

    K. Li, Y. He, Y. Wang, Y. Li, W. Wang, P . Luo, Y. Wang, L. Wang, and Y. Qiao, “Videochat: Chat-centric video understanding,” arXiv:2305.06355, 2023. 1, 4, 6

  19. [27]

    Video-llama: An instruction- tuned audio-visual language model for video understanding,

    H. Zhang, X. Li, and L. Bing, “Video-llama: An instruction- tuned audio-visual language model for video understanding,” arXiv:2306.02858, 2023. 1, 4

  20. [28]

    Pengi: An audio language model for audio tasks,

    S. Deshmukh, B. Elizalde, R. Singh, and H. Wang, “Pengi: An audio language model for audio tasks,” NeurIPS, 2024. 1, 3

  21. [29]

    Shikra: Unleashing multimodal llm’s referential dialogue magic,

    K. Chen, Z. Zhang, W. Zeng, R. Zhang, F. Zhu, and R. Zhao, “Shikra: Unleashing multimodal llm’s referential dialogue magic,” arXiv:2306.15195. 1, 9

  22. [30]

    Osprey: Pixel understanding with visual instruction tuning,

    Y. Yuan, W. Li, J. Liu, D. Tang, X. Luo, C. Qin, L. Zhang, and J. Zhu, “Osprey: Pixel understanding with visual instruction tuning,” arXiv:2312.10032. 1, 2, 9

  23. [31]

    Imagebind-llm: Multi-modality instruction tuning,

    J. Han, R. Zhang, W. Shao, P . Gao, P . Xu, H. Xiao, K. Zhang, C. Liu, S. Wen, Z. Guo et al., “Imagebind-llm: Multi-modality instruction tuning,” arXiv:2309.03905, 2023. 1, 3

  24. [32]

    Anymal: An efficient and scalable any-modality augmented language model,

    S. Moon, A. Madotto, Z. Lin, T. Nagarajan, M. Smith, S. Jain, C.-F. Yeh, P . Murugesan, P . Heidari, Y. Liu et al. , “Anymal: An efficient and scalable any-modality augmented language model,” arXiv:2309.16058, 2023. 1

  25. [33]

    Next-gpt: Any-to-any multimodal llm,

    S. Wu, H. Fei, L. Qu, W. Ji, and T.-S. Chua, “Next-gpt: Any-to-any multimodal llm,” arXiv:2309.05519, 2023. 1, 9

  26. [34]

    Large multilingual models pivot zero- shot multimodal learning across languages,

    J. Hu, Y. Yao, C. Wang, S. Wang, Y. Pan, Q. Chen, T. Yu, H. Wu, Y. Zhao, H. Zhang et al., “Large multilingual models pivot zero- shot multimodal learning across languages,” arXiv:2308.12038,

  27. [35]

    Qwen-vl: A frontier large vision-language model with versatile abilities,

    J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P . Wang, J. Lin, C. Zhou, and J. Zhou, “Qwen-vl: A frontier large vision-language model with versatile abilities,” arXiv:2308.12966, 2023. 1, 3, 4, 10

  28. [37]

    Med-flamingo: a multi- modal medical few-shot learner,

    M. Moor, Q. Huang, S. Wu, M. Yasunaga, Y. Dalmia, J. Leskovec, C. Zakka, E. P . Reis, and P . Rajpurkar, “Med-flamingo: a multi- modal medical few-shot learner,” in Machine Learning for Health (ML4H), 2023. 1, 10

  29. [38]

    Pmc-vqa: Visual instruction tuning for medical visual question answering,

    X. Zhang, C. Wu, Z. Zhao, W. Lin, Y. Zhang, Y. Wang, and W. Xie, “Pmc-vqa: Visual instruction tuning for medical visual question answering,” arXiv:2305.10415, 2023. 1, 4, 6, 10

  30. [39]

    mplug-docowl: Modularized multimodal large language model for document understanding,

    J. Ye, A. Hu, H. Xu, Q. Ye, M. Yan, Y. Dan, C. Zhao, G. Xu, C. Li, J. Tian et al. , “mplug-docowl: Modularized multimodal large language model for document understanding,” arXiv:2307.02499,

  31. [40]

    Textmonkey: An ocr-free large multimodal model for under- standing document,

    Y. Liu, B. Yang, Q. Liu, Z. Li, Z. Ma, S. Zhang, and X. Bai, “Textmonkey: An ocr-free large multimodal model for under- standing document,” arXiv:2403.04473, 2024. 1, 10

  32. [41]

    mplug-paperowl: Scientific diagram analysis with the multimodal large language model,

    A. Hu, Y. Shi, H. Xu, J. Ye, Q. Ye, M. Yan, C. Li, Q. Qian, J. Zhang, and F. Huang, “mplug-paperowl: Scientific diagram analysis with the multimodal large language model,” arXiv:2311.18248,

  33. [42]

    1 IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 15

  34. [43]

    An embodied generalist agent in 3d world,

    J. Huang, S. Yong, X. Ma, X. Linghu, P . Li, Y. Wang, Q. Li, S.-C. Zhu, B. Jia, and S. Huang, “An embodied generalist agent in 3d world,” arXiv:2311.12871, 2023. 1, 9, 10

  35. [44]

    Kosmos-2: Grounding multimodal large language models to the world,

    Z. Peng, W. Wang, L. Dong, Y. Hao, S. Huang, S. Ma, and F. Wei, “Kosmos-2: Grounding multimodal large language models to the world,” arXiv:2306.14824, 2023. 1

  36. [45]

    Appagent: Multimodal agents as smartphone users,

    Z. Yang, J. Liu, Y. Han, X. Chen, Z. Huang, B. Fu, and G. Yu, “Appagent: Multimodal agents as smartphone users,” arXiv:2312.13771, 2023. 1, 10

  37. [46]

    Cogagent: A visual language model for gui agents,

    W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Dong, M. Ding et al., “Cogagent: A visual language model for gui agents,” arXiv:2312.08914, 2023. 1, 3, 10

  38. [47]

    Mobile-agent: Autonomous multi-modal mobile device agent with visual perception,

    J. Wang, H. Xu, J. Ye, M. Yan, W. Shen, J. Zhang, F. Huang, and J. Sang, “Mobile-agent: Autonomous multi-modal mobile device agent with visual perception,” arXiv:2401.16158, 2024. 1, 10

  39. [48]

    Repro- ducible scaling laws for contrastive language-image learning,

    M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev, “Repro- ducible scaling laws for contrastive language-image learning,” in CVPR, 2023. 2, 3

  40. [49]

    Eva-clip: Improved training techniques for clip at scale,

    Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao, “Eva-clip: Improved training techniques for clip at scale,” arXiv:2303.15389, 2023. 2, 3

  41. [50]

    Eva: Exploring the limits of masked visual representation learning at scale,

    Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao, “Eva: Exploring the limits of masked visual representation learning at scale,” in CVPR, 2023. 2

  42. [51]

    Introducing our multimodal models,

    R. Bavishi, E. Elsen, C. Hawthorne, M. Nye, A. Odena, A. Somani, and S. Ta¸ sırlar, “Introducing our multimodal models,” 2023. [Online]. Available: https://www.adept.ai/blog/fuyu-8b 2

  43. [52]

    Improved baselines with visual instruction tuning,

    H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” arXiv:2310.03744, 2023. 3, 4

  44. [53]

    Monkey: Image resolution and text label are important things for large multi-modal models,

    Z. Li, B. Yang, Q. Liu, Z. Ma, S. Zhang, J. Yang, Y. Sun, Y. Liu, and X. Bai, “Monkey: Image resolution and text label are important things for large multi-modal models,” arXiv:2311.06607, 2023. 3

  45. [54]

    Mm1: Methods, analysis & insights from multimodal llm pre-training,

    B. McKinzie, Z. Gan, J.-P . Fauconnier, S. Dodge, B. Zhang, P . Dufter, D. Shah, X. Du, F. Peng, F. Weers et al. , “Mm1: Methods, analysis & insights from multimodal llm pre-training,” arXiv:2403.09611, 2024. 3, 4

  46. [55]

    Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models,

    Z. Lin, C. Liu, R. Zhang, P . Gao, L. Qiu, H. Xiao, H. Qiu, C. Lin, W. Shao, K. Chen et al. , “Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models,” arXiv:2311.07575, 2023. 3

  47. [56]

    Clap learning audio concepts from natural language supervision,

    B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, “Clap learning audio concepts from natural language supervision,” in ICASSP, 2023. 3

  48. [57]

    Imagebind: One embedding space to bind them all,

    R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V . Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” in CVPR, 2023. 3

  49. [58]

    Scaling instruction- finetuned language models,

    H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma et al. , “Scaling instruction- finetuned language models,” arXiv:2210.11416, 2022. 3, 4, 6

  50. [59]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P . Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P . Bhargava, S. Bhosale et al. , “Llama 2: Open foundation and fine-tuned chat models,” arXiv:2307.09288, 2023. 3, 4

  51. [60]

    Qwen technical report,

    J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang et al. , “Qwen technical report,” arXiv:2309.16609, 2023. 3, 4, 10

  52. [61]

    Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

    J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” arXiv:2301.12597, 2023. 3, 4

  53. [62]

    Instructblip: Towards general- purpose vision-language models with instruction tuning,

    W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P . Fung, and S. Hoi, “Instructblip: Towards general- purpose vision-language models with instruction tuning,” arXiv:2305.06500, 2023. 3, 4, 6, 7, 8

  54. [63]

    Llava-next: Improved reasoning, ocr, and world knowledge,

    H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee, “Llava-next: Improved reasoning, ocr, and world knowledge,” January 2024. [Online]. Available: https://llava-vl.github.io/ blog/2024-01-30-llava-next/ 3

  55. [64]

    An empir- ical study of scaling instruct-tuned large multimodal models,

    Y. Lu, C. Li, H. Liu, J. Yang, J. Gao, and Y. Shen, “An empir- ical study of scaling instruct-tuned large multimodal models,” arXiv:2309.09958, 2023. 3

  56. [65]

    Mobilevlm: A fast, repro- ducible and strong vision language assistant for mobile devices,

    X. Chu, L. Qiao, X. Lin, S. Xu, Y. Yang, Y. Hu, F. Wei, X. Zhang, B. Zhang, X. Wei et al. , “Mobilevlm: A fast, repro- ducible and strong vision language assistant for mobile devices,” arXiv:2312.16886, 2023. 3, 10

  57. [66]

    Mobilevlm v2: Faster and stronger baseline for vision language model,

    X. Chu, L. Qiao, X. Zhang, S. Xu, F. Wei, Y. Yang, X. Sun, Y. Hu, X. Lin, B. Zhang et al., “Mobilevlm v2: Faster and stronger baseline for vision language model,” arXiv:2402.03766, 2024. 3

  58. [67]

    Mixture-of-experts meets instruction tuning: A winning combination for large language models,

    S. Shen, L. Hou, Y. Zhou, N. Du, S. Longpre, J. Wei, H. W. Chung, B. Zoph, W. Fedus, X. Chen et al. , “Mixture-of-experts meets instruction tuning: A winning combination for large language models,” arXiv:2305.14705, 2023. 3

  59. [68]

    Mixtral of experts,

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand et al., “Mixtral of experts,” arXiv:2401.04088, 2024. 3

  60. [69]

    Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

    W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” JMLR, 2022. 3

  61. [70]

    Moe-llava: Mixture of experts for large vision-language models,

    B. Lin, Z. Tang, Y. Ye, J. Cui, B. Zhu, P . Jin, J. Zhang, M. Ning, and L. Yuan, “Moe-llava: Mixture of experts for large vision-language models,” arXiv:2401.15947, 2024. 3

  62. [71]

    End-to-end object detection with transformers,

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV, 2020. 4

  63. [72]

    X- llm: Bootstrapping advanced large language models by treating multi-modalities as foreign languages,

    F. Chen, M. Han, H. Zhao, Q. Zhang, J. Shi, S. Xu, and B. Xu, “X- llm: Bootstrapping advanced large language models by treating multi-modalities as foreign languages,” arXiv:2305.04160, 2023. 4, 6, 8, 9

  64. [73]

    Pandagpt: One model to instruction-follow them all,

    Y. Su, T. Lan, H. Li, J. Xu, Y. Wang, and D. Cai, “Pandagpt: One model to instruction-follow them all,” arXiv:2305.16355, 2023. 4, 6

  65. [74]

    Detgpt: Detect what you need via reasoning,

    R. Pi, J. Gao, S. Diao, R. Pan, H. Dong, J. Zhang, L. Yao, J. Han, H. Xu, and L. K. T. Zhang, “Detgpt: Detect what you need via reasoning,” arXiv:2305.14167, 2023. 4, 7

  66. [75]

    What matters in training a gpt4-style language model with multimodal inputs?

    Y. Zeng, H. Zhang, J. Zheng, J. Xia, G. Wei, Y. Wei, Y. Zhang, and T. Kong, “What matters in training a gpt4-style language model with multimodal inputs?” arXiv:2307.02469, 2023. 4, 7

  67. [76]

    Flamingo: a visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P . Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al., “Flamingo: a visual language model for few-shot learning,” NeurIPS, 2022. 4, 11, 12

  68. [77]

    Cogvlm: Visual expert for pretrained language models,

    W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, X. Song et al. , “Cogvlm: Visual expert for pretrained language models,” arXiv:2311.03079, 2023. 4

  69. [78]

    Llama-adapter: Efficient fine-tuning of language models with zero-init attention,

    R. Zhang, J. Han, A. Zhou, X. Hu, S. Yan, P . Lu, H. Li, P . Gao, and Y. Qiao, “Llama-adapter: Efficient fine-tuning of language models with zero-init attention,” arXiv:2303.16199, 2023. 4, 6, 8, 9

  70. [79]

    Woodpecker: Hallucination correction for multimodal large language models,

    S. Yin, C. Fu, S. Zhao, T. Xu, H. Wang, D. Sui, Y. Shen, K. Li, X. Sun, and E. Chen, “Woodpecker: Hallucination correction for multimodal large language models,” arXiv:2310.16045, 2023. 4, 9, 10, 11

  71. [80]

    From images to textual prompts: Zero-shot visual question answering with frozen large language models,

    J. Guo, J. Li, D. Li, A. M. H. Tiong, B. Li, D. Tao, and S. Hoi, “From images to textual prompts: Zero-shot visual question answering with frozen large language models,” in CVPR, 2023. 4

  72. [81]

    Caption anything: Interactive image description with diverse multimodal controls,

    T. Wang, J. Zhang, J. Fei, Y. Ge, H. Zheng, Y. Tang, Z. Li, M. Gao, S. Zhao, Y. Shan et al. , “Caption anything: Interactive image description with diverse multimodal controls,” arXiv:2305.02677,

  73. [82]

    Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions,

    D. Zhu, J. Chen, K. Haydarov, X. Shen, W. Zhang, and M. El- hoseiny, “Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions,” arXiv:2303.06594, 2023. 4, 12

  74. [83]

    mplug-owl: Modularization empowers large language models with multimodality,

    Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P . Shi, Y. Shiet al., “mplug-owl: Modularization empowers large language models with multimodality,” arXiv:2304.14178, 2023. 4, 7, 9

  75. [84]

    Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,

    W. Wang, Z. Chen, X. Chen, J. Wu, X. Zhu, G. Zeng, P . Luo, T. Lu, J. Zhou, Y. Qiaoet al., “Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,” arXiv:2305.11175,

  76. [85]

    Sharegpt4v: Improving large multi-modal models with better captions,

    L. Chen, J. Li, X. Dong, P . Zhang, C. He, J. Wang, F. Zhao, and D. Lin, “Sharegpt4v: Improving large multi-modal models with better captions,” arXiv:2311.12793, 2023. 5

  77. [86]

    Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,

    P . Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in ACL, 2018. 5

  78. [87]

    Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,

    S. Changpinyo, P . Sharma, N. Ding, and R. Soricut, “Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,” in CVPR, 2021. 5

  79. [88]

    Im2text: Describing im- ages using 1 million captioned photographs,

    V . Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing im- ages using 1 million captioned photographs,” NeurIPS, 2011. 5

  80. [89]

    Laion-5b: An open large-scale dataset for training next genera- tion image-text models,

    C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next genera- tion image-text models,” NeurIPS, 2022. 5 IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND M...

  81. [90]

    Laion coco: 600m synthetic captions from laion2b-en

    C. Schuhmann, A. Köpf, R. Vencu, T. Coombes, and R. Beau- mont, “Laion coco: 600m synthetic captions from laion2b-en.” https://laion.ai/blog/laion-coco/, 2022. 5

  82. [91]

    Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,

    J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,” in ICML, 2022. 5

  83. [92]

    Coyo-700m: Image-text pair dataset,

    M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim, “Coyo-700m: Image-text pair dataset,” https://github.com/ kakaobrain/coyo-dataset, 2022. 5

  84. [93]

    To see is to believe: Prompting gpt-4v for better visual instruction tuning,

    J. Wang, L. Meng, Z. Weng, B. He, Z. Wu, and Y.-G. Jiang, “To see is to believe: Prompting gpt-4v for better visual instruction tuning,” arXiv:2311.07574, 2023. 5, 7

  85. [94]

    Allava: Harness- ing gpt4v-synthesized data for a lite vision-language model,

    G. H. Chen, S. Chen, R. Zhang, J. Chen, X. Wu, Z. Zhang, Z. Chen, J. Li, X. Wan, and B. Wang, “Allava: Harness- ing gpt4v-synthesized data for a lite vision-language model,” arXiv:2402.11684, 2024. 5, 7

  86. [95]

    Msr-vtt: A large video descrip- tion dataset for bridging video and language,

    J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video descrip- tion dataset for bridging video and language,” in CVPR, 2016. 5

  87. [96]

    Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

    X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, “Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,” arXiv:2303.17395, 2023. 5

  88. [97]

    Training language models to follow instructions with human feedback,

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P . Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” NeurIPS, 2022. 6, 8

  89. [98]

    Opt-iml: Scaling language model instruction meta learning through the lens of generalization,

    S. Iyer, X. V . Lin, R. Pasunuru, T. Mihaylov, D. Simig, P . Yu, K. Shuster, T. Wang, Q. Liu, P . S. Koura et al. , “Opt-iml: Scaling language model instruction meta learning through the lens of generalization,” arXiv:2212.12017, 2022. 6

  90. [99]

    Mul- titask prompted training enables zero-shot task generalization,

    V . Sanh, A. Webson, C. Raffel, S. H. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, T. L. Scao, A. Rajaet al., “Mul- titask prompted training enables zero-shot task generalization,” arXiv:2110.08207, 2021. 6

  91. [100]

    Multimodal-gpt: A vision and language model for dialogue with humans,

    T. Gong, C. Lyu, S. Zhang, Y. Wang, M. Zheng, Q. Zhao, K. Liu, W. Zhang, P . Luo, and K. Chen, “Multimodal-gpt: A vision and language model for dialogue with humans,” arXiv:2305.04790,

  92. [101]

    Vqa: Visual question answering,

    S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in ICCV, 2015. 6

  93. [102]

    Deep visual-semantic alignments for generating image descriptions,

    A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in CVPR, 2015. 6

  94. [103]

    Cheap and quick: Efficient vision-language instruction tuning for large language models,

    G. Luo, Y. Zhou, T. Ren, S. Chen, X. Sun, and R. Ji, “Cheap and quick: Efficient vision-language instruction tuning for large language models,” arXiv:2305.15023, 2023. 6, 7, 8, 9

  95. [104]

    Multiinstruct: Improv- ing multi-modal zero-shot learning via instruction tuning,

    Z. Xu, Y. Shen, and L. Huang, “Multiinstruct: Improv- ing multi-modal zero-shot learning via instruction tuning,” arXiv:2212.10773, 2022. 6, 7, 8, 9

  96. [105]

    Llama-adapter v2: Parameter-efficient visual instruction model,

    P . Gao, J. Han, R. Zhang, Z. Lin, S. Geng, A. Zhou, W. Zhang, P . Lu, C. He, X. Yueet al., “Llama-adapter v2: Parameter-efficient visual instruction model,” arXiv:2304.15010, 2023. 6, 7, 8, 9

  97. [106]

    Chatbridge: Bridging modalities with large language model as a language catalyst,

    Z. Zhao, L. Guo, T. Yue, S. Chen, S. Shao, X. Zhu, Z. Yuan, and J. Liu, “Chatbridge: Bridging modalities with large language model as a language catalyst,” arXiv:2305.16103, 2023. 6, 7, 8, 9

  98. [107]

    M3it: A large-scale dataset towards multi-modal multilingual instruction tuning,

    L. Li, Y. Yin, S. Li, L. Chen, P . Wang, S. Ren, M. Li, Y. Yang, J. Xu, X. Sun, L. Kong, and Q. Liu, “M3it: A large-scale dataset towards multi-modal multilingual instruction tuning,” arXiv:2306.04387,

  99. [108]

    Self-instruct: Aligning language model with self generated instructions,

    Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi, “Self-instruct: Aligning language model with self generated instructions,” arXiv:2212.10560, 2022. 7

  100. [109]

    Gpt4tools: Teaching large language model to use tools via self- instruction,

    R. Yang, L. Song, Y. Li, S. Zhao, Y. Ge, X. Li, and Y. Shan, “Gpt4tools: Teaching large language model to use tools via self- instruction,” arXiv:2305.18752, 2023. 7, 9, 12, 13

  101. [110]

    Instructiongpt- 4: A 200-instruction paradigm for fine-tuning minigpt-4,

    L. Wei, Z. Jiang, W. Huang, and L. Sun, “Instructiongpt- 4: A 200-instruction paradigm for fine-tuning minigpt-4,” arXiv:2308.12067, 2023. 7

  102. [111]

    What makes for good visual instruc- tions? synthesizing complex visual reasoning instructions for visual instruction tuning,

    Y. Du, H. Guo, K. Zhou, W. X. Zhao, J. Wang, C. Wang, M. Cai, R. Song, and J.-R. Wen, “What makes for good visual instruc- tions? synthesizing complex visual reasoning instructions for visual instruction tuning,” arXiv:2311.01487, 2023. 7

  103. [112]

    Fine-tuning language models from human preferences,

    D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P . Christiano, and G. Irving, “Fine-tuning language models from human preferences,” arXiv:1909.08593, 2019. 8

  104. [113]

    Learning to sum- marize with human feedback,

    N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P . F. Christiano, “Learning to sum- marize with human feedback,” NeurIPS, 2020. 8

  105. [114]

    Aligning large multimodal models with factually augmented rlhf,

    Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y. Shen, C. Gan, L.-Y. Gui, Y.-X. Wang, Y. Yang et al. , “Aligning large multimodal models with factually augmented rlhf,” arXiv:2309.14525, 2023. 8, 11

  106. [115]

    Direct preference optimization: Your language model is secretly a reward model,

    R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” NeurIPS, 2023. 8

  107. [116]

    Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback,

    T. Yu, Y. Yao, H. Zhang, T. He, Y. Han, G. Cui, J. Hu, Z. Liu, H.-T. Zheng, M. Sun et al. , “Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback,” arXiv:2312.00849, 2023. 8

  108. [117]

    Silkie: Preference distillation for large visual language models,

    L. Li, Z. Xie, M. Li, S. Chen, P . Wang, L. Chen, Y. Yang, B. Wang, and L. Kong, “Silkie: Preference distillation for large visual language models,” arXiv:2312.10665, 2023. 8

  109. [118]

    Learn to explain: Multimodal reasoning via thought chains for science question answering,

    P . Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P . Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” NeurIPS,

  110. [119]

    Cider: Consensus-based image description evaluation,

    R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in CVPR, 2015. 8

  111. [120]

    Nocaps: Novel object captioning at scale,

    H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P . Anderson, “Nocaps: Novel object captioning at scale,” in ICCV, 2019. 8

  112. [121]

    From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,

    P . Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” TACL, 2014. 8

  113. [122]

    Pathvqa: 30000+ questions for medical visual question answering,

    X. He, Y. Zhang, L. Mou, E. Xing, and P . Xie, “Pathvqa: 30000+ questions for medical visual question answering,” arXiv:2003.10286, 2020. 9

  114. [123]

    A dataset of clinically generated visual questions and answers about radiology images,

    J. J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman, “A dataset of clinically generated visual questions and answers about radiology images,” Sci. Data, 2018. 9

  115. [124]

    Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering,

    B. Liu, L.-M. Zhan, L. Xu, L. Ma, Y. Yang, and X.-M. Wu, “Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering,” in ISBI, 2021. 9

  116. [125]

    Mme: A comprehensive evaluation bench- mark for multimodal large language models,

    C. Fu, P . Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, Z. Qiu, W. Lin, Z. Qiu, W. Lin et al., “Mme: A comprehensive evaluation bench- mark for multimodal large language models,” arXiv:2306.13394,

  117. [126]

    Mmbench: Is your multi-modal model an all-around player?

    Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu et al. , “Mmbench: Is your multi-modal model an all-around player?” arXiv:2307.06281, 2023. 9

  118. [127]

    Mm-vet: Evaluating large multimodal models for integrated capabilities,

    W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang, “Mm-vet: Evaluating large multimodal models for integrated capabilities,” arXiv:2308.02490, 2023. 9

  119. [128]

    Seed- bench: Benchmarking multimodal llms with generative compre- hension,

    B. Li, R. Wang, G. Wang, Y. Ge, Y. Ge, and Y. Shan, “Seed- bench: Benchmarking multimodal llms with generative compre- hension,” in CVPR, 2024. 9

  120. [129]

    Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts,

    P . Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao, “Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts,” in ICLR,

  121. [130]

    Mmmu: A massive multi- discipline multimodal understanding and reasoning benchmark for expert agi,

    X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun et al. , “Mmmu: A massive multi- discipline multimodal understanding and reasoning benchmark for expert agi,” arXiv:2311.16502, 2023. 9

  122. [131]

    Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models,

    F. Liu, T. Guan, Z. Li, L. Chen, Y. Yacoob, D. Manocha, and T. Zhou, “Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models,” in CVPR, 2024. 9

  123. [132]

    Video-chatgpt: Towards detailed video understanding via large vision and lan- guage models,

    M. Maaz, H. Rasheed, S. Khan, and F. S. Khan, “Video-chatgpt: Towards detailed video understanding via large vision and lan- guage models,” arXiv:2306.05424, 2023. 9

  124. [133]

    Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models,

    M. Ning, B. Zhu, Y. Xie, B. Lin, J. Cui, L. Yuan, D. Chen, and L. Yuan, “Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models,” arXiv:2311.16103, 2023. 9

  125. [134]

    Evaluating object hallucination in large vision-language mod- els,

    Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen, “Evaluating object hallucination in large vision-language mod- els,” arXiv:2305.10355, 2023. 9, 10

  126. [135]

    Microsoft coco: Common objects in context,

    T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P . Perona, D. Ramanan, P . Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV, 2014. 9 IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 17

  127. [136]

    Red teaming visual language models,

    M. Li, L. Li, Y. Yin, M. Ahmed, Z. Liu, and Q. Liu, “Red teaming visual language models,” arXiv:2401.12915, 2024. 9

  128. [137]

    The dawn of lmms: Preliminary explorations with gpt-4v (ision),

    Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang, “The dawn of lmms: Preliminary explorations with gpt-4v (ision),” arXiv:2309.17421. 9

  129. [138]

    On the road with gpt-4v (ision): Early explorations of visual-language model on autonomous driving,

    L. Wen, X. Yang, D. Fu, X. Wang, P . Cai, X. Li, T. Ma, Y. Li, L. Xu, D. Shang et al. , “On the road with gpt-4v (ision): Early explorations of visual-language model on autonomous driving,” arXiv:2311.05332. 9

  130. [139]

    A challenger to gpt-4v? early explorations of gemini in visual expertise,

    C. Fu, R. Zhang, H. Lin, Z. Wang, T. Gao, Y. Luo, Y. Huang, Z. Zhang, L. Qiu, G. Ye et al. , “A challenger to gpt-4v? early explorations of gemini in visual expertise,” arXiv:2312.12436. 9

  131. [140]

    Gpt4roi: Instruction tuning large language model on region-of-interest,

    S. Zhang, P . Sun, S. Chen, M. Xiao, W. Shao, W. Zhang, K. Chen, and P . Luo, “Gpt4roi: Instruction tuning large language model on region-of-interest,” arXiv:2307.03601, 2023. 9

  132. [141]

    Pink: Unveiling the power of referential comprehension for multi-modal llms,

    S. Xuan, Q. Guo, M. Yang, and S. Zhang, “Pink: Unveiling the power of referential comprehension for multi-modal llms,” arXiv:2310.00582, 2023. 9

  133. [142]

    Glamm: Pixel grounding large multimodal model,

    H. Rasheed, M. Maaz, S. Shaji, A. Shaker, S. Khan, H. Cholakkal, R. M. Anwer, E. Xing, M.-H. Yang, and F. S. Khan, “Glamm: Pixel grounding large multimodal model,” arXiv:2311.03356. 9

  134. [143]

    Ferret: Refer and ground anything anywhere at any granularity,

    H. You, H. Zhang, Z. Gan, X. Du, B. Zhang, Z. Wang, L. Cao, S.-F. Chang, and Y. Yang, “Ferret: Refer and ground anything anywhere at any granularity,” arXiv:2310.07704, 2023. 9

  135. [144]

    Lisa: Reasoning segmentation via large language model,

    X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia, “Lisa: Reasoning segmentation via large language model,” arXiv:2308.00692, 2023. 9, 13

  136. [145]

    Pointllm: Empowering large language models to understand point clouds,

    R. Xu, X. Wang, T. Wang, Y. Chen, J. Pang, and D. Lin, “Pointllm: Empowering large language models to understand point clouds,” arXiv:2308.16911, 2023. 9

  137. [146]

    Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning,

    S. Chen, X. Chen, C. Zhang, M. Li, G. Yu, H. Fei, H. Zhu, J. Fan, and T. Chen, “Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning,” arXiv:2311.18651, 2023. 9

  138. [147]

    3d-llm: Injecting the 3d world into large language models,

    Y. Hong, H. Zhen, P . Chen, S. Zheng, Y. Du, Z. Chen, and C. Gan, “3d-llm: Injecting the 3d world into large language models,” NeurIPS, 2023. 9

  139. [148]

    Generative pretraining in multimodal- ity,

    Q. Sun, Q. Yu, Y. Cui, F. Zhang, X. Zhang, Y. Wang, H. Gao, J. Liu, T. Huang, and X. Wang, “Generative pretraining in multimodal- ity,” in ICLR, 2024. 9

  140. [149]

    Anygpt: Unified multimodal llm with discrete sequence modeling,

    J. Zhan, J. Dai, J. Ye, Y. Zhou, D. Zhang, Z. Liu, X. Zhang, R. Yuan, G. Zhang, L. Li et al. , “Anygpt: Unified multimodal llm with discrete sequence modeling,” arXiv:2402.12226, 2024. 9

  141. [150]

    Jointly training large autoregressive multimodal models,

    E. Aiello, L. Yu, Y. Nie, A. Aghajanyan, and B. Oguz, “Jointly training large autoregressive multimodal models,” arXiv:2309.15564, 2023. 9

  142. [151]

    Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,

    D. Zhang, S. Li, X. Zhang, J. Zhan, P . Wang, Y. Zhou, and X. Qiu, “Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,” arXiv:2305.11000, 2023. 9

  143. [152]

    Audiopalm: A large language model that can speak and listen,

    P . K. Rubenstein, C. Asawaroengchai, D. D. Nguyen, A. Bapna, Z. Borsos, F. d. C. Quitry, P . Chen, D. E. Badawy, W. Han, E. Kharitonov et al. , “Audiopalm: A large language model that can speak and listen,” arXiv:2306.12925, 2023. 9

  144. [153]

    Modaverse: Efficiently trans- forming modalities with llms,

    X. Wang, B. Zhuang, and Q. Wu, “Modaverse: Efficiently trans- forming modalities with llms,” arXiv:2401.06395, 2024. 9

  145. [154]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P . Abbeel, “Denoising diffusion probabilistic models,” NeurIPS, 2020. 10

  146. [155]

    High-resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P . Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR, 2022. 10

  147. [156]

    Mindagent: Emergent gaming interaction,

    R. Gong, Q. Huang, X. Ma, H. Vo, Z. Durante, Y. Noda, Z. Zheng, S.-C. Zhu, D. Terzopoulos, L. Fei-Feiet al., “Mindagent: Emergent gaming interaction,” arXiv:2309.09971, 2023. 10

  148. [157]

    Embodiedgpt: Vision-language pre-training via embodied chain of thought,

    Y. Mu, Q. Zhang, M. Hu, W. Wang, M. Ding, J. Jin, B. Wang, J. Dai, Y. Qiao, and P . Luo, “Embodiedgpt: Vision-language pre-training via embodied chain of thought,” arXiv:2305.15021, 2023. 10

  149. [158]

    mplug-docowl 1.5: Unified structure learning for ocr-free document understanding,

    A. Hu, H. Xu, J. Ye, M. Yan, L. Zhang, B. Zhang, C. Li, J. Zhang, Q. Jin, F. Huang et al. , “mplug-docowl 1.5: Unified structure learning for ocr-free document understanding,”arXiv:2403.12895,

  150. [159]

    Ureader: Universal ocr-free visually-situated lan- guage understanding with multimodal large language model,

    J. Ye, A. Hu, H. Xu, Q. Ye, M. Yan, G. Xu, C. Li, J. Tian, Q. Qian, J. Zhang et al., “Ureader: Universal ocr-free visually-situated lan- guage understanding with multimodal large language model,” in EMNLP, 2023. 10

  151. [160]

    Llava-med: Training a large language-and-vision assistant for biomedicine in one day,

    C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao, “Llava-med: Training a large language-and-vision assistant for biomedicine in one day,” arXiv:2306.00890, 2023. 10

  152. [161]

    Halle-switch: Rethinking and con- trolling object existence hallucinations in large vision language models for detailed caption,

    B. Zhai, S. Yang, X. Zhao, C. Xu, S. Shen, D. Zhao, K. Keutzer, M. Li, T. Yan, and X. Fan, “Halle-switch: Rethinking and con- trolling object existence hallucinations in large vision language models for detailed caption,” arXiv:2310.01779, 2023. 10, 11

  153. [162]

    Object hallucination in image captioning,

    A. Rohrbach, L. A. Hendricks, K. Burns, T. Darrell, and K. Saenko, “Object hallucination in image captioning,” in EMNLP, 2018. 10

  154. [163]

    Evaluation and analysis of hallucination in large vision-language models,

    J. Wang, Y. Zhou, G. Xu, P . Shi, C. Zhao, H. Xu, Q. Ye, M. Yan, J. Zhang, J. Zhu et al. , “Evaluation and analysis of hallucination in large vision-language models,” arXiv:2308.15126, 2023. 10

  155. [164]

    Faithscore: Evaluating hallucinations in large vision-language models,

    L. Jing, R. Li, Y. Chen, M. Jia, and X. Du, “Faithscore: Evaluating hallucinations in large vision-language models,” arXiv:2311.01477, 2023. 10

  156. [165]

    An llm-free multi-dimensional benchmark for mllms hallucination evaluation,

    J. Wang, Y. Wang, G. Xu, J. Zhang, Y. Gu, H. Jia, M. Yan, J. Zhang, and J. Sang, “An llm-free multi-dimensional benchmark for mllms hallucination evaluation,” arXiv:2311.07397, 2023. 10

  157. [166]

    Mitigating hallucination in large multi-modal models via robust instruction tuning,

    F. Liu, K. Lin, L. Li, J. Wang, Y. Yacoob, and L. Wang, “Mitigating hallucination in large multi-modal models via robust instruction tuning,” in ICLR, 2024. 11

  158. [167]

    Mitigating object hallucinations in large vision-language models through visual contrastive decoding,

    S. Leng, H. Zhang, G. Chen, X. Li, S. Lu, C. Miao, and L. Bing, “Mitigating object hallucinations in large vision-language models through visual contrastive decoding,” in CVPR, 2024. 11

  159. [168]

    Hallucination augmented contrastive learning for multimodal large language model,

    C. Jiang, H. Xu, M. Dong, J. Chen, W. Ye, M. Yan, Q. Ye, J. Zhang, F. Huang, and S. Zhang, “Hallucination augmented contrastive learning for multimodal large language model,” arXiv:2312.06968, 2023. 11

  160. [169]

    Analyzing and mitigating object hallucination in large vision-language models,

    Y. Zhou, C. Cui, J. Yoon, L. Zhang, Z. Deng, C. Finn, M. Bansal, and H. Yao, “Analyzing and mitigating object hallucination in large vision-language models,” arXiv:2310.00754, 2023. 11

  161. [170]

    A survey for in-context learning,

    Q. Dong, L. Li, D. Dai, C. Zheng, Z. Wu, B. Chang, X. Sun, J. Xu, and Z. Sui, “A survey for in-context learning,” arXiv:2301.00234,

  162. [171]

    Chameleon: Plug-and-play compositional reasoning with large language models,

    P . Lu, B. Peng, H. Cheng, M. Galley, K.-W. Chang, Y. N. Wu, S.- C. Zhu, and J. Gao, “Chameleon: Plug-and-play compositional reasoning with large language models,” arXiv:2304.09842, 2023. 11, 12, 13

  163. [172]

    Visual programming: Composi- tional visual reasoning without training,

    T. Gupta and A. Kembhavi, “Visual programming: Composi- tional visual reasoning without training,” in CVPR, 2023. 11, 12, 13

  164. [173]

    Fan- tastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity,

    Y. Lu, M. Bartolo, A. Moore, S. Riedel, and P . Stenetorp, “Fan- tastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity,” arXiv:2104.08786, 2021. 11

  165. [174]

    Mimic-it: Multi-modal in-context instruction tuning,

    B. Li, Y. Zhang, L. Chen, J. Wang, F. Pu, J. Yang, C. Li, and Z. Liu, “Mimic-it: Multi-modal in-context instruction tuning,” arXiv:2306.05425, 2023. 11

  166. [175]

    Generative pretraining in multimodal- ity,

    Q. Sun, Q. Yu, Y. Cui, F. Zhang, X. Zhang, Y. Wang, H. Gao, J. Liu, T. Huang, and X. Wang, “Generative pretraining in multimodal- ity,” arXiv:2307.05222, 2023. 11

  167. [176]

    Towards more unified in-context visual understanding,

    D. Sheng, D. Chen, Z. Tan, Q. Liu, Q. Chu, J. Bao, T. Gong, B. Liu, S. Xu, and N. Yu, “Towards more unified in-context visual understanding,” arXiv:2312.02520, 2023. 12

  168. [177]

    Link- context learning for multimodal llms,

    Y. Tai, W. Fan, Z. Zhang, F. Zhu, R. Zhao, and Z. Liu, “Link- context learning for multimodal llms,” arXiv:2308.07891, 2023. 12

  169. [178]

    Mmicl: Empowering vision-language model with multi-modal in-context learning,

    H. Zhao, Z. Cai, S. Si, X. Ma, K. An, L. Chen, Z. Liu, S. Wang, W. Han, and B. Chang, “Mmicl: Empowering vision-language model with multi-modal in-context learning,” arXiv:2309.07915,

  170. [179]

    Hijacking context in large multi-modal models,

    J. Jeong, “Hijacking context in large multi-modal models,” arXiv:2312.07553, 2023. 12, 14

  171. [180]

    An empirical study of gpt-3 for few-shot knowledge-based vqa,

    Z. Yang, Z. Gan, J. Wang, X. Hu, Y. Lu, Z. Liu, and L. Wang, “An empirical study of gpt-3 for few-shot knowledge-based vqa,” in AAAI, 2022. 12

  172. [181]

    Multimodal few-shot learning with frozen language models,

    M. Tsimpoukelli, J. L. Menick, S. Cabi, S. Eslami, O. Vinyals, and F. Hill, “Multimodal few-shot learning with frozen language models,” NeurIPS, 2021. 12

  173. [182]

    Ot- ter: A multi-modal model with in-context instruction tuning,

    B. Li, Y. Zhang, L. Chen, J. Wang, J. Yang, and Z. Liu, “Ot- ter: A multi-modal model with in-context instruction tuning,” arXiv:2305.03726, 2023. 12

  174. [183]

    Hugging- gpt: Solving ai tasks with chatgpt and its friends in huggingface,

    Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “Hugging- gpt: Solving ai tasks with chatgpt and its friends in huggingface,” arXiv:2303.17580, 2023. 12, 13

  175. [184]

    Large language models are zero-shot reasoners,

    T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” arXiv:2205.11916,

  176. [185]

    12 IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 18

  177. [186]

    Automatic chain of thought prompting in large language models,

    Z. Zhang, A. Zhang, M. Li, and A. Smola, “Automatic chain of thought prompting in large language models,” arXiv:2210.03493,

  178. [187]

    Visual chain of thought: Bridging logical gaps with multimodal infillings,

    D. Rose, V . Himakunthala, A. Ouyang, R. He, A. Mei, Y. Lu, M. Saxon, C. Sonar, D. Mirza, and W. Y. Wang, “Visual chain of thought: Bridging logical gaps with multimodal infillings,” arXiv:2305.02317, 2023. 12

  179. [188]

    Multimodal chain-of-thought reasoning in language models,

    Z. Zhang, A. Zhang, M. Li, H. Zhao, G. Karypis, and A. Smola, “Multimodal chain-of-thought reasoning in language models,” arXiv:2302.00923, 2023. 12

  180. [189]

    Let’s think frame by frame: Evaluating video chain of thought with video infilling and prediction,

    V . Himakunthala, A. Ouyang, D. Rose, R. He, A. Mei, Y. Lu, C. Sonar, M. Saxon, and W. Y. Wang, “Let’s think frame by frame: Evaluating video chain of thought with video infilling and prediction,” arXiv:2305.13903, 2023. 12

  181. [190]

    Chain of thought prompt tuning in vision language models,

    J. Ge, H. Luo, S. Qian, Y. Gan, J. Fu, and S. Zhan, “Chain of thought prompt tuning in vision language models,” arXiv:2304.07919, 2023. 12

  182. [191]

    Visual chatgpt: Talking, drawing and editing with visual foundation models,

    C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan, “Visual chatgpt: Talking, drawing and editing with visual foundation models,” arXiv:2303.04671, 2023. 12, 13

  183. [192]

    Ddcot: Duty-distinct chain-of-thought prompting for multimodal rea- soning in language models,

    G. Zheng, B. Yang, J. Tang, H.-Y. Zhou, and S. Yang, “Ddcot: Duty-distinct chain-of-thought prompting for multimodal rea- soning in language models,” in NeurIPS, 2023. 12

  184. [193]

    Talm: Tool augmented lan- guage models,

    A. Parisi, Y. Zhao, and N. Fiedel, “Talm: Tool augmented lan- guage models,” arXiv:2205.12255, 2022. 12

  185. [194]

    Pal: Program-aided language models,

    L. Gao, A. Madaan, S. Zhou, U. Alon, P . Liu, Y. Yang, J. Callan, and G. Neubig, “Pal: Program-aided language models,” arXiv:2211.10435, 2022. 12

  186. [195]

    Toolformer: Language models can teach themselves to use tools,

    T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” arXiv:2302.04761, 2023. 12

  187. [196]

    Webgpt: Browser-assisted question-answering with human feedback,

    R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V . Kosaraju, W. Saunderset al., “Webgpt: Browser-assisted question-answering with human feedback,” arXiv:2112.09332,

  188. [198]

    Idealgpt: Iteratively decompos- ing vision and language reasoning via large language models,

    H. You, R. Sun, Z. Wang, L. Chen, G. Wang, H. A. Ayyubi, K.-W. Chang, and S.-F. Chang, “Idealgpt: Iteratively decompos- ing vision and language reasoning via large language models,” arXiv:2305.14985, 2023. 12, 13

  189. [199]

    Sus-x: Training- free name-only transfer of vision-language models,

    V . Udandarao, A. Gupta, and S. Albanie, “Sus-x: Training- free name-only transfer of vision-language models,” arXiv:2211.16198, 2022. 12

  190. [200]

    Point- clip v2: Adapting clip for powerful 3d open-world learning,

    X. Zhu, R. Zhang, B. He, Z. Zeng, S. Zhang, and P . Gao, “Point- clip v2: Adapting clip for powerful 3d open-world learning,” arXiv:2211.11682, 2022. 13

  191. [201]

    Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners,

    R. Zhang, X. Hu, B. Li, S. Huang, H. Deng, Y. Qiao, P . Gao, and H. Li, “Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners,” in CVPR, 2023. 13

  192. [202]

    Bottom-up and top-down attention for image captioning and visual question answering,

    P . Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR, 2018. 13

  193. [203]

    Deep modular co- attention networks for visual question answering,

    Z. Yu, J. Yu, Y. Cui, D. Tao, and Q. Tian, “Deep modular co- attention networks for visual question answering,” in CVPR,

  194. [204]

    Dynamic fusion with intra-and inter-modality attention flow for visual question answering,

    P . Gao, Z. Jiang, H. You, P . Lu, S. C. Hoi, X. Wang, and H. Li, “Dynamic fusion with intra-and inter-modality attention flow for visual question answering,” in CVPR, 2019. 13

  195. [205]

    Socratic models: Composing zero-shot multimodal reasoning with language,

    A. Zeng, A. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. Ryoo, V . Sindhwani, J. Lee, V . Vanhoucke et al. , “Socratic models: Composing zero-shot multimodal reasoning with language,” arXiv:2204.00598, 2022. 13

  196. [206]

    Stylenet: Generating attractive visual captions with styles,

    C. Gan, Z. Gan, X. He, J. Gao, and L. Deng, “Stylenet: Generating attractive visual captions with styles,” in CVPR, 2017. 13

  197. [207]

    Senticap: Generating image descriptions with sentiments,

    A. Mathews, L. Xie, and X. He, “Senticap: Generating image descriptions with sentiments,” in AAAI, 2016. 13

  198. [208]

    V*: Guided visual search as a core mechanism in multimodal llms,

    P . Wu and S. Xie, “V*: Guided visual search as a core mechanism in multimodal llms,” arXiv:2312.14135, 2023. 13

  199. [209]

    Least-to- most prompting enables complex reasoning in large language models,

    D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, O. Bousquet, Q. Le, and E. Chi, “Least-to- most prompting enables complex reasoning in large language models,” arXiv:2205.10625, 2022. 13

  200. [210]

    On evaluating adversarial robustness of large vision-language models,

    Y. Zhao, T. Pang, C. Du, X. Yang, C. Li, N.-M. Cheung, and M. Lin, “On evaluating adversarial robustness of large vision-language models,” arXiv:2305.16934, 2023. 14

  201. [211]

    Jailbreak in pieces: Compositional adversarial attacks on multi-modal lan- guage models,

    E. Shayegani, Y. Dong, and N. Abu-Ghazaleh, “Jailbreak in pieces: Compositional adversarial attacks on multi-modal lan- guage models,” in ICLR, 2023. 14

Pith tools

Reviewed May 16, 2026 · model on record in the stance chip above.