Pith. sign in

REVIEW 9 cited by

MoVA: Adapting Mixture of Vision Experts to Multimodal Context

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.13046 v2 pith:UOOQVITQ submitted 2024-04-19 cs.CV

classification cs.CV
keywords visionexpertsencoderimagemultimodalunderstandingabilityclip
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As the key component in multimodal large language models (MLLMs), the ability of the visual encoder greatly affects MLLM's understanding on diverse image content. Although some large-scale pretrained vision encoders such as vision encoders in CLIP and DINOv2 have brought promising performance, we found that there is still no single vision encoder that can dominate various image content understanding, e.g., the CLIP vision encoder leads to outstanding results on general image understanding but poor performance on document or chart content. To alleviate the bias of CLIP vision encoder, we first delve into the inherent behavior of different pre-trained vision encoders and then propose the MoVA, a powerful and novel MLLM, adaptively routing and fusing task-specific vision experts with a coarse-to-fine mechanism. In the coarse-grained stage, we design a context-aware expert routing strategy to dynamically select the most suitable vision experts according to the user instruction, input image, and expertise of vision experts. This benefits from the powerful model function understanding ability of the large language model (LLM). In the fine-grained stage, we elaborately conduct the mixture-of-vision-expert adapter (MoV-Adapter) to extract and fuse task-specific knowledge from various experts. This coarse-to-fine paradigm effectively leverages representations from experts based on multimodal context and model expertise, further enhancing the generalization ability. We conduct extensive experiments to evaluate the effectiveness of the proposed approach. Without any bells and whistles, MoVA can achieve significant performance gains over current state-of-the-art methods in a wide range of challenging multimodal benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A curated GPT-4o synthetic image dataset improves open-source generation models on instruction-following, surreal scenes, and multi-reference synthesis, plus two new benchmarks to measure those skills.

  2. M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks

    cs.NI 2025-08 reject novelty 6.0 of 10

    M3LLM routes each multimodal query to the semantically most suitable, wirelessly reachable vision expert using protocol-aided retrieval and a decoupled reinforcement learning agent.

  3. Multilingual Multimodal Software Developer for Code Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 7B vision-language model trained on synthetic diagram-to-code data outperforms several larger open-weight models on a new 10-language UML/flowchart code-generation benchmark.

  4. AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AutoV selects instance- and query-specific visual prompts via loss-based pairwise ranking, consistently improving LVLMs across many benchmarks with no backbone fine-tuning.

  5. GenRecal: Generation after Recalibration from Large to Small Vision-Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A learnable Recalibrator bridges different tokenizers so that small VLMs can distill knowledge from any large VLM, improving their benchmark scores.

  6. STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset

    cs.CV 2025-06 conditional novelty 6.0 of 10

    STORM is a new multi-domain ordinal-regression benchmark with coarse-to-fine Chain-of-Thought prompts that improves MLLM zero-shot visual rating, though the 'universal' claim is bounded by its five curated domains.

  7. Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    F3 adds attention-guided noise to adversarial images so that large vision-language models produce answers that are much closer to their clean-image answers.

  8. MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MINT-CoT-7B interleaves fine-grained visual tokens into each math reasoning step and reports 73.70 on MathVista-Math, 64.72 on GeoQA, and 69.6 on MMStar-Math.

  9. EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models

    cs.CL 2025-05 conditional novelty 4.0 of 10

    EvoMoE creates MoE experts as decaying averages of a single trained FFN and routes tokens with hypernetwork-generated weights, yielding small benchmark gains over MoE-LLaVA.

Pith tools