Pith. sign in

REVIEW 15 cited by

MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.07915 v3 pith:ZHEFC2KK submitted 2023-09-14 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords multi-modallearningin-contextmmiclvision-languagevlmscomplexability
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Since the resurgence of deep learning, vision-language models (VLMs) enhanced by large language models (LLMs) have grown exponentially in popularity. However, while LLMs can utilize extensive background knowledge and task information with in-context learning, most VLMs still struggle with understanding complex multi-modal prompts with multiple images, making VLMs less effective in downstream vision-language tasks. In this paper, we address the limitation above by 1) introducing vision-language Model with Multi-Modal In-Context Learning(MMICL), a new approach to allow the VLM to deal with multi-modal inputs efficiently; 2) proposing a novel context scheme to augment the in-context learning ability of the VLM; 3) constructing the Multi-modal In-Context Learning (MIC) dataset, designed to enhance the VLM's ability to understand complex multi-modal prompts. Our experiments confirm that MMICL achieves new state-of-the-art zero-shot performance on a wide range of general vision-language tasks, especially for complex benchmarks, including MME and MMBench. Our analysis demonstrates that MMICL effectively tackles the challenge of complex multi-modal prompt understanding and emerges the impressive ICL ability. Furthermore, we observe that MMICL successfully alleviates language bias in VLMs, a common issue for VLMs that often leads to hallucination when faced with extensive textual context. Our code, dataset, dataset tool, and model are available at https://github.com/PKUnlp-icler/MIC

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MentalThink: Shaping Thoughts in Mental SVG World

    cs.AI 2026-07 conditional novelty 7.0 of 10

    MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.

  2. Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

    cs.CV 2024-11 conditional novelty 7.0 of 10

    SAVs extract a sparse set of attention head outputs from a frozen large multimodal model and use them as nearest-centroid features, achieving state-of-the-art few-shot vision-language classification without finetuning.

  3. UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A six-level capability taxonomy plus UniICL-760K and a lightweight CAPM module improve unified multimodal few-shot learning and beat larger MLLMs on most understanding ICL tasks.

  4. EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A self-evolving multimodal model using continuous self-consistency rewards improves math reasoning by about 2–3% using only raw images, without labels or external reward models.

  5. Region-Level Context-Aware Multimodal Understanding

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Region-level context-aware instruction tuning with a large synthetic dataset improves MLLMs' ability to connect objects in images to their textual descriptions.

  6. True Multimodal In-Context Learning Needs Attention to the Visual Context

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 160-parameter attention-scaling method, DARA, improves true multimodal in-context learning on a new dataset, TrueMICL, that forces models to use demo images rather than copy text patterns.

  7. DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DisCo assigns each visual token to a unique concept from the caption and aligns its attention across frames, improving video MLLM accuracy and token efficiency.

  8. LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LLaVA-SP adds six spatial tokens produced by multi-scale cropping or pooling and cross-attention to MLLMs, improving 10/11 benchmarks over LLaVA-1.5 with nearly unchanged latency.

  9. Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Embodied VideoAgent augments an LLM-based video agent with persistent object memory built from egocentric video, depth, and pose, plus VLM-based memory updates, reporting gains on Ego4D-VQ3D, OpenEQA, and EnvQA.

  10. Error-driven Data-efficient Large Multimodal Model Tuning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    An error-driven teacher-student pipeline extracts a student LMM's missing skills from validation mistakes and retrieves targeted samples from a task-agnostic dataset to fine-tune it.

  11. Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Simignore improves multimodal LLM complex question answering on ScienceQA by masking image tokens whose embeddings have low cosine similarity to the text prompt.

  12. Teaching VLMs to Localize Specific Objects from In-context Examples

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Fine-tuning VLMs on video-tracking conversations with made-up object names teaches them to localize a specific object in a new image from only a few in-context examples.

  13. SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SymDPO replaces answer text in multimodal demonstrations with meaningless symbols during preference training, forcing models to use image context and improving in-context learning performance on five benchmarks.

  14. Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A single conditional generative model, trained only on RPM-style puzzles, can be repurposed via probability scoring to solve odd-one-out, analogy, and categorization tasks, with modest zero-shot transfer.

  15. Efficiently Enhancing General Agents With Hierarchical-categorical Memory

    cs.AI 2025-05 conditional novelty 4.0 of 10

    EHC, a memory-augmented tool-use agent that retrieves category-specific past experiences, outperforms the CLOVA baseline on GQA, NLVR2, MagicBrush editing, and image tagging without updating model parameters.

Pith tools