Pith. sign in

REVIEW 10 cited by

MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.07915 v3 pith:ZHEFC2KK submitted 2023-09-14 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords multi-modallearningin-contextmmiclvision-languagevlmscomplexability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Since the resurgence of deep learning, vision-language models (VLMs) enhanced by large language models (LLMs) have grown exponentially in popularity. However, while LLMs can utilize extensive background knowledge and task information with in-context learning, most VLMs still struggle with understanding complex multi-modal prompts with multiple images, making VLMs less effective in downstream vision-language tasks. In this paper, we address the limitation above by 1) introducing vision-language Model with Multi-Modal In-Context Learning(MMICL), a new approach to allow the VLM to deal with multi-modal inputs efficiently; 2) proposing a novel context scheme to augment the in-context learning ability of the VLM; 3) constructing the Multi-modal In-Context Learning (MIC) dataset, designed to enhance the VLM's ability to understand complex multi-modal prompts. Our experiments confirm that MMICL achieves new state-of-the-art zero-shot performance on a wide range of general vision-language tasks, especially for complex benchmarks, including MME and MMBench. Our analysis demonstrates that MMICL effectively tackles the challenge of complex multi-modal prompt understanding and emerges the impressive ICL ability. Furthermore, we observe that MMICL successfully alleviates language bias in VLMs, a common issue for VLMs that often leads to hallucination when faced with extensive textual context. Our code, dataset, dataset tool, and model are available at https://github.com/PKUnlp-icler/MIC

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MentalThink: Shaping Thoughts in Mental SVG World

    cs.AI 2026-07 conditional novelty 7.0 of 10

    MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.

  2. UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A six-level capability taxonomy plus UniICL-760K and a lightweight CAPM module improve unified multimodal few-shot learning and beat larger MLLMs on most understanding ICL tasks.

  3. EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A self-evolving multimodal model using continuous self-consistency rewards improves math reasoning by about 2–3% using only raw images, without labels or external reward models.

  4. Region-Level Context-Aware Multimodal Understanding

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Region-level context-aware instruction tuning with a large synthetic dataset improves MLLMs' ability to connect objects in images to their textual descriptions.

  5. True Multimodal In-Context Learning Needs Attention to the Visual Context

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 160-parameter attention-scaling method, DARA, improves true multimodal in-context learning on a new dataset, TrueMICL, that forces models to use demo images rather than copy text patterns.

  6. DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DisCo assigns each visual token to a unique concept from the caption and aligns its attention across frames, improving video MLLM accuracy and token efficiency.

  7. LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LLaVA-SP adds six spatial tokens produced by multi-scale cropping or pooling and cross-attention to MLLMs, improving 10/11 benchmarks over LLaVA-1.5 with nearly unchanged latency.

  8. Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Embodied VideoAgent augments an LLM-based video agent with persistent object memory built from egocentric video, depth, and pose, plus VLM-based memory updates, reporting gains on Ego4D-VQ3D, OpenEQA, and EnvQA.

  9. Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A single conditional generative model, trained only on RPM-style puzzles, can be repurposed via probability scoring to solve odd-one-out, analogy, and categorization tasks, with modest zero-shot transfer.

  10. Efficiently Enhancing General Agents With Hierarchical-categorical Memory

    cs.AI 2025-05 conditional novelty 4.0 of 10

    EHC, a memory-augmented tool-use agent that retrieves category-specific past experiences, outperforms the CLOVA baseline on GQA, NLVR2, MagicBrush editing, and image tagging without updating model parameters.

Pith tools