Pith. sign in

REVIEW 7 cited by

What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.20482 v1 pith:MBCAYKJO submitted 2024-10-27 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords mm-icldemonstrationmulti-modalorderingaffectfactorsin-contextinvestigate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, rapid advancements in Multi-Modal In-Context Learning (MM-ICL) have achieved notable success, which is capable of achieving superior performance across various tasks without requiring additional parameter tuning. However, the underlying rules for the effectiveness of MM-ICL remain under-explored. To fill this gap, this work aims to investigate the research question: "What factors affect the performance of MM-ICL?'' To this end, we investigate extensive experiments on the three core steps of MM-ICL including demonstration retrieval, demonstration ordering, and prompt construction using 6 vision large language models and 20 strategies. Our findings highlight (1) the necessity of a multi-modal retriever for demonstration retrieval, (2) the importance of intra-demonstration ordering over inter-demonstration ordering, and (3) the enhancement of task comprehension through introductory instructions in prompts. We hope this study can serve as a foundational guide for optimizing MM-ICL strategies in future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL

    cs.CV 2026-08 conditional novelty 6.0 of 10

    The Selection-Realization Hypothesis holds that static task vectors suffice when demonstration-induced activation changes are largely shared across queries; query-conditioned, multi-site, or routing interventions are ...

  2. MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new 2,700-task benchmark with budget, time, and distance constraints shows that even the best multimodal LLMs produce feasible plans less than 22% of the time.

  3. True Multimodal In-Context Learning Needs Attention to the Visual Context

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 160-parameter attention-scaling method, DARA, improves true multimodal in-context learning on a new dataset, TrueMICL, that forces models to use demo images rather than copy text patterns.

  4. Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

    cs.AI 2025-01 conditional novelty 6.0 of 10

    Concept shift vectors, computed as mean activation differences, can partially recover fine-tuned multimodal LLM concepts and steer model outputs without additional training.

  5. CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    CoMT is the first benchmark to ask LVLMs to produce interleaved image and text rationales, and current models perform near random on it.

  6. Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Manager aggregates multi-layer unimodal representations and improves both two-tower VLMs (ManagerTower) and MLLMs (LLaVA-OV-Manager) on 24 downstream tasks.

  7. Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Vision-language models improve little, often not at all, when given demonstrations, even when demonstrations contain explicit reasoning steps.

Pith tools