Pith. sign in

REVIEW 4 cited by

Towards Multimodal In-Context Learning for Vision & Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12736 v2 pith:AG6DKBYV submitted 2024-03-19 cs.CV

classification cs.CV
keywords modelslanguagelearningvisionvlmsabilitybenchmarksdemonstrations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

State-of-the-art Vision-Language Models (VLMs) ground the vision and the language modality primarily via projecting the vision tokens from the encoder to language-like tokens, which are directly fed to the Large Language Model (LLM) decoder. While these models have shown unprecedented performance in many downstream zero-shot tasks (eg image captioning, question answers, etc), still little emphasis has been put on transferring one of the core LLM capability of In-Context Learning (ICL). ICL is the ability of a model to reason about a downstream task with a few examples demonstrations embedded in the prompt. In this work, through extensive evaluations, we find that the state-of-the-art VLMs somewhat lack the ability to follow ICL instructions. In particular, we discover that even models that underwent large-scale mixed modality pre-training and were implicitly guided to make use of interleaved image and text information (intended to consume helpful context from multiple images) under-perform when prompted with few-shot demonstrations (in an ICL way), likely due to their lack of direct ICL instruction tuning. To enhance the ICL abilities of the present VLM, we propose a simple yet surprisingly effective multi-turn curriculum-based learning methodology with effective data mixes, leading up to a significant 21.03% (and 11.3% on average) ICL performance boost over the strongest VLM baselines and a variety of ICL benchmarks. Furthermore, we also contribute new benchmarks for ICL evaluation in VLMs and discuss their advantages over the prior art.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    BANGLAWILD is the first in-the-wild Bengali scene text benchmark with dual verbatim/standard labels, and its evaluation shows visual mis-recognition dominates errors while conjunct-related errors are nearly closed.

  2. Generalizable Object Re-Identification via Visual In-Context Prompting

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VICP uses an LLM to generate per-category visual prompts for a frozen DINOv2, enabling few-shot generalization to unseen object categories in re-identification without parameter updates.

  3. True Multimodal In-Context Learning Needs Attention to the Visual Context

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 160-parameter attention-scaling method, DARA, improves true multimodal in-context learning on a new dataset, TrueMICL, that forces models to use demo images rather than copy text patterns.

  4. HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A pipeline that converts ultrasound textbooks and reports into 45,000 reasoning QA/VQA samples improves a 7B multimodal model on self-built ultrasound benchmarks.

Pith tools