Pith. sign in

Can MLLMs Perform Text-to-Image In-Context Learning?

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The evolution from Large Language Models (LLMs) to Multimodal Large Language Models (MLLMs) has spurred research into extending In-Context Learning (ICL) to its multimodal counterpart. Existing such studies have primarily concentrated on image-to-text ICL. However, the Text-to-Image ICL (T2I-ICL), with its unique characteristics and potential applications, remains underexplored. To address this gap, we formally define the task of T2I-ICL and present CoBSAT, the first T2I-ICL benchmark dataset, encompassing ten tasks. Utilizing our dataset to benchmark six state-of-the-art MLLMs, we uncover considerable difficulties MLLMs encounter in solving T2I-ICL. We identify the primary challenges as the inherent complexity of multimodality and image generation, and show that strategies such as fine-tuning and Chain-of-Thought prompting help to mitigate these difficulties, leading to notable improvements in performance. Our code and dataset are available at https://github.com/UW-Madison-Lee-Lab/CoBSAT.

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Unified Multimodal Understanding via Byte-Pair Visual Encoding

cs.CV · 2025-06-30 · conditional · novelty 4.0

Priority-guided byte-pair encoding of quantized image patches plus curriculum training yields an 8B discrete-token MLLM competitive with continuous-embedding models on VQA and multimodal benchmarks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Unified Multimodal Understanding via Byte-Pair Visual Encoding cs.CV · 2025-06-30 · conditional · none · ref 6 · internal anchor

    Priority-guided byte-pair encoding of quantized image patches plus curriculum training yields an 8B discrete-token MLLM competitive with continuous-embedding models on VQA and multimodal benchmarks.