Pith. sign in

CREPE: Can Vision-Language Foundation Models Reason Compositionally?

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it
abstract

A fundamental characteristic common to both human vision and natural language is their compositional nature. Yet, despite the performance gains contributed by large vision and language pretraining, we find that: across 7 architectures trained with 4 algorithms on massive datasets, they struggle at compositionality. To arrive at this conclusion, we introduce a new compositionality evaluation benchmark, CREPE, which measures two important aspects of compositionality identified by cognitive science literature: systematicity and productivity. To measure systematicity, CREPE consists of a test dataset containing over $370K$ image-text pairs and three different seen-unseen splits. The three splits are designed to test models trained on three popular training datasets: CC-12M, YFCC-15M, and LAION-400M. We also generate $325K$, $316K$, and $309K$ hard negative captions for a subset of the pairs. To test productivity, CREPE contains $17K$ image-text pairs with nine different complexities plus $183K$ hard negative captions with atomic, swapping and negation foils. The datasets are generated by repurposing the Visual Genome scene graphs and region descriptions and applying handcrafted templates and GPT-3. For systematicity, we find that model performance decreases consistently when novel compositions dominate the retrieval set, with Recall@1 dropping by up to $12\%$. For productivity, models' retrieval success decays as complexity increases, frequently nearing random chance at high complexity. These results hold regardless of model and training dataset size.

fields

cs.CL 1 cs.CV 1

years

2026 1 2025 1

verdicts

CONDITIONAL 2

representative citing papers

How Do Language Models Compose Functions?

cs.CL · 2025-10-02 · conditional · novelty 6.0

LLMs solve compositional factual recall either by computing intermediates or directly, with mechanism choice correlated to translation geometry in embedding spaces.

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP

cs.CV · 2026-07-08 · conditional · novelty 5.5

Inference-time α-entmax sparsification of CLIP’s final self-attention denoises diffuse mass and lifts dense open-vocabulary segmentation and region retrieval in proportion to baseline off-class spread.

citing papers explorer

Showing 2 of 2 citing papers.

  • How Do Language Models Compose Functions? cs.CL · 2025-10-02 · conditional · none · ref 25

    LLMs solve compositional factual recall either by computing intermediates or directly, with mechanism choice correlated to translation geometry in embedding spaces.

  • Sparse Attention for Dense Open-Vocabulary Prediction in CLIP cs.CV · 2026-07-08 · conditional · none · ref 14 · internal anchor

    Inference-time α-entmax sparsification of CLIP’s final self-attention denoises diffuse mass and lifts dense open-vocabulary segmentation and region retrieval in proportion to baseline off-class spread.