Pith. sign in

REVIEW 4 cited by

CREPE: Can Vision-Language Foundation Models Reason Compositionally?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.07796 v3 pith:SNACBLR3 submitted 2022-12-13 cs.CL cs.CV

classification cs.CLcs.CV
keywords crepecompositionalitydatasetsmodelspairsproductivitysystematicitytest
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

A fundamental characteristic common to both human vision and natural language is their compositional nature. Yet, despite the performance gains contributed by large vision and language pretraining, we find that: across 7 architectures trained with 4 algorithms on massive datasets, they struggle at compositionality. To arrive at this conclusion, we introduce a new compositionality evaluation benchmark, CREPE, which measures two important aspects of compositionality identified by cognitive science literature: systematicity and productivity. To measure systematicity, CREPE consists of a test dataset containing over $370K$ image-text pairs and three different seen-unseen splits. The three splits are designed to test models trained on three popular training datasets: CC-12M, YFCC-15M, and LAION-400M. We also generate $325K$, $316K$, and $309K$ hard negative captions for a subset of the pairs. To test productivity, CREPE contains $17K$ image-text pairs with nine different complexities plus $183K$ hard negative captions with atomic, swapping and negation foils. The datasets are generated by repurposing the Visual Genome scene graphs and region descriptions and applying handcrafted templates and GPT-3. For systematicity, we find that model performance decreases consistently when novel compositions dominate the retrieval set, with Recall@1 dropping by up to $12\%$. For productivity, models' retrieval success decays as complexity increases, frequently nearing random chance at high complexity. These results hold regardless of model and training dataset size.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Punching Bag vs. Punching Person: Motion Transferability in Videos

    cs.CV 2025-07 conditional novelty 7.0 of 10

    State-of-the-art video action recognition models fail to transfer high-level motion concepts to unseen contexts, even when the new context is only a minor variation of the training data.

  2. LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.

  3. Multi-Modal Language Models as Text-to-Image Model Evaluators

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MT2IE uses a single open-source multimodal LLM to generate 20 progressively harder prompts and score image-text consistency, reproducing the 1,600-prompt GenAIBench ranking of 8 text-to-image models.

  4. Sparse Attention for Dense Open-Vocabulary Prediction in CLIP

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Replacing softmax with α-entmax in frozen CLIP's final attention layers denoises dense predictions by zeroing irrelevant token interactions, with gains proportional to baseline attention diffuseness.

Pith tools