Pith. sign in

REVIEW 3 cited by

Does CLIP Bind Concepts? Probing Compositionality in Large Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.10537 v3 pith:U2YU227R submitted 2022-12-20 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords clipcompositionalconceptscubemodelsperformancebehindbind
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale neural network models combining text and images have made incredible progress in recent years. However, it remains an open question to what extent such models encode compositional representations of the concepts over which they operate, such as correctly identifying "red cube" by reasoning over the constituents "red" and "cube". In this work, we focus on the ability of a large pretrained vision and language model (CLIP) to encode compositional concepts and to bind variables in a structure-sensitive way (e.g., differentiating "cube behind sphere" from "sphere behind cube"). To inspect the performance of CLIP, we compare several architectures from research on compositional distributional semantics models (CDSMs), a line of research that attempts to implement traditional compositional linguistic structures within embedding spaces. We benchmark them on three synthetic datasets - single-object, two-object, and relational - designed to test concept binding. We find that CLIP can compose concepts in a single-object setting, but in situations where concept binding is needed, performance drops dramatically. At the same time, CDSMs also perform poorly, with best performance at chance level.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    The accuracy of CLIP and CLIP-based visual question answering models is strongly correlated with how often the concept pair in an image appears together in pretraining captions.

  2. Position: We Need An Algorithmic Understanding of Generative AI

    cs.AI 2025-07 conditional novelty 5.0 of 10

    The paper argues for a systematic algorithmic understanding of LLMs and presents a case study suggesting that Llama models do not implement BFS or DFS on graph navigation tasks.

  3. On the rankability of visual embeddings

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Visual embeddings from CLIP and other vision encoders encode ordinal attributes along linear directions, recoverable from as few as two extreme reference images, without full supervision.

Pith tools