Concept-centric short captions and cross-modal attention pooling yield SOTA compositionality in contrastive V&L models without degrading zero-shot or retrieval performance.
Object-centric binding in contrastive language-image pretraining
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 3years
2026 3verdicts
UNVERDICTED 3representative citing papers
Introduces an information-theoretic formalization of the binding problem and a probing method to quantify binding information in deep learning model representations, tested on ViTs across challenging datasets.
Training VLMs to point via text induces serial processing that eliminates binding errors and enables compositional generalization on multi-object tasks.
citing papers explorer
-
No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models
Concept-centric short captions and cross-modal attention pooling yield SOTA compositionality in contrastive V&L models without degrading zero-shot or retrieval performance.
-
Formalizing the Binding Problem
Introduces an information-theoretic formalization of the binding problem and a probing method to quantify binding information in deep learning model representations, tested on ViTs across challenging datasets.
-
Binding Visual Features Point by Point
Training VLMs to point via text induces serial processing that eliminates binding errors and enables compositional generalization on multi-object tasks.