FoCo learns composition for zero-shot CIR via text-anchored visual aggregation and context-conditioned semantic completion trained jointly with cross-instance contrastive loss, reporting SOTA on four benchmarks.
Prompting large vision-language models for compositional reasoning.arXiv preprint arXiv:2401.11337
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CV 2verdicts
UNVERDICTED 2representative citing papers
COMPACT synthesizes compositional visual instruction data to reduce VIT training data by 90% while achieving 100.2% of full performance across eight multimodal benchmarks.
citing papers explorer
-
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
FoCo learns composition for zero-shot CIR via text-anchored visual aggregation and context-conditioned semantic completion trained jointly with cross-instance contrastive loss, reporting SOTA on four benchmarks.
-
Visual Compositional Tuning
COMPACT synthesizes compositional visual instruction data to reduce VIT training data by 90% while achieving 100.2% of full performance across eight multimodal benchmarks.