A zero-shot pipeline combining SAM2 masks, Alpha-CLIP region embeddings, and Qwen2.5-VL reranking and grounding improves mask-aware text-to-image retrieval on COCO and D3.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval
A zero-shot pipeline combining SAM2 masks, Alpha-CLIP region embeddings, and Qwen2.5-VL reranking and grounding improves mask-aware text-to-image retrieval on COCO and D3.