IMAGINE uses adaptive schema-imagery via dynamic multimodal prototypes to incorporate implicit semantics into composed video retrieval, claiming SOTA results on CVR and CIR benchmarks.
Sentence-level prompts benefit composed image retrieval
3 Pith papers cite this work. Polarity classification is still indexing.
3
Pith papers citing it
years
2026 3representative citing papers
Framework uses LLaVA for triplet generation and two-stage fine-tuning to enhance composed fashion image retrieval.
citing papers explorer
-
IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval
IMAGINE uses adaptive schema-imagery via dynamic multimodal prototypes to incorporate implicit semantics into composed video retrieval, claiming SOTA results on CVR and CIR benchmarks.
-
Exploring Multi-Modal Large Language Models and Two-Stage Fine-Tuning for Fashion Image Retrieval
Framework uses LLaVA for triplet generation and two-stage fine-tuning to enhance composed fashion image retrieval.
- Adjustable Text-Guided Backdoor Attacks with Natural-Word Triggers on Multimodal Pretrained Models