CR-CLIP adds cross-modal attention and multi-scale test augmentation to an AVION baseline, reporting 66.8% mAP and 82.1% nDCG, first place on EPIC-KITCHENS-100 multi-instance retrieval 2025.
A survey on vision transformer
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
ContextRefine-CLIP for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2025
CR-CLIP adds cross-modal attention and multi-scale test augmentation to an AVION baseline, reporting 66.8% mAP and 82.1% nDCG, first place on EPIC-KITCHENS-100 multi-instance retrieval 2025.