A multimodal retriever aligns images to text by allowing text tokens to attend to visual patches while excluding text embeddings from the trained visual representation, then combines both at scoring time.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
A multimodal retriever aligns images to text by allowing text tokens to attend to visual patches while excluding text embeddings from the trained visual representation, then combines both at scoring time.