A pre-aligned cross-modal retrieval model with a global-local Swin-style image encoder and a similarity-matrix reweighting reranker reports gains on four remote sensing image-text benchmarks.
Textrs: Deep bidirectional triplet network for matching text to remote sensing images,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
A pre-aligned cross-modal retrieval model with a global-local Swin-style image encoder and a similarity-matrix reweighting reranker reports gains on four remote sensing image-text benchmarks.