An optical-SAR visual question answering dataset with 6,008 image pairs and 1,036,694 questions is introduced, together with a text-guided fusion network that outperforms baselines on that dataset.
Mu- tual attention inception network for remote sensing visual question an- swering
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Text-Guided Coarse-to-Fine Fusion Network for Robust Remote Sensing Visual Question Answering
An optical-SAR visual question answering dataset with 6,008 image pairs and 1,036,694 questions is introduced, together with a text-guided fusion network that outperforms baselines on that dataset.