MUAN applies a stacked gated self-attention block to concatenated visual and textual tokens, jointly modeling intra-modal and inter-modal attention, and achieves top results on VQA and visual grounding benchmarks.
Multimodal deep network embedding with integrated structure and attribute information,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multimodal Unified Attention Networks for Vision-and-Language Interactions
MUAN applies a stacked gated self-attention block to concatenated visual and textual tokens, jointly modeling intra-modal and inter-modal attention, and achieves top results on VQA and visual grounding benchmarks.