A question-conditioned frequency-domain fusion module adds about 3 to 4 exact-match points over the authors' own non-frequency baseline on two medical VQA benchmarks, but the full model remains far below published state of the art.
Masked Vision and Language Pre-training with Unimodal and Multimodal Contrastive Losses for Medical Visual Question Answering
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Frequency-Domain Dual-Branch Fusion for Medical Visual Question Answering
A question-conditioned frequency-domain fusion module adds about 3 to 4 exact-match points over the authors' own non-frequency baseline on two medical VQA benchmarks, but the full model remains far below published state of the art.