A question-conditioned frequency-domain fusion module adds about 3 to 4 exact-match points over the authors' own non-frequency baseline on two medical VQA benchmarks, but the full model remains far below published state of the art.
Multi-Modal Masked Autoencoders for Medical Vision-and-Language Pre-Training
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Frequency-Domain Dual-Branch Fusion for Medical Visual Question Answering
A question-conditioned frequency-domain fusion module adds about 3 to 4 exact-match points over the authors' own non-frequency baseline on two medical VQA benchmarks, but the full model remains far below published state of the art.