SilVar fuses speech, image, and text via CLIP, Whisper, and LLaMA to perform reasoning visual question answering and object localization, but its claimed state-of-the-art results are not supported by its own tables.
Vqa: Visual question answering
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
SilVar fuses speech, image, and text via CLIP, Whisper, and LLaMA to perform reasoning visual question answering and object localization, but its claimed state-of-the-art results are not supported by its own tables.