WavFusion fuses wav2vec 2.0 audio features with text and visual features using gated cross-modal attention and a homogeneous-feature margin loss, reporting modest state-of-the-art gains on IEMOCAP and MELD.
In Proceedings of the AAAI conference on artificial intelligence, volume 30 (2016)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
WavFusion fuses wav2vec 2.0 audio features with text and visual features using gated cross-modal attention and a homogeneous-feature margin loss, reporting modest state-of-the-art gains on IEMOCAP and MELD.