WavFusion fuses wav2vec 2.0 audio features with text and visual features using gated cross-modal attention and a homogeneous-feature margin loss, reporting modest state-of-the-art gains on IEMOCAP and MELD.
In 2017 seventh international conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), pages 79–80 (2017)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
WavFusion fuses wav2vec 2.0 audio features with text and visual features using gated cross-modal attention and a homogeneous-feature margin loss, reporting modest state-of-the-art gains on IEMOCAP and MELD.