Modality-fused co-training predicts low-fluidity and low-enjoyment moments in videoconference sessions nearly as well as a fully supervised model while using only 8% of the labels.
While audio-based SSL has been widely applied in speech emotion recognition, multimodal approaches have only recently emerged [5]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience
Modality-fused co-training predicts low-fluidity and low-enjoyment moments in videoconference sessions nearly as well as a fully supervised model while using only 8% of the labels.