A decoder with modal-specialized expert groups and a learned two-level router improves noisy audio-visual speech recognition on LRS3 and MuAViC, using roughly half the active parameter count of the dense baseline.
Xls-r: Self-supervised cross-lingual speech representation learning at scale
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition
A decoder with modal-specialized expert groups and a learned two-level router improves noisy audio-visual speech recognition on LRS3 and MuAViC, using roughly half the active parameter count of the dense baseline.