A 0.1M-parameter visual frontend plus an autoregressive acoustic encoder improves online audio-visual speaker extraction by up to 0.9 dB SI-SNRi on simulated LRS3 mixtures, with roughly 0.5 dB attributable to the acoustic encoder alone.
Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Online Audio-Visual Autoregressive Speaker Extraction
A 0.1M-parameter visual frontend plus an autoregressive acoustic encoder improves online audio-visual speaker extraction by up to 0.9 dB SI-SNRi on simulated LRS3 mixtures, with roughly 0.5 dB attributable to the acoustic encoder alone.