Graph attention refinement plus overlapping label propagation yields a reported 15.94% DER on DIHARD-III without oracle VAD, though internal configuration inconsistencies and missing error bars make the SOTA claim provisional.
EEND-M2F: Masked-attention mask transformers for speaker diarization
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this paper, we make the explicit connection between image segmentation methods and end-to-end diarization methods. From these insights, we propose a novel, fully end-to-end diarization model, EEND-M2F, based on the Mask2Former architecture. Speaker representations are computed in parallel using a stack of transformer decoders, in which irrelevant frames are explicitly masked from the cross attention using predictions from previous layers. EEND-M2F is lightweight, efficient, and truly end-to-end, as it does not require any additional diarization, speaker verification, or segmentation models to run, nor does it require running any clustering algorithms. Our model achieves state-of-the-art performance on several public datasets, such as AMI, AliMeeting and RAMC. Most notably our DER of 16.07% on DIHARD-III is the first major improvement upon the challenge winning system.
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm
Graph attention refinement plus overlapping label propagation yields a reported 15.94% DER on DIHARD-III without oracle VAD, though internal configuration inconsistencies and missing error bars make the SOTA claim provisional.