A cascade of TDOA-based spatial segmentation and speaker-embedding clustering diarizes meetings without multi-channel training data and handles overlapping speech and speaker position changes.
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed system requires neither access to multi-channel training data nor prior knowledge about the number or placement of microphones. It works for both a compact microphone array and distributed microphones, with minor adjustments. Due to its superior handling of overlapping speech during segmentation, the proposed pipeline significantly outperforms the single-channel pyannote approach, both in a scenario with a compact microphone array and in a setup with distributed microphones. Additionally, we show that, unlike fully spatial diarization pipelines, the proposed system can correctly track speakers when they change positions.
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
A cascade of TDOA-based spatial segmentation and speaker-embedding clustering diarizes meetings without multi-channel training data and handles overlapping speech and speaker position changes.