Pith. sign in

REVIEW 2 cited by

Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.07272 v2 pith:LJN3FVKF submitted 2020-05-14 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords diarizationspeakerts-vadactivityapproachclustering-baseddetectioni-vectors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainly due to the limited ability to handle overlapping speech. We propose a novel Target-Speaker Voice Activity Detection (TS-VAD) approach, which directly predicts an activity of each speaker on each time frame. TS-VAD model takes conventional speech features (e.g., MFCC) along with i-vectors for each speaker as inputs. A set of binary classification output layers produces activities of each speaker. I-vectors can be estimated iteratively, starting with a strong clustering-based diarization. We also extend the TS-VAD approach to the multi-microphone case using a simple attention mechanism on top of hidden representations extracted from the single-channel TS-VAD model. Moreover, post-processing strategies for the predicted speaker activity probabilities are investigated. Experiments on the CHiME-6 unsegmented data show that TS-VAD achieves state-of-the-art results outperforming the baseline x-vector-based system by more than 30% Diarization Error Rate (DER) abs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Conditioning an SOT multi-talker ASR decoder on EEND-EDA speaker embeddings and activity information lowers WER on Libri2Mix and Libri3Mix, provided the diarization branch is accurate.

  2. Cross-attention and Self-attention for Audio-visual Speaker Diarization in MISP-Meeting Challenge

    cs.SD 2025-06 conditional novelty 4.0 of 10

    An audio-visual speaker diarization system using cross-attention and self-attention fusion achieves an 8.18% diarization error rate on the MISP 2025 Challenge, a 47.3% relative improvement over the baseline.

Pith tools