Pith. sign in

REVIEW 1 cited by

Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09142 v1 pith:GCFXRDZ7 submitted 2024-05-15 eess.AS cs.SD

classification eess.AScs.SD
keywords speakerdiarizationembeddingmodelsupervisedactivityattentioncurrent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current speaker diarization systems rely on an external voice activity detection model prior to speaker embedding extraction on the detected speech segments. In this paper, we establish that the attention system of a speaker embedding extractor acts as a weakly supervised internal VAD model and performs equally or better than comparable supervised VAD systems. Subsequently, speaker diarization can be performed efficiently by extracting the VAD logits and corresponding speaker embedding simultaneously, alleviating the need and computational overhead of an external VAD model. We provide an extensive analysis of the behavior of the frame-level attention system in current speaker verification models and propose a novel speaker diarization pipeline using ECAPA2 speaker embeddings for both VAD and embedding extraction. The proposed strategy gains state-of-the-art performance on the AMI, VoxConverse and DIHARD III diarization benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion

    cs.SD 2025-06 conditional novelty 5.0 of 10

    FusionVAD shows that simple addition or concatenation of MFCC and pre-trained model features outperforms cross-attention fusion for voice activity detection, with the best model beating Pyannote by 2.04 average DER.

Pith tools