Pith. sign in

REVIEW 3 cited by

Powerset multi-class cross entropy loss for neural speaker diarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.13025 v1 pith:O65XJBVQ submitted 2023-10-19 cs.SD cs.AIcs.CLcs.NEeess.AS

classification cs.SDcs.AIcs.CLcs.NEeess.AS
keywords diarizationmulti-labeleendformulationclassificationmulti-classneuraloverlapping
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Since its introduction in 2019, the whole end-to-end neural diarization (EEND) line of work has been addressing speaker diarization as a frame-wise multi-label classification problem with permutation-invariant training. Despite EEND showing great promise, a few recent works took a step back and studied the possible combination of (local) supervised EEND diarization with (global) unsupervised clustering. Yet, these hybrid contributions did not question the original multi-label formulation. We propose to switch from multi-label (where any two speakers can be active at the same time) to powerset multi-class classification (where dedicated classes are assigned to pairs of overlapping speakers). Through extensive experiments on 9 different benchmarks, we show that this formulation leads to significantly better performance (mostly on overlapping speech) and robustness to domain mismatch, while eliminating the detection threshold hyperparameter, critical for the multi-label formulation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FT-Boosted SV: Towards Noise Robust Speaker Verification for English Speaking Classroom Environments

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Fine-tuning pretrained speaker verification models on augmented children's speech reduces error rates in English-speaking classrooms, with ECAPA-TDNN nearly halving error on the MPT dataset.

  2. CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning

    cs.SD 2025-05 reject novelty 5.0 of 10

    A universal adversarial perturbation framework claiming to protect speech against zero-shot voice cloning by degrading cloned outputs while preserving input naturalness.

  3. Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm

    cs.SD 2025-06 conditional novelty 4.0 of 10

    Graph attention refinement plus overlapping label propagation yields a reported 15.94% DER on DIHARD-III without oracle VAD, though internal configuration inconsistencies and missing error bars make the SOTA claim pro...

Pith tools