Pith. sign in

REVIEW 2 cited by

Neural Speaker Diarization with Speaker-Wise Chain Rule

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.01796 v1 pith:4H3TCR2H submitted 2020-06-02 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords diarizationmethodnumberspeakerspeakersproposedvariableaudio
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speaker diarization is an essential step for processing multi-speaker audio. Although an end-to-end neural diarization (EEND) method achieved state-of-the-art performance, it is limited to a fixed number of speakers. In this paper, we solve this fixed number of speaker issue by a novel speaker-wise conditional inference method based on the probabilistic chain rule. In the proposed method, each speaker's speech activity is regarded as a single random variable, and is estimated sequentially conditioned on previously estimated other speakers' speech activities. Similar to other sequence-to-sequence models, the proposed method produces a variable number of speakers with a stop sequence condition. We evaluated the proposed method on multi-speaker audio recordings of a variable number of speakers. Experimental results show that the proposed method can correctly produce diarization results with a variable number of speakers and outperforms the state-of-the-art end-to-end speaker diarization methods in terms of diarization error rate.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering

    eess.AS 2025-07 conditional novelty 6.0 of 10

    A streaming Sortformer with an arrival-ordered speaker cache achieves lower diarization error than prior online systems on DIHARD III and CALLHOME, even at 0.32 second latency.

  2. Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm

    cs.SD 2025-06 conditional novelty 4.0 of 10

    Graph attention refinement plus overlapping label propagation yields a reported 15.94% DER on DIHARD-III without oracle VAD, though internal configuration inconsistencies and missing error bars make the SOTA claim pro...

Pith tools