Pith. sign in

REVIEW 1 cited by

A Review of Speaker Diarization: Recent Advances with Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.09624 v4 pith:YQNZV7ON submitted 2021-01-24 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords speakerdiarizationrecentaudiodeeplearningspeechadvancements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speaker diarization is a task to label audio or video recordings with classes that correspond to speaker identity, or in short, a task to identify "who spoke when". In the early years, speaker diarization algorithms were developed for speech recognition on multispeaker audio recordings to enable speaker adaptive processing. These algorithms also gained their own value as a standalone application over time to provide speaker-specific metainformation for downstream tasks such as audio retrieval. More recently, with the emergence of deep learning technology, which has driven revolutionary changes in research and practices across speech application domains, rapid advancements have been made for speaker diarization. In this paper, we review not only the historical development of speaker diarization technology but also the recent advancements in neural speaker diarization approaches. Furthermore, we discuss how speaker diarization systems have been integrated with speech recognition applications and how the recent surge of deep learning is leading the way of jointly modeling these two components to be complementary to each other. By considering such exciting technical trends, we believe that this paper is a valuable contribution to the community to provide a survey work by consolidating the recent developments with neural methods and thus facilitating further progress toward a more efficient speaker diarization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DNCASR: End-to-End Training for Speaker-Attributed ASR

    eess.AS 2025-06 conditional novelty 5.0 of 10

    DNCASR links speaker clustering and ASR decoders with cross-attention, achieving a 9.0% relative cpWER reduction on AMI-MDM Eval over a parallel (unlinked) system.

Pith tools