Pith. sign in

REVIEW 6 cited by

Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.13849 v2 pith:BCVXU3JH submitted 2024-11-21 eess.AS

classification eess.AS
keywords speakerdiarizationaudiodetectionrepresentationvoiceactivitiesblock
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes a novel Sequence-to-Sequence Neural Diarization (S2SND) framework to perform online and offline speaker diarization. It is developed from the sequence-to-sequence architecture of our previous target-speaker voice activity detection system and then evolves into a new diarization paradigm by addressing two critical problems. 1) Speaker Detection: The proposed approach can utilize partially given speaker embeddings to discover the unknown speaker and predict the target voice activities in the audio signal. It does not require a prior diarization system for speaker enrollment in advance. 2) Speaker Representation: The proposed approach can adopt the predicted voice activities as reference information to extract speaker embeddings from the audio signal simultaneously. The representation space of speaker embedding is jointly learned within the whole diarization network without using an extra speaker embedding model. During inference, the S2SND framework can process long audio recordings blockwise. The detection module utilizes the previously obtained speaker-embedding buffer to predict both enrolled and unknown speakers' voice activities for each coming audio block. Next, the speaker-embedding buffer is updated according to the predictions of the representation module. Assuming that up to one new speaker may appear in a small block shift, our model iteratively predicts the results of each block and extracts target embeddings for the subsequent blocks until the signal ends. Finally, the last speaker-embedding buffer can re-score the entire audio, achieving highly accurate diarization performance as an offline system. Experimental results show that ...

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering

    eess.AS 2025-07 conditional novelty 6.0 of 10

    A streaming Sortformer with an arrival-ordered speaker cache achieves lower diarization error than prior online systems on DIHARD III and CALLHOME, even at 0.32 second latency.

  2. Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

    eess.AS 2025-06 conditional novelty 6.0 of 10

    An LLM conditioned on speaker embeddings and utterance time boundaries jointly transcribes and timestamps overlapping multi-speaker speech.

  3. Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A controlled 120-configuration study of EEND-VC speaker diarization finds finetuned WavLM encoders, Conformer/Mamba decoders, and longer chunks give the biggest gains, with the best system reaching state-of-the-art on...

  4. The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

    cs.SD 2025-05 conditional novelty 5.0 of 10

    The top MISP 2025 systems achieve DER 8.09%, CER 9.48%, and cpCER 11.56%, far outperforming the provided audio-visual baselines.

  5. The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge

    eess.AS 2025-07 conditional novelty 4.0 of 10

    A challenge system combining speaker diarization, speaker embeddings, and a Qwen2.5 LLM adapter architecture reports 18.08% tcpWER on multilingual multi-speaker ASR, far below the 60.39% baseline.

  6. Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge

    eess.AS 2025-05 conditional novelty 4.0 of 10

    Extending S2SND with a channel-attention module for multi-channel audio achieves an 8.09% diarization error rate, first place in the MISP 2025 speaker diarization task.

Pith tools