Pith. sign in

REVIEW 1 cited by

Mamba-based Segmentation Model for Speaker Diarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.06459 v2 pith:Y33TFIJZ submitted 2024-10-09 cs.SD eess.AS

classification cs.SDeess.AS
keywords diarizationmambacapabilitiesmamba-basedproposedspeakerattention-basedmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mamba is a newly proposed architecture which behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are too limited. In this paper, we propose to assess the potential of Mamba for diarization by comparing the state-of-the-art neural segmentation of the pyannote pipeline with our proposed Mamba-based variant. Mamba's stronger processing capabilities allow usage of longer local windows, which significantly improve diarization quality by making the speaker embedding extraction more reliable. We find Mamba to be a superior alternative to both traditional RNN and the tested attention-based model. Our proposed Mamba-based system achieves state-of-the-art performance on three widely used diarization datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Fine-tuning WavLM on the diarization task before structured pruning yields 80% parameter removal at nearly unchanged diarization error, with 2.6x to 4x faster GPU inference.

Pith tools