REVIEW 5 cited by
SPMamba: State-space model is all you need in speech separation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Existing CNN-based speech separation models face local receptive field limitations and cannot effectively capture long time dependencies. Although LSTM and Transformer-based speech separation models can avoid this problem, their high complexity makes them face the challenge of computational resources and inference efficiency when dealing with long audio. To address this challenge, we introduce an innovative speech separation method called SPMamba. This model builds upon the robust TF-GridNet architecture, replacing its traditional BLSTM modules with bidirectional Mamba modules. These modules effectively model the spatiotemporal relationships between the time and frequency dimensions, allowing SPMamba to capture long-range dependencies with linear computational complexity. Specifically, the bidirectional processing within the Mamba modules enables the model to utilize both past and future contextual information, thereby enhancing separation performance. Extensive experiments conducted on public datasets, including WSJ0-2Mix, WHAM!, and Libri2Mix, as well as the newly constructed Echo2Mix dataset, demonstrated that SPMamba significantly outperformed existing state-of-the-art models, achieving superior results while also reducing computational complexity. These findings highlighted the effectiveness of SPMamba in tackling the intricate challenges of speech separation in complex environments.
Forward citations
Cited by 5 Pith papers
-
Cortical-SSM: A Deep State Space Model for Motor Imagery Decoding from EEG Signals
Cortical-SSM, a dual state-space architecture with wavelet-based frequency features, reports state-of-the-art motor-imagery decoding accuracy on OpenBMI, Stieger2021, and a clinical ECoG-ALS dataset.
-
TF-MossFormer: Integrating Convolution Gated Local-Global Attentions for Enhanced Time-Frequency Domain Monaural Speech Separation
A spectrogram-domain transformer combining fixed sliding-window local attention and global attention with convolutional gating reaches 24.4 dB SI-SDRi on WSJ0-2Mix at 25.4M parameters.
-
Enhancing Stereo Sound Event Detection with BiMamba and Pretrained PSELDnet
Replacing the Conformer decoder in pretrained PSELDnet with a bidirectional Mamba plus asymmetric convolution reports 39.6% versus 38.2% stereo SELD F20 on the DCASE2025 development set, using 76M versus 210M parameters.
-
EAGLE: An Efficient Global Attention Lesion Segmentation Model for Hepatic Echinococcosis
EAGLE combines Mamba-style state-space blocks with wavelet downsampling and reports 89.76% DSC on private hepatic echinococcosis CT data, but slice-level splitting and missing artifacts undermine the claim.
-
ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
A unified open-source toolkit with pretrained models for speech enhancement, separation, super-resolution, and multimodal target extraction, plus a new face-conditioned extraction model.
Discussion (0). Sign in to comment.