Pith. sign in

REVIEW 4 cited by

TF-Mamba: A Time-Frequency Network for Sound Source Localization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.05034 v2 pith:Q65XEYWG submitted 2024-09-08 eess.AS cs.SD

classification eess.AScs.SD
keywords featuresmambasoundtf-mambaexperimentsfrequencylocalizationsource
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sound source localization (SSL) determines the position of sound sources using multi-channel audio data. It is commonly used to improve speech enhancement and separation. Extracting spatial features is crucial for SSL, especially in challenging acoustic environments. Recently, a novel structure referred to as Mamba demonstrated notable performance across various sequence-based modalities. This study introduces the Mamba for SSL tasks. We consider the Mamba-based model to analyze spatial features from speech signals by fusing both time and frequency features, and we develop an SSL system called TF-Mamba. This system integrates time and frequency fusion, with Bidirectional Mamba managing both time-wise and frequency-wise processing. We conduct the experiments on the simulated and real datasets. Experiments show that TF-Mamba significantly outperforms other advanced methods. The code will be publicly released in due course.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A multi-band semantic-spatial alignment network (AV-SSAN) localizes a target sound source using a cross-instance visual prompt, achieving 16.59 degrees mean error and 71.29% accuracy on the new VGGSound-SSL benchmark.

  2. ESTM: An Enhanced Dual-Branch Spectral-Temporal Mamba for Anomalous Sound Detection

    cs.SD 2025-09 conditional novelty 5.0 of 10

    ESTM, a dual-branch Mamba with frequency/time patches and a statistical gating module, reports the best average AUC and pAUC on DCASE 2020 Task 2.

  3. Enhancing Stereo Sound Event Detection with BiMamba and Pretrained PSELDnet

    eess.AS 2025-07 conditional novelty 4.0 of 10

    Replacing the Conformer decoder in pretrained PSELDnet with a bidirectional Mamba plus asymmetric convolution reports 39.6% versus 38.2% stereo SELD F20 on the DCASE2025 development set, using 76M versus 210M parameters.

  4. Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation

    eess.AS 2025-05 conditional novelty 3.0 of 10

    A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.

Pith tools