Pith. sign in

SAM: A Mamba-2 State-Space Audio-Language Model

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We present SAM, a State-space Audio-language Model that integrates an audio encoder with a Mamba-2 backbone. SAM-2.7B achieves 21.1 mAP on AudioSet and 17.6 SPICE on AudioCaps, matching or surpassing larger 7B transformer-based models with fewer parameters. We further provide the first systematic, representation-level analysis of how SSMs interact with audio encoder outputs: (1) joint audio encoder finetuning is essential, supported by accuracy gains and observed adaptation of token representation rank and similarity across different SSM sizes; (2) despite linear scaling, SSMs benefit more from compact, information-rich audio token representations than from excessively long token sequences; and (3) incorporating instruction-following supervision substantially improves reasoning ability, boosting MMAU-Sound accuracy from 22.8 to 56.8. Through comprehensive experiments and analysis, we establish practical design principles for SSMs as strong, scalable backbones for audio-language models.

fields

cs.SD 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

SAM: A Mamba-2 State-Space Audio-Language Model

cs.SD · 2025-09-19 · conditional · novelty 4.0

A 2.7B-parameter Mamba-2 based audio captioning model (MAC) matches or surpasses larger transformer-based audio-language models on several zero-shot classification and captioning benchmarks, with additional design-space analysis.

citing papers explorer

Showing 1 of 1 citing paper.

  • SAM: A Mamba-2 State-Space Audio-Language Model cs.SD · 2025-09-19 · conditional · none · ref 2 · internal anchor

    A 2.7B-parameter Mamba-2 based audio captioning model (MAC) matches or surpasses larger transformer-based audio-language models on several zero-shot classification and captioning benchmarks, with additional design-space analysis.