Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:52:34.201226Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2509.15680.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:52:34.201226Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:52:33.994758Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-15T15:52:34.272945Z
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 34b2f0ff-f804-449f-9442-ded1a9c622d6 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Prior works have improved their ability in various ways, including large-scale QA datasets, curriculum- learning strategy, and advanced connector architectures
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b3049d6d-ce10-43d5-b88a-9c7a1b3ef7d6 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model SAM: A Mamba-2 State-Space Audio-Language Model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation de72343b-6cfe-45ed-b426-0e30fe751946 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Write an audio caption describing the sound
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7b02b90e-da14-4d97-bf44-1e642031264f · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8628da76-b2a8-4362-920d-df238e4449cf · outbound
SAM: A Mamba-2 State-Space Audio-Language Model For future work, we plan to investigate the effects of advanced connector designs that enable token mixing and Mamba-Transformer hybrid architectures in audio language modeling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 214c0e4c-cc07-4ee5-86df-bcd5e41a61d5 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 22e71d62-c33c-4e4c-be3d-142a95fef6fb · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Attention is all you need,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43f5121f-191c-48c1-8015-63f6d4246ff3 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Llama 2: Open foundation and fine-tuned chat models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 88565375-3273-408a-b3ca-b01eda865464 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Opt-iml: Scaling language model instruc- tion meta learning through the lens of generalization,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bc775082-d7d0-436f-8a60-2235e484c08c · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Qwen2.5 technical report,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bb6bbe74-1ceb-4a6d-9c84-7055db0044e3 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Listen, think, and understand,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 91ea9b0d-32ed-4b60-9e5f-5ad7a01ad702 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Transformers are ssms: generalized models and efficient algorithms through structured state space duality,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0dcacb6f-e8ef-430c-a614-aae77ac4ad9f · outbound
SAM: A Mamba-2 State-Space Audio-Language Model SALMONN: Towards generic hearing abil- ities for large language models,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 65b03824-6831-46fa-a8af-3630565e0ab0 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Audio flamingo 2: An audio-language model with long-audio understanding and expert reasoning abil- ities,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3ddb7568-6508-4a2e-9cf0-0c66f8bec392 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Qwen2-audio technical report,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fda0436f-6ef9-4067-8d70-4da53e5b895d · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Mellow: a small audio language model for reasoning,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 04448f11-cee8-45a9-bf26-48973402c580 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c41cb4-1a56-4d5f-99a8-7b92078e8638 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Eat: Self-supervised pre-training with efficient audio transformer,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a36fcd49-aaee-497e-b479-6321fa11a17d · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Ml-mamba: Efficient multi-modal large language model utilizing mamba-2,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 163366c4-b587-4510-9bbd-6f6f7330bd4f · outbound
SAM: A Mamba-2 State-Space Audio-Language Model VL-Mamba: Exploring state space models for multimodal learning,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8858a6af-f693-4b02-84f4-84432bfa7a39 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Shaking up VLMs: Comparing transformers and structured state space models for vision & language modeling,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ff6c3f51-30b5-4e3e-92a2-4974422e8aa9 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model State-Space Large Audio Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02901a8d-65a0-49d8-9265-41415d635809 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model On the parameterization and initialization of diagonal state space models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 93010108-d17c-4705-9626-6973486f9b4b · outbound
SAM: A Mamba-2 State-Space Audio-Language Model MambaPEFT: Exploring parameter-efficient fine-tuning for mamba,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d5d54761-7f60-46bf-b2e5-ff8078f73426 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Audio set: An ontology and human- labeled dataset for audio events,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2fbfffec-e55e-42ba-8277-13da51e20b4a · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Slam-aac: Enhancing audio captioning with paraphrasing augmentation and clap-refine through llms,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation abed0099-02d4-45df-b59f-0beb49cf7f2e · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Sjtu-thu automated audio captioning system for dcase 2024,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cc571dd9-6b1b-462c-a41d-4064c99f3000 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model VMamba: Visual state space model,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1259f0f5-98ae-4664-b064-f6ae607c0c7e · outbound
SAM: A Mamba-2 State-Space Audio-Language Model LoRA: Low-rank adaptation of large language models,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0e0e683d-b49e-4a34-8e5c-bb10425905cb · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Vggsound: A large-scale audio-visual dataset,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6e6f22ed-17d5-410e-8e6a-c91a6d49a426 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Parameter-efficient fine-tuning of state space models,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8e872c0a-9c4b-4afc-8540-0169fd380c27 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Flashattention-2: Faster attention with better par- allelism and work partitioning,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 80d19e94-e767-423a-8766-4e1898acf06f · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Esc: Dataset for environmental sound classifi- cation,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e4ef902a-f664-4973-bc1d-7daa16b6da8c · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Sound event detection of weakly labelled data with cnn-transformer and automatic threshold optimiza- tion,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 99d9ea46-2e78-4e77-b490-3b21099faf45 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model V ocalsound: A dataset for improving human vocal sounds recognition,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7a25311d-7ed0-4260-9e9d-e578e2de058b · outbound
SAM: A Mamba-2 State-Space Audio-Language Model BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 40f2e6fe-ba7c-45ec-a3cb-7a7580005cb0 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Fsd50k: An open dataset of human-labeled sound events,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c43b2cb6-0a08-4f2e-a77e-625e49374762 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Spice: Semantic propositional image caption evaluation,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aa8bff82-83da-4567-bc6e-24a1c925bb7e · outbound
SAM: A Mamba-2 State-Space Audio-Language Model AudioCaps: Generating captions for audios in the wild,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 23c78040-988b-40b4-8b03-40f0087be24b · outbound
SAM: A Mamba-2 State-Space Audio-Language Model 119–132, Association for Computational Lin- guistics
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ee354d0e-76f3-4c99-b887-84c7aea21df2 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Clotho: an audio captioning dataset,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c9832d8c-d146-469c-b31c-148409d8f9b9 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmen- tation,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 747ec156-4349-417e-90e9-579895ed4511 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model The effective rank: A measure of effective dimensionality,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3a38bcb5-8392-4f2a-98ec-84b8ec3ee9a0 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Open-world objectness modeling unifies novel object detection,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 364627d7-5739-4317-8afb-c9bd4fe1ce35 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Diff-erank: A novel rank-based metric for evaluating large language models,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dfbd7c2b-c1d4-4278-b932-2faf3bd15624 · outbound
SAM: A Mamba-2 State-Space Audio-Language Model Prior works on ALMs of- ten use mean pooling to reduce the length of audio tokens, mitigating the quadratic computational cost of self-attention
Reference 6144
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b3049d6d-ce10-43d5-b88a-9c7a1b3ef7d6 · inbound
SAM: A Mamba-2 State-Space Audio-Language Model SAM: A Mamba-2 State-Space Audio-Language Model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.