Pith. sign in

Paper Citation Record · LEDGER

SAM: A Mamba-2 State-Space Audio-Language Model

As of 16 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2509.15680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.15680 v3

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:52:34.201226Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:52:33.994758Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T15:52:34.272945Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 34b2f0ff-f804-449f-9442-ded1a9c622d6 · outbound

This paper cites Prior works have improved their ability in various ways, including large-scale QA datasets, curriculum- learning strategy, and advanced connector architectures.

SAM: A Mamba-2 State-Space Audio-Language Model Prior works have improved their ability in various ways, including large-scale QA datasets, curriculum- learning strategy, and advanced connector architectures

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:35.037335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:33.989045Z digest=sha256:be4be57932348726bf14828862bee2f6ff66ba6d6dd7540e9325eadd8e2156d9

Observation b3049d6d-ce10-43d5-b88a-9c7a1b3ef7d6 · outbound

This paper cites SAM: A Mamba-2 State-Space Audio-Language Model.

SAM: A Mamba-2 State-Space Audio-Language Model SAM: A Mamba-2 State-Space Audio-Language Model

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:52:34.279765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:33.994758Z digest=sha256:ddcad8545643d123e1d9a18483299dba571b9ad59a39b34cc3be3fc30666e7c7

Observation de72343b-6cfe-45ed-b426-0e30fe751946 · outbound

This paper cites Write an audio caption describing the sound.

SAM: A Mamba-2 State-Space Audio-Language Model Write an audio caption describing the sound

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:35.005633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.004897Z digest=sha256:950686f11684e4ea6a071fc2c7d8c108dac4b34c3293b72972df8a81da1d5b6e

Observation 7b02b90e-da14-4d97-bf44-1e642031264f · outbound

This paper cites an unresolved cited work.

SAM: A Mamba-2 State-Space Audio-Language Model Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:52:34.989829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.010160Z digest=sha256:5eb9f2f252e033f69236286fe812e49e7eb4c8a0687191dfe28154de3c917b2f

Observation 8628da76-b2a8-4362-920d-df238e4449cf · outbound

This paper cites For future work, we plan to investigate the effects of advanced connector designs that enable token mixing and Mamba-Transformer hybrid architectures in audio language modeling.

SAM: A Mamba-2 State-Space Audio-Language Model For future work, we plan to investigate the effects of advanced connector designs that enable token mixing and Mamba-Transformer hybrid architectures in audio language modeling

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.974028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.015396Z digest=sha256:b17e7345b34ecdff20c8d0d0d445bd0323c476a589de1abfb555f03f8fc7ba37

Observation 214c0e4c-cc07-4ee5-86df-bcd5e41a61d5 · outbound

This paper cites GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities,.

SAM: A Mamba-2 State-Space Audio-Language Model GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.889187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.043862Z digest=sha256:0632dda5ed744b49adb48176c71c00c489c675462430004f1b611398ff638462

Observation 22e71d62-c33c-4e4c-be3d-142a95fef6fb · outbound

This paper cites Attention is all you need,.

SAM: A Mamba-2 State-Space Audio-Language Model Attention is all you need,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:52:34.020228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:52:34.020228Z digest=sha256:ba1deec454adc0b24f86e6b3fda95e3925ac27b5cd401874791996be2bbb9f26

Observation 43f5121f-191c-48c1-8015-63f6d4246ff3 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models,.

SAM: A Mamba-2 State-Space Audio-Language Model Llama 2: Open foundation and fine-tuned chat models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.948110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.024935Z digest=sha256:a6434a792c9f5c13d25df0e29b484fb71af8dc72ea044502147870412b891f26

Observation 88565375-3273-408a-b3ca-b01eda865464 · outbound

This paper cites Opt-iml: Scaling language model instruc- tion meta learning through the lens of generalization,.

SAM: A Mamba-2 State-Space Audio-Language Model Opt-iml: Scaling language model instruc- tion meta learning through the lens of generalization,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.933254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.029675Z digest=sha256:3675972c8eb400ace045cb04154f752f57f8cf333d1bb09cf33cd5b9e129e81b

Observation bc775082-d7d0-436f-8a60-2235e484c08c · outbound

This paper cites Qwen2.5 technical report,.

SAM: A Mamba-2 State-Space Audio-Language Model Qwen2.5 technical report,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.918701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.034341Z digest=sha256:ba891f9762b8e8112bf96bc2c63e6e4a628311df955093742261f1c713b57890

Observation bb6bbe74-1ceb-4a6d-9c84-7055db0044e3 · outbound

This paper cites Listen, think, and understand,.

SAM: A Mamba-2 State-Space Audio-Language Model Listen, think, and understand,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.904023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.038873Z digest=sha256:b8ed774daf02e1b29b0ba2fb4654a4be0e4334c83e01bfab3d3ffaa2a1cd383e

Observation 91ea9b0d-32ed-4b60-9e5f-5ad7a01ad702 · outbound

This paper cites Transformers are ssms: generalized models and efficient algorithms through structured state space duality,.

SAM: A Mamba-2 State-Space Audio-Language Model Transformers are ssms: generalized models and efficient algorithms through structured state space duality,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.817581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.071847Z digest=sha256:0f39e1548a653021977ae106fd480018f6b6da2831acc7d33c9aba780994fd61

Observation 0dcacb6f-e8ef-430c-a614-aae77ac4ad9f · outbound

This paper cites SALMONN: Towards generic hearing abil- ities for large language models,.

SAM: A Mamba-2 State-Space Audio-Language Model SALMONN: Towards generic hearing abil- ities for large language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.875026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.048518Z digest=sha256:735d7be98966e71db669cc83502d1437bc22a26da2d7fe2460058418878319e9

Observation 65b03824-6831-46fa-a8af-3630565e0ab0 · outbound

This paper cites Audio flamingo 2: An audio-language model with long-audio understanding and expert reasoning abil- ities,.

SAM: A Mamba-2 State-Space Audio-Language Model Audio flamingo 2: An audio-language model with long-audio understanding and expert reasoning abil- ities,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.860745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.053073Z digest=sha256:be735f837e84464c9251746eafad46db4233c18697b52221ed77da2b189b2dd5

Observation 3ddb7568-6508-4a2e-9cf0-0c66f8bec392 · outbound

This paper cites Qwen2-audio technical report,.

SAM: A Mamba-2 State-Space Audio-Language Model Qwen2-audio technical report,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.845860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.057894Z digest=sha256:0611bd9c0c26d070b5f6dad3799eaaea5de041e57d14253a99f9cb7f9985500c

Observation fda0436f-6ef9-4067-8d70-4da53e5b895d · outbound

This paper cites Mellow: a small audio language model for reasoning,.

SAM: A Mamba-2 State-Space Audio-Language Model Mellow: a small audio language model for reasoning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.831690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.062451Z digest=sha256:ba8ed943eb606e4bea7e88b060982bda25ba4459f203c93bd2742dc76a2222dc

Observation 04448f11-cee8-45a9-bf26-48973402c580 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

SAM: A Mamba-2 State-Space Audio-Language Model Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:52:34.067089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:52:34.067089Z digest=sha256:65f32d2ca5cef5df1fda95681d2e18dd13296d156ceb6f41988bfc98bb72983d

Observation d0c41cb4-1a56-4d5f-99a8-7b92078e8638 · outbound

This paper cites Eat: Self-supervised pre-training with efficient audio transformer,.

SAM: A Mamba-2 State-Space Audio-Language Model Eat: Self-supervised pre-training with efficient audio transformer,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.744663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.099940Z digest=sha256:94669d57302258f95d44e38eafd93805babff59e0ac3dd37f80d164114b7a65e

Observation a36fcd49-aaee-497e-b479-6321fa11a17d · outbound

This paper cites Ml-mamba: Efficient multi-modal large language model utilizing mamba-2,.

SAM: A Mamba-2 State-Space Audio-Language Model Ml-mamba: Efficient multi-modal large language model utilizing mamba-2,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.802931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.076562Z digest=sha256:81eb8b0a6d06b597669eec1bd42e207fdc4a5adf1f17d8e5cdfbd679d6af3a3b

Observation 163366c4-b587-4510-9bbd-6f6f7330bd4f · outbound

This paper cites VL-Mamba: Exploring state space models for multimodal learning,.

SAM: A Mamba-2 State-Space Audio-Language Model VL-Mamba: Exploring state space models for multimodal learning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.788480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.081222Z digest=sha256:0864fff8496344891ec060ef5f84a2aa76ad998cb96bf034f89213c6c837af38

Observation 8858a6af-f693-4b02-84f4-84432bfa7a39 · outbound

This paper cites Shaking up VLMs: Comparing transformers and structured state space models for vision & language modeling,.

SAM: A Mamba-2 State-Space Audio-Language Model Shaking up VLMs: Comparing transformers and structured state space models for vision & language modeling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.774150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.086247Z digest=sha256:875c53b59c5579870745f4cfccce6b51be7b23890ab9d81e224bfd02c64e327f

Observation ff6c3f51-30b5-4e3e-92a2-4974422e8aa9 · outbound

This paper cites State-Space Large Audio Language Models.

SAM: A Mamba-2 State-Space Audio-Language Model State-Space Large Audio Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:52:34.090565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:52:34.090565Z digest=sha256:307a5c206932cd16b4eabdd7cf1e33e7134c0a1c46d4403e2c9b873302b27066

Observation 02901a8d-65a0-49d8-9265-41415d635809 · outbound

This paper cites On the parameterization and initialization of diagonal state space models,.

SAM: A Mamba-2 State-Space Audio-Language Model On the parameterization and initialization of diagonal state space models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.759670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.095165Z digest=sha256:9858ac45f8a8b6332ce45f3ac02528786b4b149238fa5ad5133978a92abcfea9

Observation 93010108-d17c-4705-9626-6973486f9b4b · outbound

This paper cites MambaPEFT: Exploring parameter-efficient fine-tuning for mamba,.

SAM: A Mamba-2 State-Space Audio-Language Model MambaPEFT: Exploring parameter-efficient fine-tuning for mamba,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.549328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.126370Z digest=sha256:9cf7f496434c3cd139437c074f92081116cc651582499292ab90dfe6083030d9

Observation d5d54761-7f60-46bf-b2e5-ff8078f73426 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events,.

SAM: A Mamba-2 State-Space Audio-Language Model Audio set: An ontology and human- labeled dataset for audio events,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.730225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.104304Z digest=sha256:92d448882500f08c2e424f184ec100342d60d85c36487e39cb8cfa52022d5a69

Observation 2fbfffec-e55e-42ba-8277-13da51e20b4a · outbound

This paper cites Slam-aac: Enhancing audio captioning with paraphrasing augmentation and clap-refine through llms,.

SAM: A Mamba-2 State-Space Audio-Language Model Slam-aac: Enhancing audio captioning with paraphrasing augmentation and clap-refine through llms,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.716082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.108730Z digest=sha256:5bdaff4b0adb5409ae6a89237f5ee6147746d29c539bee527a23c8dd730cbd76

Observation abed0099-02d4-45df-b59f-0beb49cf7f2e · outbound

This paper cites Sjtu-thu automated audio captioning system for dcase 2024,.

SAM: A Mamba-2 State-Space Audio-Language Model Sjtu-thu automated audio captioning system for dcase 2024,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.701876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.113012Z digest=sha256:96153819a5f2c0cfeb8600b8c18915fc2fe0caaee8f11c7e9c39933fa641f598

Observation cc571dd9-6b1b-462c-a41d-4064c99f3000 · outbound

This paper cites VMamba: Visual state space model,.

SAM: A Mamba-2 State-Space Audio-Language Model VMamba: Visual state space model,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.680464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.117522Z digest=sha256:415a05ddcacb6b9bbb7e9689a80aa8e9f8c8eea1fc0939389730a0b39c527caf

Observation 1259f0f5-98ae-4664-b064-f6ae607c0c7e · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

SAM: A Mamba-2 State-Space Audio-Language Model LoRA: Low-rank adaptation of large language models,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.564257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.122074Z digest=sha256:b5ae9a08ef253d7ea01acd62307542bedd68515ae4682a52fcd171c87f2c88c1

Observation 0e0e683d-b49e-4a34-8e5c-bb10425905cb · outbound

This paper cites Vggsound: A large-scale audio-visual dataset,.

SAM: A Mamba-2 State-Space Audio-Language Model Vggsound: A large-scale audio-visual dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.456944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.153187Z digest=sha256:455e99709a012caa814f846d609545de7f4fcf3c46b834dbdd976bb33cde6729

Observation 6e6f22ed-17d5-410e-8e6a-c91a6d49a426 · outbound

This paper cites Parameter-efficient fine-tuning of state space models,.

SAM: A Mamba-2 State-Space Audio-Language Model Parameter-efficient fine-tuning of state space models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.533480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.130716Z digest=sha256:17a96d4748da3574a6cb548933cba24dda60005db29258b95586a33ecb8783da

Observation 8e872c0a-9c4b-4afc-8540-0169fd380c27 · outbound

This paper cites Flashattention-2: Faster attention with better par- allelism and work partitioning,.

SAM: A Mamba-2 State-Space Audio-Language Model Flashattention-2: Faster attention with better par- allelism and work partitioning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.517928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.135334Z digest=sha256:36f5d638d7336b20f3175c109f7593b6f2eeb4e99d8eade211742c7436aad4ba

Observation 80d19e94-e767-423a-8766-4e1898acf06f · outbound

This paper cites Esc: Dataset for environmental sound classifi- cation,.

SAM: A Mamba-2 State-Space Audio-Language Model Esc: Dataset for environmental sound classifi- cation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.503223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.139673Z digest=sha256:3f5f68acf218c88fc6a5f1892931d810aca5d13359180eb539e3014dbb60c840

Observation e4ef902a-f664-4973-bc1d-7daa16b6da8c · outbound

This paper cites Sound event detection of weakly labelled data with cnn-transformer and automatic threshold optimiza- tion,.

SAM: A Mamba-2 State-Space Audio-Language Model Sound event detection of weakly labelled data with cnn-transformer and automatic threshold optimiza- tion,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.488360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.144392Z digest=sha256:ee0ea9d94e1be13324055a96a636746a9d82695be7db9a1dc2461598362250ee

Observation 99d9ea46-2e78-4e77-b490-3b21099faf45 · outbound

This paper cites V ocalsound: A dataset for improving human vocal sounds recognition,.

SAM: A Mamba-2 State-Space Audio-Language Model V ocalsound: A dataset for improving human vocal sounds recognition,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.473079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.148882Z digest=sha256:7314d1d882a8506d52074f7d53fbe4c614a748cac75a77ce9ad991b2fbafa401

Observation 7a25311d-7ed0-4260-9e9d-e578e2de058b · outbound

This paper cites BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models,.

SAM: A Mamba-2 State-Space Audio-Language Model BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.344218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.186999Z digest=sha256:e28c2f4c0089479c1af77e19abb0bb319ac093d99d7dbdffacf13fab3b0ed705

Observation 40f2e6fe-ba7c-45ec-a3cb-7a7580005cb0 · outbound

This paper cites Fsd50k: An open dataset of human-labeled sound events,.

SAM: A Mamba-2 State-Space Audio-Language Model Fsd50k: An open dataset of human-labeled sound events,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.439364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.158535Z digest=sha256:177435bdec1d04d651cdc7cc22cb8a8273b345ea21c0608a5c277b1e52be1aab

Observation c43b2cb6-0a08-4f2e-a77e-625e49374762 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

SAM: A Mamba-2 State-Space Audio-Language Model Spice: Semantic propositional image caption evaluation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.423460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.163297Z digest=sha256:05cc79f8e572bd69f986c5bc4a764a1f50e40b20189fdae2b6d8f609bd376781

Observation aa8bff82-83da-4567-bc6e-24a1c925bb7e · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

SAM: A Mamba-2 State-Space Audio-Language Model AudioCaps: Generating captions for audios in the wild,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.407637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.168052Z digest=sha256:85541a577459020883f4c8d39285714433e0e81c0f2f6d1959c1055ab5c26557

Observation 23c78040-988b-40b4-8b03-40f0087be24b · outbound

This paper cites 119–132, Association for Computational Lin- guistics.

SAM: A Mamba-2 State-Space Audio-Language Model 119–132, Association for Computational Lin- guistics

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.392567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.172538Z digest=sha256:cc7a494157c8f53d20a9ab2906fe4c0b6546efd4781688cff53f43896f812821

Observation ee354d0e-76f3-4c99-b887-84c7aea21df2 · outbound

This paper cites Clotho: an audio captioning dataset,.

SAM: A Mamba-2 State-Space Audio-Language Model Clotho: an audio captioning dataset,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.377042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.177321Z digest=sha256:ba8a8217bf4a978f45d8a3f31d5b05b8ac6596e575e28ea1eb66593de27a6f82

Observation c9832d8c-d146-469c-b31c-148409d8f9b9 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmen- tation,.

SAM: A Mamba-2 State-Space Audio-Language Model Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmen- tation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.360764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.182300Z digest=sha256:62585a5153194ba88672f9867ba0bef30a24c653e663da3a568066a228f9b4e6

Observation 747ec156-4349-417e-90e9-579895ed4511 · outbound

This paper cites The effective rank: A measure of effective dimensionality,.

SAM: A Mamba-2 State-Space Audio-Language Model The effective rank: A measure of effective dimensionality,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.327747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.191770Z digest=sha256:8b0d630011aa834aa977a3beeafee1a19838f5047b7d3f57afb89238418a9efb

Observation 3a38bcb5-8392-4f2a-98ec-84b8ec3ee9a0 · outbound

This paper cites Open-world objectness modeling unifies novel object detection,.

SAM: A Mamba-2 State-Space Audio-Language Model Open-world objectness modeling unifies novel object detection,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.311220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.196753Z digest=sha256:0b88a5e29344d51daf03e50a468f945e649e969324fd64b2a5dcaf4c59d80f8e

Observation 364627d7-5739-4317-8afb-c9bd4fe1ce35 · outbound

This paper cites Diff-erank: A novel rank-based metric for evaluating large language models,.

SAM: A Mamba-2 State-Space Audio-Language Model Diff-erank: A novel rank-based metric for evaluating large language models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:34.295746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:34.201226Z digest=sha256:d58d5c2f8a0c523925a6123a14544e533d2c8d01d0e9e30b689f3c8c7226ab48

Observation dfbd7c2b-c1d4-4278-b932-2faf3bd15624 · outbound

This paper cites Prior works on ALMs of- ten use mean pooling to reduce the length of audio tokens, mitigating the quadratic computational cost of self-attention.

SAM: A Mamba-2 State-Space Audio-Language Model Prior works on ALMs of- ten use mean pooling to reduce the length of audio tokens, mitigating the quadratic computational cost of self-attention

Reference 6144

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:52:35.021809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:33.999881Z digest=sha256:10fa003052e83126d0c8a916e0dfa6d2787207b8632ff940497dd0eca0ee8c4f

Pith citing papers

Observation b3049d6d-ce10-43d5-b88a-9c7a1b3ef7d6 · inbound

SAM: A Mamba-2 State-Space Audio-Language Model cites this paper.

SAM: A Mamba-2 State-Space Audio-Language Model SAM: A Mamba-2 State-Space Audio-Language Model

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:52:34.279765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:52:33.994758Z digest=sha256:ddcad8545643d123e1d9a18483299dba571b9ad59a39b34cc3be3fc30666e7c7