Pith. sign in

Paper Citation Record · LEDGER

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2501.15177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15177 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:56:54.908507Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 67c45737-4490-4377-ae36-7247300acb00 · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:54.908507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:54.908507Z digest=sha256:39e1180293a19e78d5ed3ec1e8804d401199137a53f99854c3d1b2efdaa2702b

Observation 0168ded6-1ca9-410c-b702-a6f097fbf190 · inbound

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation cites this paper.

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:11.400124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:11.400124Z digest=sha256:13a131d959d637c7b741f377ffbbb8fc9ebfc9400cf3100c3b22e0dba964a065

Observation a339adb2-3b79-47a0-9a57-8a21cead81a6 · inbound

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs cites this paper.

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:01:25.938248Z digest=sha256:22f568e44dace0bfcba1057d482fdad85bf6b102216e9fed3aef42cea72d22a6

Observation 6793b64c-7fdc-4fb7-9310-ae5fceb14b6f · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:d10731dd4799efd6df878bdf115bd868e3a70bd513f8675dccf2312cc01cc7df

Observation 5d178954-211e-4767-8bc1-1a2ab5a9ebb2 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:97d8ece2d4dc38411203522ed9bd007291525854129a94fa6a73c4c3b98e308c

Observation 9f4bed0e-2a05-402c-a285-dfbabd16d2bc · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:2ec03a5ef2bae147dd35a53b26582ad2fe23e0ddddc39bfd14c63ecf52365534

Observation 6a65ae05-1e3c-4198-9414-f910e7395f4e · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:94bdfa407feeb7b735dc3e9579021caa9dc9e048ed2a78823788973f55c91fb8

Observation 629c2bc5-ea64-466a-8b8a-032465a7be1f · inbound

Learning When to Think While Listening in Large Audio-Language Models cites this paper.

Learning When to Think While Listening in Large Audio-Language Models Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:37:34.409802Z digest=sha256:2a2dc77be867f25cf20d370338acfc0b669f5dd008de903582488e068311a728

Observation 457c8e9c-249a-4a9d-9213-9a558c55a0d6 · inbound

Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition cites this paper.

Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T20:56:59.118526Z digest=sha256:ab55761e4bca0852e99ee08521bfb21ede0cd8cebfeaea8e176657ad4aa43dd5

Observation 6d6eed2a-f84a-4875-9a06-c93df3cc29ef · inbound

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models cites this paper.

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T19:51:59.265579Z digest=sha256:52645f02b946a5d194d130bc65e1bd0c85326bae506968dc887b6bc49fffaaef

Observation bca6fe96-5ef7-486b-a6d4-e155b5dd30a4 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 180

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:b0a638698df2b4496bea8fb8ced524b78cc529934398d356187bd4a167a281b3

Observation 250fabc9-f870-4cf2-8d21-bcc649aa5c8f · inbound

When the Same Musical Knowledge Forgets Differently: A Clean Probe of Pathway-Dependent Forgetting cites this paper.

When the Same Musical Knowledge Forgets Differently: A Clean Probe of Pathway-Dependent Forgetting Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:50:48.712481Z digest=sha256:491ea662e43534ca533305f19a7fc4a3d0281f21740e56c3545a2a45888cd04f

Observation 2b42ec49-bb0c-4847-86ec-5eb5e652407a · inbound

ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation cites this paper.

ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:17:37.444210Z digest=sha256:256be08dfa922076c27513856ff980fb99b1b7c6adcf2fdac612b93f58f5cea7

Observation 2b321af5-938c-4c76-be6f-ee2c3b5c379b · inbound

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? cites this paper.

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T13:22:12.541923Z digest=sha256:f8800b1067d5ae38df8c82fac8f6e903cfdb9e7b84274e5b9d5911b954e0f76e

Observation 09ee30bd-04f3-4e16-827e-6e2bcd8dd040 · inbound

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios cites this paper.

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:07.899293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T07:36:34.307652Z digest=sha256:08ad21f76119312b7c0200e9a7c8e2602a9ebb8fb7ff4d21b643f8cd0f83f7ca