Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:36:19.603067Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 100 of 168 outbound references and 16 inbound Pith citation observations for arXiv:2501.15177.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:36:19.603067Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:50:44.994154Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
100 of 168 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 12b7cf0a-724e-4c01-8d39-5bf644a579bd · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Pengi: An Audio Language Model for Audio Tasks,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b524c830-54c3-40a3-827d-038fd56e5599 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey CLAP Learn- ing Audio Concepts from Natural Language Supervision,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 666fa1d9-21c6-420e-b5dd-d18ed62c0a76 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Bridging language gaps in audio-text retrieval,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a083652-315d-41aa-adee-c55addd7e074 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Audioldm: text-to-audio generation with latent diffusion models,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6452b7c5-7446-4034-8e18-5dfaca868a5f · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38faad7a-526e-45fd-ad24-e100c5311819 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Separate what you describe: Language-queried audio source separation,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b35920c-04cd-48fa-8849-0aa5d7244e9a · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30671dc0-f696-4464-b101-c7bda7e0e60b · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Speechgpt: Empowering large language models with intrinsic cross- modal conversational abilities,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8e074c-1f63-45ae-bdc7-79a6890c39c2 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Audio Dataset,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68d9703-36cb-40df-ad40-0095180aadfa · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Audio caption: Listen and tell,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22dd4c4a-e67b-4063-912c-c09c2820cbe7 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey A Survey of Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 388ed2cb-d343-40c8-9560-cfcf09fa9a56 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Recent advances in speech language models: A survey,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5207cc1-6394-4a24-acb1-e3f238b866c6 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey WavChat: A Survey of Spoken Dialogue Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdfca223-bb90-4a11-a52f-1d01963f68d1 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Audio retrieval with natural language queries: A benchmark study,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 909caf73-5985-4ea7-aae3-1927fd21afd7 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Beyond the Status Quo: A Con- temporary Survey of Advances and Challenges in Audio Captioning,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72b3fb2f-b8f9-41c8-bf11-0d499b3b2db8 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Recent Advances in Direct Speech-to-text Translation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c12161ed-ee77-4d57-8fa2-de2e17ddc279 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Audio-Language Datasets of Scenes and Events: A Survey
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 865c371f-ee60-450d-b19a-33590dfb8747 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey AudioCaps: Generating Cap- tions for Audios in The Wild,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 272fa887-9511-4073-87a6-21af1a5f006d · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Clotho: an Audio Captioning Dataset,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcc22394-cf69-4339-8753-1ddf6522f958 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Clotho-aqa: A crowdsourced dataset for audio question answering,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7669c81f-aaea-4b18-8b5e-9872b341ad8b · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdf349f5-12d0-4697-9651-57b26868c376 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Beats: audio pre-training with acoustic tokenizers,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5bdd9c-c85b-44b5-9c17-643280d377b0 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Listen, think, and understand,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 617bebe3-aa22-4e97-be69-52a77f10cc99 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Automated data augmentation for audio classification,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 501235d7-29b7-4738-b5d6-53e4e6723a46 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Compa: Addressing the gap in compositional reasoning in audio- language models,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f8ff340-5ee6-4ed8-9486-b1b70ef402f1 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Finetuned language models are zero-shot learners,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f3efb4-5131-4b49-8fdf-e5fe6f7be77c · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7573674a-1a43-400a-adce-1d9c637b4657 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea22914-42a0-42f2-aee2-95f6a75f5005 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Multimodal contrastive learning with limoe: the language-image mix- ture of experts,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5184169b-4978-4c4d-b578-96338c72d7f1 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Unifying vision- language representation space with single-tower transformer,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d48753c6-8634-4af9-9e5c-81ec798f5a0e · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7380048-7629-4ea8-b5bf-68ab5d21c936 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 040a6ff7-9330-459f-b740-7260e9ead14e · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Mint: Boosting audio-language model via multi-target pre-training and instruction tuning,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33a9f8c5-90a4-4232-a731-8e67e4b5bb0c · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey CALM: Contrastive Aligned Audio-Language Multirate and Multimodal Representations
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 98870a65-85c1-4129-aa9b-7fd0613d8de3 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey A survey on llm- based multi-agent systems: workflow, infrastructure, and challenges,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cba960e-0e83-4e47-9272-c42de57830cf · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Representation Learning with Contrastive Predictive Coding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abace351-ec70-4724-ae05-06aeafb05997 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Masked autoencoders that listen,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2896c677-f2a7-441c-bcf8-1bdc4d7d9862 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Palm: Scaling language modeling with pathways,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad1d089-3e04-408f-a422-119f17b10928 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey WavCaps: A ChatGPT-Assisted Weakly- Labelled Audio Captioning Dataset for Audio-Language Multimodal Research,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a6d20bb-6d94-4f44-92c5-ce5f52e95f16 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Automatic tagging using deep convolutional neural networks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 970b31af-b834-4338-9cf9-a0ac4d70b7de · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Cnn architectures for large-scale audio classification,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1397e390-83d7-425a-ab1f-128ca3480ded · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Panns: Large-scale pretrained audio neural networks for audio pattern recognition,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49c473b8-02d0-4bd0-a49d-eca87b0c000c · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92efebeb-f5a6-4db8-9b90-db6aa4311d4b · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Hubert: Self-supervised speech representation learning by masked prediction of hidden units,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79dbc15c-ebac-49f2-b5aa-66368b9a1ea8 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Shortcut learning in deep neural networks,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6df1d3ff-f31c-43f1-9a2e-31c6d7607c78 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Robust speech recognition via large-scale weak supervi- sion,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b7ac7c-80b5-4c0a-ad74-bda9894d9955 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey AST: Audio Spectrogram Transformer
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac7ea6b-dbf1-4da9-9afb-1811b95ccf11 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bd65bdd-f3fc-461e-9d40-c69a1bfcacd5 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Natural language supervision for general-purpose audio representations,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187e7e93-3dae-424a-a998-97168e30c4bf · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Masked autoencoders are scalable vision learners,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eea972aa-1ba7-4598-a556-a9c68713d9fd · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Towards audio language modeling -- an overview
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47566490-aea1-49e4-b5be-c86ec1c6ac23 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Soundstream: An end-to-end neural audio codec,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03918a5a-9086-48fb-97a0-1f67737b79cf · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Seanet: A multi- modal speech enhancement network,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb47107-a10d-48bf-acac-d05b1466d082 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey High fidelity neural audio compression,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 330c9ed0-1917-4338-a7f0-fd49768d9df1 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Improving language understanding by generative pre- training,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8323bd2c-b489-48ba-ad17-f49e69bba502 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Language models are unsupervised multitask learners,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74c341f4-e9ed-427a-82c3-f04b8804ed39 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Language models are few-shot learners,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82daa610-5e93-4b78-ad2e-25de64fc7f0d · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey GPT-4 Technical Report
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36780942-25f6-4e18-80b0-1068d38eb8bc · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Training language models to follow instructions with human feedback,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcbb9986-bba0-4f20-86ad-db16fd443e57 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Introducing chatgpt,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aeffd84-c60b-4393-9b31-2d9e83f883c4 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey LLaMA: Open and Efficient Foundation Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 299797a9-998c-4840-a08a-2f2a92b311cf · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99251b4c-530f-454e-867b-430fad32b952 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Alpaca: A strong, replicable instruction- following model,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d80bc2-d6c9-4f72-9113-1d130bf95c5b · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Vicuna: An open- source chatbot impressing gpt-4 with 90%* chatgpt quality,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b9f0ec1-ae64-443b-9233-611d280a9d11 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Qwen Technical Report
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f29d03a1-47aa-47f9-9491-4f614bb0c2d6 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Qwen2 Technical Report
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d31559-ec7b-4a40-b734-55bf9459816a · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Qwen2.5-Coder Technical Report
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b414b4f5-5841-4e44-a16c-ab54138d0a32 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13f57039-dc2d-4b1c-a2d9-a10f4c0e30bc · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey OPT: Open Pre-trained Transformer Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740b64c8-9108-478f-8406-1f4f03f156c8 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89a93bc4-2756-4807-ae74-19e76f58de54 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Large Language Models: A Survey
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68363f9-eaa8-461b-ae82-639d20d11720 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Blat: Bootstrapping language-audio pre-training based on audioset tag-guided synthetic data,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51b6a62d-17c9-4dcf-9318-ee19bc357c65 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Audio-text models do not yet leverage natural language,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ee11212-2e75-41ef-b2c8-09d2919a1c4c · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Advancing multi-grained align- ment for contrastive language-audio pre-training,
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21946010-80d4-4e66-87ac-4a68cbbb744c · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Contrastive audio- language learning for music,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c3e113f-e52e-45f6-8609-8e7b05ff030e · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Flap: Fast language-audio pre-training,
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d2687ad-813e-4b7c-87f8-1d8251663d1b · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2a3d6981-0b78-432b-80bf-01e144a37c99 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Cacophony: An Improved Contrastive Audio-Text Model
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d1b034-1e2a-47cd-8c08-7cba1fa32087 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Learning transferable visual models from natural language supervision,
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4edf8561-fc46-4df4-a5af-0f20fa2d5d75 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Masked modeling duo: Learning representations by encouraging both networks to model the input,
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe95b06-abf1-49eb-aefc-0a0b884f8b1e · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Coca: Contrastive captioners are image-text foundation mod- els,
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d412e7-936a-4dcd-b220-da495dc210e8 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a540f9f8-f850-4a4d-aa07-8b4d219d5e4b · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c711961-8b0e-4533-a395-969bcc99c33e · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad64020b-cff5-4485-9922-e18c42653483 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed5aba6-9d71-4d4c-99ba-af1d401b03b2 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 716b2f3e-1cb8-48b2-81ae-e1da5beb5dfd · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c825c1-24f3-412d-a28c-85cac637a18f · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Audiolm: A language modeling approach to audio generation,
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d3ed5b-a6b8-42e5-b69e-895b69e92a5d · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey MusicLM: Generating Music From Text
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fdffaf9-56c2-448c-9ddb-78391af87148 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Mustango: Toward controllable text-to-music generation,
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c4a3b66-d69f-492d-b3ee-03fee6f069ec · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Audiogen: Textually guided audio generation,
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd797f66-df8b-4775-ad84-c7ec8f535c84 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Diffsound: Discrete diffusion model for text-to-sound generation,
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dab5646-809a-4a06-abaa-e6080e616efb · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey CLIPSep: Learning text-queried sound separation with noisy unlabeled videos,
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 253a674d-a8d9-481b-af07-8ceb909f0402 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Separate Anything You Describe
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0127a8f9-7991-4dde-8c71-f659a66aca71 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Opensep: Leveraging large language models with textual inversion for open world audio separation,
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b22e6b7-09f4-412e-8555-cd466df31e48 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01761d08-92c4-4551-a165-da75e24932b0 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Exploring text-queried sound event detection with audio source separation,
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fdf5a94-3b79-45a9-8da6-5b0923330a3c · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b1f541-060f-41ed-8f1c-ad0228b1741b · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 098e0395-4feb-42f1-8c84-937ecb6c59c4 · outbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Lp-musiccaps: Llm-based pseudo music captioning,
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9181726-0b0f-4d86-974f-52685268cfe8 · inbound
Learning Sparsity for Effective and Efficient Music Performance Question Answering Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67c45737-4490-4377-ae36-7247300acb00 · inbound
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0168ded6-1ca9-410c-b702-a6f097fbf190 · inbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a339adb2-3b79-47a0-9a57-8a21cead81a6 · inbound
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6793b64c-7fdc-4fb7-9310-ae5fceb14b6f · inbound
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5d178954-211e-4767-8bc1-1a2ab5a9ebb2 · inbound
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9f4bed0e-2a05-402c-a285-dfbabd16d2bc · inbound
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6a65ae05-1e3c-4198-9414-f910e7395f4e · inbound
A Survey of Audio Reasoning in Multimodal Foundation Models Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 629c2bc5-ea64-466a-8b8a-032465a7be1f · inbound
Learning When to Think While Listening in Large Audio-Language Models Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 457c8e9c-249a-4a9d-9213-9a558c55a0d6 · inbound
Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6d6eed2a-f84a-4875-9a06-c93df3cc29ef · inbound
GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bca6fe96-5ef7-486b-a6d4-e155b5dd30a4 · inbound
Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 180
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 250fabc9-f870-4cf2-8d21-bcc649aa5c8f · inbound
When the Same Musical Knowledge Forgets Differently: A Clean Probe of Pathway-Dependent Forgetting Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2b42ec49-bb0c-4847-86ec-5eb5e652407a · inbound
ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2b321af5-938c-4c76-be6f-ee2c3b5c379b · inbound
AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 09ee30bd-04f3-4e16-827e-6e2bcd8dd040 · inbound
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.