Pith. sign in

Paper Citation Record · LEDGER

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

As of 4 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 100 inbound Pith citation observations for arXiv:2311.07919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.07919 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T18:57:28.666194Z

measured 137 of 137 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 100 of 120 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T03:03:32.342985Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:27:36.579792Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact21
  • verified fuzzy13
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f82b18c7-0b96-499e-bba2-fcb3970209e0 · outbound

This paper cites Spice: Semantic propositional image caption evaluation.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Spice: Semantic propositional image caption evaluation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.885441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:029e0e0cc27bac1181b21a374f9930e177ca5d0a52f2e3b844e58ebeccd5d676

Observation 0b34eb66-5d5e-4197-b53c-291d3fa9631a · outbound

This paper cites PaLM 2 Technical Report.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models PaLM 2 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.779045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:f23029e24b674dfdcd2b9abf335d8d2ed51aa385b0e689e7ec93c653dfdbc590

Observation 1a9a4096-7be8-487e-bd11-85d4f50bebe4 · outbound

This paper cites SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.714811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:8a2e71ca92d84f5a15c1e71eec5afdb90a85c0a7b734fdfba58d75d2c05d6b79

Observation 686a88c1-2cb5-4cc3-9743-96b07e7caff0 · outbound

This paper cites Qwen Technical Report.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Qwen Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.736051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:d69fb1dd114eb8566e2d44a8a268855538e841dd5950a13e911dc0af8644af50

Observation 978cba6e-4754-4698-89bd-496af42f7de5 · outbound

This paper cites AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.831322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:4519f3c21a89b6158a00999f40cc72a8f785f5a7b19232458917eb527f07d904

Observation 40075b4d-b472-4db9-9b1e-be1236df4437 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.801899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:c6d1ade95192ebfe47ba015bf6617c9d08add2e452abb167272286da85a3f748

Observation c2e4f870-f674-4337-aaa5-83566ad92f9f · outbound

This paper cites SpeechNet: A Universal Modularized Model for Speech Processing Tasks.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models SpeechNet: A Universal Modularized Model for Speech Processing Tasks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.808161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:9fc664a5cb74970d8f743a78110648d41e37c8f3182823589b22885c81fbca93

Observation 0389ca6e-f17d-4aec-b5e2-8381638cda49 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.817947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:e9d21586cef5b0518cdd413a6694149562573224a47051c4be90f2fd8dfbf48f

Observation 513cee5e-03c7-4c91-98bd-c697f2efda83 · outbound

This paper cites High Fidelity Neural Audio Compression.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models High Fidelity Neural Audio Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:49:52.306785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:bf7a4f8a0088e44578e95d6e9ef77a5c23eed594298bf08045ef3ac07872a0a5

Observation 527b4f02-96cd-4cc8-b586-ec04dde22436 · outbound

This paper cites Clotho: an audio captioning dataset.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Clotho: an audio captioning dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.835232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:6ed379ff9adc7ccb2dd468d366c359f957e978be629adc341912b9de44c2a9a7

Observation 68ff7646-e1bf-41d7-bcfa-203ecfe6a788 · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.720600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:81d4f0f6f12c5155631150b4e58469b8d03218d8195cc9fe06f802ed535a2992

Observation 6f5b030d-787a-4213-8656-d5a8ac082472 · outbound

This paper cites CLAP: Learning Audio Concepts From Natural Language Supervision.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models CLAP: Learning Audio Concepts From Natural Language Supervision

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.731393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:614e1c8c82c3b06cd0b3fbd0238a01ad49c91824c7777a2101c88f05e517a5e9

Observation 8797769d-2f3c-4613-9ead-011e07268843 · outbound

This paper cites Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.838846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:b45d2101ebc48f3c82d5343ab70ae12e91779221c5b6c2ddcc080ae3eb9dc515

Observation 3ad2844e-c9f5-4714-bff0-1a4b7959d5fd · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.775426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:e83adcdf4c52325466ffe30d37d7826e89fdcb89942cdccaa5c13f066af16686

Observation 36888e76-61f8-40ff-8187-3bcc31335b98 · outbound

This paper cites Vocalsound: Adatasetforimprovinghumanvocalsoundsrecognition.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Vocalsound: Adatasetforimprovinghumanvocalsoundsrecognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.845340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:5fd441fab92b7c3a47341c3f3932c6c0a7dc48a3befb272ff6a809b2bc2d07b8

Observation f7b85189-ba7f-4e93-8573-4e11a561165d · outbound

This paper cites author Zhou, A.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models author Zhou, A

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.701257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:917c139f0dc3be32c2d4f9af07c0cb26affbb5c0cf2735bef5f11ba54efbc138

Observation 2af46fc4-917a-4c63-8f0e-57d6da47ab1e · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.783179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:1300a987f904c555b20c54cfa4a3d3e100230f874a894d021ebd85062fd14465

Observation fe563662-f13d-4f5b-860d-d8d85889f490 · outbound

This paper cites AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.789265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:7ac7d5dbadeb671c0f1957b6f39ed03669ac6a54e58ab56926a473712e19014b

Observation 06221470-b5d6-497b-bdc2-9c210a4d8566 · outbound

This paper cites CochlScene: Acquisition of acoustic scene data using crowdsourcing.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models CochlScene: Acquisition of acoustic scene data using crowdsourcing

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.796789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:684352815c052b8042193ed650064e59a362224aef2505bf67b71343a2de1721

Observation 59b741bb-9023-4ffb-849d-3efec1ae8228 · outbound

This paper cites an unresolved cited work.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-12T18:57:28.849814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:2d71a83d6b6c62b358d8d76791d0947ce73a5bc5766ef73fe0f48a8c782219ce

Observation 85bb758f-2015-4066-8751-0af96520986f · outbound

This paper cites Clotho-aqa: A crowdsourceddatasetforaudioquestionanswering.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Clotho-aqa: A crowdsourceddatasetforaudioquestionanswering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.859905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:6e171b83a651982d30a843a6cc97b6fadf16770c7d58f26a12f3f78bff09051d

Observation 0e5f935e-91f7-42dd-90d9-b091d4afbb92 · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.813519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:d9c7e2e1fcb75153941adf2b1fbc53f9aaad59fd8f49c89a71318b660899c4ed

Observation 8f2a679e-c6bd-4afd-850c-f503f410178f · outbound

This paper cites Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.823612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:24806d32e44488f562a13f327b4a978448fbab01e4294a1fbeddc8b167775645

Observation 86c9db3c-e837-44ff-a9fc-598806601471 · outbound

This paper cites Montreal forced aligner: Trainable text-speech alignment using kaldi.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Montreal forced aligner: Trainable text-speech alignment using kaldi

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.863538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:c7888152e9a0e2c449cd8d2742995dcc4eca9250489b5e1bd60ae7dae9ce7256

Observation 9138ff39-1341-42d4-9330-8c7a9faac660 · outbound

This paper cites DCASE2017 challenge setup: Tasks, datasets and baseline system.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models DCASE2017 challenge setup: Tasks, datasets and baseline system

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.866991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:84bcb4b854cee37cac502f7708babd9401f28dddf1471f185a83053ed6fd4d12

Observation 7d1d4f61-5064-42a0-8e1d-3aa78034ab6a · outbound

This paper cites Librispeech: AnASRcorpusbasedon public domain audio books.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Librispeech: AnASRcorpusbasedon public domain audio books

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.870304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:fdf535f4581ce28c0b7cfc0953faf88f95d017ab934b14f23d969df489a21119

Observation 24f53383-474b-4b6d-b5df-d15bf63ea882 · outbound

This paper cites Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.873606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:ac99e1ffd4014b13ba29f9edd9fa9f1e4ffd78b14155235202ff6d93535ffc27

Observation 4578af68-d155-46d8-8b0e-6e1618754984 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.726199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:6dd326e689d5cc2ec984a2d327de2b2b3c3f2367cc4f978e6a65ff99ffe5de80

Observation 1b0ec48a-8858-49fd-8f34-0c73b749ae2a · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversations.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models MELD: A multimodal multi-party dataset for emotion recognition in conversations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.876807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:a77072c03ddcd0f0a96066cf05a3c68b26af8b7a6c4477e488cbf013173ab225

Observation 5fe6b423-28aa-4c97-9865-085fe94460da · outbound

This paper cites Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.880806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:2f0223397ea47e8877a20d376c73f349fdcd89225f94fa5c0a2ddc9be749b6bf

Observation fb10e4d5-8f2d-43af-8bde-5bdd7fc47cd6 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:06:45.816188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:c48a65710e4dbe2ef4bcd3b1c6ec1b550a6878d6975c4ba5a547d80fb33deb8b

Observation 2bcc33b7-1fb3-4b57-ad7c-82fe00b902f9 · outbound

This paper cites LLaSM: Large Language and Speech Model.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models LLaSM: Large Language and Speech Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.748529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:b84b7b630a84b73ba9e4e9d6cf56dd28797f38e75cfd356eded25ce2ecc7fae5

Observation d1b73c67-50be-45a3-9491-94e911eddeed · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Emu: Generative Pretraining in Multimodality

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:22:11.462164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:b0d24a600df0d4b507aa418ec0584c92ad8b13699bc99343887281e0ed3d3d5d

Observation 0877f768-621a-4a4f-a3aa-e5cdd7c7ac10 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:57:28.757812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:1d4ca47dfd73270c128a68099542f39c7d5de13fadf3766464d84e074eaa83cc

Observation 492184fe-5f82-4b06-a1a1-3b6b33fe2ea8 · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.762806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:b15385d032867f3a77d2ff7f5794d7b8307450a4074f6cf378e396dd4f4e3a98

Observation e09f75fc-0890-4b05-b290-52f56213f9f7 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.769454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:ef5245fdeafac9fac8ff02ba2f02e81c4d3e54e8c91a02ba45cba233a3b61caf

Observation fe3be170-510f-402f-8ce1-f13f025a6ccb · outbound

This paper cites Whisper-large-v2 Qwen-audio 1st-stage LLM init.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models Whisper-large-v2 Qwen-audio 1st-stage LLM init

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:57:28.827623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:041daaae3d190caf6313a88b4f497cf76b6d7dbb6066d21438a958bbcdbf1255

Pith citing papers

Observation 1fce556b-0194-4d5f-9565-dc42f0d4f564 · inbound

Qwen2-Audio Technical Report cites this paper.

Qwen2-Audio Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T02:14:45.371564Z digest=sha256:d094eb10c187dfc0d4692bc5d0ea311b6ed3785f600c179db1d7be345826811e

Observation 23e023fc-61b4-45fd-a4c8-b6b1a7d5011d · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:50:13.957823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:4959ea82c1c01f82250cafb4e66df6e39bff5b9ff755893c14a47b0596290e39

Observation af400f69-26f3-42aa-8581-4de09686fd5e · inbound

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot cites this paper.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:53:47.586781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:5e40910c115d2e40b7004dbb87a01c5e644ade3f629b9b41ee9f6667e8e322b7

Observation 5a11210e-e87e-4a6c-b8c7-d7f1f4be233d · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:53:26.220877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:4ea10f64797066fad4cd1d1b90ed83f06871b217ba5a2823af3d4da9aa66d577

Observation fee4a322-b101-4a6d-8353-d3e47ddd677e · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 199

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:18:53.363845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:551e677af8a2c452d81b6d58cf008a132bec55c3aded943220a05455f74cd890

Observation f7f4e2d5-eef8-4afe-ac6e-be2886e2d488 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T20:45:08.215895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:e87471953fb8dd9a888007ec40015aa61060d90c4a6a25f02e3230a7c29c06ac

Observation 069f554c-bd79-4cfd-a636-58f85cf251ce · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:77e8be1990bab4105c52d15f99bdcc6ea214617b7d614ca84df35d8f4ceec5b4

Observation a68178db-90bf-4a1b-a8a9-5ee404ebe3b1 · inbound

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping cites this paper.

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:41:36.647489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T13:38:17.640841Z digest=sha256:719188f1852809f723b296e18d54d8fba0b0f7e33a7a5169f42122dafdc185cc

Observation 8b25ced2-3437-4776-af5b-64709d4dc2ab · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.391139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:7281159304cb4e41b3faa5e39e1c60758ad6b984b01f0a28aa8693e9c15bf59b

Observation 2f947a66-9d89-4f63-8508-05c68e9a5cdb · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:55:40.434915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:7e703eb788e08e3f56929920d7ced6609aba1e2150d24cc3e6ebea344aff0ceb

Observation bb4ccade-bec5-4864-aff7-30c54bb5815c · inbound

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games cites this paper.

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:02:16.689598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T12:01:42.681135Z digest=sha256:dbbbff9dd2797124d1eafa0d3056bac81ec09d448678f3c17257a718d63e3431

Observation a85c5273-a4ad-4bb3-83da-cfd9fe8cb236 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:59:50.980061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:6580924198419c0f910dde39e892a012c209915bd10851a26e2b79ea9d0c8852

Observation 5ba525c0-e879-4e1c-ac0c-dfc514a71efe · inbound

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks cites this paper.

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:41:59.607161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-19T02:41:52.996457Z digest=sha256:afc60337505eb62deaa8738103d6bd149f1da85c59669ca638efbb6f0136bc46

Observation 9459bfa7-63e8-4767-b177-81c377a0642e · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:24:23.487552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:ac1c7eac0a02886e2a9086716569aa15d078422087ace76cea9bbf1daf194715

Observation dd84d51e-7157-4d6f-a6f6-17f81ad08e06 · inbound

Direct Simultaneous Translation Activation for Large Audio-Language Models cites this paper.

Direct Simultaneous Translation Activation for Large Audio-Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:36.817162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:823b46ef6477647d8adea3a5675a04a432c65438bfb39e7620b7cfed998f96e5

Observation a15dbc66-93b9-4f32-825e-c029cb124bea · inbound

GraphMend: Code Transformations for Fixing Graph Breaks in PyTorch 2 cites this paper.

GraphMend: Code Transformations for Fixing Graph Breaks in PyTorch 2 Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:06:35.002150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T16:04:29.215393Z digest=sha256:234a850b4f86e2b3219ed399276ece7fda923c905cf5c8d519b6af93fbecf01c

Observation f87db4fa-bc05-4d2f-8cab-d25c11610b5e · inbound

Qwen3-Omni Technical Report cites this paper.

Qwen3-Omni Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T00:20:37.406351Z digest=sha256:aeaf974d0e540e0920e93b9c8942272cd969c8ec575417953e1c81e60e051595

Observation e78fa995-c65f-4887-8e7c-cdeffa3ff2e6 · inbound

Investigating Modality Contribution in Audio LLMs for Music cites this paper.

Investigating Modality Contribution in Audio LLMs for Music Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:50:43.452848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T22:46:49.051113Z digest=sha256:d394f62ed1ddfc73ea51bfa8ece7f82d4d727bf458237900c96948b8e3805f64

Observation a0e35937-98ce-43e6-bdbc-85b965b995a4 · inbound

End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering cites this paper.

End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:45:24.311935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-17T22:44:49.949759Z digest=sha256:22cd14d69dd0ffeaf0badf215c59663aca1c696282011881135d82fef849b097

Observation 5a4f7985-4e58-499b-94de-6d9b360f85cf · inbound

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages cites this paper.

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:18:57.209538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T03:15:04.685150Z digest=sha256:8f70267d3f039df0a3957f6f1803eab311dbb71004ecbe876184a88f0ed31931

Observation b2fca029-712f-4cfe-9d69-cc6c5f6b7842 · inbound

Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching cites this paper.

Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:48:53.066189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:48:53.066189Z digest=sha256:ad79bd1d51b386cd57f367358e48bc6df8c1e5b33d16affbfe12123e05d1e5a5

Observation 938d5d20-137b-438b-92bb-7c6fc392fde4 · inbound

MOSS Transcribe Diarize Technical Report cites this paper.

MOSS Transcribe Diarize Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:49:50.056919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:49:50.056919Z digest=sha256:24eb848cdf745281db74f2b59a2988cbaf9dedb39615463b78c235afaf3a15cb

Observation 82d211a8-aba7-478b-aa7b-03ef17f646a1 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T12:01:59.150857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:01:59.150857Z digest=sha256:82ebc5113f62ae33abd52d8e86d54a4ea5bd020e148f546fba37eee14b851ba4

Observation b4e00556-f75b-4533-98b7-ea35e4ccfbd6 · inbound

AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering cites this paper.

AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:07:58.679402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T14:04:17.935630Z digest=sha256:47f043f4d6d8b3aa199b52d23d020a7d9ce24c85ca7ccf60d43cda146a375564

Observation 915656ab-cbcd-4b88-b8c6-f189a0119684 · inbound

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer cites this paper.

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:42.307716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:42.307716Z digest=sha256:5ad63e076a765d7e708519379af834b9b2df1afe2f7108861d4ff8dd52040e57

Observation 736bbc75-f1be-4855-a73d-8c2fb4eef428 · inbound

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity cites this paper.

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:50.902443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:50.902443Z digest=sha256:40bb87dc9814a2d2e96dd55ddf72313e2cb2a5c756926d6734c5c6efd8cf95d8

Observation 4eaee3f0-f2e7-4f2a-89ec-0f84c995de76 · inbound

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling cites this paper.

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:41:12.014149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T15:40:07.421894Z digest=sha256:3a1da3fdf56d9e15545649d05df8109fe795cefc4bbd6dd4e2753d5d43f70a3b

Observation dec17be9-5a17-4f83-abe7-6e3b2f56bacb · inbound

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision cites this paper.

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T18:42:05.405658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:42:05.405658Z digest=sha256:9d4a2290c7315f3b84c769a414cee44273f0b3d0f2b99c741c2e45b86cdd26b3

Observation c879398c-5410-45c8-b011-e0edb2ffcbaa · inbound

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips cites this paper.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:524a0237b6e8891bd8879ebb4386f46f38c64c0d95b93dee8f5e93ba5ce8d26e

Observation 290cb93c-7d3d-4bcb-bfbd-7e3b0930fc6b · inbound

Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition cites this paper.

Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:22:08.670559Z digest=sha256:4b0141c18994b6759bc519eb1d60d2f7bae8ca9d8ab10357b0cbe62143b7dab4

Observation e9a965b2-0d0a-4ef0-9396-243c375e80d4 · inbound

Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training cites this paper.

Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:31:50.670807Z digest=sha256:114592cb91beb515c946cb5a8cddc7d7d301179042ce67669aa5636f4010baac

Observation bd9ea5f8-f59b-491e-ad1c-ff72371925c7 · inbound

Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS cites this paper.

Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:08:56.412939Z digest=sha256:eb39842efdcbdba04cdc87766884ab406ce3bfe352cf3bea7f309e0fcca264cf

Observation c749619a-078b-4728-8ab2-c3c25b7369e6 · inbound

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models cites this paper.

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:24:27.118694Z digest=sha256:bfcf83fc1af156f9771058e560733b3ddd00a007c479a141de66a8f12b703bf1

Observation c3034d23-7de3-4507-864f-79a5b5ad6212 · inbound

SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding cites this paper.

SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T13:54:52.275011Z digest=sha256:62d33c052829a27a80a70bdb24df2921e6c7eafb5ba07fbbc9e6f06efb75c10e

Observation b52494af-c2e3-4ba0-853e-437c93d78287 · inbound

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt cites this paper.

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T12:37:17.843365Z digest=sha256:9c666ae3e56d9827005ac0dd049d3e671d279421d0815674268ae6041e0b9102

Observation 8379af5e-6e66-4f3f-ad8f-1aa9fea82c16 · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:9a12473df3a49a5f265a0d4b62edffd955f3adc05b16581cd1bdd133dc44e9fe

Observation 7547d7e1-6daa-43f2-8e63-084d53cb4443 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:d93d79caafa9d12e7d3ff8de740e2dedf517afc29f67686130c92e2222344d52

Observation 5b6f4c32-1675-40bd-987f-f73297d3e4fa · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:f4dd7cf146cc564ad62d86078140c78e160527bc1208a87a6cce15cfc001ed81

Observation 70ea96b3-229c-4098-bf72-f96e84c1f355 · inbound

Qwen3.5-Omni Technical Report cites this paper.

Qwen3.5-Omni Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:11:22.402552Z digest=sha256:c5045d90f4fb8178b63ae3485ff3eda106f6ee704794ace2dd32c1431cc59a95

Observation d0052591-51b9-4b3f-9b14-905b6495f83f · inbound

TinyMU: A Compact Audio-Language Model for Music Understanding cites this paper.

TinyMU: A Compact Audio-Language Model for Music Understanding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:17:23.740979Z digest=sha256:abba30bddeaf61e8271951a58fd4b2add26bbc89cfed4f4878677a3982a64fe5

Observation 95e3d43a-fc71-4bc5-87b6-3b5faa34f5a9 · inbound

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval cites this paper.

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T03:23:40.993175Z digest=sha256:f5961e16960cd7bda85e801b7bccc232c2e05996fa5f69afac63ce559fb704d2

Observation 2edfa591-0bad-47f3-b02a-3ae6c8c974dc · inbound

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models cites this paper.

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T01:34:54.375266Z digest=sha256:dc9ef29e36413bce3790187c2b129c1de47c1d8dfa31147b2bd205bd1529e51e

Observation e5ad04c3-579a-43ea-aa50-76833e525b9c · inbound

Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages cites this paper.

Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 193

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T00:34:30.978387Z digest=sha256:df46f439277313ef6399f7f5447065d5e67f9f133df81e0523e18be3e9983e1a

Observation fda59e30-64c1-4c59-bba0-294afcb332f5 · inbound

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence cites this paper.

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T23:03:01.356594Z digest=sha256:04cfbb9d4840b060cfe8471e85a62e932bda6b037b044be91b67f5b356018129

Observation 2c5cbd56-3b87-4777-9b2b-6e83b530d87e · inbound

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions cites this paper.

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T13:25:50.524448Z digest=sha256:1f066377c08fc861ddfdf177c5323b075346b07dd167b213e8f524adc66f5149

Observation abfbf2bf-15ea-4e16-bd20-bb0be9f09530 · inbound

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition cites this paper.

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T18:06:47.859998Z digest=sha256:7a4c820147d54701a1764f266a05386e248a53b65d968b0447906e80327d119f

Observation bdaf2b62-393c-44b1-885c-7505b775a31c · inbound

AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition cites this paper.

AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T07:25:42.870223Z digest=sha256:7f6b27959851bb2c7914cadbc64f42cdf6973fd270c388fae3523a70151a9e59

Observation 9e68c7f7-ce71-4e29-b9ae-f44abb081e5c · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:db6da9351081c394405e4a7b53d77b835319fed9a782bdf4d424b9aadd3b9136

Observation 5729abf1-bc37-4611-9f1e-e94c51ac38bf · inbound

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes cites this paper.

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-11T01:14:43.744389Z digest=sha256:9e19d1099f78ce6f69aac32723e3eaaf3e1d4900e8f768ee003f73ad5c9fac4d

Observation f5cc76f9-0d30-4fda-a649-90c29cfca4d9 · inbound

Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration cites this paper.

Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-12T05:15:55.193062Z digest=sha256:83b231856b20aa7ed4dacb6825c2a95f499af59f6b84574e4a0f05329bfd4cdd

Observation 620ffce7-c856-4ca7-9bb0-3c8f10fadcb3 · inbound

NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating cites this paper.

NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:17:35.034442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-14T18:16:22.762405Z digest=sha256:a5409a12949c40509a0fa077061408df3aaf90fa4722956c1e1f87a6808c6143

Observation cdc99437-c98d-415e-b494-73f7e9b21483 · inbound

SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning cites this paper.

SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:24:55.762374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T03:20:37.488507Z digest=sha256:067627c56f38d2d0e02fea0681129aac558b498cfd661d3470a2346a5a380b93

Observation df4a1cc0-e0b4-4445-bf82-31cf959c3e95 · inbound

Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction cites this paper.

Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:58:13.582853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-20T10:57:40.997128Z digest=sha256:d9f2580f911ec6858ffd1fdd28df182cab78dd7591b48f0a884e9c298d6ae23e

Observation 1e53ada0-156c-4bd4-97c7-dbbb9a207388 · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:13:11.904848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:73b776148c5d0ae9f3621821d706259c4b1e1c6de73266599a432fd5aebb75d5

Observation 9e8fa346-5ab2-42b9-a605-1bb28993f172 · inbound

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training cites this paper.

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T07:18:07.283584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-20T07:13:29.041056Z digest=sha256:e93ed670f6e75379b013a364c50fe4a574fc680b8ea018c3c3a2024fb85c78a3

Observation 4f00b568-ccd7-42fa-9273-d5de3b870bea · inbound

AffectVerse: Emotional World Models for Multimodal Affective Computing cites this paper.

AffectVerse: Emotional World Models for Multimodal Affective Computing Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:48:05.771479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T06:46:33.612905Z digest=sha256:c4a72c51ade8508bdd0c8ff27ca796b289824ddb0dc39b1aff3416028c4409eb

Observation eb3712d9-5bb2-4d85-8236-08592f036432 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:49.032020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:29887fddf8db823ab8a8de02fa0e535e857ebceb0d9e9928340da59b80232b46

Observation d0601d4d-da6c-4519-8b2d-1c2ceed0618d · inbound

Codec-Robust Attacks on Audio LLMs cites this paper.

Codec-Robust Attacks on Audio LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:44:00.638652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T06:43:52.735211Z digest=sha256:006159a3ded804f19d2615c54cf9156713dae3b66d41ef8c68ba6dbf9ba8917d

Observation 479df41f-2545-4208-89ef-c1e4ca2aa2a3 · inbound

Codec-Robust Attacks on Audio LLMs cites this paper.

Codec-Robust Attacks on Audio LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:45:23.736609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-25T05:44:46.831360Z digest=sha256:09d715bf8e85e531ca22d353206584a6c7fc8dae7b7d8f33747e6b21c1567602

Observation 7f4ccaf0-1710-4abd-8708-534d46d21743 · inbound

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods cites this paper.

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:40:53.973928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T01:38:39.499793Z digest=sha256:0b62853b6f5135aa311cc64401730dc101a4a3fa40f001179684c7fdf6a72c82

Observation 5c25e27e-88b1-45a4-872a-19bd9b4dd46c · inbound

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods cites this paper.

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:44:57.896841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T17:37:57.477175Z digest=sha256:c98ff0de36fa47deb5a14f5e44be763949ab69216c35b316543a71f787677f27

Observation 740e0a2d-bb52-478b-b45b-c7f7d49cb0e8 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.933661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:6a8b4e23eb565d5e31d0f9a8a3df8caa15c087c5d48180a106e0608523cd4990

Observation da3ac5f8-e845-4de2-bfd5-97485eec1690 · inbound

Learning When to Think While Listening in Large Audio-Language Models cites this paper.

Learning When to Think While Listening in Large Audio-Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:43:50.761327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T18:37:34.409802Z digest=sha256:8af9ccafd74c920148dffbc87d37d69021e3afcde57cee972cb96a2842738ed0

Observation 0bc2763b-f789-42ea-9194-ebae95b4c9f7 · inbound

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation cites this paper.

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T11:43:23.632863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T11:40:46.165547Z digest=sha256:b5e33b479398916755aab3b73a6431e62270c598e605ab679095087a703802f3

Observation 9d928bee-1a71-4611-b74a-b293c28cb711 · inbound

Decoding Strategies for Diffusion-Based ASR: A Systematic Evaluation of Confidence-Based Thresholding cites this paper.

Decoding Strategies for Diffusion-Based ASR: A Systematic Evaluation of Confidence-Based Thresholding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T05:53:09.221791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T05:45:19.762023Z digest=sha256:142f5026a7c2379e53f42e4c93e2fb68fe541d28b1470b330b6c38a744d1ec2c

Observation aefe8a3d-a213-4bab-a83d-2b41b781e83d · inbound

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion cites this paper.

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-28T20:52:37.927179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T20:44:20.190064Z digest=sha256:7e7db85f31f26368d83f8440812e4c4d7bb360b21dc606102e80cedd01c0f551

Observation 179085e3-992f-4d2a-ae0a-f132a28e8a89 · inbound

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors cites this paper.

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:32:35.160558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T19:23:19.422735Z digest=sha256:d8b2ccb4d74e2bcb5449de716c4f816bac5d586dcee88396c556d261898575cf

Observation f4ca3108-7a9c-4749-b2ad-8a07fd214955 · inbound

MOSS-Audio Technical Report cites this paper.

MOSS-Audio Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:25.088047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:84f8b10a122a85bf649799a65312e03df35f21783cbda92c768a8b44953b48be

Observation 8849e7a8-7d9b-4497-bb3f-0ad7da2baa51 · inbound

SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification cites this paper.

SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:06:39.874784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T08:34:11.270002Z digest=sha256:e8419bcdca2b561d2015503cf0c73e2ec65f64d1c46ddc90cc0aa97fc629f7cd

Observation 723b62a9-8d87-4d49-b458-0387ac1675a5 · inbound

SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification cites this paper.

SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:24:37.761323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T11:24:22.023188Z digest=sha256:3b05a22d88838db4ce68809a217b766e6113d40d2878b94730dda97cc1acd4a9

Observation e2bb7787-a114-40f1-8466-be63fe23eef8 · inbound

Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention cites this paper.

Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:46:46.222824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T06:41:25.592270Z digest=sha256:14e2c4d14be7b581a9fbc486049e2b6ede1fee984c9a09885e9dc3d10c6ad977

Observation eb536c04-8d0a-4304-9e23-2ec0aa406ed2 · inbound

UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning cites this paper.

UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:16:53.506262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T04:22:13.478562Z digest=sha256:2e605299231d048b1b533270dda7056e9de9445f2f5a57de5e96a15c85d320e1

Observation a5855a8f-e10c-4cb7-bb7c-fab514158adb · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T10:46:52.393805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:2200b2894ae691fa48810ad8f0762c640d39fb81c1cacb62abfd423644133042

Observation 932b0ab3-6afd-45de-b862-4d56d7a095e1 · inbound

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models cites this paper.

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:47:20.020338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T21:16:07.007759Z digest=sha256:0ee90d49d277972c483871f242c9e593c71eef4872fb8fb9e41796f7f5dde204

Observation 3176a1fb-2ea5-4502-bd1c-55e08d7d8913 · inbound

Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation cites this paper.

Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:07:08.932270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T23:00:07.144841Z digest=sha256:33738ddff1f25857ae4b18bba7cfc37744b180db86959081fd53212cfad309f8

Observation 2eae3c81-fd6b-4d33-881c-12a91f9888d2 · inbound

Titans-as-a-Layer: Test-Time Memory for Conversational Speech Emotion Recognition cites this paper.

Titans-as-a-Layer: Test-Time Memory for Conversational Speech Emotion Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:07:26.859912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T18:28:44.652813Z digest=sha256:d10669816398166bd010a55ad8672d881cc8472922c1fd68602113f3e690e414

Observation 4ba04366-f80d-42f3-b828-93d0309bccac · inbound

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs cites this paper.

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.605027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T16:26:32.594986Z digest=sha256:11a9e227f4b13e09d803777e246d63bae06ea3094dfb91fb2eb9621123c021e7

Observation c241abff-92c0-4a8e-9fa7-a78e8cf487c2 · inbound

A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales cites this paper.

A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.830670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T16:33:41.230574Z digest=sha256:8988d8e908699cbd2d07329b3a74f785a5ec7339816a2c350aac7e819e6972a0

Observation e4925622-4525-4f07-9bf9-2d1dc972c69e · inbound

Speech Encoder Fusion for LLM-based Automatic Speech Recognition cites this paper.

Speech Encoder Fusion for LLM-based Automatic Speech Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T07:57:44.685719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T11:32:58.335783Z digest=sha256:de2d8be6f8d9f6efdf8eb30bd06e5f944fd371787f775b457bfa5886e8de6e28

Observation b6ca04b8-b71d-4992-bd17-b1f28879ad73 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:15:05.345453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:0e3842128b34a0b12a3e8dd73bb23b638bc1c456a3ce618c78563a984262b8f3

Observation 9452a331-52ee-4638-9edb-3d048f2d355a · inbound

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark cites this paper.

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T07:17:43.937486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T12:11:56.629002Z digest=sha256:f137479eab0c834b35a0d9ac16eb0484b61dc3b6b599702d1bc6de370d10eaa2

Observation f1b06a76-9818-43c9-900a-f37b5ca039e7 · inbound

DeceptionX: From Multimodal Evidence to Explainable Deception Detection cites this paper.

DeceptionX: From Multimodal Evidence to Explainable Deception Detection Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:27:40.118590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T13:15:49.796647Z digest=sha256:ce3d44ac350ebb93911dbb854363ef35a82137b6d69ea0e7877e6d1c61ccc9fe

Observation 9b7135e4-56ed-4edb-95a9-bb132db28181 · inbound

DeceptionX: From Multimodal Evidence to Explainable Deception Detection cites this paper.

DeceptionX: From Multimodal Evidence to Explainable Deception Detection Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T03:03:32.342985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:03:32.342985Z digest=sha256:254ef62553aa9c318f354f9ec2b1045a7ed39b588e2e8f722484d0fda8d09e1d

Observation ac7ff34f-1db1-4410-b224-25bed8213b59 · inbound

Continuous Audio Thinking for Large Audio Language Models cites this paper.

Continuous Audio Thinking for Large Audio Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:47:18.010086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T21:52:49.839901Z digest=sha256:cf7d597f3e03d473e02361536caae4a00f910b3665134b6543c995c751c498ee

Observation 619b0362-c950-408b-be0b-f8183fec4b63 · inbound

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era cites this paper.

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:19:44.206175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T11:54:08.457573Z digest=sha256:75345a9cfee6022613dc06115486fb677cca5ae8d7d64869cd7a02153f1089de

Observation 0b2d104a-d72a-4026-8a0a-2d0483a7cfd6 · inbound

Uncertainty-based Debiasing and Unlearning for Decontamination cites this paper.

Uncertainty-based Debiasing and Unlearning for Decontamination Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T12:49:52.436103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-26T05:57:44.576411Z digest=sha256:f07e963f1d4831a74b5d88194374788116e9e46e727c875e4825eea85b295060

Observation 5ea7e87f-1ab4-41c0-9425-0c9ebd3b1236 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 141

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:20:06.521202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:cb0b56c8dc55fe3485793f502fec678108e834f80dc2872f43d851ec06fbbeac

Observation 4fb3ee76-5ee9-4a24-9bed-c477ba7f6356 · inbound

Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models cites this paper.

Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:30:07.197249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-25T20:08:33.322532Z digest=sha256:1f3582c80615cf71e6830f58b9c527bb71635db49c99fe46d3ba73fb501c753a

Observation 4d7d6bd8-907e-4ffd-8ad1-d78f8f3a39b2 · inbound

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs? cites this paper.

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs? Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:30:08.192339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-25T20:03:10.858349Z digest=sha256:4c25ad49629e365531d1c13b4648dbaa3810b8e778b12090fd5fd4e7f651f168

Observation 9df4c430-c924-4e36-9570-2ed681fabf5e · inbound

wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval cites this paper.

wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-04T14:39:58.377508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T02:49:35.819155Z digest=sha256:1a55a4ced266e527ac66aed05144dbb3afbadf94d47cf78bcdc82a7dd52ac07d

Observation 83072bef-a268-4225-a828-3e1b98b061ed · inbound

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy cites this paper.

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:51.453211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T05:06:09.216428Z digest=sha256:4415163f027a603ab25a7af3a78bb1d926f0613b8726dcf2b8df2dfc5f6cc1a7

Observation a1f64d49-bfc4-47c2-8d5c-48b413637a29 · inbound

How to Leverage Synthetic Speech for LLM-Based ASR Systems? cites this paper.

How to Leverage Synthetic Speech for LLM-Based ASR Systems? Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T12:54:40.667112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T09:35:03.660671Z digest=sha256:c947ac9242a4674c38a2ace629f374309092c686487a084c00a569766863e6b0

Observation 89de6386-c342-448a-951d-515551432d5f · inbound

How to Leverage Synthetic Speech for LLM-Based ASR Systems? cites this paper.

How to Leverage Synthetic Speech for LLM-Based ASR Systems? Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T11:11:08.878850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:11:08.878850Z digest=sha256:6eae6fa61d2e9f6614d0b1a9b22d4fa6f252206cd37ce6b9210c05aade2e6140

Observation 0891cb16-4399-4f0e-b866-501bbb99ede7 · inbound

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs cites this paper.

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:21.799793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T07:26:38.380118Z digest=sha256:7c707b9566ff239a991f5eaaa0fd22ae450b209d7dc1d9a15b644f23771693f1

Observation cc259b6a-0406-45d8-868c-c62268c37d7d · inbound

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models cites this paper.

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T11:55:42.994017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-01T03:31:42.611472Z digest=sha256:d6d5e1a53bbc758739d39bf9a69344afc645622ef647bdcb2473d405e15d0484

Observation b8215101-0e5e-467c-8a2a-491459e6a730 · inbound

CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds cites this paper.

CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T03:03:59.776103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T03:03:59.776103Z digest=sha256:26706ba56edadfee3bfc23a94a9871c11a37890b5b2ea80366c685f5c89cf26e

Observation 2f60d2a8-ff7e-4bd3-a57e-289062cf1888 · inbound

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding cites this paper.

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T19:34:49.358453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:34:49.358453Z digest=sha256:215933836a3436ada85522c5fe447a51cdfb1369265337d98d546d5c94afec68

Observation ad3f1393-442d-4c5a-b34a-89707874ac6f · inbound

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding cites this paper.

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:43:33.132312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:43:33.132312Z digest=sha256:b68269cbdd0423b1bed8979a84626399f4864f60bfaaf34d8845128902574e92

Observation 2c0c3017-8736-467a-85ae-40de1a7a8f12 · inbound

Context-Aware ASR for Mandarin Technical Lectures cites this paper.

Context-Aware ASR for Mandarin Technical Lectures Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T09:26:54.286790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T09:26:54.286790Z digest=sha256:e06705887f2dc055ccac5fc31ce3445979df8b2d4ceba85dc58ab645f72a1189

Observation 71b2450c-44af-4992-b2b3-d8adbe4b8090 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 136

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.353380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:d47b4b6afd2302e5079fee23f4f30df6e20a9241474dc980d33d71438dc39885