Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T13:05:29.813707Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 6 inbound Pith citation observations for arXiv:2606.01802.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T13:05:29.813707Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T08:35:56.402445Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-07T14:53:55.822643Z
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation af73f21c-68b5-4ae8-9f05-e1997837d6fb · outbound
MOSS-Audio Technical Report MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f448aabb-33d2-4c3d-9271-026bb43374f2 · outbound
MOSS-Audio Technical Report Audio set: An ontology and human-labeled dataset for audio events
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae8c38fa-9e02-4ccc-9d8c-db4a3f5fe9b0 · outbound
MOSS-Audio Technical Report A udio C aps: Generating captions for audios in the wild
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42c08a71-b4c5-4add-817f-8cc304eda715 · outbound
MOSS-Audio Technical Report IEEE/ACM Transactions on Audio, Speech, and Language Processing 29, 1161–1173 (2021).https://doi.org/10.1109/TASLP
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 68aa2f45-0b4c-4c0e-aaa6-a43f8b2c6865 · outbound
MOSS-Audio Technical Report Robust speech recognition via large-scale weak supervision
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ca3108-7a9c-4749-b2ad-8a07fd214955 · outbound
MOSS-Audio Technical Report Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9af80082-ed7b-4cdd-beda-7d9c93e5f27a · outbound
MOSS-Audio Technical Report SALMONN: T owards generic hearing abilities for large language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed21e055-1ae8-4962-8023-f7d1eb8b0350 · outbound
MOSS-Audio Technical Report Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dadd2c6-621d-4838-83a5-e8afc13519c4 · outbound
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 80a6f3c8-f54f-46e1-b2fc-7ba9f9949b4e · outbound
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa56ac27-c8be-475c-b64d-18ff7d03ab88 · outbound
MOSS-Audio Technical Report Sakshi, Oriol Nieto, Ra- mani Duraiswami, and Dinesh Manocha
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c0ed9ed-f01f-4924-a57e-cdbe6b4f47c5 · outbound
MOSS-Audio Technical Report Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cbf4515f-2c82-412e-82a8-fcc8532e758e · outbound
MOSS-Audio Technical Report Enhancing temporal understanding in audio question answer- ing for large audio language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c3e04d-08df-463a-af71-f8221bc326d2 · outbound
MOSS-Audio Technical Report Sakshi, Utkarsh T yagi, Sonal Kumar, Ashish Seth, Ramaneswaran Selvakumar, Oriol Nieto, Ramani Duraiswami, Sreyan Ghosh, and Dinesh Manocha
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1785c1d-8c56-4343-a04d-f990d3fcfe49 · outbound
MOSS-Audio Technical Report MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2fa4aeee-6e4e-44ed-b021-194bd073135b · outbound
MOSS-Audio Technical Report Joint Audio and Speech Understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8cf30ed0-c42a-477a-9c38-10e81a1c5d40 · outbound
MOSS-Audio Technical Report DeepStack: Deeply stacking visual tokens is surprisingly simple and effective for large multimodal models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7bcb2bb-69a6-48ba-85d5-6edbed107792 · outbound
MOSS-Audio Technical Report MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9a4d7d6-2427-414e-a2be-34b3cf0d5bda · outbound
MOSS-Audio Technical Report Transformer Transducer: A Streamable Speech Recognition Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29492b91-fa32-486a-bd0b-e733ab0bd189 · outbound
MOSS-Audio Technical Report Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82c14792-db0e-40de-9be4-3c637d99cb18 · outbound
MOSS-Audio Technical Report SUPERB: Speech processing Universal PERformance Benchmark
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9155490e-5eca-45c3-b645-19c8d3a8447b · outbound
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c031f99c-4eee-4423-9a11-79cc20a86d16 · outbound
MOSS-Audio Technical Report MOSS Transcribe Diarize Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85af66fd-6d64-46cf-82ed-a73c719d329e · outbound
MOSS-Audio Technical Report BEATs: Audio pre-training with acoustic tokenizers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a674584b-8f8f-40dc-9bf0-fbe39c94a07e · outbound
MOSS-Audio Technical Report Effective Pre-Training of Audio Transformers for Sound Event Detection
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 991c5787-464c-4103-a247-2600900de08c · outbound
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 435d8d96-5cc5-4fa0-bbd8-6286f3af106c · outbound
MOSS-Audio Technical Report Fun-ASR technical report
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3540138f-e7be-41da-b26b-03a4953c4ced · outbound
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 58e70b04-d5df-4ef7-be84-d9fd491a4d04 · outbound
MOSS-Audio Technical Report Bag of tricks for efficient text classification
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 180dc482-029c-4068-b5b0-9eb89575ada1 · outbound
MOSS-Audio Technical Report FastText.zip: Compressing text classification models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d5aab964-c959-454b-9230-7efd00127a8f · outbound
MOSS-Audio Technical Report Scaling speech technology to 1,000+ languages
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89d43869-c876-4363-b6ea-e8f5e9620074 · outbound
MOSS-Audio Technical Report Scaling speech technology to 1,000+ languages
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a03b87e3-1d1c-434f-a5c0-672390f50972 · outbound
MOSS-Audio Technical Report Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfec829-f9cd-47b5-b323-07c90c8155e2 · outbound
MOSS-Audio Technical Report Leveraging self- supervised learning for speaker diarization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c8970c1-e9b3-4385-96df-fe2f77c33b87 · outbound
MOSS-Audio Technical Report Xiangyang Chen, Shuzhao Li, Xiuwen Zhu, Yongfan Chen, Fan Yang, Cheng Fang, Lin Qu, Xiaoxiao Xu, Hu Wei, and Minggang Wu
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cabe4df8-1509-4859-98be-f8d12619485b · outbound
MOSS-Audio Technical Report Listening be- tween the frames: Bridging temporal gaps in large audio-language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f1b8ab3-fdad-4dd4-b7e7-880150545f25 · outbound
MOSS-Audio Technical Report Bryan, Zeyu Jin, and Justin Salamon
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 544e3d97-bbb5-4a42-b50c-e3fec440bbf8 · outbound
MOSS-Audio Technical Report Music Flamingo: Scaling music understanding in audio language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 570a94d8-0279-42fb-b1c4-ed2ab0745ded · outbound
MOSS-Audio Technical Report Sakshi, Jaehyeon Kim, Wei Ping, Rafael Valle, Dinesh Manocha, and Bryan Catanzaro
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d99664-cffd-4375-9bab-48efff57341d · outbound
MOSS-Audio Technical Report Approximate note transcription for the improved identification of difficult chords
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13dc7cf9-557f-4237-953d-65435b0cbd77 · outbound
MOSS-Audio Technical Report Beatnet: Crnn and particle filtering for online joint beat downbeat and meter tracking
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46cd78cd-555b-4402-b3fe-5dfd99eb7535 · outbound
MOSS-Audio Technical Report madmom: a new Python Audio and Music Signal Processing Library
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5c39b61-9ab2-42cf-bf52-15183feb8203 · outbound
MOSS-Audio Technical Report Essentia: an open-source library for sound and music analysis
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62dfaf62-8a9e-4991-8e20-170792b8da2b · outbound
MOSS-Audio Technical Report Recent developments in openSMILE, the Munich open -source multimedia feature extraction toolkit
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96a6de44-74cc-458e-b677-f25f64246bed · outbound
MOSS-Audio Technical Report Codified audio language modeling learns useful representa- tions for music information retrieval
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892d7e75-acc7-493f-b0f1-01096d5d30e0 · outbound
MOSS-Audio Technical Report SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 731d1eee-03d8-462a-b9b5-bc1eae320375 · outbound
MOSS-Audio Technical Report SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a3a7cb2a-6403-4ac3-afa6-4eb6e0ac3465 · outbound
MOSS-Audio Technical Report Unified Speech-Text Pre-training for Speech Translation and Recognition
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cc14cf4-5788-4e9a-b665-2d17d061b13e · outbound
MOSS-Audio Technical Report SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 30a14c5d-4d17-4ed7-8736-72fb790a3865 · outbound
MOSS-Audio Technical Report Spirit-lm: Interleaved spoken and written language model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96592c46-965e-4138-bf16-40c353f910c1 · outbound
MOSS-Audio Technical Report Moshi: a speech-text foundation model for real-time dialogue
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 314060b0-211a-4bc5-82a2-6dc28975b8ea · outbound
MOSS-Audio Technical Report Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 363b0d7f-2e34-4f19-b7e8-1c3d336660ec · outbound
MOSS-Audio Technical Report GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 181163d2-d0d8-453f-8e9a-015a68694417 · outbound
MOSS-Audio Technical Report Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4b3731c-80aa-49c6-84d1-d6fe04227c59 · outbound
MOSS-Audio Technical Report Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 971d39a6-a5f5-47b7-909f-c7eb6657c1bf · outbound
MOSS-Audio Technical Report Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3992ce51-9e1f-41ea-9f6e-6ae2ebcbacb9 · outbound
MOSS-Audio Technical Report CLAP: Learning audio concepts from natural language supervision
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cddd868-3b8f-4fc0-b0e5-10c8b458dd5a · outbound
MOSS-Audio Technical Report Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb0502f-60fc-4dec-b079-65cbe03b1652 · outbound
MOSS-Audio Technical Report High Fidelity Neural Audio Compression
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9483aace-6f76-4ce4-8758-b80fa5199566 · outbound
MOSS-Audio Technical Report High-fidelity audio compression with improved rvqgan
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c02b7ec1-3b85-4b62-a36a-ec76c0345d7d · outbound
MOSS-Audio Technical Report SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1065d7de-b651-49af-9d75-51695c2fc376 · outbound
MOSS-Audio Technical Report Codec does matter: Exploring the semantic shortcoming of codec for audio language model
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91131f25-74cd-4d44-9320-fcf040698176 · outbound
MOSS-Audio Technical Report SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bec92328-9791-408e-9a7c-b75b49aadede · outbound
MOSS-Audio Technical Report TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0e53b4a1-9a58-4b03-aca8-204bec188cd6 · outbound
MOSS-Audio Technical Report The interspeech 2026 audio encoder capability challenge for large audio lan- guage models, 2026
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85c0811a-7d11-4253-898c-821335d34d10 · outbound
MOSS-Audio Technical Report MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72f844c2-04dc-4cec-9949-df97b3a813a5 · inbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation MOSS-Audio Technical Report
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 765abb12-078e-4eb1-b78c-613e5d6aab34 · inbound
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing MOSS-Audio Technical Report
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 259ddbf8-222e-4dfc-9a60-ca007a266265 · inbound
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing MOSS-Audio Technical Report
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c3c4844-4d0a-458c-9cdd-2ee5f9de746b · inbound
SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations MOSS-Audio Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ce1bda-0752-4599-90f8-53709420fa9e · inbound
Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering MOSS-Audio Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ab4d448-f56a-435f-9ae3-44456c282ad9 · inbound
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens MOSS-Audio Technical Report
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.