Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T15:40:07.421894Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2603.05094.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T15:40:07.421894Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-15T15:40:07.421894Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-15T15:41:12.087219Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c2ad3e17-146d-47eb-8c94-d0335a3f926a · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c1d7ee02-2725-4d0a-b589-b43a669cf98f · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 58dc4f1c-1a71-4e0c-bdd4-4c66a4cf9c8a · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling acoustic long-tail
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3a18b3d-3762-497f-a18e-4ec2015e42aa · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eaabe514-74a4-4f48-ace0-10d9ab9592d6 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling To preserve speech-free soundmarks, clips where both ASRs yield empty outputs bypass the text check
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5c7abf1-20f0-4e0f-b76f-ef565e77dd8e · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61b305de-86fc-4728-85e5-ce1b2c649505 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Hello everyone
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ba5f6592-67c2-4563-a3e3-aafa29c606b3 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Experimental Setup Implementation Details:The proposed model, Tai-LALM, is developed as a localized adaptation of the DeSTA 2.5-Audio framework
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 944ae7f0-7b37-4398-97ba-a9f7769ac5c6 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Architectural scal- ing alone is insufficient for robust sound-to-meaning ground- ing without localized acoustic semantics
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e02ad1cd-4a20-4dcf-8ac0-0ab1f5aaf742 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Beyond the Taiwanese context, this pipeline offers a method for regional adaptation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ab9292e-e197-4519-b3b5-3d214061f0da · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling The perfor- mance gains underscore the necessity of the VGC pipeline for robust training-time curation and Dual-ASR arbitration for sta- bilizing inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b65cae82-50fa-4474-8a71-11e63c0c172f · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 75e9f46a-c54b-45ad-b32d-60bdc1d41b2e · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Dynamic-SUPERB Phase-2: A collaboratively expanding benchmark for measuring the capa- bilities of spoken language models with 180 tasks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8dc8a962-a94b-49f4-a55e-d0b1629820e1 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Listen, think, and understand
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e0a623f6-c828-4464-b9e0-9fc2fc9bb675 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling SALMONN: Towards generic hearing abilities for large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4eaee3f0-f2e7-4f2a-89ec-0f84c995de76 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cce735d0-dced-49b6-86da-d08e9df6816c · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling CultureLLM: Incorporating cultural differences into large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3adab25d-ddcf-42b0-b2ae-5dffbb57b0ea · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Universal paralinguistic speech representations using self-supervised con- formers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a00cccf8-d019-448b-8617-d4d74dfe1569 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b4a85159-a69d-4ed1-a07f-9f381714c7cb · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04407e70-5fd3-4939-afd3-cddb23e862ff · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Mitigating subgroup dis- parities in multi-label speech emotion recognition: A pseudo- labeling and unsupervised learning approach
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0110998a-97a3-4756-a826-22975c002534 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling AudioSet: An ontology and human-labeled dataset for audio events
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cd723bb6-25f5-4c47-808b-76a4968ca288 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Lib- riSpeech: An ASR corpus based on public domain audio books
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31f3d6d9-cfe5-4661-a66a-4cf0dea0c4cd · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9bc4780a-01bc-4901-8c04-2fb1c7fb6497 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling AudioGen: Tex- tually guided audio generation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da662ac1-43f3-41c4-bf05-78d6facf5437 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90cd461a-39cf-41a4-a45f-a9b9926f70e0 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b8819a61-e5ed-4733-8bf8-cacc9813cd48 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Ke- Speech: An open source speech dataset of Mandarin and its eight subdialects
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0863d271-bea5-4dde-8292-44d96213b520 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling LESS: Large language model enhanced semi-supervised learning for speech foundational models using in-the-wild data
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b00ad799-b7d4-4046-a998-7c2f3cda719e · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Training language mod- els to follow instructions with human feedback
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation efa01878-ef1b-4e10-8478-cd0d23bff92e · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Data-centric lessons to improve speech-language pretraining
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 300b3fd7-8048-4c67-ae5b-09cced34f132 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Reducing ob- ject hallucination in large audio-language models via audio-aware decoding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 075cd9f5-c2ee-428d-8f81-d231c6423f2a · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Beyond transcription: Mechanistic interpretability in ASR
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 46975b0f-9fff-4e39-bb9c-6e1e06c911ac · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling TAU: A benchmark for cultural sound understanding beyond semantics
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63112cbf-69af-4d20-b282-69d48e16c237 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling DeSTA2.5-Audio: Toward general- purpose large audio language model with self-generated cross- modal alignment
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fb14dec9-9e30-49d3-ab98-5dc0eda6dfb4 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Attention- passing models for robust and data-efficient end-to-end speech translation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df6cac8d-acff-4af8-b471-8312b88f3a36 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Lost in transcription, found in distribution shift: Demystifying hallucination in speech foundation models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eafd01da-5065-411d-8b0c-f6a289d81f4a · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Teaching audio-aware large language models what does not hear: Mitigating hallucinations through synthesized negative samples
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2086a53c-cd28-4c44-8134-1640c39e5fea · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd2de4d1-4754-464e-be92-2f2d992f940f · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Robust speech recognition via large-scale weak su- pervision
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 814d604a-bab6-4ee4-a3bd-e5d0fba97762 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c05e7224-bbf4-4475-ab49-064a0f8c7e9c · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d1ab2d00-a676-43e9-bbfa-04079f2a6d73 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Qwen2.5-Omni Technical Report
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 33308ca8-b00c-4034-be1d-b4d7af0aec01 · outbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling Qwen2-Audio Technical Report
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c2ad3e17-146d-47eb-8c94-d0335a3f926a · inbound
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.