Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:02:41.528122Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.01881.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:02:41.528122Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 77fb404d-c569-4f15-a722-967f86973136 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models FSD50K: An Open Dataset of Human-Labeled Sound Events
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7211c381-a8d4-40ec-9e22-b8567e9480d9 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e2a3dbc-9498-453c-a23b-34a3e2ccf11a · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6187c0fc-5fe1-4a16-ad07-e5a4d9c68382 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models InIEEE Au- tomatic Speech Recognition and Understanding Workshop, ASRU2025,Honolulu,HI,USA,December6-10,2025,1–4
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b09f8e2-b07b-4cbf-ad50-41d679f06ce7 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models In Lacerda, F., ed.,18th Annual Conference of the Inter- national Speech Communication Association, Interspeech 2017, Stockholm, Sweden, August 20-24, 2017, 2616–2620
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f5ce39-792b-48b8-87a2-61e208bc2c30 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b4f6fc5-f7b4-4187-9e5a-c35004e175f5 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Sakshi, S.; Tyagi, U.; and Kumar, S
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98553ed0-0e79-4d3e-a5ce-dc733f5d1faa · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Qwen3-Omni Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc62cbbc-a02d-4d4d-bb71-bace030c9f48 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Qwen3.5-Omni Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b09b3f9-95d8-4749-8218-152e34f768e0 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Tong, S.; Li, X.; and Wang, Y
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f03065c-a706-42ed-8410-2f5b6d2d8c21 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Wang, B.; Zou, X.; and Lin, G
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0e38b8-5a61-44bb-a818-8123a40a4ac6 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f91bf7a1-04a2-4871-b173-21ed1dd452e9 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c834345b-e7e9-450f-a683-d09f2176cb2e · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Audio-Mind: An Auditable Agentic Framework for Audio Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb13be1e-37ad-453a-a48e-e22c2db5df73 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Xie,Z.;Lin,M.;andLiu,Z.2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be4a0b8-7f58-41f1-a43b-7f2bb47ef9ac · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models CoRR, abs/2606.15141
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13256ed0-1e9f-4e28-bbfc-217dd82fbb6a · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8bcbcec-dc56-4ace-9a1a-ea4a8567596b · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d71e7252-bb08-46c4-9f06-f0a42d6c960f · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models LibriMix: An Open-Source Dataset for Generalizable Speech Separation
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 724f824a-7e93-4787-920e-4654dc3fc791 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dcc0451-20e1-44a7-b907-4a2f4d95a0fc · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c752120-f0a1-4d5f-b55f-f14617d2bbb0 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132df9a7-37fb-47bd-a385-bfcccba9b1e1 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adbdab8b-0935-4d1e-abf1-c57a5fb4f5e9 · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models KimiTeam;Ding,D.;andJu,Z.2025
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d5a1be-9bd3-4ea1-8cc1-7032541fc61f · outbound
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models CoRR, abs/2602.10439
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.