Pith. sign in

Paper Citation Record · LEDGER

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

As of 5 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.01881.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01881 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:02:41.528122Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77fb404d-c569-4f15-a722-967f86973136 · outbound

This paper cites FSD50K: An Open Dataset of Human-Labeled Sound Events.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models FSD50K: An Open Dataset of Human-Labeled Sound Events

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.771164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.771164Z digest=sha256:8e0bff4db93d5b1766cce44bbf303e390ea64f30909b5228344e5eb87193b79b

Observation 7211c381-a8d4-40ec-9e22-b8567e9480d9 · outbound

This paper cites Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.932166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.932166Z digest=sha256:debf0d952b9e5497e0e9865f42cdc6b89250ee6b4632e5647f2176aa04215054

Observation 2e2a3dbc-9498-453c-a23b-34a3e2ccf11a · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.013510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.013510Z digest=sha256:dfe699eca56fef437ee79ebeeb93ef9a5f9fa312cf9c5479bbc113ea698b7bf5

Observation 6187c0fc-5fe1-4a16-ad07-e5a4d9c68382 · outbound

This paper cites InIEEE Au- tomatic Speech Recognition and Understanding Workshop, ASRU2025,Honolulu,HI,USA,December6-10,2025,1–4.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models InIEEE Au- tomatic Speech Recognition and Understanding Workshop, ASRU2025,Honolulu,HI,USA,December6-10,2025,1–4

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.053261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.053261Z digest=sha256:56036f2630643d694bb0bb96becca1b9fbce1fd92d585156041cb20a9205357f

Observation 8b09f8e2-b07b-4cbf-ad50-41d679f06ce7 · outbound

This paper cites In Lacerda, F., ed.,18th Annual Conference of the Inter- national Speech Communication Association, Interspeech 2017, Stockholm, Sweden, August 20-24, 2017, 2616–2620.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models In Lacerda, F., ed.,18th Annual Conference of the Inter- national Speech Communication Association, Interspeech 2017, Stockholm, Sweden, August 20-24, 2017, 2616–2620

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.067683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.067683Z digest=sha256:34b1df732ad17445129700120bcc2e224f0b55b7da7b1642c1b2f00c2d412f0b

Observation 34f5ce39-792b-48b8-87a2-61e208bc2c30 · outbound

This paper cites Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.107158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.107158Z digest=sha256:0d822b977aa5e5b52f8c6114a465a9d071d7ab18c8173811aaa1c1270ca4630e

Observation 1b4f6fc5-f7b4-4187-9e5a-c35004e175f5 · outbound

This paper cites Sakshi, S.; Tyagi, U.; and Kumar, S.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Sakshi, S.; Tyagi, U.; and Kumar, S

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.148943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.148943Z digest=sha256:bb7534ad0cd6c110b41f8e5237938304621b4af710832c5b91f20643c7517a4e

Observation 98553ed0-0e79-4d3e-a5ce-dc733f5d1faa · outbound

This paper cites Qwen3-Omni Technical Report.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Qwen3-Omni Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.221713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.221713Z digest=sha256:fdb9464bfe979840cd1cffec4e0ad54b0327aaf8ef93b176af5044929eb86b3f

Observation bc62cbbc-a02d-4d4d-bb71-bace030c9f48 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Qwen3.5-Omni Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.256146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.256146Z digest=sha256:cfe72e464ac0ce4637ecfd5553838b764fd080a6d44b1ceb5c214c6fc7c75f32

Observation 4b09b3f9-95d8-4749-8218-152e34f768e0 · outbound

This paper cites Tong, S.; Li, X.; and Wang, Y.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Tong, S.; Li, X.; and Wang, Y

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.293091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.293091Z digest=sha256:2c4cdd351418774b48a3dfea487ee3e6afe9b52954fd3ea7a89f337ad2bde6ee

Observation 3f03065c-a706-42ed-8410-2f5b6d2d8c21 · outbound

This paper cites Wang, B.; Zou, X.; and Lin, G.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Wang, B.; Zou, X.; and Lin, G

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.335720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.335720Z digest=sha256:d13de26fbdef30f5d2b550815af35e30f510bdf5d4a8e049da32050bcf82dd08

Observation 6f0e38b8-5a61-44bb-a818-8123a40a4ac6 · outbound

This paper cites an unresolved cited work.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.353060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.353060Z digest=sha256:c78047a72ef345e78af35d7327da123f5f4d1706906b72349cec0f41f9514d62

Observation f91bf7a1-04a2-4871-b173-21ed1dd452e9 · outbound

This paper cites MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.392890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.392890Z digest=sha256:2d17ee848b68eb9557408ef0d485c4c89be40afb4f3fd22f1ebbd6d7098e7e0c

Observation c834345b-e7e9-450f-a683-d09f2176cb2e · outbound

This paper cites Audio-Mind: An Auditable Agentic Framework for Audio Understanding.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Audio-Mind: An Auditable Agentic Framework for Audio Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.428446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.428446Z digest=sha256:735cf20bae7fc7b99fb6f6128f2d53d8979a6a7139bc79da756b29e39500cb41

Observation eb13be1e-37ad-453a-a48e-e22c2db5df73 · outbound

This paper cites Xie,Z.;Lin,M.;andLiu,Z.2025.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Xie,Z.;Lin,M.;andLiu,Z.2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.470759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.470759Z digest=sha256:25a4a71f5d1a501174f71ef82b59bf675a741409a8f855538262c6fa98ca1a81

Observation 9be4a0b8-7f58-41f1-a43b-7f2bb47ef9ac · outbound

This paper cites CoRR, abs/2606.15141.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models CoRR, abs/2606.15141

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.491746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.491746Z digest=sha256:5f731fe24199ae0678c550fbb82f849fab46232560bea4c4e3198620e1d5bd1a

Observation 13256ed0-1e9f-4e28-bbfc-217dd82fbb6a · outbound

This paper cites AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.653124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.653124Z digest=sha256:e9230f70798fd9cf6664671d0ec86c67d44df1d630083d26bdef1eaf091ad09b

Observation c8bcbcec-dc56-4ace-9a1a-ea4a8567596b · outbound

This paper cites Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.885771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.885771Z digest=sha256:32a40417fb66a8782a15e71f351fe13dc977583ca0c2ad666d384ffda7f78e78

Observation d71e7252-bb08-46c4-9f06-f0a42d6c960f · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.730450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.730450Z digest=sha256:a98fc221279dda5016a13cf430d74c4afe18214711a2b8962a053c9ab919b8bb

Observation 724f824a-7e93-4787-920e-4654dc3fc791 · outbound

This paper cites AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.804727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.804727Z digest=sha256:4d7d79e8cb53b820befa8feffb632f7f61fde71feb18f209a419eb66b8af4867

Observation 4dcc0451-20e1-44a7-b907-4a2f4d95a0fc · outbound

This paper cites ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.969960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.969960Z digest=sha256:56048f6d0af37c46f3dd0478018957f33b6606ec6798e6e846c4864c63e78aaa

Observation 9c752120-f0a1-4d5f-b55f-f14617d2bbb0 · outbound

This paper cites LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.528122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.528122Z digest=sha256:98f1ebba54c34f6d8c580b3efd3c9a0a860e1e0947468c2ddac1f5a8c4f2d3f6

Observation 132df9a7-37fb-47bd-a385-bfcccba9b1e1 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.181470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.181470Z digest=sha256:ec2b2b6a58ad6d831cf6dd5291e9da212335be438eed9c051176da3b73449d53

Observation adbdab8b-0935-4d1e-abf1-c57a5fb4f5e9 · outbound

This paper cites KimiTeam;Ding,D.;andJu,Z.2025.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models KimiTeam;Ding,D.;andJu,Z.2025

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.849020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.849020Z digest=sha256:0cb90d082190b70f94aecf6a38de5c27599e9209c634b60445193066d35d4ef4

Observation a4d5a1be-9bd3-4ea1-8cc1-7032541fc61f · outbound

This paper cites CoRR, abs/2602.10439.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models CoRR, abs/2602.10439

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:40.693596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:40.693596Z digest=sha256:f5e4b3349940b03ab918198fee7c338e68307280708e7e59ec2c39ffc2d0b64e

Pith citing papers

No inbound Pith citation observations are available.