Pith. sign in

Paper Citation Record · LEDGER

Audio-Mind: An Auditable Agentic Framework for Audio Understanding

As of 7 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2605.28480.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28480 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T10:03:54.653164Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:02:41.428446Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact39
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1506217-13ea-4ee1-b44f-b84d1c492850 · outbound

This paper cites Qwen3-VL Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.942116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:d320c8067d1d2d2a5a0cb06f09fc3b7883c308b8bcbf5e85a5cf4c2c5eae4552

Observation 7f46dda7-5abc-4b0b-9806-0ebd148e07b2 · outbound

This paper cites WhisperX: Time-Accurate Speech Transcription of Long-Form Audio.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding WhisperX: Time-Accurate Speech Transcription of Long-Form Audio

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.548499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:74774a41c651d882736487b282873daa381154e9f3c2b234ef43c3c522e5e6b6

Observation 329522a9-4391-477f-bf17-e9651b0ed5a3 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:a08a030660d6f3a7647d621648a918cd63186f063fe5ba78b2753267d531c8ea

Observation cf248996-2560-46a7-8d1f-25da4e8dc723 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.551269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:6196db247a89a86d34bb3ac8b494f0e7f9337fd8ff0201bead7a4cdb740925b7

Observation b0753620-f81b-494e-b3a4-2db185ec4dcb · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:0c7bb6b4c7a2ae8f182e66ac764c3376053fb54cc6dcded316e1f09e52c38611

Observation 9d196c85-a767-40bc-a470-8570b1c3169c · outbound

This paper cites Qwen2-Audio Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Qwen2-Audio Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.588585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:7e43648f994cb9105d10f8c8a6d1ec400842374e140432de09fa3ba27b191b08

Observation f6e9c809-2583-4e2d-858f-43685f97e682 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.930336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:6e4f92c8221dd90d779698f613f079f92c03cef3122c817bf95b5390c61b94a7

Observation c94f2c74-304e-43c0-a0a7-0e9a07fef51c · outbound

This paper cites Recent Advances in Speech Language Models: A Survey.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Recent Advances in Speech Language Models: A Survey

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.598343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:0a9733c28acad5fe2a64ad11669234c3dd196b86e1bc6c95d6116e68de73ee76

Observation 787d3b99-3b5d-4bd7-83b8-9d284ab7cf94 · outbound

This paper cites Kimi-Audio Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Kimi-Audio Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.935066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:d688589431b049d9b4c823774d54e12c5968cc3c99ccb05361d1cc8142d535f9

Observation ccd6e72f-21c0-4627-a178-cd4a6c480c30 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.625234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:be84ad1858fcda0710b9f0de51b307ba30d4bd4265632e73b61872cc25c13359

Observation 7f417745-1396-4451-9c96-88e4e4df356e · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:b4645db2020db160043a77a216f5895abca53eb19a31364a7d38551b8a83b383

Observation 9948a2fc-0017-4122-8a73-fc1160340dc7 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:9f0d682e286f05492631a1d9fef24abe74dc5b83f7bc8ff0e5deccb08b882952

Observation a674a3bd-4aa2-4ed3-bf00-2038385cfcbe · outbound

This paper cites Efficient and generalizable speaker diarization via structured pruning of self-supervised models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Efficient and generalizable speaker diarization via structured pruning of self-supervised models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.581607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:9998f8183b3b98c7bafe345a290d3fadbbc02e523e32ffde0a33c33302f4e534

Observation 6105dafd-069c-4ad2-bbf2-f004becb9545 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.925272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:d52d7d2693e387bebca41866f3790e415414260dfc6665fff037f5b146de16b3

Observation 856faebb-ddd4-4d10-8861-fc1523e065cb · outbound

This paper cites GPT-4o System Card.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding GPT-4o System Card

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.927951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:4f0fd0bbd380ab8313ec321f63edd66fa7485a8a90dd150df65490c03e3ec02a

Observation e44bec0d-72aa-4f24-8a42-d703c3f2990d · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:b520862ec2243827a56795dd1ea319c3d29476056a0daa33e62fa0aa8c7f051d

Observation 1834ea9a-084a-42d3-8884-b32c1d7595a6 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.584186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:7c26cc0ad7b34f895341606d88739eb919752a669fcb6e5b84646eb22284b1e5

Observation 2c8b6ede-7e34-42c0-9af4-3eb310b479da · outbound

This paper cites Chord Label Personalization through Deep Learning of Integrated Harmonic Interval-based Representations.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Chord Label Personalization through Deep Learning of Integrated Harmonic Interval-based Representations

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.922805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:b774da4c95bb0c48d6d5ab6d908df718e1537e6a3eeded80adc9f2845abe4a8e

Observation 91bd9b4c-97f0-4bf3-b81b-9634cd6dba38 · outbound

This paper cites Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.573738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:41cd5046ddd197fccb4991134da1699b255476a36caa64fc84442303638e54f5

Observation 107c5050-9c44-47d3-9984-f432ec791da5 · outbound

This paper cites MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.567135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:ad9d53ca20eb28be9665750a7a0d0fab85cc01ecfac6e133a9c5653329b837d5

Observation d10481d1-be7e-4d14-ac7e-5fecb6c3ad85 · outbound

This paper cites Audio-Maestro: Enhanc- ing large audio-language models with tool-augmented reasoning,.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Audio-Maestro: Enhanc- ing large audio-language models with tool-augmented reasoning,

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.571518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:0ba52b5a312bfddd9fad9b04e7e6d0e553481c926c2ba2247f4f2aa1bc3fd453

Observation f915d858-d9f7-49f3-9296-81babc305094 · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.578490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:2f5fbb0d10848d3bdafbd721260000590758ad071b7016ecd0c8c88a85de90af

Observation b1752b84-5199-4337-a201-24af2e5789c6 · outbound

This paper cites The interspeech 2026 audio reasoning chal- lenge: Evaluating reasoning process quality for audio reasoning models and agents.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding The interspeech 2026 audio reasoning chal- lenge: Evaluating reasoning process quality for audio reasoning models and agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.586530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:41117f2f2b19fb448c2001e4c551aac6fdc616fa301013c2ffbe399f0f5f20dd

Observation 5940a835-3ead-4561-9d4c-b83f3b107d76 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:01fbef95ca7b382fa2b6bd14440c6a5fd51d9c17178dcbdc4eeadeb14f04868b

Observation 021fa0e0-ea9a-4152-ad90-2388694b91a2 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.595640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:a38f17d28f0a10f70e8126655aa50673d65975f58b1c320e22c9346b696e5d90

Observation 6e02d4d5-1fba-4931-8dd9-f269d5b180e6 · outbound

This paper cites Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:13:17.553901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:3fab441c8f17e5e3e4f86dd698b7bdf4f36c01f1716741a2662d964349be5c15

Observation d8c259d6-6d82-4ef9-940e-bd8ec60704dd · outbound

This paper cites A survey on speech large language models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding A survey on speech large language models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.556558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:a4a4fbc1199d782ef4cc50fbb4ef3da7337788e10021bd5382b65e2eda54141e

Observation 5dcd32bd-4df6-424b-9551-05497ca559c9 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:d26b7bff00d37e93f71d6efd571b2af27d54a90c836100af04aa9af37698dceb

Observation 5384afd3-e3d9-4f2b-b910-6fcf3d05e811 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Qwen3.5-Omni Technical Report

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.559143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:9f29376acf766e78426b74af69678fee833c36101ea5ee3cdf23e25679bb8c08

Observation 5224713f-bdac-42ea-abb5-cff0f84df98a · outbound

This paper cites Audiogenie-reasoner: A training- free multi-agent framework for coarse-to-fine audio deep reasoning.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Audiogenie-reasoner: A training- free multi-agent framework for coarse-to-fine audio deep reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.561988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:66d7b8390b84fd6bc96b15b85da8bdc8f31b10fe022998101737abfedaca4ce3

Observation 47ce616d-5194-43da-b624-434a76da46b0 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T10:13:17.564501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:a82350c7dcbe2ee6ed8f10120c15632d7560a43d5a51890f157e144f1af12606

Observation 82219fe5-09c2-46ae-8452-afcaa9b840b4 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 32

Resolution
verified exact
doi, observed 2026-06-29T10:13:17.575295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:5804902d48ad5fca92b30e81feef1e5359a90ca6c23bb26639e09430b39b39c9

Observation ce687dc8-6093-4861-8d1a-28f5153af99d · outbound

This paper cites Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.611327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:7cbdf086cadad764ac4f4e79914bd1e05f6a08d2cb96a78c134eefac4e034255

Observation 94dc971e-2020-4cbc-9a94-bb5a7c3dfd48 · outbound

This paper cites Qwen3-ASR Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Qwen3-ASR Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.569247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:609ca1d36f60b25910c20b420a13ebfbd82e94b5e715669fcb5ad341b2bf4f64

Observation b943b762-204c-45c0-9770-8229f2d793b4 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T10:13:17.605795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:ef0e347f85ed5fa21fba5448228d07875f9d4fe7961ea2c7702ea33f05a434ff

Observation 16113a10-cb66-4f05-8e04-cdfd9c3cd49d · outbound

This paper cites Step-audio-r1 technical report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Step-audio-r1 technical report

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.614085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:52928b1905cf27af8edeb5ec268b17aaffe7b1101756f080a7178a59107627f4

Observation ba763aba-a68c-4075-aaad-79bb67eb5773 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:504247e08e24667c0afdb80b05277be32f1e275bc96dd4ca0cc2743f812f7acc

Observation 6a673bc6-0ab6-430a-a9d1-76761ab82065 · outbound

This paper cites Autagent: A reinforcement learning framework for tool-augmented audio reasoning.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Autagent: A reinforcement learning framework for tool-augmented audio reasoning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.591027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:2dd82cbe5bd991f3e093c8b50e5d420238921ba661741c262ae1d5b49ef5cf9f

Observation b3c730db-8447-417a-b87e-592e8131b0c0 · outbound

This paper cites Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.600602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:4208b42a51d9400e0e76d50a42888a13ea7428a51d12df99f25c884c64aa8ef1

Observation 56905657-2e58-441a-afda-e53445f1bd55 · outbound

This paper cites Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.603410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:924658a2d481dec5341bb9555937dd50dca5ea9de9978cb3a307ccac3e985abf

Observation d8ff21f0-a445-46e5-875e-9c9043bc14fe · outbound

This paper cites MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.619641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:7fa6962d4a4d63a8766917e0ad1bf7a033984d62692d3ef2e24842154ac2d542

Observation c4b9eb55-a699-4a37-855f-038fa2baac62 · outbound

This paper cites Au- diotoolagent: An agentic framework for audio-language mod- els.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Au- diotoolagent: An agentic framework for audio-language mod- els

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.622541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:3fd521128dc75ac62cf82a72c46a56526e8e9a5ab932bbb841214f3b9c25be6d

Observation 74408ede-d0f8-4bbe-b9df-84d9d21dea95 · outbound

This paper cites Step-Audio 2 Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Step-Audio 2 Technical Report

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.937590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:8172db146160a09d4fc01ad8e0be00748658ac415556f8e034fdb64e34b3cff5

Observation 0fb953db-3e5a-4cbb-86a1-f05b8102e72e · outbound

This paper cites Thinking with sound: Audio chain-of-thought enables multimodal reasoning in large audio-language models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Thinking with sound: Audio chain-of-thought enables multimodal reasoning in large audio-language models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.617023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:5ec9cd4cd57aff91df53a02aa87fb34c04395be6dc0d0d64fa5158d3f5a27ad9

Observation 7072e4bb-2f5e-4f97-851b-e1bded280374 · outbound

This paper cites Qwen3-Omni Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Qwen3-Omni Technical Report

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.939878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:4a8590b750e77640cfe486b357035bb668ca2e0680266c735a140d68d7aaf33b

Observation 0edb7d8a-b937-42d0-a225-51d7743eb8b3 · outbound

This paper cites FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.609016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:42d1c4d545fcff2cea1cdf8d5fff8f015797a94a4f8a25d3a57fd27de89b22f6

Observation 89ada80d-cd91-48a5-af2e-9229bf82fc58 · outbound

This paper cites Fireredasr2s: A state-of-the-art industrial-grade all-in-one automatic speech recognition system.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Fireredasr2s: A state-of-the-art industrial-grade all-in-one automatic speech recognition system

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.627925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:3fe88eb82dffd03445081e01d8aebe82304e34370a618c94f6f67c3a27800b5e

Observation d0c51ded-a237-4e37-8afa-e567bec4833c · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.630607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:52ae867bc2ce5db6cafc18e7d02d69bf9a54a184da68cd98e106997207311bbd

Observation e6993045-ba14-481a-b61d-1fe74d5dc1dc · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding ReAct: Synergizing Reasoning and Acting in Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.593173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:72e3adbffbf496b2dd119d1a2d81bc2fd4e5698d2ddc8f5941f456ba8c7da537

Observation 657e8447-a686-4bea-aaa3-e04ba9014f86 · outbound

This paper cites Mimo-audio: Audio language models are few-shot learners.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Mimo-audio: Audio language models are few-shot learners

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.932817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:fdd130d5559ff2cd94043fd6a65ca7fe9e677db8f9a8471244dd0ad9e7e01946

Observation c57bac54-7fe5-4cd1-a4f8-b6857bf31b18 · outbound

This paper cites online" 'onlinestring :=.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding online" 'onlinestring :=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:dc4698dcef3f8ec0706a980edd3c561303aa6bd4fb855e7578932b752abd0d66

Observation 5a0372af-caf5-4143-ab5b-f5fca9d7c233 · outbound

This paper cites write newline.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding write newline

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:fd20020f925d2789c6bef63ed36425ba70d5bd8719dba8bd3f72c9a38f65e9e2

Pith citing papers

Observation b931952c-42e1-4b70-a83b-9a03bd5b5a6c · inbound

Voice Memory for Agentic Speech Recognition cites this paper.

Voice Memory for Agentic Speech Recognition Audio-Mind: An Auditable Agentic Framework for Audio Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.362798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.362798Z digest=sha256:d834be58cf4fa7e3c31ce6ed9c1eef185a46545ea3f58b3707727ac53f84262d

Observation c834345b-e7e9-450f-a683-d09f2176cb2e · inbound

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models cites this paper.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Audio-Mind: An Auditable Agentic Framework for Audio Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.428446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.428446Z digest=sha256:8e816a015b903179daee8e2ac549bf222efade642eaadaa1497630bf4eb321a4