Pith. sign in

Paper Citation Record · LEDGER

Audio-Mind: An Auditable Agentic Framework for Audio Understanding

As of 14 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2605.28480.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28480 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T10:03:54.653164Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:02:41.428446Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact39
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1506217-13ea-4ee1-b44f-b84d1c492850 · outbound

This paper cites Qwen3-VL Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.942116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:8b0de54e77dfcbeb11c5500b93b09ee5f85853aeb968283900e3eba8a37ac482

Observation 7f46dda7-5abc-4b0b-9806-0ebd148e07b2 · outbound

This paper cites WhisperX: Time-Accurate Speech Transcription of Long-Form Audio.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding WhisperX: Time-Accurate Speech Transcription of Long-Form Audio

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.548499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:5167938530dcfa9f13627604bc1d63255cae8405c3b360aaf138e21517211d7d

Observation 329522a9-4391-477f-bf17-e9651b0ed5a3 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:3c1a62254d3937c3185cd2c885bda8f87627b9c925d18e48edbf514a519f82b0

Observation cf248996-2560-46a7-8d1f-25da4e8dc723 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.551269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:d805a5f0d68827ebf216d424730d6bf186212a796f628de52df7da11afa3110d

Observation b0753620-f81b-494e-b3a4-2db185ec4dcb · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:27c38352e26dca6fbeb0fe2e228d07e17c61ad4a82592aee095768bce714bf56

Observation 9d196c85-a767-40bc-a470-8570b1c3169c · outbound

This paper cites Qwen2-Audio Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Qwen2-Audio Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.588585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:dcff1214629a7ed6a42b410d919ee8254801b15b9beddb1a2b988378d360e85c

Observation f6e9c809-2583-4e2d-858f-43685f97e682 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.930336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:4ad221eaa661cfecdebe9b8101f3cf01fdb30ef7bbf29c752cd6467637f4b3a0

Observation c94f2c74-304e-43c0-a0a7-0e9a07fef51c · outbound

This paper cites Recent Advances in Speech Language Models: A Survey.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Recent Advances in Speech Language Models: A Survey

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.598343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:6cdb905c663366efe1b2dfee44caa0a9d3a6f22fdea1605cce9d918d2333aa83

Observation 787d3b99-3b5d-4bd7-83b8-9d284ab7cf94 · outbound

This paper cites Kimi-Audio Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Kimi-Audio Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.935066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:1346c4c063d67bc1038af6489085ec1b8a83098506752539ec768367291c9c89

Observation ccd6e72f-21c0-4627-a178-cd4a6c480c30 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.625234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:577a776174a200ff7a8664a39baff5065e817676011a59438d188b845a8b0fd9

Observation 7f417745-1396-4451-9c96-88e4e4df356e · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:6ea6459b80fba949fcf96989b7129e262af25d80c0114e53a50eeafe6d9d7a97

Observation 9948a2fc-0017-4122-8a73-fc1160340dc7 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:202179772ed912ef5f7617917ae8287e5cdb53fbb8d00544b2da2c05d6063fe4

Observation a674a3bd-4aa2-4ed3-bf00-2038385cfcbe · outbound

This paper cites Efficient and generalizable speaker diarization via structured pruning of self-supervised models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Efficient and generalizable speaker diarization via structured pruning of self-supervised models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.581607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:1c3177ba07ccb58a884ee6c294be15093e2060db9f2a11e32c4c6cf383b25f70

Observation 6105dafd-069c-4ad2-bbf2-f004becb9545 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.925272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:aad42c3d504b4564359ac03e9f70a97389f1a24aaefce647671cd15c57236b82

Observation 856faebb-ddd4-4d10-8861-fc1523e065cb · outbound

This paper cites GPT-4o System Card.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding GPT-4o System Card

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.927951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:a07f0bdf7b40e7aceeedb29e44ac05e55792361614aa293d03f51d36e62c19df

Observation e44bec0d-72aa-4f24-8a42-d703c3f2990d · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:8dcccf5bf51459a70329430e56bf3cb118b8b423727b3c989587822c36059280

Observation 1834ea9a-084a-42d3-8884-b32c1d7595a6 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.584186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:0c98520d7fe9e41b0be155e122a53c42efe794ef702692b94cfad86e27baca35

Observation 2c8b6ede-7e34-42c0-9af4-3eb310b479da · outbound

This paper cites Chord Label Personalization through Deep Learning of Integrated Harmonic Interval-based Representations.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Chord Label Personalization through Deep Learning of Integrated Harmonic Interval-based Representations

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.922805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:c7492a8280d87d097636dcc997038292bd68c2b0677f17c10e729e49f196b75f

Observation 91bd9b4c-97f0-4bf3-b81b-9634cd6dba38 · outbound

This paper cites Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.573738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:cf39cc4230c941bfdbc7bb24e857c20ca5649f1387434efba7a3d512a33a6740

Observation 107c5050-9c44-47d3-9984-f432ec791da5 · outbound

This paper cites MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.567135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:15732450b1a28577fac75150a3d534b51d4d824c3471bd6fbd491229810f01c7

Observation d10481d1-be7e-4d14-ac7e-5fecb6c3ad85 · outbound

This paper cites Audio-Maestro: Enhanc- ing large audio-language models with tool-augmented reasoning,.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Audio-Maestro: Enhanc- ing large audio-language models with tool-augmented reasoning,

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.571518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:17e713c05ba7383b74df71c227546a814deac7ee9cc1822d81e14dba950a1481

Observation f915d858-d9f7-49f3-9296-81babc305094 · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.578490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:15d0768de70e455b8238be31ad1ba2ba17d80a3b8264d3fb21c014ae57e24491

Observation b1752b84-5199-4337-a201-24af2e5789c6 · outbound

This paper cites The interspeech 2026 audio reasoning chal- lenge: Evaluating reasoning process quality for audio reasoning models and agents.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding The interspeech 2026 audio reasoning chal- lenge: Evaluating reasoning process quality for audio reasoning models and agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.586530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:2ec5bc949a589033b979880bc2057f9466b741c0e60a944cb880d37f5c97ca71

Observation 5940a835-3ead-4561-9d4c-b83f3b107d76 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:2d929b63795efaeedc1e08e081dcc8b3a911e4b0d5fcb10cb1a7c5b45473f45c

Observation 021fa0e0-ea9a-4152-ad90-2388694b91a2 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.595640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:7d414a3057e360ea6b4a45128d143f47c3907f89e9712498319160dd06d93ba1

Observation 6e02d4d5-1fba-4931-8dd9-f269d5b180e6 · outbound

This paper cites Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:13:17.553901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:b05991f27214e71b2a6f9c1a50c7af0323b560ab692dc6e7ae713ce697cac89f

Observation d8c259d6-6d82-4ef9-940e-bd8ec60704dd · outbound

This paper cites A survey on speech large language models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding A survey on speech large language models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.556558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:8ab265b8d28fa7923a78994c5b879653ddbe4b74a984d1babe8ff4af8b73eb7a

Observation 5dcd32bd-4df6-424b-9551-05497ca559c9 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:9a1373c61fb4c2f0a372850d62320eed468f0ee8f7c09cf594ee8df13f10f7ea

Observation 5384afd3-e3d9-4f2b-b910-6fcf3d05e811 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Qwen3.5-Omni Technical Report

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.559143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:1827cf2580d7e27e9fd91a0e284dd088c3e42622b49fd7e301dfeecd6461771b

Observation 5224713f-bdac-42ea-abb5-cff0f84df98a · outbound

This paper cites Audiogenie-reasoner: A training- free multi-agent framework for coarse-to-fine audio deep reasoning.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Audiogenie-reasoner: A training- free multi-agent framework for coarse-to-fine audio deep reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.561988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:be5d06e6cfcfc141e8c25ed27b29481545f2d4e8ed95bfb43617d6f3ac456ee2

Observation 47ce616d-5194-43da-b624-434a76da46b0 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T10:13:17.564501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:b13483caa356def5be1d96ebffaf8bb2d99291cf3fc7a8b05e0337723f63a775

Observation 82219fe5-09c2-46ae-8452-afcaa9b840b4 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 32

Resolution
verified exact
doi, observed 2026-06-29T10:13:17.575295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:9b1049a2caf54c38c730d217e7bd7687e9ca5e389e9e84f7075721944b9c48e1

Observation ce687dc8-6093-4861-8d1a-28f5153af99d · outbound

This paper cites Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.611327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:52f4ced00166e5fe2cc86e4a53ff57a5781aac3e24edd8ba643429070ceb29e7

Observation 94dc971e-2020-4cbc-9a94-bb5a7c3dfd48 · outbound

This paper cites Qwen3-ASR Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Qwen3-ASR Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.569247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:3812635b9b025a6797bc414df178a1e7582764b1035c405ece14f0b21d2ef3ad

Observation b943b762-204c-45c0-9770-8229f2d793b4 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T10:13:17.605795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:966165252415cb5eb5e98879774eb5ed6e4407c50a1095b4deeb0146bcf78e7f

Observation 16113a10-cb66-4f05-8e04-cdfd9c3cd49d · outbound

This paper cites Step-audio-r1 technical report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Step-audio-r1 technical report

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.614085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:cc7d636550bc89cdc7165e38d9f7bc76fcbab0cc59b0c02f64801f6e9d9a9d48

Observation ba763aba-a68c-4075-aaad-79bb67eb5773 · outbound

This paper cites an unresolved cited work.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:51981670ed1a2c85f1ce10dbc6c0424419dddf640be40b7918a478d43319ccbc

Observation 6a673bc6-0ab6-430a-a9d1-76761ab82065 · outbound

This paper cites Autagent: A reinforcement learning framework for tool-augmented audio reasoning.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Autagent: A reinforcement learning framework for tool-augmented audio reasoning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.591027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:b48688fe870157707114468318c195eb03ddbeb655ed6c980a2fa9e0abc39655

Observation b3c730db-8447-417a-b87e-592e8131b0c0 · outbound

This paper cites Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.600602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:81ca9c1ba2e9bd6e52a6e912731f60cf6438601532def5879e34b48b056055c3

Observation 56905657-2e58-441a-afda-e53445f1bd55 · outbound

This paper cites Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.603410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:72b68ab97d9fba1f7f9c4377051644a1b0bf7295d3982a0d11e418a3d2003818

Observation d8ff21f0-a445-46e5-875e-9c9043bc14fe · outbound

This paper cites MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.619641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:98e858907b974d44db2fb51049964911d3097ef9645684304a1d53a870b46859

Observation c4b9eb55-a699-4a37-855f-038fa2baac62 · outbound

This paper cites Au- diotoolagent: An agentic framework for audio-language mod- els.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Au- diotoolagent: An agentic framework for audio-language mod- els

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.622541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:54d4dd7f0a82f51f1166d459e94601ca5426660bf1db3e24813a17c80c8f5e23

Observation 74408ede-d0f8-4bbe-b9df-84d9d21dea95 · outbound

This paper cites Step-Audio 2 Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Step-Audio 2 Technical Report

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.937590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:cc3c20b8f1bb8a762165941973a21556d0c31f33048b908be066e4e81796ab02

Observation 0fb953db-3e5a-4cbb-86a1-f05b8102e72e · outbound

This paper cites Thinking with sound: Audio chain-of-thought enables multimodal reasoning in large audio-language models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Thinking with sound: Audio chain-of-thought enables multimodal reasoning in large audio-language models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.617023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:f4972a5e2d64a86de91b5b7173dab774e1bf332675167b1b6c442c453d49bfe8

Observation 7072e4bb-2f5e-4f97-851b-e1bded280374 · outbound

This paper cites Qwen3-Omni Technical Report.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Qwen3-Omni Technical Report

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.939878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:68122c188ab920b7feac22ef29fe0000577299665a8c6122cab79f3c04994014

Observation 0edb7d8a-b937-42d0-a225-51d7743eb8b3 · outbound

This paper cites FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.609016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:a3ab3962936a760c2d6566349e39123338f18321e44809c61bb20964c7e2824f

Observation 89ada80d-cd91-48a5-af2e-9229bf82fc58 · outbound

This paper cites Fireredasr2s: A state-of-the-art industrial-grade all-in-one automatic speech recognition system.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Fireredasr2s: A state-of-the-art industrial-grade all-in-one automatic speech recognition system

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.627925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:bfdf0f99b0b086e9a09885f6395204aa2b9cfb3e59033323361c612d845e2aa1

Observation d0c51ded-a237-4e37-8afa-e567bec4833c · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.630607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:e8835f5879572c6db963f9d7907156a89246b098387e5a64a0334f720e950ef2

Observation e6993045-ba14-481a-b61d-1fe74d5dc1dc · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding ReAct: Synergizing Reasoning and Acting in Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.593173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:b0e7891cbb4961aa2f59fe0d92afd72073c7d1a25976f56d7d6274ddccde25e7

Observation 657e8447-a686-4bea-aaa3-e04ba9014f86 · outbound

This paper cites Mimo-audio: Audio language models are few-shot learners.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Mimo-audio: Audio language models are few-shot learners

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.932817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:7e06264b10de0dceb7f67dc0374204d05f6b544eb18d2093e96e5e7a343821e2

Observation c57bac54-7fe5-4cd1-a4f8-b6857bf31b18 · outbound

This paper cites online" 'onlinestring :=.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding online" 'onlinestring :=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:daf741a703e70966805c9c6777d40ac7986c78109675bd7e944cfe7cabd3ef13

Observation 5a0372af-caf5-4143-ab5b-f5fca9d7c233 · outbound

This paper cites write newline.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding write newline

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T10:03:54.653164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:6f45e178dd7983051be32c22b39640552d2dac4cf2b84150c9c8be7e69ece3ae

Pith citing papers

Observation b931952c-42e1-4b70-a83b-9a03bd5b5a6c · inbound

Voice Memory for Agentic Speech Recognition cites this paper.

Voice Memory for Agentic Speech Recognition Audio-Mind: An Auditable Agentic Framework for Audio Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.362798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.362798Z digest=sha256:87b3451e8c8092fd39fd0b6de2a7d9adefb9e5e751ce05df1263d6b297a4d36d

Observation c834345b-e7e9-450f-a683-d09f2176cb2e · inbound

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models cites this paper.

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models Audio-Mind: An Auditable Agentic Framework for Audio Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:02:41.428446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:02:41.428446Z digest=sha256:23daa6a222e40c34b20230c2c6c27b13752d5eeee0be789b963442d9eca32926