Pith. sign in

Paper Citation Record · LEDGER

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs

As of 4 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2509.08031.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08031 v3

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T18:07:01.257126Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact10
  • verified fuzzy7
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 438bb6eb-afa4-4118-9e01-f38ceee8c266 · outbound

This paper cites AHELM: A Holistic Evaluation of Audio-Language Models.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs AHELM: A Holistic Evaluation of Audio-Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:11:43.185900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:428fee5f92e86ee00490030aee7654febd86163099b2f81adacf7977cdb53fec

Observation 624181be-7626-47fa-9fab-f13bf5e31ff2 · outbound

This paper cites Evaluation of real-time tran- scriptions using end-to-end asr models.arXiv preprint arXiv:2409.05674.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Evaluation of real-time tran- scriptions using end-to-end asr models.arXiv preprint arXiv:2409.05674

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:11:43.180239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:bab0874c0deff277d7c64242445ad5f08599c150062a6d72fa932e64c87933d8

Observation b204fc5c-f951-48d3-9111-17386e35fffc · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:11:43.130147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:bd8ee3d5689cef4c5d7da289f2869b7f52fee25eacf239dbc1bbb79030a0d1bf

Observation 40b264ad-83e4-4711-9cf1-c13c21202f61 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:11:43.175119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:44fb21a934dffebd65a799754d913885d5ac1ecaf8e11d7e5d7810212a84f4c4

Observation 098c0c90-6348-4246-8634-a424d1abd924 · outbound

This paper cites Recent Advances in Speech Language Models: A Survey.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Recent Advances in Speech Language Models: A Survey

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:11:43.169892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:848bd8eef610190db1665f369319e1327243727f5c6cde08c0fb52f2d535f534

Observation 289cde76-5e4e-4b72-909c-70570a587388 · outbound

This paper cites Kimi-Audio Technical Report.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Kimi-Audio Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:11:43.164703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:e1adeffcb16f67b269711270ad5e4b2aca7ab082100a2540cc5cbb8b5d9b88ce

Observation 3b5a1ef1-58fd-4ad9-bfb2-c0eae2d5eef4 · outbound

This paper cites How numerical precision affects arithmetical reasoning capabilities of llms.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs How numerical precision affects arithmetical reasoning capabilities of llms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:12:47.433644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:a22bb91f69c88ae7009cfbd34100d73db94873e6a9e63faab93d4251120225ef

Observation f86e5007-4d89-4a6e-85d8-77cfdacf9837 · outbound

This paper cites Dynamic-superb phase-2: A collab- oratively expanding benchmark for measuring the capabilities of spoken language models with 180 tasks.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Dynamic-superb phase-2: A collab- oratively expanding benchmark for measuring the capabilities of spoken language models with 180 tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:12:47.430983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:6e34064f8bb5ac2aef683c7001214a40b61b79293af83c2c20022d44539c088d

Observation 662da643-4ff2-4908-b0df-431df5ff5a86 · outbound

This paper cites Voxtral.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Voxtral

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:11:43.160420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:0af9217ecb0e09fda9253d760b54c01ee275db93bd076c658ba91c89c1cb9f3f

Observation c0dc75f9-f8f4-41be-9e44-9da96f94fff7 · outbound

This paper cites MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:11:43.155180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:c96699639525171246ee86f979afc6006ef54a89291894732830e3fca6630fa2

Observation 6a2628a2-4e66-45a8-9c12-468c126252d1 · outbound

This paper cites A survey on speech large language models.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs A survey on speech large language models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:11:43.136599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:35736d0489845d4aa801318c93d8ff31745602de7aed56fcc58364acc5ff5fc3

Observation 99db9f7f-4cbf-4370-bab9-cfde87796100 · outbound

This paper cites Joint speech recognition and speaker diariza- tion via sequence transduction.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Joint speech recognition and speaker diariza- tion via sequence transduction

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:12:47.427740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:9b88fed8560c32031d02ab06bd15b8535d4d0147e6e00054fc8d3acf715f4792

Observation 1a822000-00e9-4d2d-8b92-bf9b370133f0 · outbound

This paper cites Versa: A versatile evaluation toolkit for speech, audio, and music.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Versa: A versatile evaluation toolkit for speech, audio, and music

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:12:47.424523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:794bb3aa211c947475a0de5b2b87c85c059fc61e113b865cda4b17102f100426

Observation 25bef7fc-d493-4d89-ab0d-846b97359280 · outbound

This paper cites In: Zong, C., Xia, F., Li, W., Navigli, R.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs In: Zong, C., Xia, F., Li, W., Navigli, R

Reference 14

Resolution
malformed identifier
doi_truncated, observed 2026-05-18T18:11:42.482554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:f6b5af9fc72bf4a6531fb917fa9659d0bb20d05bbae2dc730ea2700a438c67a7

Observation 0c231d1e-31af-4e30-981f-bf0469d9b7ff · outbound

This paper cites Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:12:47.421850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:2340e5e29874431177df6f7df0c276919fdfdcd888b8c4dd373e7135f23f661a

Observation d9c93039-74e3-47ef-a318-586598aad692 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Qwen2.5-Omni Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:11:43.150673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:1dece5b022562e9c2bc77f4558edede3fc6a5d9ed0aca0a3f1f8d94dd08cbd35

Observation 703cab93-17bd-4556-9661-de930d227615 · outbound

This paper cites Air-bench: Benchmarking large audio-language models via generative comprehension.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Air-bench: Benchmarking large audio-language models via generative comprehension

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:12:47.418077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:90edb5901c01d482726bb30d0cf41ffe990fa7a10d44c23c3701dd08d4eced36

Observation 20f76a52-e733-4cab-ae8b-84181a8953b9 · outbound

This paper cites Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T18:12:47.415066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:b03413939ac10f40e57acd693ba3f5abe769efe707f427c9078d584e0b58b13a

Observation 33896ca0-5b44-4b21-9cc2-7610012defdb · outbound

This paper cites NO INSIGHT.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs NO INSIGHT

Reference 19

Resolution
metadata mismatch
doi, observed 2026-05-18T18:11:42.486423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:3e77854f2a4b2bc806466348e2565fa65a2c16bc57a7be13deb4bab97d8bf361

Observation 054cc496-1d59-4e3e-8e16-1c7faa9956a9 · outbound

This paper cites X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:11:43.145790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:031266f8c23f0014eae01aa07078d0181c865dbdf57b1aa0815a295a38907f03

Observation 3798f55a-a09f-48af-812b-9e6ac8ad2946 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Instruction-Following Evaluation for Large Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:11:43.140551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:04e2fe5b3fdbd1a51dc21b0dd52ed0d639e8479078c9ccca6478f36226406fc2

Pith citing papers

No inbound Pith citation observations are available.