Pith. sign in

Paper Citation Record · LEDGER

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.20166.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20166 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:40:25.940622Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 592cc0c8-1bcf-43e5-bbb7-9ba80a2972ea · outbound

This paper cites Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.870899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.870899Z digest=sha256:711abd7343703451437ddbe6312593e3cfb04a36b4b62cf67bb9cbfc274a8c01

Observation 370586df-40cc-4efa-b777-f9a491e488a4 · outbound

This paper cites MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.881784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.881784Z digest=sha256:cb2f2ece9a7ce4598174965abe39f6ad81d30a60d9f9de815678793fe4843359

Observation 6c96c274-9eb9-4085-849a-a5c4f52338d5 · outbound

This paper cites Visplay: Self-evolving vision-language models from images.arXiv preprint arXiv:2511.15661,.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Visplay: Self-evolving vision-language models from images.arXiv preprint arXiv:2511.15661,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.887648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.887648Z digest=sha256:299d14830de1615dde72a576ef1fcccf639c2e6684c97d9a215ab6c53d95786e

Observation fc197c8c-d538-4e0f-b32b-15e76bae5ea2 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.890871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.890871Z digest=sha256:efe7bb5df103824918bf36815f79d6685ac5855d3c0eac8b04d5efe4c68f2f0e

Observation db2dece3-a474-4384-933b-90fc9997a4e1 · outbound

This paper cites MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.893951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.893951Z digest=sha256:39073f04422bfac56a3efe2cb0336eb5458c6ea8e0fb2750f13da0db24508136

Observation 5924d362-6311-4e55-906c-f7e5c3a8e008 · outbound

This paper cites Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.897788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.897788Z digest=sha256:42be3d6b032f832d82d7f96949679e514271fd60a6f0cf6d6fff1c565f0020ee

Observation 181d56f9-4e1d-42aa-bf9b-f1cf57ac083f · outbound

This paper cites Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119, 2025a.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119, 2025a

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.901360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.901360Z digest=sha256:84e85b6a8e405ae0b6f63cbca82d5e21a4e7e7229e047a86c1cfbbf8a8fc7762

Observation 8135aef0-bdd2-4814-9e97-2451bffa676e · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.903803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.903803Z digest=sha256:aad1b012f977a75ecfd61ce6ebeb1bbe04087c169db5d5cfedc8413526aabbfe

Observation 2a201464-23cb-4b0c-a63f-bc6bc9e4ed20 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.906733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.906733Z digest=sha256:1adcdea920dcf0c2d59396abc8cfbb23d50197c5301cfa29abcd65803a534c7a

Observation d7c1d744-ed98-43e1-b926-9227f81be223 · outbound

This paper cites Can speech llms think while listening?arXiv preprint arXiv:2510.07497,.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Can speech llms think while listening?arXiv preprint arXiv:2510.07497,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.909339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.909339Z digest=sha256:7495c60526b0ba1958a2bab5016eeb8fc0ec26bb19e6fe4122a99033988484cc

Observation 7c5a4885-3cc0-4171-99a7-266986784d2a · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Salmonn: Towards generic hearing abilities for large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.912584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.912584Z digest=sha256:9d8f054dfee33e20ac0d1c924301f2f08487160db947fec088ae5a243d7aab08

Observation e88fcc04-f999-4af0-baa1-b380b00b4370 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Qwen3.5-Omni Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.915711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.915711Z digest=sha256:5bc35de254335bf96049ee1c7b01e8dfe03728d01c768601cb78a3b9fd6b2141

Observation 6f0549ba-ab6f-4375-8d8a-1f07ed819873 · outbound

This paper cites Emotion- thinker: Prosody-aware reinforcement learning for explainable speech emotion reasoning.arXiv preprint arXiv:2601.15668,.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Emotion- thinker: Prosody-aware reinforcement learning for explainable speech emotion reasoning.arXiv preprint arXiv:2601.15668,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.920896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.920896Z digest=sha256:513881ee0cdac3f10206a4e0054a3d3477e30d41f9b8165329e25a1b6e7847ad

Observation d83a353f-3aaf-4cc8-b0c2-97dc00bd9bec · outbound

This paper cites Vision-zero: Scalable vlm self-improvement via strategic gamified self-play.arXiv preprint arXiv:2509.25541, 2025a.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Vision-zero: Scalable vlm self-improvement via strategic gamified self-play.arXiv preprint arXiv:2509.25541, 2025a

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.923609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.923609Z digest=sha256:d84318828d5505e8c10b8f1d9b6e5ead0ace271a58cd378efc67c0bf8822b26c

Observation a84d212c-93e7-43a9-bf2a-459438127293 · outbound

This paper cites Qwen3-Omni Technical Report.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Qwen3-Omni Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.926154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.926154Z digest=sha256:c070d224691d4392497c81dc8e47a7a8bbbfae0b62b2544ace9e799cc38c71ad

Observation 57dbc71f-bea7-413d-ad3c-41516d32857d · outbound

This paper cites Spell: Self-play reinforcement learning for evolving long-context language models.arXiv preprint arXiv:2509.23863, 2025b.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Spell: Self-play reinforcement learning for evolving long-context language models.arXiv preprint arXiv:2509.23863, 2025b

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.929340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.929340Z digest=sha256:123a704437fa77bb56f92588d660a5a79721708c4aedb7b23dffda34ed9363cc

Observation 6751f093-cfd9-40b0-8119-0ced94544a48 · outbound

This paper cites AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.936762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.936762Z digest=sha256:1591ff904fb3231deb876d345eaf8d9394b775bd5323c3d7b529e641e981ae86

Observation 8c4bc8d0-ba15-4f49-8267-24d31f53ef37 · outbound

This paper cites We treat the chosen and rejected audios as an audio contrast pair (aref , avar).

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning We treat the chosen and rejected audios as an audio contrast pair (aref , avar)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.940622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.940622Z digest=sha256:10a68b4057bbf427e5dd431f697ac8224d6b380dc4cf3dc71711279f119a5536

Observation ebb9399c-db7c-4ad2-a17d-9402744aee6b · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.933165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.933165Z digest=sha256:83ce675e5be336d92988ac6f8918944874d53cbd1a71c649426cb409bda6353f

Observation 9a3897d3-6906-4378-925d-6e2ecde66d6b · outbound

This paper cites Incentivizing consistent, effective and scalable reasoning capability in audio llms via reasoning process rewards.arXiv preprint arXiv:2510.20867,.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Incentivizing consistent, effective and scalable reasoning capability in audio llms via reasoning process rewards.arXiv preprint arXiv:2510.20867,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.884823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.884823Z digest=sha256:9af4fa2939892fc2891567ad0fa586a954de435c85d89addd4c71f2dd588f682

Observation 3125e7d7-0f57-4742-88d6-3ef66a0012b2 · outbound

This paper cites Qwen2-Audio Technical Report.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Qwen2-Audio Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.877689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.877689Z digest=sha256:b166748ae6f70dd182b3df1ada2bb4e3be110025dd3b183bc4c03b68d81c4ee2

Observation 8200c25c-853e-43bf-bc2e-27fd8bdcd96a · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.874679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.874679Z digest=sha256:5da1e3f07fe31e35ab1ef015c410d639cc92f5310be77f0d5c9abbe0df86f786

Observation d9c5e5d8-0658-4001-9652-35f4f55975a1 · outbound

This paper cites Game-rl: Synthesizing multimodal verifiable game data to boost vlms’ general reasoning.arXiv preprint arXiv:2505.13886,.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning Game-rl: Synthesizing multimodal verifiable game data to boost vlms’ general reasoning.arXiv preprint arXiv:2505.13886,

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.918429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.918429Z digest=sha256:497f3867722259c4a094f747c481e8b129bae322a6a8951a5207e0d35889cbe5

Pith citing papers

No inbound Pith citation observations are available.