Pith. sign in

Paper Citation Record · LEDGER

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2510.05478.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.05478 v3

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:23:20.431056Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:23:16.842205Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31a6a2d0-840a-457c-b667-49799b97ceba · outbound

This paper cites advantage collapse.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning advantage collapse

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.702602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.702602Z digest=sha256:c250b26781a5eca41d9d614b9707ae38885fd4b0e3fe27d57f45df59f61c7833

Observation 88954ec1-b658-48c6-96e2-cf676a9b9362 · outbound

This paper cites AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.842205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.842205Z digest=sha256:5af1bd4431ce9314b3d443ab7f2e5dbcc4746bcfa486cdb4c6da18bdc3b95577

Observation 4ab987d7-7b65-4965-87a7-698d29b843c9 · outbound

This paper cites 𝑟#…RewardsAdvantage 𝐴! 𝐴.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning 𝑟#…RewardsAdvantage 𝐴! 𝐴

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.003452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.003452Z digest=sha256:1cd98621f174768f91f5c62a5c1e2ef953bbbbabb820e0e1f3012710461dcb0a

Observation bdf2ce7f-821e-4f41-bc9f-89f130a6d0ab · outbound

This paper cites {question}Please choose the answer from the follow- ing options:{choice string}. Output the final answer in<answer> </answer>.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning {question}Please choose the answer from the follow- ing options:{choice string}. Output the final answer in<answer> </answer>

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.135795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.135795Z digest=sha256:1a4f5b4a74dc4f49aa09e74bb28c691333829b3064da808c3e3fa8eb151ff6d4

Observation 316b11d5-8a9e-40db-aa51-15869eabe27d · outbound

This paper cites Our framework establishes a closed-loop learning process by generating confidence-weighted pseudo-labels from majority voting to guide policy optimization with GRPO.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Our framework establishes a closed-loop learning process by generating confidence-weighted pseudo-labels from majority voting to guide policy optimization with GRPO

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.263471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.263471Z digest=sha256:63b711a08c73c3be4ba7610f32a27d9b45be8501a0da13397b24ef8346062e9a

Observation 589131f7-a8e1-4d51-9804-9259c77714ec · outbound

This paper cites Listen, Think, and Understand.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Listen, Think, and Understand

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.396777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.396777Z digest=sha256:a8c8c8d5f4e5d55e0bd1746d9db2994068b390785fdf3b1dd904a0d8085169e2

Observation 24c589b2-21f9-4b24-8e9c-b95b9da91b0f · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Salmonn: Towards generic hearing abilities for large language models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.559550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.559550Z digest=sha256:8d209a97de1ee690497c956c633abae59a9a50cdcb4f404a28d9839045749f1a

Observation b69095e2-c45c-4a89-b52e-57faec7ab2ac · outbound

This paper cites Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.704277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.704277Z digest=sha256:b7d858c343ea97b2dac7addaf04e5fc4af4cc7b95f3c8c1d1f2af729cf93c2eb

Observation 0fec4e2b-0834-4458-92cd-27227a7beb4b · outbound

This paper cites Audio flamingo 2: An audio-language model with long-audio understanding and expert reasoning abilities,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Audio flamingo 2: An audio-language model with long-audio understanding and expert reasoning abilities,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.776171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.776171Z digest=sha256:c63e77bec77066e3583ede94771a431250e357f2e7777e093481cfa5f5a27776

Observation f0a6f0d2-4a32-4b43-8924-e64f340f3f43 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.876357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.876357Z digest=sha256:8257eb9bb611b283b52e493d2760878535c4eeb37f4a0d41ddea1be77c9155bf

Observation 89840c29-ad96-42f3-868b-da8d84e3413b · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.996500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.996500Z digest=sha256:8ccfd6d1d384e3d9ca6f412b972c4c7dde01070bd117b8ea184143f450d71b60

Observation 93cfa4e8-ca69-49d3-8e57-caa48c531417 · outbound

This paper cites Qwen2-Audio Technical Report.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Qwen2-Audio Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.145873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.145873Z digest=sha256:711f2c0032fd1890e022e5b0fa462697b0e982f1a8bc3aab32969f7b46e28c9f

Observation 513d9acc-d37d-431a-9187-d31604016403 · outbound

This paper cites Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.440292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.440292Z digest=sha256:b7c240a5c6566119f0012c85030494dbad20fa93e5324b3b905f9325c0cdc244

Observation a186adcf-1b91-40f4-8702-c0d913298660 · outbound

This paper cites Omni-r1: Do you really need audio to fine-tune your audio llm?,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Omni-r1: Do you really need audio to fine-tune your audio llm?,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.597964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.597964Z digest=sha256:1e1f1cf3e3d508cbcd387caae17da0fe56088773634d6dde6148fa675102fbb9

Observation bcb41631-a86a-47b7-a51d-857cad3553ce · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.772322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.772322Z digest=sha256:bd1fd74e9285bba8ec5985b61029c15d3315de1c6b5d26d4f8da45cbe1d19d38

Observation 2d3bb884-ac91-4cdb-8693-dbc4baaa8db6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.925637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.925637Z digest=sha256:827d3e8c110d905ca46c1724a6fa8432299e34b86789d1f8dc8691df56628c9e

Observation 0110b7e2-d07c-4589-b4cf-4ab7830a3201 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.056503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.056503Z digest=sha256:44b29de9b06fd220ca629ae2bbda09c5545553604e138879405b6220d80a16f2

Observation 8424f054-81a8-4cd0-8309-8e0bf1930b9c · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.218335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.218335Z digest=sha256:e4305e0c6216365300679aecc91fb37b9bd75be7545cefd6ccdc020322b679a2

Observation f721397d-e7b9-4569-98c5-5247d43af320 · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.304868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.304868Z digest=sha256:6ebb9b34fb895b4c10735f27549019691d43533c0bf9a092fd22e41168ea3466

Observation b2cef981-2ec9-4d57-ad06-77b058848d37 · outbound

This paper cites MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.414877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.414877Z digest=sha256:d9fdb31181c34d972c662431c7d817dab2916879569efc85e3f96b445b821f21

Observation ef63f2d9-aef0-4e24-8243-9876e2f2ab91 · outbound

This paper cites Desta: Enhancing speech language models through descrip- tive speech-text alignment,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Desta: Enhancing speech language models through descrip- tive speech-text alignment,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.542550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.542550Z digest=sha256:7f4575910704715e3c0cb99670615d5979b8b83af3258f3a758403d9344c6c02

Observation 704371f3-759e-4cf2-8beb-2747351e0162 · outbound

This paper cites Developing instruction-following speech language model without speech instruction-tuning data,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Developing instruction-following speech language model without speech instruction-tuning data,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.717992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.717992Z digest=sha256:07e8c8ddba2fa1cb4f457aefed2ff4b2faca9a83cfdad9a0307992d7fa152a51

Observation 513f4a43-d9e6-47ef-96c5-66ec8faf365b · outbound

This paper cites Desta2.5- audio: Toward general-purpose large audio language model with self-generated cross-modal alignment,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Desta2.5- audio: Toward general-purpose large audio language model with self-generated cross-modal alignment,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.880966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.880966Z digest=sha256:b0dbb2864f1bfd1a25c373d1f4c9b872150d8c30a799da2a5e40ea9ba7fc24ca

Observation fbe99f8e-e2ae-4851-85b9-d275d025b811 · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.055373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.055373Z digest=sha256:73210b0abd21da28a13b3aef441e5bc6d6b5a72bcceb1b7726a2f1bba93cc1d6

Observation c28a2971-a4d6-4140-b04f-4f1f68ed129d · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.230179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.230179Z digest=sha256:5666e701870e5a0a6ec26d7824b76a6bbf31883d3d0270a6cb6c8809a166c721

Observation 62b33e03-d5e7-46f3-a8d3-77a9e3b3abb4 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.346056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.346056Z digest=sha256:731b260ec03f7bab5047ba068fe6379dae0f3683c8f13a3cbc0230cb24d52581

Observation 9ada16e7-107a-49cb-adf2-49ca8fd00ad0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Qwen2.5-Omni Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.431056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.431056Z digest=sha256:1c1454c6d79e19ed9fca1ebbbe2d6ef395f8a6df4d14201f9c9506719aba26a0

Pith citing papers

Observation 88954ec1-b658-48c6-96e2-cf676a9b9362 · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.842205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.842205Z digest=sha256:5af1bd4431ce9314b3d443ab7f2e5dbcc4746bcfa486cdb4c6da18bdc3b95577

Observation 6751f093-cfd9-40b0-8119-0ced94544a48 · inbound

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning cites this paper.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.936762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.936762Z digest=sha256:428f79f4b64c8bda9de331789f80eeb4f3e59da99042db33cc5a1c0ffda4a6c4