Pith. sign in

Paper Citation Record · LEDGER

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

As of 22 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2510.05478.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.05478 v3

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:23:20.431056Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:23:16.842205Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31a6a2d0-840a-457c-b667-49799b97ceba · outbound

This paper cites advantage collapse.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning advantage collapse

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.702602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.702602Z digest=sha256:f6ecffeed3db2414690771eb0682af31469fa102429713b627250367a8b161f3

Observation 88954ec1-b658-48c6-96e2-cf676a9b9362 · outbound

This paper cites AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.842205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.842205Z digest=sha256:8ea0c7c9832290f716a83b5d6fc77ad288259e968e194ac3030bd94c01057e3b

Observation 4ab987d7-7b65-4965-87a7-698d29b843c9 · outbound

This paper cites 𝑟#…RewardsAdvantage 𝐴! 𝐴.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning 𝑟#…RewardsAdvantage 𝐴! 𝐴

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.003452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.003452Z digest=sha256:dadeb0b19756d5754b3ef01b2841f872a7771ba5d251e45ef6a9a3d1ae3d35bb

Observation bdf2ce7f-821e-4f41-bc9f-89f130a6d0ab · outbound

This paper cites {question}Please choose the answer from the follow- ing options:{choice string}. Output the final answer in<answer> </answer>.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning {question}Please choose the answer from the follow- ing options:{choice string}. Output the final answer in<answer> </answer>

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.135795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.135795Z digest=sha256:bc7372ca6969663026f997dda6878d4a50626b339dcd98e62e9f4b3f26aac3e3

Observation 316b11d5-8a9e-40db-aa51-15869eabe27d · outbound

This paper cites Our framework establishes a closed-loop learning process by generating confidence-weighted pseudo-labels from majority voting to guide policy optimization with GRPO.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Our framework establishes a closed-loop learning process by generating confidence-weighted pseudo-labels from majority voting to guide policy optimization with GRPO

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.263471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.263471Z digest=sha256:1378a05f43b6ca79967a9433bfbb4a558dc557ef96bad1f65525571f6e19eb3b

Observation 589131f7-a8e1-4d51-9804-9259c77714ec · outbound

This paper cites Listen, Think, and Understand.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Listen, Think, and Understand

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.396777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.396777Z digest=sha256:049d91dee63e745efe9caa2be470736a0062c688dee3bae3693722395f056706

Observation 24c589b2-21f9-4b24-8e9c-b95b9da91b0f · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Salmonn: Towards generic hearing abilities for large language models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.559550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.559550Z digest=sha256:4d12ffc552620f11b06f598be299ba91496a45c3e8cf76665a596992e9e1fcf5

Observation b69095e2-c45c-4a89-b52e-57faec7ab2ac · outbound

This paper cites Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.704277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.704277Z digest=sha256:17ae8428ee43e33d050325c89e6473ec5bafc08a06fb1c8bdcbb610355660bf0

Observation 0fec4e2b-0834-4458-92cd-27227a7beb4b · outbound

This paper cites Audio flamingo 2: An audio-language model with long-audio understanding and expert reasoning abilities,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Audio flamingo 2: An audio-language model with long-audio understanding and expert reasoning abilities,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.776171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.776171Z digest=sha256:e0215ae5d01dafeb0f223a91c8e5e6bae98f078826e345a561a566b99e695625

Observation f0a6f0d2-4a32-4b43-8924-e64f340f3f43 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.876357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.876357Z digest=sha256:fb8a57284b3249e676d1405ea4670b880d5aa8e1f89f904f46a0477d9a5b2784

Observation 89840c29-ad96-42f3-868b-da8d84e3413b · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:17.996500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:17.996500Z digest=sha256:bdc9b0a7725d6549d2d636e4536b38de0579e5e6b72bc1a43f54c75013f34dce

Observation 93cfa4e8-ca69-49d3-8e57-caa48c531417 · outbound

This paper cites Qwen2-Audio Technical Report.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Qwen2-Audio Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.145873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.145873Z digest=sha256:0d97c9668908535bda725f4a1d163162a698f39bec4bea93986b122734b3777f

Observation 513d9acc-d37d-431a-9187-d31604016403 · outbound

This paper cites Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.440292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.440292Z digest=sha256:3bbf6e3e5f486b9da277515bf12fff60f6c2bdefb643cb31dc17237402dc0b41

Observation a186adcf-1b91-40f4-8702-c0d913298660 · outbound

This paper cites Omni-r1: Do you really need audio to fine-tune your audio llm?,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Omni-r1: Do you really need audio to fine-tune your audio llm?,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.597964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.597964Z digest=sha256:3319994b5a6620f4799fe9db11a33a1dabb1718b117f9a09b0f143bf952a2e22

Observation bcb41631-a86a-47b7-a51d-857cad3553ce · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.772322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.772322Z digest=sha256:d26bfbc827aa622d7a5129fd1dc1277a58b874548277555e0b908467dc2aef1f

Observation 2d3bb884-ac91-4cdb-8693-dbc4baaa8db6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:18.925637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:18.925637Z digest=sha256:9841ee88fc7ea647532c328d2598cb5d53f724607da68b8e1beb26aae2af95dd

Observation 0110b7e2-d07c-4589-b4cf-4ab7830a3201 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.056503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.056503Z digest=sha256:fe3f606e59719b6cd877a922fe2d4585462a63e3feed8efbca22dd3cd5a7039b

Observation 8424f054-81a8-4cd0-8309-8e0bf1930b9c · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.218335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.218335Z digest=sha256:f7da54d0d1b4871a7737651f3d87742119cdebc718f247a42364c6c30e1bf6f3

Observation f721397d-e7b9-4569-98c5-5247d43af320 · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.304868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.304868Z digest=sha256:d25df9754f235d403ce5010f9b6b835a336c9532ad9db16cd23e37b1b5a65eeb

Observation b2cef981-2ec9-4d57-ad06-77b058848d37 · outbound

This paper cites MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.414877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.414877Z digest=sha256:fa69bdf2917daf8e02094a65299bf8f5d6fbd59ff8937bc626f8d967ca14335d

Observation ef63f2d9-aef0-4e24-8243-9876e2f2ab91 · outbound

This paper cites Desta: Enhancing speech language models through descrip- tive speech-text alignment,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Desta: Enhancing speech language models through descrip- tive speech-text alignment,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.542550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.542550Z digest=sha256:93817ace0962ed8a822ca93181f80dc7fe1e1ccae558b6a545374105f9636037

Observation 704371f3-759e-4cf2-8beb-2747351e0162 · outbound

This paper cites Developing instruction-following speech language model without speech instruction-tuning data,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Developing instruction-following speech language model without speech instruction-tuning data,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.717992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.717992Z digest=sha256:8a1f9d1449d7d947c32fe590cb26a78b29468a043396db4fcae493af12fae50d

Observation 513f4a43-d9e6-47ef-96c5-66ec8faf365b · outbound

This paper cites Desta2.5- audio: Toward general-purpose large audio language model with self-generated cross-modal alignment,.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Desta2.5- audio: Toward general-purpose large audio language model with self-generated cross-modal alignment,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:19.880966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:19.880966Z digest=sha256:a51ad447b4a98b105644e9ee6c1a7def65855f3403480de750fb5836d06fb59f

Observation fbe99f8e-e2ae-4851-85b9-d275d025b811 · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.055373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.055373Z digest=sha256:34db04f48b1f66351720006ca386ed9ac5c3b33fa1cfba2a55ab1fbea44ac37a

Observation c28a2971-a4d6-4140-b04f-4f1f68ed129d · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.230179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.230179Z digest=sha256:2e691a008e61da6e9182088ef34702d50a84614635e206f40de4f410c2a93a71

Observation 62b33e03-d5e7-46f3-a8d3-77a9e3b3abb4 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.346056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.346056Z digest=sha256:f91cc8683c2e4a6b9aba6a58180f934eb6d5bba082af2a33f420a666653f7ab5

Observation 9ada16e7-107a-49cb-adf2-49ca8fd00ad0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Qwen2.5-Omni Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.431056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.431056Z digest=sha256:756902dcc724ddb550fa3df5715e24b87bee55d1596ef41b3962bf5d2e6aa83f

Pith citing papers

Observation 88954ec1-b658-48c6-96e2-cf676a9b9362 · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:16.842205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:16.842205Z digest=sha256:8ea0c7c9832290f716a83b5d6fc77ad288259e968e194ac3030bd94c01057e3b

Observation 6751f093-cfd9-40b0-8119-0ced94544a48 · inbound

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning cites this paper.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.936762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.936762Z digest=sha256:1591ff904fb3231deb876d345eaf8d9394b775bd5323c3d7b529e641e981ae86