Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:23:20.431056Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2510.05478.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:23:20.431056Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:23:16.842205Z
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 31a6a2d0-840a-457c-b667-49799b97ceba · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning advantage collapse
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88954ec1-b658-48c6-96e2-cf676a9b9362 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ab987d7-7b65-4965-87a7-698d29b843c9 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning 𝑟#…RewardsAdvantage 𝐴! 𝐴
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf2ce7f-821e-4f41-bc9f-89f130a6d0ab · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning {question}Please choose the answer from the follow- ing options:{choice string}. Output the final answer in<answer> </answer>
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 316b11d5-8a9e-40db-aa51-15869eabe27d · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Our framework establishes a closed-loop learning process by generating confidence-weighted pseudo-labels from majority voting to guide policy optimization with GRPO
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 589131f7-a8e1-4d51-9804-9259c77714ec · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Listen, Think, and Understand
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24c589b2-21f9-4b24-8e9c-b95b9da91b0f · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Salmonn: Towards generic hearing abilities for large language models,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69095e2-c45c-4a89-b52e-57faec7ab2ac · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fec4e2b-0834-4458-92cd-27227a7beb4b · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Audio flamingo 2: An audio-language model with long-audio understanding and expert reasoning abilities,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0a6f0d2-4a32-4b43-8924-e64f340f3f43 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89840c29-ad96-42f3-868b-da8d84e3413b · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93cfa4e8-ca69-49d3-8e57-caa48c531417 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Qwen2-Audio Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513d9acc-d37d-431a-9187-d31604016403 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a186adcf-1b91-40f4-8702-c0d913298660 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Omni-r1: Do you really need audio to fine-tune your audio llm?,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcb41631-a86a-47b7-a51d-857cad3553ce · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning TTRL: Test-Time Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d3bb884-ac91-4cdb-8693-dbc4baaa8db6 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0110b7e2-d07c-4589-b4cf-4ab7830a3201 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8424f054-81a8-4cd0-8309-8e0bf1930b9c · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f721397d-e7b9-4569-98c5-5247d43af320 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2cef981-2ec9-4d57-ad06-77b058848d37 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef63f2d9-aef0-4e24-8243-9876e2f2ab91 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Desta: Enhancing speech language models through descrip- tive speech-text alignment,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 704371f3-759e-4cf2-8beb-2747351e0162 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Developing instruction-following speech language model without speech instruction-tuning data,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513f4a43-d9e6-47ef-96c5-66ec8faf365b · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Desta2.5- audio: Toward general-purpose large audio language model with self-generated cross-modal alignment,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe99f8e-e2ae-4851-85b9-d275d025b811 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28a2971-a4d6-4140-b04f-4f1f68ed129d · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b33e03-d5e7-46f3-a8d3-77a9e3b3abb4 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ada16e7-107a-49cb-adf2-49ca8fd00ad0 · outbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Qwen2.5-Omni Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88954ec1-b658-48c6-96e2-cf676a9b9362 · inbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6751f093-cfd9-40b0-8119-0ced94544a48 · inbound
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.