Pith. sign in

Paper Citation Record · LEDGER

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

As of 19 August 2026, this Paper Citation Record lists 100 of 233 outbound references and 0 inbound Pith citation observations for arXiv:2607.21550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21550 v1

Coverage vector

measured 100 of 233 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:12:17.572218Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 233 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e5715848-40de-4ea2-9290-1d6634e81f78 · outbound

This paper cites Aho and Jeffrey D.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Aho and Jeffrey D

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:06.756641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:06.756641Z digest=sha256:27e94b3c340d22197e45cdc540e7028ed0122a2b77f4e06a369ffe060a0de8c3

Observation f327697c-0e8f-45b2-a46d-03dadc3ddf77 · outbound

This paper cites an unresolved cited work.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:06.834197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:06.834197Z digest=sha256:88a2f12b0768e1f4e57151cdf75c0b8aec00e6badc714f27131fd4613f070926

Observation 4cc9e917-5c4a-4b69-9c69-98fd68785d60 · outbound

This paper cites Chandra and Dexter C.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Chandra and Dexter C

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:06.938263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:06.938263Z digest=sha256:58845606ba78b6e4c817bd931314abd5fcaf337de2c17b49dfc7b0578809c91b

Observation 088b67c6-abff-41f9-a533-6c3f636bb6e1 · outbound

This paper cites Scalable training of.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Scalable training of

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:07.093622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:07.093622Z digest=sha256:d7d1435c2a5cafed93ce59d43b11ab39cc9e555cc413ec0b577e026ec7165354

Observation 0917e93f-2422-4f06-a968-480cd27278d7 · outbound

This paper cites an unresolved cited work.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:07.260687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:07.260687Z digest=sha256:941fe621b20810a02f050b7dcae0f7f27b93e8c95c0878d2377b4271f8b33984

Observation d7bd3ee8-a727-4a3e-b53e-803113813988 · outbound

This paper cites Tetreault , title =.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Tetreault , title =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:07.416008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:07.416008Z digest=sha256:8e481158cad750083786bf2d157a23734ac92a140f8aa2f990634f6b1690d21e

Observation 6aea00fd-e092-4ee4-8bcb-6a038f717ab9 · outbound

This paper cites A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:07.593193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:07.593193Z digest=sha256:d2f3ceb8a63edd9de62c7e3f515e5dea56eb23740667e11a5e8174748c78bced

Observation b8185f40-55c2-416a-9de7-913595996a98 · outbound

This paper cites Scaling Learning Algorithms Towards.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Scaling Learning Algorithms Towards

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:07.727557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:07.727557Z digest=sha256:850c26cc44e8d0c94fdb7d509d403435a03fffe83f34e3778d53d8453ab7f82a

Observation e25d93eb-e36e-4909-9e71-0980a8d6e671 · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment and Osindero, Simon and Teh, Yee Whye , journal =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:07.854999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:07.854999Z digest=sha256:352ef3cb20925300d2e85507c2ea54da6fb4622616868a599424d33e81ab12bf

Observation 83cc3342-d0dd-4fe6-949c-13595d840cd0 · outbound

This paper cites Step-Audio 2 Technical Report.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Step-Audio 2 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:07.986457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:07.986457Z digest=sha256:8d224f1e67273de0afd7043b71f16ec8879fcbb465f919648a75cdbf729818f2

Observation 46fb2046-e84d-4c07-840c-83b750e7cd33 · outbound

This paper cites SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:08.148931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:08.148931Z digest=sha256:1f57472fb5cbfc3cf998a66175aa16af572cfb107b336cdbd63f535e95bfeeee

Observation 72a440c8-4c3a-4ec3-aedd-53308aac27f0 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:08.294139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:08.294139Z digest=sha256:3a3979dd9b447ddc39dca0a35ad740934136f0a9899fdb46b5e59db26af602ca

Observation 101d21f0-8e8a-42fe-a32c-51224e641643 · outbound

This paper cites 2016 , publisher=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment 2016 , publisher=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:08.454212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:08.454212Z digest=sha256:865de6dfc86973527a9f51e764f10c53c5d9ac239fe4ebd90d15961cff8cad8c

Observation bee5254f-8fec-4d1d-9913-05f4bb2b7f41 · outbound

This paper cites an unresolved cited work.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:08.633748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:08.633748Z digest=sha256:8e4cad4289ec3e302a1e9e01fe9068d952d78c2a5da69c1447b821aa0987759d

Observation ee2d21fb-74c6-4651-8323-41a72afe571f · outbound

This paper cites International conference on machine learning , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment International conference on machine learning , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:08.772091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:08.772091Z digest=sha256:471b198c0d96bf2e0ea0e4933db852fa8ceb562ad4b7a911747db6a57c5cbed6

Observation 485d7e01-2773-4c0f-8bfa-e4f75d2375fc · outbound

This paper cites https://cdn.openai.com/gpt-4o-system-card.pdf , year=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment https://cdn.openai.com/gpt-4o-system-card.pdf , year=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.055716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.055716Z digest=sha256:04082fe3d4bd21806c41479fe931ca21f42181e62f0c33ba58f4546b87ea7250

Observation 78a30ecc-3368-46ef-89e9-581d48552d13 · outbound

This paper cites https://openai.com/index/chatgpt-can-now-see-hear-and-speak/ , year=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment https://openai.com/index/chatgpt-can-now-see-hear-and-speak/ , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.199962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.199962Z digest=sha256:449de74d5647513b7c66a70661a98cfe20614851f622dad52b87b6a5f46ef086

Observation a9f3aa32-2b0a-4df9-8681-485e09de9725 · outbound

This paper cites 2009 , publisher=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment 2009 , publisher=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.320404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.320404Z digest=sha256:019f87dd61f383db3de94be1c918c46620267fcca5a01e5c709c09d8875b721e

Observation 5eea0196-5cd2-429f-bdc8-c28a9be84a81 · outbound

This paper cites 2016 , publisher=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment 2016 , publisher=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.455539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.455539Z digest=sha256:2d3b7b2a81d8fa2c4ff6c48ed128affdb03c0c6e31b7cca6666d8d421628e522

Observation ee87b208-b527-4a46-ab69-651fec2e2632 · outbound

This paper cites International Journal of Speech Technology , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment International Journal of Speech Technology , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.517358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.517358Z digest=sha256:947c22b476d1e2e7fa7bf77561fd8f71e308f78ea3074a9caa948deaeda72ad5

Observation a6a840f9-ef49-49e9-b96b-65268a27ee67 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.599563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.599563Z digest=sha256:6d6edd4f64b8bd5385ed351ea5109f13395a2a2d584cda17d50b06f932bbb2db

Observation 4d49bade-5a8e-4f9a-8d24-84856afda852 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.684733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.684733Z digest=sha256:74b52ef94da0df3079a27558f40ac6f9f792b0a2c9b86751b9195f191a803543

Observation 6722b941-deab-4175-8dfe-6088a097031f · outbound

This paper cites 2025 , eprint=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment 2025 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.767021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.767021Z digest=sha256:4c86a21e06a10c3067f0b26df30366adb3b8374efeeecc1f268a378960406b08

Observation ab6ad535-b454-4a7a-b44e-43a4407ff4bc · outbound

This paper cites 2025 , eprint=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment 2025 , eprint=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.847222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.847222Z digest=sha256:6e6c351c1cee1ac2acd69e6c52b1ac9328db16ec0ec40c8f00a007ee646055df

Observation 20155239-a0ad-491b-a478-9f6901150394 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.927490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.927490Z digest=sha256:26b4d1ed30dc3c64c10978ebe75e1ac052e2fd507d75f20dceca304b9d76107b

Observation f64c19ce-d65f-420d-8d49-364838d3480b · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:09.988524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:09.988524Z digest=sha256:d56af864da305e8639038de440bd4c0a6427e9d2db6550154a9e1a779b9ff50a

Observation 4d351f54-86e3-4725-925c-88f850bc0b3b · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.070955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.070955Z digest=sha256:a3a8116564e7c36f0ba9370bc76229088c30303bd686ac1128f81eed966e33c7

Observation 4977d279-9efb-4c3e-85ac-e0f92bdb0309 · outbound

This paper cites Acoustics, speech, and signal processing, ieee international conference on , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Acoustics, speech, and signal processing, ieee international conference on , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.153355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.153355Z digest=sha256:59ba1ec6112a7e11739bc26b13fe5628c371422f93c444e62901b84a13d9cc09

Observation 246322d0-02e8-47f8-a425-78bdd04036be · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.212235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.212235Z digest=sha256:9fbc16f1ddb16603fd0677091b66b1af7e47219dea60a0741d8bd097564b9870

Observation 52efee45-f724-4463-8a43-d4f929878baf · outbound

This paper cites , author=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment , author=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.290516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.290516Z digest=sha256:9328d646f0a067c569d80876a19233d202f94010ea375941a2f5a7ac39417282

Observation f0fefe6b-67f9-407f-8c10-6c53607a2e75 · outbound

This paper cites Language resources and evaluation , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Language resources and evaluation , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.349619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.349619Z digest=sha256:44c96528888c982b67dd1d53d99fedbc51bbb824bae9eacf74a5ae969cbfd796

Observation acaaa15d-3a5b-44be-9078-d94072db7188 · outbound

This paper cites Proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (SIGDIAL) , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the 15th annual meeting of the special interest group on discourse and dialogue (SIGDIAL) , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.413911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.413911Z digest=sha256:167b3f23360d81188b54a81672e1b0e38a289e18fa59b4181a10f718da3295fd

Observation 256e3517-1a63-4ca0-b573-407c72110a94 · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.474782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.474782Z digest=sha256:4ca749c433ebfd81227c8a32d32df01b01aa3d11e286065ba1de0a542464ca4f

Observation 0a077164-05fb-4d99-bc22-0a90c183f1c2 · outbound

This paper cites ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.532059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.532059Z digest=sha256:226b8cb5b0f4a8e8a4b9ae5f0f578daa7b625f1ed122d842846d7bc30f2d64c1

Observation 8cac6bd0-c1bc-4bf6-8b32-02864f55f387 · outbound

This paper cites Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.622660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.622660Z digest=sha256:4c32dc60e2da174c1b4eaf1d2baaf95eb7a5be194355d5dca35600a4b6a5ea91

Observation 8851f57e-099b-40a7-a95a-cc9680a588c1 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Advances in Neural Information Processing Systems , volume=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.726722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.726722Z digest=sha256:f7b9bfed4b02825d6b63b3b48c2f61389c111f9c27de8c8a2e32cb95d025df78

Observation 2b21b3ca-b6b4-40c7-8dc5-4c487328d496 · outbound

This paper cites E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.901933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.901933Z digest=sha256:4a319ff9d5cdd215219f92c82a58681f327efc949840da6928bc06b9b85928dc

Observation 64971b66-3ff6-48c4-847d-96dcacdeb01c · outbound

This paper cites PCWorld Retrieved October , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment PCWorld Retrieved October , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.069973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.069973Z digest=sha256:fb749e1fa198dc7ce9647ab32655f2f4d987628e27782e2e46597c3f359c8e97

Observation 5e51c39d-defd-4a26-86bf-4cf89f94e98d · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.242181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.242181Z digest=sha256:125b137025de377f236ddd2f2e2a6a53652956aa244cf6fcab5e3100309314d2

Observation 57453f28-eb6e-402b-a255-ad737c671313 · outbound

This paper cites Medical reference services quarterly , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Medical reference services quarterly , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.398365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.398365Z digest=sha256:aca97d7b5abaec8853492ea0559bed50b1f41610dc1ba7c345683e8b27dcecd3

Observation e98932ef-6fc6-4774-b43e-c1653e5faa6c · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.459753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.459753Z digest=sha256:7201ce1f59175d3eed077b3e38ae2965fd160b3a2e1a118aee1baea28a79279f

Observation 5d988626-77ea-4a0a-a71e-28ddacfb1f1b · outbound

This paper cites International Conference on Machine Learning , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment International Conference on Machine Learning , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.546538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.546538Z digest=sha256:1d17fcddaa79ec542da980cd08308b97ee14a65fe50380c943a59eab435c0860

Observation a74e8fc7-d053-4e79-b5e4-9a50930558b1 · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Advances in Neural Information Processing Systems , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.628642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.628642Z digest=sha256:186a15031b57411dd5f87cb0678946e1ca800256e1e12e8821f2a23cc58e122f

Observation 3474efc5-4f5c-4a75-8ad1-1ee3d4184eae · outbound

This paper cites International conference on machine learning , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment International conference on machine learning , pages=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.732058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.732058Z digest=sha256:3d5b8c8b212dc67f2388f0c2891db22cde0aa28bedabdd52ef8bf3fc0cce7196

Observation b0730a86-9db4-4bf7-982e-c9ac644c3e76 · outbound

This paper cites Proceedings of the 40th annual meeting of the Association for Computational Linguistics , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the 40th annual meeting of the Association for Computational Linguistics , pages=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.930013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.930013Z digest=sha256:8c682f11a29409a9ce70a3d82e9a02057376b1e0ae57ce6956e9a03a2e027f32

Observation e9bae5a8-7636-41fd-a00f-b72ba6ed2f8b · outbound

This paper cites Text summarization branches out , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Text summarization branches out , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.971081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.971081Z digest=sha256:8a79f3b7c3ced39d46baaef4e8d3522a6ae160b6aa2271542be04c4cdf6d9bc9

Observation c7d358ef-e263-4686-bcde-1e1eaf98190c · outbound

This paper cites Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization , pages=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.092018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.092018Z digest=sha256:1a35d5d8e6d0f737fe75771321b9220baf8c7ee11df6d30596c623d982a51710

Observation 94c15fc3-4315-405e-9ecd-a143f547e081 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment BERTScore: Evaluating Text Generation with BERT

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.186994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.186994Z digest=sha256:9c202b0061d3bab9a90373089d7038e817e64d0ebd1a89c4ed3b6c8c1377fa2f

Observation 47553334-8e5f-485f-9fc3-c4c39eb394af · outbound

This paper cites an unresolved cited work.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.289894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.289894Z digest=sha256:b71a74ce64039afd9ee64bddddd52c2c29ae57d0d5395301962bdd3d3db8a95d

Observation 0a05c3c2-ccb1-40ed-bff3-ee5e2d174595 · outbound

This paper cites Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.382374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.382374Z digest=sha256:cc17a3e4230f666d54fd57d5f36bb3430e0ef764cfada04b74616abcafe9777d

Observation bf01ea21-adfc-4f0d-ac04-56124e9a977b · outbound

This paper cites an unresolved cited work.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.481058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.481058Z digest=sha256:63ee332e4ae714c518c29eebfe42cfcacdbc53212e9681135f4176aefd9632e4

Observation 88afe302-0d6e-4d8a-81c7-d889ef53df20 · outbound

This paper cites an unresolved cited work.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.556830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.556830Z digest=sha256:42595da3e062cc357e8be0d65b4c9a40e248bf6a24b22ab7c13bfc6433fe28e5

Observation c737f141-3e85-4f47-8da6-15215ba5abff · outbound

This paper cites The Llama 3 Herd of Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment The Llama 3 Herd of Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.619135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.619135Z digest=sha256:4123636144e75ba8a5ad1056e645106bb784fc73b080db4bdb0481d70d5cad75

Observation ac4750f5-df8d-4cf0-b632-758959a915dc · outbound

This paper cites AIR -Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment AIR -Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.722309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.722309Z digest=sha256:d7cc7b876e0df283d318826f81f477cbdd9ae8527ba68ee967bb0efa79cce0c0

Observation 34e00dab-3979-45ff-a610-3e9cc738bbae · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.789704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.789704Z digest=sha256:1cd5cd2c06d42fca5b1ce6cf32bce9196b433abab2fcd98c82c5bb74a071f587

Observation 8c4c97bd-8c9e-4f18-9c4c-a053c5fea72c · outbound

This paper cites ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.851050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.851050Z digest=sha256:2c0bda9e82d8c0b594d053b16936e225ce11273e1b388905b21e5ba1c89a70ed

Observation eeac6aa0-6d14-4af2-a656-689ff120777a · outbound

This paper cites SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.911134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.911134Z digest=sha256:ed09c2ec590700d4393cf368ab3fd567a72c91d0a706ef30a678e10ade7e7218

Observation d65ae3e8-d898-4b20-bd84-c976d9391b04 · outbound

This paper cites EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.988342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.988342Z digest=sha256:bc6385808a0602e41fd307fe286e7530b283f292ba1d5bafcf4a219dda486e64

Observation 537b5764-7b62-40b8-b487-ac214c907a15 · outbound

This paper cites 2024 , eprint=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment 2024 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.058354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.058354Z digest=sha256:5c38dfccc46006dbc6020064e5386b76eb96ec048b51f36f7c3366985bcfcf2a

Observation 788183f2-e31e-4971-a137-16cb22e0c543 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.127289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.127289Z digest=sha256:c3542ba92c31b09ae083fe86b0a50bf2856ccd1925aece65d0abe7e6b88a4a56

Observation 85185671-4be5-4501-b00f-36aa0bf3d94e · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment The Thirteenth International Conference on Learning Representations , year=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.218676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.218676Z digest=sha256:d2023194151f7ec06ba6f523cf18c19e455842c0cbef9eec00177c209b57ba43

Observation da57ffd8-7705-4d92-8a51-a7d843c25581 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.300557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.300557Z digest=sha256:d52987c7b577b58e564c89396ae8882becf0eed024304e224515fc66c65df2f7

Observation 3ade6424-1a5d-4c06-a5bf-5b2ab43eb11f · outbound

This paper cites On the Audio Hallucinations in Large Audio-Video Language Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment On the Audio Hallucinations in Large Audio-Video Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.375070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.375070Z digest=sha256:f25107bd5c46c333aefbe2938c0b6bd2567333a9613faa21bfbd5235cfe782bc

Observation 91e29d06-1446-4d36-972a-27d491729775 · outbound

This paper cites arXiv preprint arXiv:2503.02318 , year=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment arXiv preprint arXiv:2503.02318 , year=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.484221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.484221Z digest=sha256:3168304c5b14e63324091eac1d9a467458ad6f727923b43aa08454c7e93f7e28

Observation 5ad67401-fa29-4e19-ae9d-06f0da398954 · outbound

This paper cites Proceedings of the 32nd ACM International Conference on Multimedia , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the 32nd ACM International Conference on Multimedia , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.561062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.561062Z digest=sha256:1903fd2bc74f6670896dfbf33ea7aecf119b4aa1c24de4e69ddaf31173f704a8

Observation 847507f7-43bc-4b2e-9f8a-cc64315c3893 · outbound

This paper cites doi:10.5281/zenodo.1251982 , url =.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment doi:10.5281/zenodo.1251982 , url =

Reference 68

Resolution
verified exact
doi, observed 2026-08-01T07:14:10.687930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-01T07:12:13.767987Z digest=sha256:d72e3796afc7f23670683daf4da2b5676f506022e68d410c00bae52ac59e876b

Observation 54777cc6-4c33-474c-ad29-200716bdb6b1 · outbound

This paper cites Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems , pages=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.978068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.978068Z digest=sha256:9ad4a52a22284b9548332b47680eba47615ed715a3929b5642fbc5028be3b6aa

Observation 3de9232a-1ecc-4aae-bd1c-cf3263bb4b4c · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Evaluating Object Hallucination in Large Vision-Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:14.152450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:14.152450Z digest=sha256:b9cc9f0e5356b86654ea8e923347b3b5ccab8bf5a4c98433fdec4622045286a9

Observation 4225fb6d-d4d5-47c5-b176-a716dd65f93b · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the AAAI Conference on Artificial Intelligence , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:14.375866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:14.375866Z digest=sha256:0d608f5fa894ce51a451f76f24d78a2b0c0a5267e796f57874db3a5b60524f2c

Observation 6d6df894-ae6d-4756-b326-2297a6a1578d · outbound

This paper cites A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:14.607895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:14.607895Z digest=sha256:3da0f3d8c2296ac4767bf656dfd2b5da3f3481be5e5ce304ac57c20f8718f23a

Observation 9c5e98ce-e9c2-481b-9d03-7c4b010f9b61 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Transactions of the Association for Computational Linguistics , volume=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:14.822627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:14.822627Z digest=sha256:2be2775fd245ff91cadb7b232b1986d361208a4b0b6aa6f2d9c16ac8785ff5fc

Observation 15e0ed74-aff8-4d29-ba38-0b9a098f85cb · outbound

This paper cites CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:14.955221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:14.955221Z digest=sha256:222fc3d6c0557dd44609bfdfd317b3d79ba7ecd5024b5a1b130ee60547f64b2a

Observation 2266438c-c857-43bc-9626-7a2a9a070e10 · outbound

This paper cites ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:15.101215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:15.101215Z digest=sha256:bb6ebf1c228b2650ae3b8e83164e4b08463bf6a8823661343c2f4806e98ff410

Observation 27fa3c65-bbcb-4518-8b61-c2db9a331c32 · outbound

This paper cites 2024 Conference of the International Speech Communication Association (INTERSPEECH) , year=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment 2024 Conference of the International Speech Communication Association (INTERSPEECH) , year=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:15.271713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:15.271713Z digest=sha256:0e2ce6d4883e341f1e673164277f63a20f0300fb9f346996498128f6bb23dea6

Observation 798d44f9-d625-476e-b436-eb0ef3435f70 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:15.461952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:15.461952Z digest=sha256:884874046df85c95a37d8b4f20f1d6b3ea1c68577583ccb2a546439f9c4b807a

Observation 424ba466-d2a2-41b6-846d-a9a8d457e801 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:15.632025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:15.632025Z digest=sha256:2ea94a5b473087a627a6a0d3a7a9ed37c946b22739c30da3a9274cdbb8367928

Observation e5c0c47f-6b82-4139-b0e8-2c95d94bf9d0 · outbound

This paper cites AudioBench: A Universal Benchmark for Audio Large Language Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment AudioBench: A Universal Benchmark for Audio Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:15.816142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:15.816142Z digest=sha256:6bfe5c9ed9be38a5607d45bb592ce47374eed83ea4fa21cdfb362b1d759823a0

Observation aa1a1529-2aa6-46ad-8c87-452f3da9e328 · outbound

This paper cites Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:15.903633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:15.903633Z digest=sha256:920694399708e6572e3a7fb09c27739dbe6ec877b7285769d746464a6fcd2cae

Observation b1511d11-6966-4856-8b76-de8a36588d31 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.034175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.034175Z digest=sha256:f7c843b4d7c41e605bcd3650e3bc0439e00de57441256d0903c7ccaca22cfa86

Observation 4125ddc8-3dcc-4bea-9ca8-136e363c02fe · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.095101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.095101Z digest=sha256:19f2b413a4b6275d2bd29cca88d9aa036d61a05f9ca3e2d77ed4a3d57ad2e5ba

Observation 4375b592-f9ba-4093-8b3d-6fcc451728dc · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.146281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.146281Z digest=sha256:18024324f4e811b9598cda9c0656e4c0108cae3e41bdce2058d9fa783e65af45

Observation ae8db77a-1ced-4a00-901d-ca31dd04de34 · outbound

This paper cites Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.200386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.200386Z digest=sha256:79a60dcf19de72dab6e847f9e1dea557f26ef79dc17e357cee72fea35b7ba9f1

Observation 53654071-4067-4275-b19c-607b39c4e9d1 · outbound

This paper cites International Journal of Computer Applications , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment International Journal of Computer Applications , volume=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.341950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.341950Z digest=sha256:88b6df924799513deb5e1aa655993a605a2740ad0ed5a1317db6c00d43b8d08e

Observation 3720dd43-b85a-41ec-b2c4-d263f33ec55a · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.390007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.390007Z digest=sha256:dd0971d89745932f09b76009e79a194025b77c576e1a3b805db9899c51253dd0

Observation b380fad3-7681-4a7d-9d9f-5faba4309e90 · outbound

This paper cites 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) , pages=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.437302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.437302Z digest=sha256:0e37a1f58e0f23a7da5aa0ad1ad4db737bf17640a765f4cd75f87f236638501b

Observation 599b2daa-9040-461b-83ca-e872bee9b881 · outbound

This paper cites Proceedings of the 58th annual meeting of the association for computational linguistics , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the 58th annual meeting of the association for computational linguistics , pages=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.510602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.510602Z digest=sha256:09e57cea1a683235c1d56209ffb0205d1a3d020931c56c3b73e4cbb8f861c11c

Observation 18352f4d-a462-4945-a5a9-133ff75fc6eb · outbound

This paper cites Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.661745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.661745Z digest=sha256:d8995f030d83fde216652faab4710242903d38278ddaf1947b4c46f45135c9ce

Observation a7b89f91-fc62-4b2f-a55c-280fec50ed11 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.774057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.774057Z digest=sha256:8adb12c435b7018122f67d4a12ba44de3c1fe68eb2830e70f5ae2a110d381329

Observation a077d62c-96f8-46b0-b0c1-5554bd719916 · outbound

This paper cites L-Eval: Instituting Standardized Evaluation for Long Context Language Models.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.866244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.866244Z digest=sha256:5991f966e1c620a53539eed090211e2de65c815829fd46c887e5079efb4909d3

Observation b849c910-6f92-4919-b7d6-13fa323e0dbd · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.961984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.961984Z digest=sha256:e3fef7b9efe77446abbd235c363be6f21b6b2d77a47c075b4d0cebec8f2b86c2

Observation 9d781903-13b5-4533-9198-444390d591e6 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.070993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.070993Z digest=sha256:5408d7d08f3492fe965a11fa7dc12e585a5c3f5b166a1386d1a49ea45ee1194a

Observation ea645a66-611d-4f17-8f19-4fe61d3755fd · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.158490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.158490Z digest=sha256:63239577c5ded3013017cd847144be4d64c432e5d7d026e7b079b48a1732ffa0

Observation b6e0f166-6d11-4580-b458-f1320e3f0227 · outbound

This paper cites , booktitle =.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment , booktitle =

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.255053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.255053Z digest=sha256:0be796a807c1a4e4b98c85f869a1de7bc20d19d7d26bde81742e5826d179eef8

Observation 8f6cf091-33d3-41e8-8353-c4a1b3287f16 · outbound

This paper cites Proceedings of the 23rd ACM international conference on Multimedia , pages=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Proceedings of the 23rd ACM international conference on Multimedia , pages=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.414459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.414459Z digest=sha256:4566dabe0e5347692e241778101b88e8c26e0b7116389181d8c32f1773ce2406

Observation ff021f35-3403-4f2a-b417-8604fee4d2f8 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.523618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.523618Z digest=sha256:1daa9469fe7dda6b13b1addb7def041c00be0a74e274e9a1e7a6484e12a8c8bb

Observation bb4a6421-3771-41ba-8c79-f93640cae73a · outbound

This paper cites Neurocomputing , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Neurocomputing , volume=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.561993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.561993Z digest=sha256:2b575a51bc4f3ac75f9bc5818eae5cc8f3a107f28ed232ec01d5e0861d293e39

Observation c3ba17b6-882f-4947-a654-8af4a9f0dad9 · outbound

This paper cites Long Range Arena: A Benchmark for Efficient Transformers.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Long Range Arena: A Benchmark for Efficient Transformers

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.564271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.564271Z digest=sha256:9e872a4ece6d5b632c05573ab728599015b0df6b99e8ea0b61528e6943bff8b1

Observation ada37a26-5187-4f53-b72a-338eaae07299 · outbound

This paper cites Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.567045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.567045Z digest=sha256:3c61a8876ae67f2ec8056d97e82e4e3e1fd2c2baadf2092a48ab492d5cb0ccc2

Observation 149b6ae8-df44-46f2-b627-27211285b6a6 · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.569658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.569658Z digest=sha256:45184ef540ff9ee730c0eaab6b3eb44b1b3017cb48bffdca1ddd44879aa1d888

Observation 82a98992-e04b-46ad-8857-1e8923c54ebd · outbound

This paper cites ACM Transactions on Information Systems , volume=.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment ACM Transactions on Information Systems , volume=

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.572218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.572218Z digest=sha256:4d583ef387e25b51b4d202cfb30efe4c1ec3f7b4a22e7cf4cdad8a1bd5c1f9f7

Pith citing papers

No inbound Pith citation observations are available.