Pith. sign in

Paper Citation Record · LEDGER

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 4 inbound Pith citation observations for arXiv:2506.08967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08967 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:16.518476Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T05:59:50.900436Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T05:59:51.133845Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10cb4c00-fb42-4a17-8643-93c8cb711898 · outbound

This paper cites Claude 3.5 sonnet.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Claude 3.5 sonnet

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:20.303920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:11.339879Z digest=sha256:add6239f0fb591492dee82df9210670f0d1b7617fce434bb0e843e87c6f4fe00

Observation abd46843-edbf-447c-9d90-54b740e55788 · outbound

This paper cites ML4CO-KIDA: Knowledge Inheritance in Dataset Aggregation.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model ML4CO-KIDA: Knowledge Inheritance in Dataset Aggregation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.413339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.413339Z digest=sha256:d20d186d89a187c91838113672459b70a229c0fb8b3a84eba7c9965499882625

Observation c04f6727-a3cd-412a-8844-60e880dc20d2 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.529192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.529192Z digest=sha256:e74a1b0fd8155e2047b0128ac60f81ab77a668dd2a14be7f91616ccd38a1d123

Observation d0a5322d-cb5c-40cb-a5c9-1180a7413fd3 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.622724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.622724Z digest=sha256:c0a492eebafee752f9988bd612072ec2da61c796140add826f253a63b46da101

Observation a22e50bd-1de3-49eb-aca5-215897c7b636 · outbound

This paper cites Qwen2-Audio Technical Report.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Qwen2-Audio Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.715744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.715744Z digest=sha256:85a161022a452447bd8d205302f74db0f9b5d679b09eea260e9a02947ac764d8

Observation 6c0ef7c8-88f3-49f6-a9a9-767451b04f45 · outbound

This paper cites Simple and controllable music generation.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Simple and controllable music generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:20.146037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:11.799823Z digest=sha256:8287faabf15875648b9ef53ac6a79c357dc58382a1338297d4a650ec88002cce

Observation 41695169-4cb5-48d0-8092-c55a769324ee · outbound

This paper cites Recent Advances in Speech Language Models: A Survey.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Recent Advances in Speech Language Models: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.869766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.869766Z digest=sha256:b5de0be9cf2950a1bb44a7f0222f943a456b0bf7e9bf70737a4359eeb53bcae6

Observation 98a60a68-68ec-4beb-b36f-3992a44c7334 · outbound

This paper cites Pengi: An audio language model for audio tasks.Advances in Neural Information Processing Systems, 36:18090–18108, 2023.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Pengi: An audio language model for audio tasks.Advances in Neural Information Processing Systems, 36:18090–18108, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.991060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:11.955150Z digest=sha256:5893a637158a27bed394f76c921eff644fc22d059fd75bf99c878170be6c6f24

Observation 04ae4ddb-c8ae-4042-b533-bbee52d4e131 · outbound

This paper cites Kimi-Audio Technical Report.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Kimi-Audio Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.052740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.052740Z digest=sha256:cd0dca0eab4ca4872db079748783abeea9aacd85f7a3921adc689418ee1f9c7f

Observation 054ed7e2-d761-4403-9623-b840e7a9ba1e · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.174526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.174526Z digest=sha256:92bee6d8675ccf1d4a31f39822c4905c647ece09ce2f3afd9743d2f720605244

Observation 1daf89f6-b755-43eb-ba41-ef48f608a2c6 · outbound

This paper cites Emo-dpo: Controllable emo- tional speech synthesis through direct preference optimization.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Emo-dpo: Controllable emo- tional speech synthesis through direct preference optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.785059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:12.244985Z digest=sha256:c53be5f369bed7b2b7f2ad48307c1bf2632fc1885350e9173359f733407ce7f1

Observation debb1561-7cd4-450d-94b9-44605811f8d0 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.369009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.369009Z digest=sha256:e7e3377e64b245f5f75511e7e64d587a1df5e462966b96cda055ccb5d168a510

Observation a942c027-9629-45eb-a593-242adb83c897 · outbound

This paper cites Joint audio and speech understanding.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Joint audio and speech understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.651919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:12.467905Z digest=sha256:89f0b579f8dc8ec190555bc3ab719e1b9e4a13af0a6bdecfee28c2bac8ebea4a

Observation b205fbe5-d36f-4836-bd3b-f5605229ace5 · outbound

This paper cites Gemini 2.0 pro.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Gemini 2.0 pro

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.537131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:12.579313Z digest=sha256:724a65a2848d84f810f5e098491cb48dd7ddd110f0a51adba5068bbf535d8534

Observation f127b7f0-c905-4dec-8d96-fee8ce8bed7d · outbound

This paper cites The Llama 3 Herd of Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.674168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.674168Z digest=sha256:8d92f7efa58fd9ab82d8fea078f41dab4c7bc1b7183ec42e94c1eeb686124753

Observation 6d71bed5-e07f-464b-8b25-890a7a7847c6 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.752879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.752879Z digest=sha256:0c33a7395c1634b0a5bc5625a72225781617618d9bcc2b139a6bf6a6fcc5719b

Observation fa979691-538a-4d86-a543-e88cc561bffa · outbound

This paper cites Deep residual learning for image recognition.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Deep residual learning for image recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.402351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:12.827575Z digest=sha256:16e63a1e75e5b36a76d6f4539abc25748c3e7eaf71699c3de5248d2ba885c7d9

Observation 82e4a2a7-0850-40ef-9200-49d35d6dd197 · outbound

This paper cites DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.935432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.935432Z digest=sha256:cc5fcfeacb96bdd981e93d9a2e35b21391d00a103dc7df52bdadfc9b207b70dd

Observation 2ecf714d-a632-43b9-948e-0244e0216d2f · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.012840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.012840Z digest=sha256:732de9889aef968c7cdf1e54144c9202b1e70c26d3e80bcf840d8d5a8ab60e40

Observation e0ab10ad-e0c6-4cc2-9cb0-80ac60a65f79 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.099690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.099690Z digest=sha256:69ea9520f3ec00cb4262278e4332c9d796a47beb0cbfba8e38aeb859d1851941

Observation 4f417131-a65b-427e-8cfc-91cb7a895586 · outbound

This paper cites GPT-4o System Card.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.194856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.194856Z digest=sha256:1699c30b317af40ebbf58ab173d2fd52c8e13b4108fafb07c9a1c913c79283c8

Observation 8dfec203-8715-435f-9432-4cfbb285ceb1 · outbound

This paper cites OpenAI o1 System Card.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model OpenAI o1 System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.295965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.295965Z digest=sha256:a47b6b25466f66335bd8a30655c5b5e3a789954260f3944582d1671228328ff9

Observation 681a5cdb-7ba7-4382-9e1f-b936d7d04130 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model WavChat: A Survey of Spoken Dialogue Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.393147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.393147Z digest=sha256:d4d1a03efcca8a21ef1c273c7d36f3f4d93873d8d91a1544761fc683463f85a7

Observation 3b51b3cd-7c60-4af5-b140-ae3258ec7077 · outbound

This paper cites An llm compiler for parallel function calling.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model An llm compiler for parallel function calling

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.310225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:13.488079Z digest=sha256:02dd67a4fe7bc34ff671485fce751143024572b7dd0e180766b518585ea37cf3

Observation 4784308c-b4a6-411a-855a-1dc0db1d4d6d · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.573676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.573676Z digest=sha256:8a692f14ab3c9767428d311408201fb667633b54c271718c2d42d980e7609421

Observation 4d3e3944-1c43-49b2-a452-e36b1b2eda44 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.651263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.651263Z digest=sha256:4d5f67205f882b5487cc5431c5ac501ce99e0dd046df07f450d3d8e242b232a7

Observation fa47b73e-e297-4d3f-a577-85bf14e82af4 · outbound

This paper cites Flexkbqa: A flexible llm-powered framework for few-shot knowledge base question answering.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Flexkbqa: A flexible llm-powered framework for few-shot knowledge base question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:19.137327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:13.754388Z digest=sha256:7b4ff5713669d26e16e7d4a532ec1153d8090fdc197033ae0ae582b6a528c272

Observation 945b7c88-6470-46fa-989f-fd6bb4e79769 · outbound

This paper cites Merging models with fisher-weighted averaging.Advances in Neural Information Processing Systems, 35:17703–17716, 2022.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Merging models with fisher-weighted averaging.Advances in Neural Information Processing Systems, 35:17703–17716, 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.853463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.853463Z digest=sha256:4bd3100cedb246f104b199c5f9cf9ba168cbe3baa6451eb7665f243f11c88475

Observation a3ac23ad-a909-40b9-81fe-b2e8d0973b32 · outbound

This paper cites Using an llm to help with code understanding.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Using an llm to help with code understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.950462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.950462Z digest=sha256:ddbe7d43fde049a32dba6cec6aa674c41847287cc872aba276b194718c815767

Observation 6c60b368-3315-473a-bc21-6a6e4f930483 · outbound

This paper cites The role of paralinguistic cues in social life.ANALYSIS OF MODERN SCIENCE AND INNOVATION, 1(2):215–218, 2024.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model The role of paralinguistic cues in social life.ANALYSIS OF MODERN SCIENCE AND INNOVATION, 1(2):215–218, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.931107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:14.045361Z digest=sha256:b6764cfe420da62c403888f8244ae64ceb9bca3c08de8795e34d40eedddf45bb

Observation 2905d971-6f5a-4ca6-b4ba-96df5210b1ce · outbound

This paper cites A survey on speech large language models.arXiv preprint arXiv:2410.18908, 2024.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model A survey on speech large language models.arXiv preprint arXiv:2410.18908, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.141710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.141710Z digest=sha256:8f8a42bb2059ea1337f6cf70eb39dc45e27f93f6a1d9f268e5f9f812f487cf57

Observation cd1534b0-420e-4502-8a87-676fc3828313 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.234520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.234520Z digest=sha256:c3241539283895d79fd0b8ef996a6c9b899896171ee324d315debd3612b5e3b8

Observation 7f2b0cb2-a3ab-483b-adf7-e8e47cc6535a · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.333366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.333366Z digest=sha256:e7da80939cfff5235f79308ecabcc7c2838dec1ecbd90cfc57e7bf989b7a67fa

Observation ddd32cd2-70b9-49a6-85db-d6918228071c · outbound

This paper cites How to debug code with github copilot.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model How to debug code with github copilot

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.762496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:14.429648Z digest=sha256:c2bff9641c870c8eba5e8b28ba411c5f20326f5d266158e51b2ad05e888133c4

Observation 0450f964-3514-4b32-b6d4-1193c5f26a67 · outbound

This paper cites Paralinguistics in speech and language—state-of-the-art and the challenge.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Paralinguistics in speech and language—state-of-the-art and the challenge

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.596275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:14.503974Z digest=sha256:4a7e6bc226622f3bc3d346b386fa57b3a5c80eee9a1bba2c9191ee4c32e59e81

Observation 1f9684bb-3fd8-44ed-ba66-efc89785ae4c · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.618551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.618551Z digest=sha256:07a4b38d948cc42f0bbad100b2e892ebf4b19c602b5443bde01b928c0fe818d6

Observation 2c510c4f-08a8-4eb0-a358-fcd62f230dd2 · outbound

This paper cites Stepeval-audio-360.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Stepeval-audio-360

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.443217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:14.752046Z digest=sha256:f0e679cf8115e0618ae6eefc6f3507057e43c3f5becc7d361df1bc95e7a257e2

Observation 95470915-60f4-4c1e-8152-ab198d7116a5 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:14.876677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:14.876677Z digest=sha256:5da1f9a5a85cae358fd66e545a1a5b42eab019ff07a387fba7da4647c9d86aa0

Observation 81d875f8-529e-42f4-811b-7d2830f9f202 · outbound

This paper cites Preference alignment improves language model-based tts.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Preference alignment improves language model-based tts

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.297584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:14.945288Z digest=sha256:89c938bbb6ab486fe41a115547cdf366faca5cd61b9d7b4f4e8afc13f3a8530b

Observation d1195625-63a9-4c2a-b0c9-f18331437a73 · outbound

This paper cites Lami: Large language models for multi-modal human-robot interaction.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Lami: Large language models for multi-modal human-robot interaction

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:18.113523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:15.071882Z digest=sha256:c164492f7054fcf9bcb7c5674ebf75b61844974f47ba2722c376991c5328a721

Observation 907f9a29-8e0a-4a27-99f6-cd12850217ed · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.171753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.171753Z digest=sha256:019e21582435be47acd1960117d650f63d10fe10fb22ce2a1bd39867cedf6d73

Observation 41321313-bb77-41b8-856f-84ad7fcbe144 · outbound

This paper cites Reinforcement Learning for LLM Post-Training: A Survey.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Reinforcement Learning for LLM Post-Training: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.250044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.250044Z digest=sha256:1c33e15a2c5b0c139ce72dd5beb96101fb8d81a553dcfaa263e4d087acaea126

Observation 66595b68-8294-47cb-94a7-89faf54611fd · outbound

This paper cites Model soups: aver- aging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Model soups: aver- aging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.331162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.331162Z digest=sha256:c5d2ff58e81c537a1e41c595807a929890278e62fc9c6a2a00262e219291e920

Observation 962da2ef-045a-4b6e-b459-fe49ff99f826 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.425974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.425974Z digest=sha256:61dea7f5653293c70f38436ce288c3f6625be4c540278fbdfd0b433a7c738e11

Observation 07125a7f-1975-4d32-a803-21d414475f7b · outbound

This paper cites When search engine services meet large language models: visions and challenges.IEEE Transactions on Services Computing, 2024.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model When search engine services meet large language models: visions and challenges.IEEE Transactions on Services Computing, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:17.895203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:15.557323Z digest=sha256:a273d890a9ec9a51e9d642a888a8d4b4c16f2434e01eb800d06a708b419b9e68

Observation 7f1bb81c-d213-49f6-b752-269636b914d0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Qwen2.5-Omni Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.651302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.651302Z digest=sha256:457b85a5352663d73ed41a469206e0d1dfa5ed30358552b9efb3d197a7fd1d9b

Observation bd288521-e2f9-45f7-b906-647636d89695 · outbound

This paper cites Uniaudio: Towards universal audio generation with large language models.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Uniaudio: Towards universal audio generation with large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:17.676913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:15.749688Z digest=sha256:89b5c7d9245653cdeae82c3f9799d9d01892f9eaa9df3995cea2e72e131ca585

Observation 5eade042-b8c6-4f85-aa5f-faaa4851b64b · outbound

This paper cites Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.837924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.837924Z digest=sha256:d4511a146b0244fae72a763d808ced2763b111232f1c4c2dd5961567c1271180

Observation 6364a05c-56a0-4e4d-a795-8ba0b4bfed1d · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.913730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.913730Z digest=sha256:ad63c53d0c452a7a2eaa66e3fef61a5b55c669e155c595cf44a1e179d0302aae

Observation fb0bdad5-184f-4cb4-a36d-75ce3e54ac98 · outbound

This paper cites Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.976166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.976166Z digest=sha256:7fad08aaa3cd922f3995b9d35f4adeb0a3f3dffe71cf5c2d4068f1283d03723a

Observation 0d49440e-e10f-4e6a-9510-fbd1db9a59de · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.062120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.062120Z digest=sha256:c7c27786ab9b1e1b70368e2cf7d42d2cca18012535b11a0a1f15a6ddde31362b

Observation 359d011e-1276-4af3-b3e7-480e3cd1ae5c · outbound

This paper cites SpeechAlign: Aligning Speech Generation to Human Preferences.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model SpeechAlign: Aligning Speech Generation to Human Preferences

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.158400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.158400Z digest=sha256:fdb6614cd44190796cb63f496daeddab942936897ac3914d0c7e13f35fd96b26

Observation 2e36a9ec-f727-4c09-a3e8-6fcb7d392395 · outbound

This paper cites Vistorybench: Comprehensive benchmark suite for story visualiza- tion.arXiv preprint arXiv:2505.24862, 2025.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Vistorybench: Comprehensive benchmark suite for story visualiza- tion.arXiv preprint arXiv:2505.24862, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.265675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.265675Z digest=sha256:251b58fa22472f34bb5f43b0ac91747c12ad980f1bd9e7ae1de82110485279a2

Observation 4e967e3e-1770-489c-9c03-e9a1b3228968 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.Advances in Neural Information Processing Systems, 36:50117–50143, 2023.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Toolqa: A dataset for llm question answering with external tools.Advances in Neural Information Processing Systems, 36:50117–50143, 2023

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:16.429765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:16.429765Z digest=sha256:c86d2d8221ff4ab90da1adc7d12f06d280b90eb41feab7c5b90da3572d96a730

Observation a601aaa1-bf4b-4671-b545-3f298fe43b31 · outbound

This paper cites Pre-trained language model based ranking in baidu search.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Pre-trained language model based ranking in baidu search

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:17.385454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:03:16.518476Z digest=sha256:e3302b16188742848382d5946516e7fe559e2978eefcdbc1f339b1b24f89e526

Pith citing papers

Observation 77ab2a72-895c-410e-9eb3-653860b0be7d · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.136199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:c3be6df7d5f0df475e7974d5447a45d78a14b2dafb5edfba58eca50fe7cdf0e4

Observation ac1ddd35-444d-4687-88c5-1ac7ceb4cc4e · inbound

WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training cites this paper.

WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:10:09.183579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:06:07.355633Z digest=sha256:12edad596681ad4f39dad2f340d6277339e7b79d0ad8932a3d2327f7f3326343

Observation 86c944ae-9249-46eb-8c34-1251a9201072 · inbound

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints cites this paper.

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:05.205344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T03:09:41.001660Z digest=sha256:b4ab0249098831379ce03cfe8b1cd8a1d30208c528585ac79704874f5e6b64cc

Observation ab9d91c2-f249-4a97-9649-5b3a42591fff · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.014102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:114d614d28eee05a1f701846b327d849e667cbad08b60a47405ba7b77ed17f64