Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T13:39:48.225482Z
Paper Citation Record · LEDGER
As of 2 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 35 inbound Pith citation observations for arXiv:2502.11946.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T13:39:48.225482Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:43:34.049064Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T02:49:25.012133Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 140d72a0-7b80-44a1-a3cd-11115e1b2599 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction The Method of Paired Comparisons , author=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 59504f3a-02fa-4189-8681-69058398f4db · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction International conference on machine learning , pages=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation db7516ef-eed4-463e-9032-585405001c8b · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction ICASSP 23 , year=
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7999e265-947b-48c0-8b19-8ee96d91f9a3 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction 2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA) , pages=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation caefa079-983b-414a-b880-a257b96ff91e · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 09d7c2d1-71a6-44f1-ba0d-78e57b6748b4 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e6957376-f545-4975-a7a1-1d0188dc2d06 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4c593f38-50e3-4bd9-b64d-ea8488ed5415 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction 2024 , howpublished =
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 11452c23-5220-4e88-a6f6-ef9c92a2b030 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction 2024 , howpublished =
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5af1c320-77b5-4d62-9907-d3a655eb9244 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction 2024 , howpublished =
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0024d1c8-e6d0-4bcc-9ca5-e0bd5e3fc405 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e92b9024-1312-4650-859e-7071f8a5f753 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction \ Terry, M E
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 39d98dc5-3309-4596-8421-75a62949bf42 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction APACrefauthors \ 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b8366d02-27ad-4e39-8c1c-979a30e28278 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 852657ce-7c19-473e-95bf-10933b2c1e2d · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5161f6ab-9758-4d17-9c3e-4ba99d038d7a · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Qwen2-Audio Technical Report
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 62792d3a-d3f9-48b5-badd-2459f82d3802 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction SpeechVerse: A Large-scale Generalizable Audio Language Model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e650777f-ee34-4217-a15e-44188e6de56f · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Moshi: a speech-text foundation model for real-time dialogue
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d59f2916-ea4e-4f3d-a6e4-fbe15deea9f2 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 353f164c-5839-4f53-9a37-9d6825c69551 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1625c794-0f28-4b4f-9768-e6e77f92aca3 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction The Llama 3 Herd of Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2d3ae2fc-a0b0-4740-8024-10f7dbc5bf8e · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7d3c2227-45f5-494e-b5d2-7890fd337c21 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction LUCY: Linguistic Understanding and Control Yielding Early Stage of Her
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 828da4f7-53b1-4efe-aa1f-31d64759fe8c · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation eb985253-43d0-4221-9e6f-15dd2dab079f · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction WavLLM: Towards Robust and Adaptive Speech Large Language Model
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3d6e2987-5ff5-4967-b155-b0902bf75b4d · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1c1befd4-35e1-49c3-86e6-7597a76b290d · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction GPT-4o System Card
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d578991a-7be3-4bf3-a85b-9d47a8c05ba0 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Cao, Y
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 72d6ebc0-f803-40cb-830e-3e6e6ebe8195 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction ARCON: Advancing Auto-Regressive Continuation for Driving Videos
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1fe309e4-2f44-4648-913b-fde8697162f6 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Spirit LM: Interleaved Spoken and Written Language Model
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 305b1f51-acbd-419f-be71-38d19996f661 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Kim, J W
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1bc57ac3-dbca-4126-941b-3dbca3c8a708 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Massa, F
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 357ca484-add2-458f-aa01-810c9ea87b08 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Proximal Policy Optimization Algorithms
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 724f732a-8be3-4d29-860c-d7261582ab66 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction APACrefauthors \ 2024 1
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d233337b-4296-40e9-81d4-9167ccf7b470 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction APACrefauthors \ 2024 2
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation eef0cf8f-ad35-490b-9f52-6c2b723ccec9 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4b4fd7f8-9bd0-4b4d-b44b-0a7df5b37aa7 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b6ccf608-f3df-4855-b006-c7740879e3e6 · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d15e40ce-8a13-48cd-95d1-f464427993df · outbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Disttrain: Addressing model and data heterogeneity with disaggregated training for multimodal large language models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 13ac8ad6-f421-41ce-b01f-54bfe7b17bbc · inbound
Kimi-Audio Technical Report Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ab37e7fe-3520-4d25-a911-e918eb5b194c · inbound
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a90d7b25-65a4-4383-92f1-9f8da6f8cceb · inbound
Step-Audio 2 Technical Report Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e1b44793-28aa-422b-870b-ede372a11e0b · inbound
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9cfbed63-0060-4540-9caa-d89e840fbd21 · inbound
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7a3b9a6b-048d-4fbd-94c0-3d777db17cb0 · inbound
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation fbe14310-0d2a-48b2-ae6a-962ebab34970 · inbound
ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 104e889b-1350-4146-acf2-dbed95c018de · inbound
Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation fa3793dc-dee9-444c-8054-be4de0c92e79 · inbound
Same Words, Different Judgments: How Preferences Vary Across Modalities Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a8a81817-877a-4a6b-9868-775f19da4368 · inbound
Bridging What the Model Thinks and How It Speaks: Self-Aware Speech Language Models for Expressive Speech Generation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8a072de6-a6ef-45d3-aa97-e0e955af4523 · inbound
VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65ab0000-4558-4b3b-bc61-87cf5914c5db · inbound
Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation cdd9973b-928b-4ddb-ade5-a614c7dfdb0e · inbound
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 74e2921e-86fe-472d-9b9f-c95e31ff23d8 · inbound
Sema: Semantic Transport for Real-Time Multimodal Agents Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 86600281-9dfa-478c-b03d-9c1f35a28d3f · inbound
ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 72d91187-ece0-4e0d-8c42-5fb3f5bbe05d · inbound
ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 72c7adcb-7c33-4a2f-b2c8-e3c7d78203c3 · inbound
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c1ee0ef2-4d5a-470d-b565-8b193a0b988c · inbound
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d66889ff-1e86-49a8-ad0a-88aae65164a4 · inbound
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 51020bc1-d6ab-477c-b772-dc7bf98b6e5f · inbound
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a6edd245-4b5b-4bd0-a826-576966cff970 · inbound
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 292a4020-df8c-4849-b154-45e59cd1802e · inbound
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bad9d523-5735-4474-afb4-04a70cb742b1 · inbound
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b8d0dbc3-3147-49a9-84a8-58631c65090c · inbound
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 909eb336-2e95-4c26-9119-bc8976c9d00b · inbound
BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 6105dafd-069c-4ad2-bbf2-f004becb9545 · inbound
Audio-Mind: An Auditable Agentic Framework for Audio Understanding Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d4b3731c-80aa-49c6-84d1-d6fe04227c59 · inbound
MOSS-Audio Technical Report Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5c44acbd-621d-4bbf-b340-5a12e43a4a9a · inbound
Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 989c9785-7c98-4ffd-9537-92a5aa516087 · inbound
M*: A Modular, Extensible, Serving System for Multimodal Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation da34a01a-b062-45f0-a7d8-25496182aff1 · inbound
A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1f1ee91c-9617-45ad-a0f8-c85f5226a99a · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 214
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bb9728ae-77b5-48a0-b022-301d27390afb · inbound
Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1fb345c3-36e5-48f3-8748-467b166f3ba4 · inbound
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5608d3f-e015-4d6e-8d61-6f737ab51a14 · inbound
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1511d11-6966-4856-8b76-de8a36588d31 · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.