{"work":{"id":"03de0b74-a423-4e22-be1b-76d588429220","openalex_id":null,"doi":null,"arxiv_id":"2502.11946","raw_key":null,"title":"Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction","authors":null,"authors_text":"Ailin Huang, Boyong Wu, Bruce Wang, Chao Yan, Chen Hu, Chengli Feng","year":2025,"venue":"cs.CL","abstract":"Real-time speech interaction, serving as a fundamental interface for human-machine collaboration, holds immense potential. However, current open-source models face limitations such as high costs in voice data collection, weakness in dynamic control, and limited intelligence. To address these challenges, this paper introduces Step-Audio, the first production-ready open-source solution. Key contributions include: 1) a 130B-parameter unified speech-text multi-modal model that achieves unified understanding and generation, with the Step-Audio-Chat version open-sourced; 2) a generative speech data engine that establishes an affordable voice cloning framework and produces the open-sourced lightweight Step-Audio-TTS-3B model through distillation; 3) an instruction-driven fine control system enabling dynamic adjustments across dialects, emotions, singing, and RAP; 4) an enhanced cognitive architecture augmented with tool calling and role-playing abilities to manage complex tasks effectively. Based on our new StepEval-Audio-360 evaluation benchmark, Step-Audio achieves state-of-the-art performance in human evaluations, especially in terms of instruction following. On open-source benchmarks like LLaMA Question, shows 9.3% average performance improvement, demonstrating our commitment to advancing the development of open-source multi-modal language technologies. Our code and models are available at https://github.com/stepfun-ai/Step-Audio.","external_url":"https://arxiv.org/abs/2502.11946","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-04T02:49:25.012133+00:00","pith_arxiv_id":"2502.11946","created_at":"2026-05-09T06:25:45.999395+00:00","updated_at":"2026-07-04T02:49:25.012133+00:00","title_quality_ok":true,"display_title":"Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction","render_title":"Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction"},"hub":{"state":{"work_id":"03de0b74-a423-4e22-be1b-76d588429220","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":26,"external_cited_by_count":null,"distinct_field_count":9,"first_pith_cited_at":"2025-04-25T15:31:46+00:00","last_pith_cited_at":"2026-07-01T03:02:31+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T07:19:43.772928+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":3},{"context_role":"baseline","n":2},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":2},{"context_polarity":"baseline","n":2},{"context_polarity":"support","n":1},{"context_polarity":"use_method","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}