Pith. sign in

REVIEW 17 cited by

WavChat: A Survey of Spoken Dialogue Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.13577 v2 pith:AZVUUSCO submitted 2024-11-15 eess.AS cs.CLcs.LGcs.MMcs.SD

classification eess.AScs.CLcs.LGcs.MMcs.SD
keywords dialoguespokenmodelssystemsspeechtechnologiescascadedinteraction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in spoken dialogue models, exemplified by systems like GPT-4o, have captured significant attention in the speech domain. Compared to traditional three-tier cascaded spoken dialogue models that comprise speech recognition (ASR), large language models (LLMs), and text-to-speech (TTS), modern spoken dialogue models exhibit greater intelligence. These advanced spoken dialogue models not only comprehend audio, music, and other speech-related features, but also capture stylistic and timbral characteristics in speech. Moreover, they generate high-quality, multi-turn speech responses with low latency, enabling real-time interaction through simultaneous listening and speaking capability. Despite the progress in spoken dialogue systems, there is a lack of comprehensive surveys that systematically organize and analyze these systems and the underlying technologies. To address this, we have first compiled existing spoken dialogue systems in the chronological order and categorized them into the cascaded and end-to-end paradigms. We then provide an in-depth overview of the core technologies in spoken dialogue models, covering aspects such as speech representation, training paradigm, streaming, duplex, and interaction capabilities. Each section discusses the limitations of these technologies and outlines considerations for future research. Additionally, we present a thorough review of relevant datasets, evaluation metrics, and benchmarks from the perspectives of training and evaluating spoken dialogue systems. We hope this survey will contribute to advancing both academic research and industrial applications in the field of spoken dialogue systems. The related material is available at https://github.com/jishengpeng/WavChat.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dual Information Speech Language Models for Emotional Conversations

    cs.CL 2025-08 conditional novelty 7.0 of 10

    A dual-adapter design with equivalence replacement regularization lets frozen LLMs perceive both paralinguistic and linguistic information from speech for emotional conversation.

  2. Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Lychee-FD resolves modality interference in full-duplex spoken language models by separating acoustic and semantic parameters in deep layers and adding a dense semantic alignment channel, achieving state-of-the-art pe...

  3. TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

    cs.CL 2026-07 unverdicted novelty 6.0 of 10

    TurnNat introduces a likelihood-based automatic evaluation method for turn-taking naturalness in dyadic spoken dialogues using a causal prediction model and a human-validated perturbation benchmark.

  4. A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff

    cs.IT 2026-04 unverdicted novelty 6.0 of 10

    Synset-based reconstruction and synonymous variational inference are claimed to derive the distributional divergence in RDP and unify it with classical rate-distortion theory.

  5. Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Soft phoneme-similarity labels in multi-task training improve verbatim phoneme recognition and yield two new pronunciation error metrics.

  6. OpusLM: A Family of Open Unified Speech Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new open family of speech-language models trained on public data matches or beats prior systems across ASR, TTS, and text benchmarks.

  7. Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Using Mimi neural codec features with label-delayed training reduces endpoint cutoff errors by 42.7% (single-stream) and 37.5% (two-stream) at 160 ms median latency.

  8. BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A speech-and-bias contrastive retrieval framework with homophone-aware curriculum learning scales contextual ASR biasing to 200,000 entries while improving B-WER on LibriSpeech.

  9. PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs

    cs.SD 2025-05 conditional novelty 6.0 of 10

    An LLM prompted with transcripts, laughter, turn-taking, backchannel, and emotion statistics predicts Big Five personality traits from two-person phone calls, agreeing with human ratings more closely than three text-o...

  10. X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

    cs.LG 2026-07 conditional novelty 5.0 of 10

    X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.

  11. SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM

    cs.CV 2025-11 conditional novelty 5.0 of 10

    SMART, an audio-enhanced MLLM with shot-aware token compression, reports new state-of-the-art moment retrieval accuracy on Charades-STA and QVHighlights.

  12. ChipChat: Low-Latency Cascaded Conversational Agent in MLX

    eess.AS 2025-08 conditional novelty 5.0 of 10

    ChipChat is a fully on-device cascaded voice agent that reports about 920 ms total latency using streaming ASR, a state-action LLM, streaming TTS, and a vocoder.

  13. StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding

    cs.SD 2025-06 conditional novelty 5.0 of 10

    StreamFlow achieves streaming speech token decoding with 180 ms first-packet latency by using hierarchical block-wise attention masks in a DiT flow matching model.

  14. Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Step-Audio-AQAA, a 130B end-to-end audio language model using dual-codebook tokens, text-audio interleaving, masked DPO and weight merging, is claimed to outperform Kimi-Audio and Qwen-Omni on the authors' StepEval-Au...

  15. Speechless: Speech Instruction Training Without Speech for Low Resource Languages

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Fine-tuning an LLM on text instructions converted to Whisper semantic tokens enables it to understand spoken instructions at inference, bypassing TTS and speech instruction data.

  16. Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    Under matched settings, continuous SSL speech features generally outperform discrete tokens on six spoken language understanding tasks in SpeechLLMs.

  17. Breaking the Barriers of Text-Hungry and Audio-Deficient AI

    cs.SD 2025-06 reject novelty 4.0 of 10

    A proposed audio-native translation framework called MAST with fractional diffusion is described, but no evidence is given that it produces working translations.

Pith tools