LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

Dong Yu; Hao Zhang; Meng Yu; Rilin Chen; Vinay Kothapally; Weiwei Li

arxiv: 2502.14145 · v3 · pith:OTWV6PFHnew · submitted 2025-02-19 · 💻 cs.CL · eess.AS

LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

Hao Zhang , Weiwei Li , Rilin Chen , Vinay Kothapally , Meng Yu , Dong Yu This is my paper

classification 💻 cs.CL eess.AS

keywords dialoguefull-duplexsemanticreal-timespokensystemswhileaccuracy

0 comments

read the original abstract

Achieving full-duplex communication in spoken dialogue systems (SDS) requires real-time coordination between listening, speaking, and thinking. This paper proposes a semantic voice activity detection (VAD) module as a dialogue manager (DM) to efficiently manage turn-taking in full-duplex SDS. Implemented as a lightweight (0.5B) LLM fine-tuned on full-duplex conversation data, the semantic VAD predicts four control tokens to regulate turn-switching and turn-keeping, distinguishing between intentional and unintentional barge-ins while detecting query completion for handling user pauses and hesitations. By processing input speech in short intervals, the semantic VAD enables real-time decision-making, while the core dialogue engine (CDE) is only activated for response generation, reducing computational overhead. This design allows independent DM optimization without retraining the CDE, balancing interaction accuracy and inference efficiency for scalable, next-generation full-duplex SDS.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning
eess.AS 2026-03 unverdicted novelty 7.0

FLAIR enables spoken dialogue AI to conduct continuous latent reasoning while perceiving speech through recursive latent embeddings and an ELBO-based finetuning objective.
The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning
eess.AS 2026-03 unverdicted novelty 4.0

FLAIR enables simultaneous latent reasoning during speech input in full-duplex dialogue models via recursive latent embeddings and an ELBO-based training objective without added latency.