REVIEW 9 cited by
A Full-duplex Speech Dialogue Scheme Based On Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present a generative dialogue system capable of operating in a full-duplex manner, allowing for seamless interaction. It is based on a large language model (LLM) carefully aligned to be aware of a perception module, a motor function module, and the concept of a simple finite state machine (called neural FSM) with two states. The perception and motor function modules operate in tandem, allowing the system to speak and listen to the user simultaneously. The LLM generates textual tokens for inquiry responses and makes autonomous decisions to start responding to, wait for, or interrupt the user by emitting control tokens to the neural FSM. All these tasks of the LLM are carried out as next token prediction on a serialized view of the dialogue in real-time. In automatic quality evaluations simulating real-life interaction, the proposed system reduces the average conversation response latency by more than threefold compared with LLM-based half-duplex dialogue systems while responding within less than 500 milliseconds in more than 50% of evaluated interactions. Running an LLM with only 8 billion parameters, our system exhibits an 8% higher interruption precision rate than the best available commercial LLM for voice-based dialogue.
Forward citations
Cited by 9 Pith papers
-
PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue
Playback-aligned context repair lifts referent anchoring after user interruptions from 25.0% to 96.3% on a new 108-case full-duplex voice benchmark.
-
Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals
A multi-modal (text, audio, video) model trained on a newly collected 210-hour conversation dataset predicts turn-taking and backchannel actions with F1 about 0.81 and 0.91.
-
RoboEgo System Card: An Omnimodal Model with Native Full Duplexity
RoboEgo combines a full-duplex audio/text backbone with vision and action heads to achieve an 80 ms theoretical response granularity, reporting competitive quality and better responsiveness than Qwen2.5-Omni in a smal...
-
LUCY: Linguistic Understanding and Control Yielding Early Stage of Her
LUCY uses curated synthetic data to train a Mini-Omni-style speech model that responds to both what you say and how you say it, makes natural short replies, and calls external tools.
-
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
A single aligned multimodal model, MinMo, achieves strong or state-of-the-art results on speech recognition, translation, emotion recognition, voice generation, and full-duplex dialogue while preserving the underlying...
-
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
The authors integrate streaming perception, compressed long-term memory, and a reasoning model into one open-source system, reporting SOTA open-source results on several video and audio benchmarks.
-
Toward Embodied AGI: A Review of Embodied AI and the Road Ahead
Embodied AI today sits between Level 1 and Level 2 on a new five-level roadmap toward all-purpose humanlike robots.
-
Real-Time Textless Dialogue Generation
A streaming, textless dialogue model predicts turn-taking actions every 160 ms and generates speech units, improving naturalness over cascaded systems at the cost of lower semantic coherence.
-
WavChat: A Survey of Spoken Dialogue Models
WavChat categorizes spoken dialogue models into cascaded and end-to-end paradigms and surveys speech representations, training strategies, streaming, duplex interaction, datasets, and evaluation benchmarks.
Discussion (0). Continue with ORCID to comment.