Pith. sign in

REVIEW 9 cited by

A Full-duplex Speech Dialogue Scheme Based On Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19487 v2 pith:KKEHSO7C submitted 2024-05-29 cs.CL

classification cs.CL
keywords dialoguesystemallowingfull-duplexfunctioninteractionlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present a generative dialogue system capable of operating in a full-duplex manner, allowing for seamless interaction. It is based on a large language model (LLM) carefully aligned to be aware of a perception module, a motor function module, and the concept of a simple finite state machine (called neural FSM) with two states. The perception and motor function modules operate in tandem, allowing the system to speak and listen to the user simultaneously. The LLM generates textual tokens for inquiry responses and makes autonomous decisions to start responding to, wait for, or interrupt the user by emitting control tokens to the neural FSM. All these tasks of the LLM are carried out as next token prediction on a serialized view of the dialogue in real-time. In automatic quality evaluations simulating real-life interaction, the proposed system reduces the average conversation response latency by more than threefold compared with LLM-based half-duplex dialogue systems while responding within less than 500 milliseconds in more than 50% of evaluated interactions. Running an LLM with only 8 billion parameters, our system exhibits an 8% higher interruption precision rate than the best available commercial LLM for voice-based dialogue.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

    cs.SD 2026-08 conditional novelty 6.0 of 10

    Playback-aligned context repair lifts referent anchoring after user interruptions from 25.0% to 96.3% on a new 108-case full-duplex voice benchmark.

  2. Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A multi-modal (text, audio, video) model trained on a newly collected 210-hour conversation dataset predicts turn-taking and backchannel actions with F1 about 0.81 and 0.91.

  3. RoboEgo System Card: An Omnimodal Model with Native Full Duplexity

    cs.AI 2025-06 conditional novelty 5.0 of 10

    RoboEgo combines a full-duplex audio/text backbone with vision and action heads to achieve an 80 ms theoretical response granularity, reporting competitive quality and better responsiveness than Qwen2.5-Omni in a smal...

  4. LUCY: Linguistic Understanding and Control Yielding Early Stage of Her

    cs.CL 2025-01 conditional novelty 5.0 of 10

    LUCY uses curated synthetic data to train a Mini-Omni-style speech model that responds to both what you say and how you say it, makes natural short replies, and calls external tools.

  5. MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A single aligned multimodal model, MinMo, achieves strong or state-of-the-art results on speech recognition, translation, emotion recognition, voice generation, and full-duplex dialogue while preserving the underlying...

  6. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

    cs.CV 2024-12 conditional novelty 5.0 of 10

    The authors integrate streaming perception, compressed long-term memory, and a reasoning model into one open-source system, reporting SOTA open-source results on several video and audio benchmarks.

  7. Toward Embodied AGI: A Review of Embodied AI and the Road Ahead

    cs.AI 2025-05 accept novelty 4.0 of 10

    Embodied AI today sits between Level 1 and Level 2 on a new five-level roadmap toward all-purpose humanlike robots.

  8. Real-Time Textless Dialogue Generation

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A streaming, textless dialogue model predicts turn-taking actions every 160 ms and generates speech units, improving naturalness over cascaded systems at the cost of lower semantic coherence.

  9. WavChat: A Survey of Spoken Dialogue Models

    eess.AS 2024-11 conditional novelty 4.0 of 10

    WavChat categorizes spoken dialogue models into cascaded and end-to-end paradigms and surveys speech representations, training strategies, streaming, duplex interaction, datasets, and evaluation benchmarks.

Pith tools