Pith. sign in

REVIEW 7 cited by

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.17799 v2 pith:Z6CMF4JK submitted 2024-10-23 cs.CL cs.AIcs.SDeess.AS

classification cs.CLcs.AIcs.SDeess.AS
keywords dialoguefull-duplexconversationomniflattensystemsbackboneend-to-endmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and natural interactions in full-duplex dialogue systems remains a significant challenge, especially considering human conversation dynamics such as interruptions, backchannels, and overlapping speech. In this paper, we introduce a novel End-to-End GPT-based model OmniFlatten for full-duplex conversation, capable of effectively modeling the complex behaviors inherent to natural conversations with low latency. To achieve full-duplex conversation capabilities, we propose a multi-stage post-training scheme that progressively adapts a text large language model (LLM) backbone into a speech-text dialogue LLM, capable of generating text and speech in real time, without modifying the architecture of the backbone LLM. The training process comprises three stages: modality alignment, half-duplex dialogue learning, and full-duplex dialogue learning. In all training stages, we standardize the data using a flattening operation, which enables unifying the training methods and the GPT backbone across different modalities and tasks. Our approach offers a simple modeling technique and a promising research direction for developing efficient and natural end-to-end full-duplex spoken dialogue systems. Audio samples of dialogues generated by OmniFlatten can be found at this web site (https://omniflatten.github.io/).

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant

    cs.SD 2025-06 conditional novelty 6.0 of 10

    SOVA-Bench is a new evaluation framework for speech LLMs covering knowledge, recognition, linguistic and paralinguistic understanding, and semantic and acoustic generation quality.

  2. MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A joint CTC and token-and-duration transducer keyword spotter with frame-skipping and CDC-Last score fusion reports better recall on Snips, MobvoiHotwords, and LibriKWS-20 while decoding 1.47 to 1.63 times faster.

  3. SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A speech-to-speech language model uses channel fusion of a streaming encoder and codec tokens to handle barge-in and turn-taking without speech pretraining, showing improved metrics over Moshi at 0.6 kbps.

  4. Towards a Japanese Full-duplex Spoken Dialogue System

    cs.CL 2025-06 conditional novelty 5.0 of 10

    J-Moshi, the first public Japanese full-duplex spoken dialogue model, is built from Moshi and outperforms a Japanese dGSLM baseline on naturalness and meaningfulness.

  5. RoboEgo System Card: An Omnimodal Model with Native Full Duplexity

    cs.AI 2025-06 conditional novelty 5.0 of 10

    RoboEgo combines a full-duplex audio/text backbone with vision and action heads to achieve an 80 ms theoretical response granularity, reporting competitive quality and better responsiveness than Qwen2.5-Omni in a smal...

  6. Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Staged ASR-to-text-response-to-TTS training makes open end-to-end spoken dialogue systems trainable on 300 hours of public human-human data and more coherent than one-step speech-to-speech models.

  7. OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction

    cs.CL 2025-05 conditional novelty 5.0 of 10

    OmniCharacter is a speech-language role-playing agent that generates character-specific voice responses with low latency, trained on a new 10K-dialogue, 135K-audio dataset of 20 game characters.

Pith tools