Pith. sign in

REVIEW 27 cited by

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.00774 v5 pith:7XOWG4QH submitted 2024-11-01 cs.SD cs.AIcs.CLeess.AS

classification cs.SDcs.AIcs.CLeess.AS
keywords freeze-omniabilitydatadialoguespeechspeech-to-speechtrainingfrozen
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought impressive experience to users. Researchers have recently proposed several multi-modal LLMs in this direction that can achieve user-agent speech-to-speech conversations. This paper proposes a novel speech-text multimodal LLM architecture called Freeze-Omni. Our main contribution is that the speech input and output modalities can be easily connected to a textual LLM while keeping the LLM's parameters frozen throughout the training process. We design a three-stage training strategy for modeling both the speech input and output, enabling Freeze-Omni to obtain speech-to-speech conversation ability using text-speech paired data (such as ASR and TTS data) and only 60,000 multi-round text Q&A data on 8 GPUs. Moreover, we can effectively ensure that the intelligence of the Freeze-Omni in the speech modality is at the same level compared with that in the text modality of its backbone LLM, while achieving low latency end-to-end spoken response. In addition, we also designed a method to achieve duplex dialogue ability through multi-task training, giving Freeze-Omni a more natural style of dialogue ability between users and agents. In summary, Freeze-Omni holds great potential to conduct speech-to-speech dialogue based on a multimodal LLM under the condition of a frozen LLM, avoiding the catastrophic forgetting problem caused by limited data and training resources.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?

    cs.CL 2026-05 conditional novelty 7.0 of 10

    Current full-duplex speech systems largely fail to follow explicit turn-taking instructions: the best of six models scores 64.4% adherence on the new Instruct-FD benchmark, and proactive behaviors like interruption an...

  2. The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

    eess.AS 2026-03 unverdicted novelty 7.0 of 10

    FLAIR enables spoken dialogue AI to conduct continuous latent reasoning while perceiving speech through recursive latent embeddings and an ELBO-based finetuning objective.

  3. ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

    cs.CL 2025-07 conditional novelty 7.0 of 10

    A speech language model trained only on audiobooks with explicit word-level prosody tokens displays emerging abilities in prosody-controlled generation, emphasis and emotion understanding, and prosodic consistency acr...

  4. EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A new four-branch benchmark, EmoDialogue dataset, and GRPO-trained evaluator (EmoS) claim near-human emotional intelligence scoring for spoken language models.

  5. PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

    cs.SD 2026-08 conditional novelty 6.0 of 10

    Playback-aligned context repair lifts referent anchoring after user interruptions from 25.0% to 96.3% on a new 108-case full-duplex voice benchmark.

  6. M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    M3-DuplexBench, a bilingual English/Japanese benchmark over chat and multi-turn QA with teacher-forced full-context evaluation, reveals clear model-, language-, and domain-dependent gaps in full-duplex spoken dialogue...

  7. AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues

    cs.MM 2025-10 conditional novelty 6.0 of 10

    Current omni-modal LLMs underperform on audio-visual emotional reasoning, and automatic scores diverge from human perceptual judgments; AV-EMO-Reasoning provides a benchmark to measure this.

  8. FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems

    eess.AS 2025-07 conditional novelty 6.0 of 10

    FD-Bench is a new LLM/TTS/ASR-based benchmark for full-duplex spoken dialogue systems, and applying it to Moshi, Freeze-omni, and VITA-1.5 shows all three struggle with frequent interruptions and noise.

  9. Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction

    eess.AS 2025-07 conditional novelty 6.0 of 10

    A causal audio language model with continuous-valued tokens and masked next-token prediction matches diffusion-based text-to-audio quality with smaller, streamable models.

  10. SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant

    cs.SD 2025-06 conditional novelty 6.0 of 10

    SOVA-Bench is a new evaluation framework for speech LLMs covering knowledge, recognition, linguistic and paralinguistic understanding, and semantic and acoustic generation quality.

  11. SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A speech-to-speech language model uses channel fusion of a streaming encoder and codec tokens to handle barge-in and turn-taking without speech pretraining, showing improved metrics over Moshi at 0.6 kbps.

  12. SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SALMONN-omni is a standalone full-duplex speech LLM that interleaves continuous speech and text embeddings in one model and uses special thinking tokens to learn turn-taking and barge-in behavior.

  13. Contrastive Learning for Task-Independent SpeechLLM-Pretraining

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Contrastive pre-training that aligns speech and text across all model layers beats ASR-based pre-training and, with 10% of task data, matches or exceeds specialized models on translation and question answering.

  14. Raon-Speech Technical Report

    cs.CL 2026-04 conditional novelty 5.5 of 10

    A 9B SpeechLM trained on 1.38M hours of English/Korean data plus a full-duplex chat extension trained on 119K hours of time-aligned dialogue outperform same-size audio models on speech tasks and FDB turn-taking metrics.

  15. Sharp spectral estimates for free boundary problems arising in plasma physics

    math.AP 2026-04 unverdicted novelty 5.0 of 10

    For a constrained superlinear free-boundary plasma model, the non-local first eigenvalue σ₁ is always positive on balls in every dimension N≥2, despite lacking a general Faber–Krahn property.

  16. FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

    cs.SD 2025-09 conditional novelty 5.0 of 10

    A full-duplex voice system with streaming personalized VAD and semantic end-of-turn detection reports fewer false barge-ins and latencies near commercial benchmarks.

  17. Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

    cs.AI 2025-06 conditional novelty 5.0 of 10

    Stream-Omni uses CTC-based layer-dimension mapping to align speech with text, achieving vision, speech, and text interaction in one 8B model trained on 23,000 hours of speech.

  18. TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment

    cs.CL 2025-06 conditional novelty 5.0 of 10

    TESU-LLM shows that a frozen LLM can answer spoken queries after training only a 13M-parameter projector on text, using SeamlessM4T's shared speech-text encoder.

  19. LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A series of Qwen2.5-based speech chatbots (0.5B to 14B) that stream speech output with a CosyVoice2-style autoregressive decoder and outperform prior SpeechLMs using only 200K synthetic multi-turn dialogues.

  20. MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A single aligned multimodal model, MinMo, achieves strong or state-of-the-art results on speech recognition, translation, emotion recognition, voice generation, and full-duplex dialogue while preserving the underlying...

  21. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

    eess.AS 2024-12 conditional novelty 5.0 of 10

    A 0.5B spoken dialogue model trained end-to-end in one stage, with grouped semantic tokens for faster generation and text-only history for multi-turn dialogue.

  22. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

    cs.CV 2024-12 conditional novelty 5.0 of 10

    The authors integrate streaming perception, compressed long-term memory, and a reasoning model into one open-source system, reporting SOTA open-source results on several video and audio benchmarks.

  23. BoSS: Beyond-Semantic Speech

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Current spoken-language models perform poorly on a new five-task evaluation of beyond-semantic speech signals, including dialect, emotion, age, and non-verbal cues.

  24. Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners

    cs.SD 2024-12 conditional novelty 4.0 of 10

    A speech-to-speech dialogue model that predicts mel-spectrograms with flow matching, jointly trained with discrete text tokens, achieves lower WER than a discrete speech token baseline.

  25. AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM

    eess.AS 2024-12 conditional novelty 4.0 of 10

    AlignFormer, a CTC-guided dynamic-window adapter, lets a frozen LLM trained on ASR data alone achieve high instruction-following rates on zero-shot speech translation and question answering.

  26. WavChat: A Survey of Spoken Dialogue Models

    eess.AS 2024-11 conditional novelty 4.0 of 10

    WavChat categorizes spoken dialogue models into cascaded and end-to-end paradigms and surveys speech representations, training strategies, streaming, duplex interaction, datasets, and evaluation benchmarks.

  27. JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

    cs.SD 2026-08 conditional novelty 3.0 of 10

    JoyAI-Talker claims a modular Thinker-Talker-plus-Duplex speech LLM that preserves text reasoning while adding empathetic, expressive, full-duplex interaction.

Pith tools