Pith. sign in

REVIEW 24 cited by

VibeVoice Technical Report

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.19205 v1 pith:PBEMVMG4 submitted 2025-08-26 cs.CL cs.AIcs.SDeess.AS

VibeVoice Technical Report

classification cs.CL cs.AIcs.SDeess.AS
keywords speechvibevoicecontinuousdatadiffusionlong-formmodelnovel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively generating latent vectors via diffusion. To enable this, we introduce a novel continuous speech tokenizer that, when compared to the popular Encodec model, improves data compression by 80 times while maintaining comparable performance. The tokenizer effectively preserves audio fidelity while significantly boosting computational efficiency for processing long sequences. Thus, VibeVoice can synthesize long-form speech for up to 90 minutes (in a 64K context window length) with a maximum of 4 speakers, capturing the authentic conversational ``vibe'' and surpassing open-source and proprietary dialogue models.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

    eess.AS 2026-06 unverdicted novelty 7.0

    Mel-LLM shows an LLM can achieve competitive ASR by directly ingesting pre-processed Mel spectrogram patches through a linear projection layer.

  2. HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

    eess.AS 2026-06 unverdicted novelty 7.0

    HoliDubber introduces a patch-based autoregressive diffusion transformer for joint text-guided synthesis of speech and ambient audio in video dubbing, with a new benchmark showing outperformance over prior speech-only...

  3. ATIR: Towards Audio-Text Interleaved Contextual Retrieval

    cs.SD 2026-04 unverdicted novelty 7.0

    Defines ATIR task and benchmark for mixed audio-text queries; MLLM model with token compression shows substantial gains over strong baselines.

  4. MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech

    eess.AS 2026-04 unverdicted novelty 7.0

    MINT-Bench is a new benchmark using hierarchical taxonomy, multi-stage data pipeline, and hybrid evaluation to assess instruction-following TTS systems, revealing major gaps in compositional and paralinguistic controls.

  5. ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

    cs.CV 2025-12 unverdicted novelty 7.0

    ViBES introduces a speech-language-behavior model using modality-specific transformer experts that jointly generates dialogue and 3D body actions, showing gains over separate co-speech and text-to-motion baselines on ...

  6. ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

    cs.SD 2026-07 conditional novelty 6.5

    Hierarchical multi-prompt representation generation plus generalized flow matching yields high-quality single-stage waveform diffusion from 12.5 Hz latents and efficient LDM TTS.

  7. Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

    eess.AS 2026-07 conditional novelty 6.0

    Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...

  8. ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching

    eess.AS 2026-07 conditional novelty 6.0

    A latent flow-matching dialog TTS generates speech in a 4x-compressed 25 Hz space, cutting peak memory about 11x and inference time about 2.2x while keeping UTMOS competitive.

  9. LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

    eess.AS 2026-06 unverdicted novelty 6.0

    Mel-LLM shows that LLMs can process Mel spectrograms directly for competitive ASR performance without a dedicated speech encoder, with limited degradation versus encoder-based versions when using multimodal initializa...

  10. dots.tts Technical Report

    cs.SD 2026-06 unverdicted novelty 6.0

    dots.tts reports SOTA benchmark results on Seed-TTS-Eval and other tests via continuous latent-space autoregressive modeling with three listed innovations and code release.

  11. CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

    cs.SD 2026-06 unverdicted novelty 6.0

    CleanCodec reframes audio tokenization as a selective information bottleneck to encode only perceptually important features at 12.5 tokens per second, outperforming prior codecs in efficiency, speaker similarity, and ...

  12. SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

    eess.AS 2026-05 unverdicted novelty 6.0

    SwanVoice is a zero-shot TTS system for 1-4 speakers that reports higher richness and hierarchy scores than open-source baselines on monologue and dialogue tasks via mixed training and DiffusionNFT post-training.

  13. StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration

    cs.CV 2026-05 unverdicted novelty 6.0

    StreamChar decouples LLM-based orchestration from DiT denoising to achieve real-time long-horizon streaming character audio-video generation with reduced drift and misalignment.

  14. Qwen3-TTS Technical Report

    cs.SD 2026-01 unverdicted novelty 6.0

    Qwen3-TTS delivers state-of-the-art multilingual TTS performance with 3-second voice cloning, description control, and ultra-low-latency streaming via dual tokenizers and a dual-track LM architecture trained on over 5...

  15. Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

    eess.AS 2026-07 conditional novelty 5.0

    A lightweight face adapter plus soft-tuning aligns face embeddings to a frozen StyleTTS 2 style space, yielding natural zero-shot face-to-speech and language-agnostic transfer to Spanish.

  16. ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching

    eess.AS 2026-07 unverdicted novelty 5.0

    ZipL-Dialog cuts peak GPU memory 11.22× and speeds inference 2.23× for multi-minute zero-shot dialog TTS by doing conditional flow matching in a 4× compressed latent space while keeping perceptual naturalness.

  17. Qwen-Audio-VAE Technical Report

    eess.AS 2026-07 conditional novelty 5.0

    A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.

  18. MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data

    cs.SD 2026-06 unverdicted novelty 5.0

    MagpieTTS-LF enables coherent long-form TTS via three inference-time innovations without any retraining on long-form data.

  19. VoxCPM2 Technical Report

    cs.SD 2026-06 unverdicted novelty 5.0

    VoxCPM2 scales hierarchical continuous-latent speech modeling to 2B parameters and over 2M hours of multilingual data, unifying voice cloning, style control, and continuation in one backbone with open release.

  20. F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

    cs.SD 2026-06 unverdicted novelty 5.0

    F3-Tokenizer adapts audio autoencoder latents with noise-regularized bottleneck (channel normalization and stochastic perturbation) and a representation encoder (RQ-MTP plus frozen-LLM supervision) to support both hig...

  21. On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation

    cs.SD 2026-04 unverdicted novelty 5.0

    Joint-marginal alignment plus adaptive weighting in speech VAE distillation yields the best combined performance on reconstruction, understanding, and generation tasks.

  22. UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training

    cs.CL 2026-05 unverdicted novelty 4.0

    UNIQUE enables efficient top-k sparse attention in LLMs by using a mean-plus-std page importance score and a soft-mask training approach, achieving up to 11.4x kernel speedup while preserving performance.

  23. PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

    cs.SD 2026-05 unverdicted novelty 4.0

    PilotTTS achieves lowest WER 1.50% (en) and CER 0.87% (zh) plus highest speaker similarity on Seed-TTS Eval using a Q-Former conditioned autoregressive architecture and a released multi-stage open data pipeline.

  24. On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation

    cs.SD 2026-04 unverdicted novelty 4.0

    Joint-marginal distillation with adaptive weighting is reported as the best overall alignment loss for speech VAEs across reconstruction, understanding, and generation.