Pith. sign in

REVIEW 6 cited by

LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02897 v2 pith:NDK3VL7D submitted 2024-06-05 cs.SD eess.AS

classification cs.SDeess.AS
keywords audiolow-latencytext-to-speechzero-shotautoregressivecodebooklanguagelivespeech
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Prior works have demonstrated zero-shot text-to-speech by using a generative language model on audio tokens obtained via a neural audio codec. It is still challenging, however, to adapt them to low-latency scenarios. In this paper, we present LiveSpeech - a fully autoregressive language model-based approach for zero-shot text-to-speech, enabling low-latency streaming of the output audio. To allow multiple token prediction within a single decoding step, we propose (1) using adaptive codebook loss weights that consider codebook contribution in each frame and focus on hard instances, and (2) grouping codebooks and processing groups in parallel. Experiments show our proposed models achieve competitive results to state-of-the-art baselines in terms of content accuracy, speaker similarity, audio quality, and inference speed while being suitable for low-latency streaming applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.

  2. Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

    cs.LG 2025-05 conditional novelty 6.0 of 10

    SMLLE generates speech frame-by-frame using a Transducer for streaming semantic tokens plus a fully autoregressive mel-spectrogram model, reaching quality close to sentence-level zero-shot TTS.

  3. SpeakStream: Streaming Text-to-Speech with Interleaved Data

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A decoder-only TTS trained on force-aligned interleaved text-speech chunks generates audio after a few words, achieving ~30ms TTS latency and WER comparable to non-streaming.

  4. MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model

    eess.AS 2025-01 conditional novelty 6.0 of 10

    A small 70M-parameter codec-based TTS model, combining hierarchical decoding and several stabilization tricks, matches or beats far larger models on expressive reference cloning, especially in speaker similarity.

  5. Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

    eess.AS 2024-12 conditional novelty 6.0 of 10

    Training a decoder-only LM on fixed-ratio interleaved text and speech tokens yields a simple zero-shot streaming TTS, with a 1:3 text-to-speech chunk ratio keeping WER within about 8 percent relative of non-streaming.

  6. BanglaFake: Constructing and Evaluating a Specialized Bengali Deepfake Audio Dataset

    cs.SD 2025-05 reject novelty 5.0 of 10

    A new Bengali deepfake audio dataset is introduced, but its size figures contradict each other and the paper provides no detector evaluation to substantiate the benchmark claim.

Pith tools