Pith. sign in

REVIEW 6 cited by

ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.07333 v1 pith:UHKJ2V3C submitted 2024-01-14 cs.CL cs.AIcs.SDeess.AS

classification cs.CLcs.AIcs.SDeess.AS
keywords audioella-vphonemetokensacousticlanguagemodelsynthesized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation. However, existing methods still have some limitations: 1) repetitions, transpositions, and omissions in the output synthesized speech due to limited alignment constraints between audio and phoneme tokens; 2) challenges of fine-grained control over the synthesized speech with autoregressive (AR) language model; 3) infinite silence generation due to the nature of AR-based decoding, especially under the greedy strategy. To alleviate these issues, we propose ELLA-V, a simple but efficient LM-based zero-shot text-to-speech (TTS) framework, which enables fine-grained control over synthesized audio at the phoneme level. The key to ELLA-V is interleaving sequences of acoustic and phoneme tokens, where phoneme tokens appear ahead of the corresponding acoustic tokens. The experimental findings reveal that our model outperforms VALL-E in terms of accuracy and delivers more stable results using both greedy and sampling-based decoding strategies. The code of ELLA-V will be open-sourced after cleanups. Audio samples are available at https://ereboas.github.io/ELLAV/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A compact zero-shot TTS that applies discrete flow matching with separate prediction heads for prosody and acoustic tokens, reporting near-best quality, best prosody/energy metrics, and up to 25.8x faster inference.

  2. Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Compressed-to-fine language modeling improves speech token prediction by retaining prompt and local tokens while compressing long-range token spans into compact summaries.

  3. Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

    cs.LG 2025-05 conditional novelty 6.0 of 10

    SMLLE generates speech frame-by-frame using a Transducer for streaming semantic tokens plus a fully autoregressive mel-spectrogram model, reaching quality close to sentence-level zero-shot TTS.

  4. CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech

    cs.SD 2025-01 conditional novelty 6.0 of 10

    A new large-scale dataset and codec taxonomy show that codec re-synthesized speech, especially balanced by decoder type, trains detectors that catch codec-based deepfake speech better than traditional anti-spoofing training.

  5. MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model

    eess.AS 2025-01 conditional novelty 6.0 of 10

    A small 70M-parameter codec-based TTS model, combining hierarchical decoding and several stabilization tricks, matches or beats far larger models on expressive reference cloning, especially in speaker similarity.

  6. Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance

    cs.SD 2025-02 conditional novelty 5.0 of 10

    Koel-TTS combines ASR/SV-based preference alignment (DPO/RPO) with classifier-free guidance to improve zero-shot TTS intelligibility, speaker similarity, and naturalness.

Pith tools