Pith. sign in

REVIEW 18 cited by

Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.03926 v1 pith:U7OBJS3L submitted 2023-03-07 cs.CL cs.AIcs.SDeess.AS

classification cs.CLcs.AIcs.SDeess.AS
keywords languagespeechcross-lingualvall-ecodectargetacousticforeign
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a cross-lingual neural codec language model, VALL-E X, for cross-lingual speech synthesis. Specifically, we extend VALL-E and train a multi-lingual conditional codec language model to predict the acoustic token sequences of the target language speech by using both the source language speech and the target language text as prompts. VALL-E X inherits strong in-context learning capabilities and can be applied for zero-shot cross-lingual text-to-speech synthesis and zero-shot speech-to-speech translation tasks. Experimental results show that it can generate high-quality speech in the target language via just one speech utterance in the source language as a prompt while preserving the unseen speaker's voice, emotion, and acoustic environment. Moreover, VALL-E X effectively alleviates the foreign accent problems, which can be controlled by a language ID. Audio samples are available at \url{https://aka.ms/vallex}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  2. OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    cs.SD 2026-02 conditional novelty 6.0 of 10

    A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.

  3. DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A compact zero-shot TTS that applies discrete flow matching with separate prediction heads for prosody and acoustic tokens, reporting near-best quality, best prosody/energy metrics, and up to 25.8x faster inference.

  4. XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation

    eess.AS 2025-08 conditional novelty 6.0 of 10

    XEmoRAG synthesizes Thai speech with emotions cloned from Chinese reference audio by retrieving matching Thai prompts and aligning prosody with flow matching, outperforming baseline TTS in emotion similarity and intel...

  5. Next Tokens Denoising for Speech Synthesis

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.

  6. Differentiable Reward Optimization for LLM based TTS system

    cs.SD 2025-07 conditional novelty 6.0 of 10

    DiffRO optimizes codec-based TTS models directly on differentiable token-level rewards, improving WER and enabling zero-shot emotion control.

  7. OpusLM: A Family of Open Unified Speech Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new open family of speech-language models trained on public data matches or beats prior systems across ASR, TTS, and text benchmarks.

  8. Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

    cs.LG 2025-05 conditional novelty 6.0 of 10

    SMLLE generates speech frame-by-frame using a Transducer for streaming semantic tokens plus a fully autoregressive mel-spectrogram model, reaching quality close to sentence-level zero-shot TTS.

  9. VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.

  10. Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A unified multi-trait speech benchmark, Vox-Profile, evaluates foundation models on age, sex, accent, emotion, voice quality, fluency, and expressiveness, and demonstrates downstream uses in ASR analysis and speech ge...

  11. CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

    eess.AS 2025-08 conditional novelty 5.0 of 10

    CLEAR is a zero-shot TTS model that autoregressively predicts compact continuous audio latents with a per-token rectified flow head, reaching 1.88% WER on LibriSpeech Subset-B with an RTF of 0.29 and a 96 ms streaming delay.

  12. SecureSpeech: Prompt-based Speaker and Content Protection

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A dual anonymization pipeline that uses an LLM to replace sensitive entities and a prompt-driven TTS to generate speech with a new, unrelated voice.

  13. Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Zero-shot TTS for Sanskrit, two Konkani dialects, Maithili, and Kurukh is achieved by matching shared phone labels and parsing rules to each language's phonotactics.

  14. DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation

    cs.SD 2025-06 reject novelty 5.0 of 10

    DS-TTS adds a second MFCC-based style encoder and a length-adaptive variance adapter to a StyleSpeech-style TTS model, reporting higher speaker similarity but not lower WER than two strong baselines.

  15. Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes

    cs.SD 2025-05 conditional novelty 5.0 of 10

    ADD-GP, a Gaussian Process classifier with XLS-R speech embeddings, adapts to unseen TTS models with as few as 5 samples and achieves state-of-the-art low error rates on the new LibriFake benchmark.

  16. Voice Adaptation for Swiss German

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning XTTS-v2 on about 5,000 hours of weakly labeled Swiss podcast audio produces a voice adaptation model that renders Standard German text in seven Swiss German dialect regions with near-reference quality in h...

  17. Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection

    cs.SD 2026-07 conditional novelty 4.0 of 10

    A component-wise ensemble of four fine-tuned anti-spoofing models with margin fusion and calibration achieves 0.7828 macro-F1 on the ESDD2 test set, ranking 5th of 31.

  18. Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A speculative decoding variant with a tolerance-based acceptance rule speeds up CosyVoice 2 inference by 1.4x while keeping subjective quality comparable.

Pith tools