REVIEW 18 cited by
Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a cross-lingual neural codec language model, VALL-E X, for cross-lingual speech synthesis. Specifically, we extend VALL-E and train a multi-lingual conditional codec language model to predict the acoustic token sequences of the target language speech by using both the source language speech and the target language text as prompts. VALL-E X inherits strong in-context learning capabilities and can be applied for zero-shot cross-lingual text-to-speech synthesis and zero-shot speech-to-speech translation tasks. Experimental results show that it can generate high-quality speech in the target language via just one speech utterance in the source language as a prompt while preserving the unseen speaker's voice, emotion, and acoustic environment. Moreover, VALL-E X effectively alleviates the foreign accent problems, which can be controlled by a language ID. Audio samples are available at \url{https://aka.ms/vallex}.
Forward citations
Cited by 18 Pith papers
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.
-
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching
A compact zero-shot TTS that applies discrete flow matching with separate prediction heads for prosody and acoustic tokens, reporting near-best quality, best prosody/energy metrics, and up to 25.8x faster inference.
-
XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation
XEmoRAG synthesizes Thai speech with emotions cloned from Chinese reference audio by retrieving matching Thai prompts and aligning prosody with flow matching, outperforming baseline TTS in emotion similarity and intel...
-
Next Tokens Denoising for Speech Synthesis
Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.
-
Differentiable Reward Optimization for LLM based TTS system
DiffRO optimizes codec-based TTS models directly on differentiable token-level rewards, improving WER and enabling zero-shot emotion control.
-
OpusLM: A Family of Open Unified Speech Language Models
A new open family of speech-language models trained on public data matches or beats prior systems across ASR, TTS, and text benchmarks.
-
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
SMLLE generates speech frame-by-frame using a Transducer for streaming semantic tokens plus a fully autoregressive mel-spectrogram model, reaching quality close to sentence-level zero-shot TTS.
-
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.
-
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
A unified multi-trait speech benchmark, Vox-Profile, evaluates foundation models on age, sex, accent, emotion, voice quality, fluency, and expressiveness, and demonstrates downstream uses in ASR analysis and speech ge...
-
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
CLEAR is a zero-shot TTS model that autoregressively predicts compact continuous audio latents with a per-token rectified flow head, reaching 1.88% WER on LibriSpeech Subset-B with an RTF of 0.29 and a 96 ms streaming delay.
-
SecureSpeech: Prompt-based Speaker and Content Protection
A dual anonymization pipeline that uses an LLM to replace sensitive entities and a prompt-driven TTS to generate speech with a new, unrelated voice.
-
Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages
Zero-shot TTS for Sanskrit, two Konkani dialects, Maithili, and Kurukh is achieved by matching shared phone labels and parsing rules to each language's phonotactics.
-
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
DS-TTS adds a second MFCC-based style encoder and a length-adaptive variance adapter to a StyleSpeech-style TTS model, reporting higher speaker similarity but not lower WER than two strong baselines.
-
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
ADD-GP, a Gaussian Process classifier with XLS-R speech embeddings, adapts to unseen TTS models with as few as 5 samples and achieves state-of-the-art low error rates on the new LibriFake benchmark.
-
Voice Adaptation for Swiss German
Fine-tuning XTTS-v2 on about 5,000 hours of weakly labeled Swiss podcast audio produces a voice adaptation model that renders Standard German text in seven Swiss German dialect regions with near-reference quality in h...
-
Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection
A component-wise ensemble of four fine-tuned anti-spoofing models with margin fusion and calibration achieves 0.7828 macro-F1 on the ESDD2 test set, ranking 5th of 31.
-
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
A speculative decoding variant with a tolerance-based acceptance rule speeds up CosyVoice 2 inference by 1.4x while keeping subjective quality comparable.
Discussion (0). Sign in to comment.