Pith. sign in

Minimax- speech: Intrinsic zero-shot text-to-speech with a learnable speaker encoder

12 Pith papers cite this work. Polarity classification is still indexing.

12 Pith papers citing it

citation-role summary

dataset 1

citation-polarity summary

years

2026 10 2025 2

roles

dataset 1

polarities

use dataset 1

representative citing papers

Qwen3-TTS Technical Report

cs.SD · 2026-01-22 · unverdicted · novelty 6.0

Qwen3-TTS delivers state-of-the-art multilingual TTS performance with 3-second voice cloning, description control, and ultra-low-latency streaming via dual tokenizers and a dual-track LM architecture trained on over 5 million hours of data.

Qwen3-Omni Technical Report

cs.CL · 2025-09-22 · unverdicted · novelty 6.0

Qwen3-Omni is a unified multimodal model that achieves open-source SOTA on 32 of 36 audio and audio-visual benchmarks and overall SOTA on 22 without degrading performance on text, image, or video relative to single-modal Qwen counterparts.

Step-Audio 2 Technical Report

cs.CL · 2025-07-22 · unverdicted · novelty 6.0

Step-Audio 2 integrates a latent audio encoder, reasoning-centric reinforcement learning, and discrete audio token generation into language modeling to deliver state-of-the-art performance on audio understanding and conversational benchmarks.

VoxCPM2 Technical Report

cs.SD · 2026-06-05 · unverdicted · novelty 5.0

VoxCPM2 scales hierarchical continuous-latent speech modeling to 2B parameters and over 2M hours of multilingual data, unifying voice cloning, style control, and continuation in one backbone with open release.

Voxtral TTS

cs.AI · 2026-03-26 · unverdicted · novelty 5.0

Voxtral TTS produces expressive multilingual speech from 3-second reference audio with a hybrid autoregressive-plus-flow-matching architecture and a new VQ-FSQ tokenizer, achieving 68.4% win rate over ElevenLabs in human evaluations.

citing papers explorer

Showing 12 of 12 citing papers.

  • MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control cs.SD · 2026-04-23 · unverdicted · none · ref 4

    MAGIC-TTS is the first TTS system with explicit token-level duration and pause control that improves timing accuracy while preserving natural quality when controls are absent.

  • Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors cs.SD · 2026-06-17 · unverdicted · none · ref 52

    ScenA generates multi-speaker audio scenes by conditioning a flow-matching foundation model on reference voices and natural language prompts, using a high-noise-biased timestep schedule to prevent reference shortcut.

  • EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cs.CL · 2026-06-08 · unverdicted · none · ref 32

    EmoInstruct-TTS uses Emotion2embed and an Instruction-Conditioned Emotion Flow Model (ICE-Flow) to generate acoustically grounded emotion representations from free-form instructions and integrate them into an LLM-based TTS pipeline.

  • RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cs.SD · 2026-05-21 · conditional · none · ref 31

    By training flow-matching TTS to avoid augmented repeat/skip latent trajectories, RobustSpeechFlow cuts Seed-TTS-eval WER from 1.44 to 1.38 and improves CER on a new multilingual benchmark.

  • OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cs.CL · 2026-04-01 · unverdicted · none · ref 51

    OmniVoice introduces a diffusion language model-style non-autoregressive TTS system that directly maps text to multi-codebook acoustic tokens, scaling zero-shot synthesis to over 600 languages with SOTA results on multilingual benchmarks using 581k hours of open data.

  • Qwen3-TTS Technical Report cs.SD · 2026-01-22 · unverdicted · none · ref 25

    Qwen3-TTS delivers state-of-the-art multilingual TTS performance with 3-second voice cloning, description control, and ultra-low-latency streaming via dual tokenizers and a dual-track LM architecture trained on over 5 million hours of data.

  • Qwen3-Omni Technical Report cs.CL · 2025-09-22 · unverdicted · none · ref 33

    Qwen3-Omni is a unified multimodal model that achieves open-source SOTA on 32 of 36 audio and audio-visual benchmarks and overall SOTA on 22 without degrading performance on text, image, or video relative to single-modal Qwen counterparts.

  • Step-Audio 2 Technical Report cs.CL · 2025-07-22 · unverdicted · none · ref 81

    Step-Audio 2 integrates a latent audio encoder, reasoning-centric reinforcement learning, and discrete audio token generation into language modeling to deliver state-of-the-art performance on audio understanding and conversational benchmarks.

  • VoxCPM2 Technical Report cs.SD · 2026-06-05 · unverdicted · none · ref 39

    VoxCPM2 scales hierarchical continuous-latent speech modeling to 2B parameters and over 2M hours of multilingual data, unifying voice cloning, style control, and continuation in one backbone with open release.

  • Voxtral TTS cs.AI · 2026-03-26 · unverdicted · none · ref 20

    Voxtral TTS produces expressive multilingual speech from 3-second reference audio with a hybrid autoregressive-plus-flow-matching architecture and a new VQ-FSQ tokenizer, achieving 68.4% win rate over ElevenLabs in human evaluations.

  • PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis cs.SD · 2026-05-26 · unverdicted · none · ref 11

    PilotTTS achieves lowest WER 1.50% (en) and CER 0.87% (zh) plus highest speaker similarity on Seed-TTS Eval using a Q-Former conditioned autoregressive architecture and a released multi-stage open data pipeline.

  • Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling cs.CV · 2026-04-26 · unreviewed · ref 29