Pith. sign in

REVIEW 4 major objections 5 minor 24 cited by

Dialogue TTS goes live: FireRedTTS-2 streams multi-speaker speech turn by turn

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-05 11:57 UTC pith:FQM4XTKB

load-bearing objection FireRedTTS-2 is a credible systems paper with a genuinely new tokenizer/dual-transformer combination, but the headline podcast gains rest on a diarized-baseline comparison that needs a same-pipeline recheck before the SOTA claim is taken at face value. the 4 major comments →

arxiv 2509.02020 v2 pith:FQM4XTKB submitted 2025-09-02 cs.SD eess.AS

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

classification cs.SD eess.AS
keywords text-to-speechmulti-speaker dialogue generationstreaming speech tokenizerdual-transformerinterleaved text-speech modelingpodcast generationinteractive chatemotion control
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

FireRedTTS-2 is a long-form, streaming text-to-speech system that generates multi-speaker dialogue such as podcasts and chatbot responses in real time, sentence by sentence, rather than requiring the full dialogue script upfront and producing one inseparable mixed track. The paper claims that its combination of a low-frame-rate semantically enriched speech tokenizer and a dual-transformer architecture trained on interleaved speaker-labeled text and speech enables stable synthesis, reliable speaker switching, and prosody that tracks long-range conversational context. If correct, this would let voice assistants, chatbots, and podcast tools synthesize natural back-and-forth speech with under-100-millisecond first-packet latency, and adapt emotion from implicit context without explicit emotion labels.

Core claim

The paper presents FireRedTTS-2, a TTS system that models multi-speaker dialogue as a single interleaved sequence of speaker-labeled text and aligned speech tokens, and generates it autoregressively turn by turn. The central claim is that this design, together with a custom 12.5 Hz streaming speech tokenizer that injects semantic supervision, produces more stable synthesis, more reliable speaker transitions, and more context-coherent prosody than existing dialogue TTS systems. In zero-shot podcast generation it reports lower WER/CER, higher speaker similarity, and lower mel-cepstral distortion than MoonCast, ZipVoice-Dialogue, and MOSS-TTSD, and its fine-tuned version is judged as natural as

What carries the argument

The text-speech interleaved sequence with speaker tags (e.g. "[S1]<text><audio>[S2]<text><audio>"), modeled by a dual-transformer: a large decoder-only backbone predicts the first-layer speech token at each step, and a smaller decoder completes the remaining token layers using both the backbone's hidden states and the predicted first-layer token. This replaces the delay-pattern parallel-head design, giving the model complete access to prior context and cutting first-packet latency to under 100 ms. The speech tokenizer operates at 12.5 Hz (half the usual 50 Hz), with Whisper-derived semantic features concatenated to acoustic features and discretized by a 16-layer RVQ, which the paper argues s

Load-bearing premise

The head-to-head outperformance claim assumes that diarizing the competing mixed-track systems and then scoring each speaker's separated audio gives a fair and accurate measurement of their intelligibility and speaker similarity.

What would settle it

Re-run the podcast comparison on the same dialogue-zh and dialogue-en sets but score the mixed-track baselines with a speech recognizer and speaker-embedding extractor that operates directly on the mixed audio, or hand-segment the baselines before scoring; if their WER/CER and SIM no longer trail FireRedTTS-2, the claimed improvements over those systems collapse. As a second check, measure first-packet latency on a live run with a stopwatch-verified streaming pipeline to confirm the under-100 ms claim.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Dialogue TTS can be deployed in truly interactive settings: a chatbot or voice assistant can begin speaking a response in under 100 ms without waiting for the full turn's text.
  • Speaker identity and turn structure stay reliable across long multi-turn conversations, which existing mixed-track dialogue systems handle poorly.
  • Prosody carries over context across turns, so a podcast or conversation sounds coherent rather than like separately synthesized utterances.
  • Emotion can be controlled implicitly from conversational context after light fine-tuning, removing the need to annotate every utterance with emotion labels.
  • A single system covers monologue voice cloning, chat, and podcast production with the same tokenizer and backbone.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The 12.5 Hz tokenizer with semantic injection is a separable contribution: even without the dual-transformer dialogue model, monologue TTS could adopt this tokenizer to cut inference compute and latency while keeping intelligibility and voice similarity.
  • The interleaved text-speech format generalizes beyond two-speaker podcasts; extending the training data to more speakers and longer contexts should scale the approach to conference calls or multi-participant live narration.
  • Because the model emits speech tokens that decode directly to waveform, it could be paired with a streaming text LLM to build an end-to-end spoken-dialogue agent where prosody and emotion respond to the live conversation state, not just a fixed script.
  • The emotion-inference result suggests a testable extension: whether the model can generalize to emotion shifts within a single turn or to emotions not in the six-category training set, which the current 30-case-per-emotion evaluation does not cover.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. FireRedTTS-2 presents a long-form, streaming text-to-speech (TTS) system for multi-speaker dialogue, combining a newly designed 12.5 Hz streaming speech tokenizer with a dual-transformer architecture that operates on text–speech interleaved sequences. The tokenizer uses semantic injection from a Whisper encoder and RVQ quantization; the backbone transformer predicts first-layer speech tokens while a smaller decoder predicts the remaining layers. The system is evaluated for speech-tokenizer reconstruction quality (Table 1), zero-shot voice cloning on Seed-TTS-eval (Table 2), emotion control in an interactive chat setting (Table 3), and zero-shot podcast generation against MoonCast, ZipVoice-Dialogue, and MOSS-TTSD (Table 4). The paper claims superior intelligibility, speaker-turn reliability, and perceived naturalness for podcast generation, as well as sentence-by-sentence generation with under-100 ms first-packet latency and emotion inference from implicit context.

Significance. If substantiated, the paper makes a meaningful engineering contribution: a low-frame-rate tokenizer with explicit semantic supervision, a dual-transformer that avoids delay-pattern latency, and an interleaved text–speech format that enables streaming multi-speaker dialogue generation. The objective results in Tables 1, 2, and 4 directionally support the main claims, and the comparison against several recent open-source dialogue systems is valuable. The paper also ships practical details on training scales and SFT recipes. However, the headline outperformance claim rests on a measurement asymmetry in Table 4, and several central claims (latency, emotion control, CMOS preference) lack the experimental support needed to be fully convincing.

major comments (4)
  1. [§4.4, Table 4] The comparison against MoonCast, ZipVoice-Dialogue, and MOSS-TTSD is asymmetric: FireRedTTS-2 outputs per-utterance speech and is scored directly, whereas the baselines produce a single mixed track and are first diarized with Pyannote. Diarization errors (missed turns, cross-talk, speaker confusion) systematically inflate WER/CER and depress SIM/MCD for the baselines but not for FireRedTTS-2. The margins are modest on dialogue-zh (CER 2.08 vs. 2.93; MCD 7.99 vs. 8.32), so even a small diarization-induced degradation could explain the gap. The paper does not report diarization accuracy on the test sets, nor does it provide a same-pipeline check (e.g., mixing FireRedTTS-2 outputs and diarizing them before scoring, or using oracle segmentation for the baselines). This is load-bearing for the central claim of outperforming state-of-the-art dialogue systems.
  2. [Tables 1–4, especially §4.4] No error bars, confidence intervals, or significance tests are reported for any objective or subjective metric. The CMOS scores in Table 4 are differences of 0.13–0.31 relative to 0.0, and without the number of raters or a confidence interval these could be within listener noise. Similarly, WER/CER and SIM differences are small (e.g., Table 1: WER 2.16 vs. 2.26; Table 4: SIM 0.753 vs. 0.736 on dialogue-zh). Given that the paper's conclusions are comparative ('surpasses', 'highest', 'lowest'), the lack of statistical support weakens the force of these claims. The authors should provide bootstrap confidence intervals or significance tests, and at minimum report rater counts and agreement for subjective tests.
  3. [§5 and Abstract] The paper repeatedly claims 'first-packet latency under 100 ms' (Abstract, Section 5), but no latency measurement or experimental protocol is reported anywhere in Section 2.2 or Section 4. The dual-transformer architecture plausibly reduces latency relative to delay-pattern decoding, but the specific under-100 ms figure is unsupported. Either provide a latency measurement (e.g., time-to-first-packet on the podcast or chat setup) or remove the quantitative claim.
  4. [§4.3, Table 3] The emotion-control evaluation reports accuracy for six emotions after fine-tuning on a 15-hour single-speaker corpus, but there is no baseline comparison (e.g., the post-trained model without chat fine-tuning, a model with explicit emotion labels, or an existing chat TTS system). Without such a baseline, the claim that FireRedTTS-2 'infers and adjusts emotion from implicit contextual cues' is not directly demonstrated. Additionally, emotion labels are manually assigned, and no inter-annotator agreement is reported, so the reliability of the 76.7–93.3% accuracies is unclear.
minor comments (5)
  1. [§2.1, Table 1] The tokenizer comparison would be easier to interpret if the bitrate were listed consistently: 'BPS' is given, but for XY-Tokenizer the WER is '-' and for some rows only PESQ-NB is shown. Please clarify why XY-Tokenizer lacks WER and whether the missing entries are intentional.
  2. [§3.3] Typo: 'utlizes' should be 'utilizes'. Also 'Y ocos' in §2.1 appears as a formatting artifact; ensure the vocoder name is rendered correctly.
  3. [§4.4] The paper uses inconsistent names for the same system: 'ZipVoice-Dialogue' in the abstract and Table 4, but 'ZipV oice-Dialog' at first mention in §4.4. Please standardize.
  4. [§4.4, Figure 4] The naturalness preference test against ground truth reports 28% 'Win' and 28% 'Even', but no details are given about the number of raters, the number of samples rated, or the interface. This information is needed to assess the robustness of the 56% 'match or exceed' conclusion.
  5. [§2.2, Eq. (1)] The loss weighting is described in the text but the sentence 'To improve training efficiency, we optimize the decoder transformer only on 1/8 of the speech segments' could be clarified: does this mean the decoder loss is computed on a randomly selected 1/8 of segments per step, and is the same mask applied to L_text? Please specify the masking procedure.

Circularity Check

0 steps flagged

No circular derivation; the evaluation is measured against external baselines, and self-citations are not load-bearing.

full rationale

FireRedTTS-2 is an empirical systems paper. The speech tokenizer and TTS model are trained with cross-entropy losses (Eq. 1) on large external speech corpora, and the headline podcast results in Table 4 are measured against MoonCast, ZipVoice-Dialogue, and MOSS-TTSD on WER/CER, SIM, MCD, and CMOS. These targets are not defined in terms of FireRedTTS-2's own outputs, nor are parameters fitted to the benchmark and then repackaged as predictions. The dual-transformer and interleaved text-speech format are adopted from external prior work [21,25], and the curriculum training follows external systems [13,19]. Citations to the authors' own FireRedTTS and FireRedTTS-1S are used only as one comparison baseline and for the general observation that objective TTS metrics are imperfect; that observation is not load-bearing for the central outperformance claim. The main measurement caveat—that mixed-track baselines are diarized with Pyannote while FireRedTTS-2 is scored directly (Section 4.4)—is a fairness/threat-to-validity concern, not a circular reduction: it does not make any result true by construction. No equation, fitted input, or self-citation chain forces the reported outcome, so no significant circularity is present.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 0 invented entities

The ledger reflects what the central comparison actually rests on: hand-set loss weights, a tokenizer design rate, a quantizer configuration, and several domain assumptions about feature sufficiency, interleaving, and evaluation fairness. No new physical or conceptual entities are introduced; all novel components are implemented artifacts rather than unexplained postulates.

free parameters (5)
  • lambda_text = 0.01
    Chosen in Eq. (1) to weight the text cross-entropy loss; no tuning study is reported.
  • lambda_decoder = 0.6
    Chosen in Eq. (1) to weight the decoder loss; no tuning study is reported.
  • decoder_training_fraction = 1/8
    The decoder transformer is optimized on only one eighth of speech segments for efficiency; no ablation shows the impact on final quality.
  • tokenizer_frame_rate = 12.5 Hz
    Halved from the common 50 Hz to shorten sequences; a hand-set design choice at the core of the approach.
  • RVQ configuration = 16 layers, 2048 code entries
    Architectural capacity choice for the speech tokenizer; no ablation compares it to alternative quantizer sizes.
axioms (4)
  • domain assumption Whisper encoder features are a sufficient semantic representation for stable text-to-token learning.
    Tokenizer design in Section 2.1 concatenates Whisper semantic features with acoustic features; no independent validation that another semantic encoder would not change the results.
  • domain assumption Interleaved speaker-labeled text and speech tokens preserve enough context for coherent prosody.
    Core modeling claim in Section 2.2; evaluated only indirectly via MCD and CMOS, with no ablation that isolates the interleaving.
  • domain assumption Diarizing mixed-track baselines with Pyannote before computing WER/CER and SIM does not bias the comparison against them.
    Section 4.4 applies Pyannote only to MoonCast, ZipVoice-Dialogue, and MOSS-TTSD; if diarization errors inflate their error rates or lower similarity, the claim of superiority is not established.
  • domain assumption Objective metrics (WER/CER, SIM, MCD, CMOS) are valid proxies for intelligibility, speaker-turn reliability, and naturalness.
    Section 4 uses these metrics throughout, while citing reference [2] that objective metrics can mislead; the tension is not resolved.

pith-pipeline@v1.4.0-alltime-deepseek-medium · 9010 in / 14123 out tokens · 139623 ms · 2026-08-05T11:57:35.578601+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot." pith.science (2026). https://pith.science/paper/FQM4XTKB

@misc{pith2026250902020,
  author       = {Pith},
  title        = {Pith review of: FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FQM4XTKB}},
  note         = {Machine review of arXiv:2509.02020}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Current dialogue generation approaches typically require the complete dialogue text before synthesis and produce a single, inseparable speech containing all voices, making them unsuitable for interactive chat; moreover, they suffer from unstable synthesis, inaccurate speaker transitions, and incoherent prosody. In this work, we present FireRedTTS-2, a long-form streaming TTS system for multi-speaker dialogue generation, delivering stable, natural speech with reliable speaker switching and context-aware prosody. A new 12.5Hz streaming speech tokenizer accelerates training and inference, extends maximum dialogue length, encodes richer semantics to stabilize text-to-token modeling and supports high-fidelity streaming generation for real-time applications. We adopt a text-speech interleaved format, concatenating speaker-labeled text with aligned speech tokens in chronological order, and model it with a dual-transformer: a large decoder-only transformer predicts tokens at the first layer, and a smaller one completes subsequent layers. Experimental results show that FireRedTTS-2 integrates seamlessly with chat frameworks and, with minimal fine-tuning, produces emotionally expressive speech guided by implicit contextual cues. In podcast generation, it surpasses existing systems including MoonCast, Zipvoice-Dialogue, and MOSS-TTSD in objective intelligibility, speaker-turn reliability, and perceived naturalness with context-consistent prosody. Our demos are available at https://fireredteam.github.io/demos/firered_tts_2.

Figures

Figures reproduced from arXiv: 2509.02020 by Feiyu Shen, Fenglong Xie, Junjie Li, Kun Xie, Xu Tang, Yao Hu.

Figure 1
Figure 1. Figure 1: An overview of FireRedTTS-2, including: (a) a new speech tokenizer with a 12.5Hz frame [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Integration of FireRedTTS-2 into interactive chat scenarios. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Zero-shot podcast generation of FireRedTTS-2. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Subjective preference results between FireRedTTS-2 fine-tuned on two podcast speakers [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

    eess.AS 2026-06 unverdicted novelty 8.0

    WavTTS is the first raw-waveform diffusion TTS model using DiT flow matching and multi-scale mel supervision that approaches SOTA latent zero-shot performance while beating prior end-to-end models.

  2. Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis

    cs.SD 2026-06 unverdicted novelty 7.0

    Sarashina2.2-TTS achieves SOTA kanji reading accuracy via data scaling and Joyo-kanji-targeted synthesis, introduces the Joyo Kanji Yomi Benchmark and Kana-CER metric, and shows stable cross-lingual performance.

  3. Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

    eess.AS 2026-05 unverdicted novelty 7.0

    GibbsTTS combines a training-free kinetic-optimal scheduler with finite-step moment correction in MI-DFM to deliver top naturalness and strong speaker similarity in zero-shot TTS.

  4. CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation

    cs.SD 2026-04 unverdicted novelty 7.0

    CapTalk unifies single-utterance and dialogue voice design via utterance- and speaker-level captions plus a hierarchical variational module for stable timbre with adaptive expression.

  5. DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech

    eess.AS 2026-07 conditional novelty 6.5

    A LoRA-plus-convolution adaptation converts an AR TTS backbone into a confidence-ordered discrete diffusion model that improves WER and speed on limited data.

  6. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  7. Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

    eess.AS 2026-07 conditional novelty 6.0

    Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...

  8. ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching

    eess.AS 2026-07 conditional novelty 6.0

    A latent flow-matching dialog TTS generates speech in a 4x-compressed 25 Hz space, cutting peak memory about 11x and inference time about 2.2x while keeping UTMOS competitive.

  9. HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

    eess.AS 2026-06 unverdicted novelty 6.0

    HPRO uses a differentiable HD-Emo codec to extract separate content and style tokens and progressively aligns frame-, word-, and sentence-level rewards to improve emotional expressiveness in TTS while preserving intel...

  10. EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

    cs.CL 2026-06 unverdicted novelty 6.0

    EmoInstruct-TTS uses Emotion2embed and an Instruction-Conditioned Emotion Flow Model (ICE-Flow) to generate acoustically grounded emotion representations from free-form instructions and integrate them into an LLM-base...

  11. TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

    cs.SD 2026-06 unverdicted novelty 6.0

    TLDR groups codec tokens into patches for patch-level autoregressive modeling in pretrained TTS systems, yielding 1.8x speedup and 75% KV-cache reduction at patch size 4.

  12. dots.tts Technical Report

    cs.SD 2026-06 unverdicted novelty 6.0

    dots.tts reports SOTA benchmark results on Seed-TTS-Eval and other tests via continuous latent-space autoregressive modeling with three listed innovations and code release.

  13. SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

    eess.AS 2026-05 unverdicted novelty 6.0

    SwanVoice is a zero-shot TTS system for 1-4 speakers that reports higher richness and hierarchy scores than open-source baselines on monologue and dialogue tasks via mixed training and DiffusionNFT post-training.

  14. SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis

    eess.AS 2026-05 unverdicted novelty 6.0

    SemaVoice adds SFM-guided alignment to refine continuous speech representations in autoregressive TTS, reporting 1.71% English WER on Seed-TTS and competitiveness with open-source SOTA.

  15. From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation

    cs.CL 2026-05 unverdicted novelty 6.0

    S2ST-Omni 2 uses typology-informed hierarchical encoding, gated Dual-CTC, and typology-aware prompting to improve multilingual S2ST over flat-label baselines on CVSS-C, with gains in low-data regimes.

  16. TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

    cs.CL 2026-04 unverdicted novelty 6.0

    TTS-PRISM defines a 12-dimensional perceptual schema, builds a targeted diagnostic dataset via adversarial synthesis and expert labels, and tunes an end-to-end model that outperforms generalist LLMs in human alignment...

  17. Qwen3-TTS Technical Report

    cs.SD 2026-01 unverdicted novelty 6.0

    Qwen3-TTS delivers state-of-the-art multilingual TTS performance with 3-second voice cloning, description control, and ultra-low-latency streaming via dual tokenizers and a dual-track LM architecture trained on over 5...

  18. ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching

    eess.AS 2026-07 unverdicted novelty 5.0

    ZipL-Dialog cuts peak GPU memory 11.22× and speeds inference 2.23× for multi-minute zero-shot dialog TTS by doing conditional flow matching in a 4× compressed latent space while keeping perceptual naturalness.

  19. FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation

    eess.AS 2026-06 unverdicted novelty 5.0

    FlashTTS delivers a streaming TTS system using multi-track input processing and X-pred mean flow matching to reach 325 ms latency in two function evaluations while retaining zero-shot voice cloning.

  20. VoxCPM2 Technical Report

    cs.SD 2026-06 unverdicted novelty 5.0

    VoxCPM2 scales hierarchical continuous-latent speech modeling to 2B parameters and over 2M hours of multilingual data, unifying voice cloning, style control, and continuation in one backbone with open release.

  21. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 accept novelty 5.0

    MLLM-enabled video translation is usefully framed as three roles—Semantic Reasoner, Expressive Performer, and Visual Synthesizer—rather than a cascade of ASR, MT, TTS, and lip-sync.

  22. FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

    cs.SD 2025-09 conditional novelty 5.0

    A full-duplex voice system with streaming personalized VAD and semantic end-of-turn detection reports fewer false barge-ins and latencies near commercial benchmarks.

  23. PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

    cs.SD 2026-05 unverdicted novelty 4.0

    PilotTTS achieves lowest WER 1.50% (en) and CER 0.87% (zh) plus highest speaker similarity on Seed-TTS Eval using a Q-Former conditioned autoregressive architecture and a released multi-stage open data pipeline.

  24. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 4.0

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

Reference graph

Works this paper leans on

37 extracted references · 17 canonical work pages · cited by 22 Pith papers

  1. [1]

    Fireredtts: A foundation text-to-speech framework for industry-level generative speech applications

    Hao-Han Guo, Yao Hu, Kun Liu, Fei-Yu Shen, Xu Tang, Yi-Chen Wu, Feng-Long Xie, Kun Xie, and Kai-Tuo Xu. Fireredtts: A foundation text-to-speech framework for industry-level generative speech applications. arXiv preprint arXiv:2409.03283, 2024

  2. [2]

    Fireredtts-1s: An upgraded streamable foundation text-to-speech system

    Hao-Han Guo, Yao Hu, Fei-Yu Shen, Xu Tang, Yi-Chen Wu, Feng-Long Xie, and Kun Xie. Fireredtts-1s: An upgraded streamable foundation text-to-speech system. arXiv preprint arXiv:2503.20499, 2025

  3. [3]

    Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens

    Zhihao Du, Qian Chen, Shiliang Zhang, Kai Hu, Heng Lu, Yexin Yang, Hangrui Hu, Siqi Zheng, Yue Gu, Ziyang Ma, et al. Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens. arXiv preprint arXiv:2407.05407, 2024

  4. [4]

    Cosyvoice 2: Scalable streaming speech synthesis with large language models

    Zhihao Du, Yuxuan Wang, Qian Chen, Xian Shi, Xiang Lv, Tianyu Zhao, Zhifu Gao, Yexin Yang, Changfeng Gao, Hui Wang, et al. Cosyvoice 2: Scalable streaming speech synthesis with large language models. arXiv preprint arXiv:2412.10117, 2024

  5. [5]

    Cosyvoice 3: Towards in-the-wild speech generation via scaling-up and post-training

    Zhihao Du, Changfeng Gao, Yuxuan Wang, Fan Yu, Tianyu Zhao, Hao Wang, Xiang Lv, Hui Wang, Chongjia Ni, Xian Shi, et al. Cosyvoice 3: Towards in-the-wild speech generation via scaling-up and post-training. arXiv preprint arXiv:2505.17589, 2025

  6. [6]

    Indextts: An industrial-level controllable and efficient zero-shot text-to-speech system

    Wei Deng, Siyi Zhou, Jingchen Shu, Jinchao Wang, and Lu Wang. Indextts: An industrial-level controllable and efficient zero-shot text-to-speech system. arXiv preprint arXiv:2502.05512, 2025

  7. [7]

    Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens

    Xinsheng Wang, Mingqi Jiang, Ziyang Ma, Ziyu Zhang, Songxiang Liu, Linqin Li, Zheng Liang, Qixi Zheng, Rui Wang, Xiaoqin Feng, et al. Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens. arXiv preprint arXiv:2503.01710, 2025

  8. [8]

    F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching

    Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, and Xie Chen. F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching. arXiv preprint arXiv:2410.06885, 2024

  9. [9]

    E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts

    Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker, Canrun Li, Chung-Hsien Tsai, Zhen Xiao, Hemin Yang, Zirun Zhu, Min Tang, Xu Tan, et al. E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts. In 2024 IEEE Spoken Language Technology Workshop (SLT), pages 682–689. IEEE, 2024. 8

  10. [10]

    Step-audio: Unified understanding and generation in intelligent speech interaction

    Ailin Huang, Boyong Wu, Bruce Wang, Chao Yan, Chen Hu, Chengli Feng, Fei Tian, Feiyu Shen, Jingbei Li, Mingrui Chen, et al. Step-audio: Unified understanding and generation in intelligent speech interaction. arXiv preprint arXiv:2502.11946, 2025

  11. [11]

    Audiogpt: Understanding and generating speech, music, sound, and talking head

    Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al. Audiogpt: Understanding and generating speech, music, sound, and talking head. arXiv preprint arXiv:2304.12995, 2023

  12. [12]

    Podagent: A comprehensive framework for podcast generation

    Yujia Xiao, Lei He, Haohan Guo, Fenglong Xie, and Tan Lee. Podagent: A comprehensive framework for podcast generation. pages 23923–23937, 2025

  13. [13]

    Covomix: Advancing zero-shot speech generation for human-like multi-talker conversations

    Leying Zhang, Yao Qian, Long Zhou, Shujie Liu, Dongmei Wang, Xiaofei Wang, Midia Yousefi, Yanmin Qian, Jinyu Li, Lei He, et al. Covomix: Advancing zero-shot speech generation for human-like multi-talker conversations. In Advances in Neural Information Processing Systems, volume 37, pages 100291–100317, 2024

  14. [14]

    Covomix2: Advancing zero-shot dialogue generation with fully non-autoregressive flow matching

    Leying Zhang, Yao Qian, Xiaofei Wang, Manthan Thakker, Dongmei Wang, Jianwei Yu, Haibin Wu, Yuxuan Hu, Jinyu Li, Yanmin Qian, et al. Covomix2: Advancing zero-shot dialogue generation with fully non-autoregressive flow matching. arXiv preprint arXiv:2506.00885, 2025

  15. [15]

    Mooncast: High-quality zero-shot podcast generation

    Zeqian Ju, Dongchao Yang, Jianwei Yu, Kai Shen, Yichong Leng, Zhengtao Wang, Xu Tan, Xinyu Zhou, Tao Qin, and Xiangyang Li. Mooncast: High-quality zero-shot podcast generation. arXiv preprint arXiv:2503.14345, 2025

  16. [16]

    Github - nari-labs/dia: A tts model capable of generating ultra-realistic dialogue in one pass.,

  17. [17]

    Parakeet, 2024

    Jordan Darefsky, Ge Zhu, and Zhiyao Duan. Parakeet, 2024

  18. [18]

    Zipvoice-dialog: Non-autoregressive spoken dialogue generation with flow matching

    Han Zhu, Wei Kang, Liyong Guo, Zengwei Yao, Fangjun Kuang, Weiji Zhuang, Zhaoqing Li, Zhifeng Han, Dong Zhang, Xin Zhang, et al. Zipvoice-dialog: Non-autoregressive spoken dialogue generation with flow matching. arXiv preprint arXiv:2507.09318, 2025

  19. [19]

    Text to spoken dialogue generation

    OpenMOSS Team. Text to spoken dialogue generation. 2025

  20. [20]

    Vibevoice technical report

    Zhiliang Peng, Jianwei Yu, Wenhui Wang, Yaoyao Chang, Yutao Sun, Li Dong, Yi Zhu, Weijiang Xu, Hangbo Bao, Zehua Wang, et al. Vibevoice technical report. arXiv preprint arXiv:2508.19205, 2025

  21. [21]

    Crossing the uncanny valley of conversa- tional voice., 2025

    Johan Schalkwyk, Ankit Kumar, Dan Lyth, Sefik Emre Eskimez, Zack Hodari, Cinjon Resnick, Ramon Sanabria, Raven Jiang, and the Sesame team. Crossing the uncanny valley of conversa- tional voice., 2025. [Accessed: 12-08-2025]

  22. [22]

    Codec does matter: Exploring the semantic shortcoming of codec for audio language model

    Zhen Ye, Peiwen Sun, Jiahe Lei, Hongzhan Lin, Xu Tan, Zheqi Dai, Qiuqiang Kong, Jianyi Chen, Jiahao Pan, Qifeng Liu, et al. Codec does matter: Exploring the semantic shortcoming of codec for audio language model. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 25697–25705, 2025

  23. [23]

    Llasa: Scaling train-time and inference-time compute for llama-based speech synthesis

    Zhen Ye, Xinfa Zhu, Chi-Min Chan, Xinsheng Wang, Xu Tan, Jiahe Lei, Yi Peng, Haohe Liu, Yizhu Jin, Zheqi Dai, et al. Llasa: Scaling train-time and inference-time compute for llama-based speech synthesis. arXiv preprint arXiv:2502.04128, 2025

  24. [24]

    Speechtokenizer: Unified speech tokenizer for speech large language models

    Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu. Speechtokenizer: Unified speech tokenizer for speech large language models. arXiv preprint arXiv:2308.16692, 2023

  25. [25]

    Moshi: a speech-text foundation model for real-time dialogue

    Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. Moshi: a speech-text foundation model for real-time dialogue. arXiv preprint arXiv:2410.00037, 2024

  26. [26]

    Robust speech recognition via large-scale weak supervision

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In International conference on machine learning, pages 28492–28518. PMLR, 2023. 9

  27. [27]

    Soundstream: An end-to-end neural audio codec

    Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. Soundstream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021

  28. [28]

    V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis

    Hubert Siuzdak. V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis. In The Twelfth International Conference on Learning Representations, 2023

  29. [29]

    Scaling transformers for low-bitrate high-quality speech coding

    Julian D Parker, Anton Smirnov, Jordi Pons, CJ Carr, Zack Zukowski, Zach Evans, and Xubo Liu. Scaling transformers for low-bitrate high-quality speech coding. In The Thirteenth International Conference on Learning Representations, 2024

  30. [30]

    Simple and controllable music generation

    Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. Simple and controllable music generation. In Advances in Neural Information Processing Systems, volume 36, pages 47704–47720, 2023

  31. [31]

    Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors

    Imtiaz Ahmed, Sadman Islam, Partha Protim Datta, Imran Kabir, Naseef Ur Rahman Chowdhury, and Ahshanul Haque. Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors. Authorea Preprints, 2025

  32. [32]

    Seed-tts: A family of high-quality versatile speech generation models

    Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, et al. Seed-tts: A family of high-quality versatile speech generation models. arXiv preprint arXiv:2406.02430, 2024

  33. [33]

    Maskgct: Zero-shot text-to-speech with masked generative codec transformer

    Yuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng, Haotian Guo, Jiachen Zheng, Qiang Zhang, Xueyao Zhang, Shunsi Zhang, and Zhizheng Wu. Maskgct: Zero-shot text-to-speech with masked generative codec transformer. arXiv preprint arXiv:2409.00750, 2024

  34. [34]

    Qwen3 technical report

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025

  35. [35]

    Powerset multi-class cross entropy loss for neural speaker diarization

    Alexis Plaquet and Hervé Bredin. Powerset multi-class cross entropy loss for neural speaker diarization. In Proc. INTERSPEECH 2023, 2023

  36. [36]

    pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe

    Hervé Bredin. pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe. In Proc. INTERSPEECH 2023, 2023. 10

  37. [2025]

    [Accessed: 12-08-2025]