REVIEW 4 major objections 5 minor 44 references
A single production TTS system can simultaneously deliver top content accuracy, speaker similarity, naturalness, controllability, multilingual coverage, and robustness to degraded input audio.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-07-31 23:30 UTC pith:GMPLL7XY
load-bearing objection A genuine, mostly solid production-TTS systems paper whose headline SOTA claim is one notch too strong for its own tables; the controllability evidence is the weakest load-bearing piece. the 4 major comments →
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim, on the paper's own terms, is that Qwen-Audio-3.0-TTS combined a 12.5 Hz supervised speech tokenizer with an enlarged 59,049-entry codebook — halving the frame rate of its predecessor while preserving content and speaker information — and a five-stage progressive training pipeline: independent LM and FM pretraining; joint training where the flow-matching model conditions on continuous LM hidden states rather than only discrete token IDs, with high-quality data annealing; LM reinforcement learning with a composite reward and a differentiable corrective branch; FM acoustic-robustness training on corrupted prompts with a frozen LM; and FM reinforcement learning with an SDE-bas
What carries the argument
The central machinery is the two-part generation stack: a 12.5 Hz supervised speech tokenizer (a causal encoder with a 10-dimensional finite scalar quantization bottleneck of 59,049 entries, trained with multi-task supervision from ASR, language identification, emotion recognition, audio-event detection, speaker analysis, and general audio analysis) that produces semantic tokens at half the previous frame rate; and a five-stage progressive training pipeline. The load-bearing design choice is that the flow-matching acoustic renderer conditions on continuous hidden states from the language model, not just discrete tokens, so the FM can access information lost in quantization while the autoregr
Load-bearing premise
The claim of state-of-the-art hinges on the validity of the paper's self-constructed evaluation protocol — particularly the LLM judge for instruction-following (Section 4.4.4, agreeing with humans only 70% of the time with a conservative bias, p = 0.007) and the native-speaker dialect evaluation (Section 4.6.1) that includes no external baseline systems; if those judgments are biased toward this model, the headline performance claim would weaken even though the architecture a
What would settle it
Re-run the 440-case instruction-following benchmark with human listeners as primary judges on the same criteria and compare win rates against the closest competitor; if humans prefer the competitor or if human–LLM agreement falls to chance levels on the disputed criteria, the instruction-following advantage would not survive. Likewise, run the 20-dialect prompts through several competitive multilingual systems and have native speakers score dialect authenticity blind; if the gap narrows or reverses, the dialect coverage claim would be unsupported.
If this is right
- If the claims hold, a single production TTS can synthesize in sixteen languages and twenty Chinese dialect regions, follow free-style natural-language instructions plus fine-grained inline tags, generate up to three-minute passages in one pass, and clone voices from noisy or reverberant reference audio without an explicit denoising mode.
- Halving the token rate to 12.5 Hz while scaling the codebook roughly ninefold cuts autoregressive decoding steps with no loss in downstream TTS quality, pointing to a general recipe for low-latency speech language models.
- Feeding continuous LM hidden states into the flow-matching renderer, rather than only discrete tokens, is claimed to alleviate the information bottleneck of token-only interfaces and to let the FM loss shape LM representations.
- The two-stage speaker adaptation protocol, with joint LM plus FM adaptation followed by FM-only refinement, improves content consistency across all four reported target speakers.
- The robustness training integrates prompt enhancement into the cloning path, so degraded-input quality no longer depends on an inference-time denoiser.
Where Pith is reading between the lines
- The success of a 12.5 Hz tokenizer with a larger codebook suggests that future systems may push to even lower frame rates, such as 6.25 Hz, trading a small amount of similarity for further latency gains — a testable extension the paper does not run.
- The reported 70% agreement between the LLM judge and human raters on instruction-following, together with a conservative bias (p = 0.007), implies that instruction-following evaluations may need human arbitration as systems improve; the disputed criteria could be re-scored by humans to check whether the ranking survives.
- The leaderboard rank-one claim has a confidence interval that overlaps the runner-up, so "statistically tied at the top" is the more defensible public framing until more paired samples accumulate.
- The robustness results come from real noisy and reverberant prompts with no paired clean references; this design highlights that future TTS benchmarks should include degradation-matched enrollment conditions rather than only synthetic corruption.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents Qwen-Audio-3.0-TTS, a production-oriented TTS system combining a 12.5 Hz supervised speech tokenizer, a five-stage progressive LM/FM training pipeline, natural-language and fine-grained inline-tag controllability, 16-language/20-dialect coverage, one-pass long-form synthesis, and acoustic robustness. The central scientific claim is that the system achieves state-of-the-art or 'strongest aggregate' performance across SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness evaluations, and that it ranks first on the independent Artificial Analysis TTS leaderboard. The paper includes tokenizer ablations, extensive objective comparisons, some arena-style human evaluation, and a self-constructed deployment-oriented benchmark (Qwen-Audio-TTS-Eval).
Significance. If the empirical claims were fully supported, this would be a significant systems contribution: it demonstrates that one architecture can combine low-frame-rate discrete tokens, continuous hidden-state conditioning, RL-based training, controllable instruction following, multilingual/dialect coverage, and robustness to degraded prompts. The training pipeline is technically interesting, the tokenizer ablation is useful, and the independent leaderboard result is a genuine external data point. However, several load-bearing evaluation choices currently leave the headline SOTA claims under-supported. The underlying system is plausible, but the evidence as presented is not yet proportionate to the abstract's claims.
major comments (4)
- [§4.4.4, Table 10] The headline 'free-style instruction following' SOTA is scored exclusively by Gemini-2.5-Pro, whose calibration on this benchmark is weak: 70.0% single-attribute agreement, 64.9% recall, conservative positive-rate bias (McNemar p=0.007), and 56.7% criterion-level agreement on complex instructions. The authors disclose these limitations but do not show that the judge's bias is system-independent; no human ranking validates the reported margins (e.g., English Overall 80.45 vs CosyVoice3's 64.09). Because controllability is a central contribution, this claim is not established. I request a pre-registered human A/B or human rating on the same 440 cases, or evidence that LLM-judge errors are uncorrelated with system identity.
- [§4.6.1, Table 11] The dialect synthesis evaluation contains no external baselines: Table 11 reports only Qwen-Audio-3.0-TTS's absolute scores (Dialect Authenticity 3.639, Pronunciation Accuracy 3.935, Prosodic Naturalness 3.680). Native-speaker mean scores without a comparison system cannot support a competitive 'robust deployment coverage' claim for 20 dialect regions. Please add at least one strong multilingual/dialect baseline (e.g., a CosyVoice-family or commercial system) and report inter-annotator agreement (e.g., Fleiss' κ).
- [Abstract, §4.2, §4.4.2] 'State-of-the-art performance on many reported dimensions or the strongest aggregate results' is not supported by the authors' own tables. In Table 3, Qwen-Audio-3.0-TTS has test-en WER 1.54 vs 1.30 (Dots.TTS) and 1.24 (Qwen3-TTS), and test-hard CER 7.00 vs 6.04 (LongCat). In Table 8, all zh CER is 2.22 vs 0.54 (VoxCPM2) and all en WER is 5.00 vs 3.20 (VoxCPM2). The phrase 'strongest aggregate' is never defined or computed. Please define the aggregate metric a priori and with uncertainty, or temper the abstract to name the specific dimensions on which the model is best.
- [Tables 3–10] No confidence intervals or significance tests are reported for the core comparative tables, yet many 'best' claims rest on small margins (e.g., Table 4 Thai WER 1.45 vs MiniMax-Speech 1.65; Table 6 claiming 'best in eight' of twelve transfer directions without error bars). Without measures of uncertainty, 'best' vs 'statistically tied' cannot be distinguished. Add bootstrap CIs or pairwise significance tests for the comparisons that support the abstract's SOTA claims.
minor comments (5)
- [§3.3] Qwen-Audio-TTS-Eval is a self-constructed benchmark; please state whether it will be released and provide the detailed annotation instructions, since several headline claims depend on it.
- [§4.4.1, Table 7] Text normalization is also scored by Gemini-2.5-Pro, but unlike §4.4.4 no human-agreement calibration is reported for this task. Please add the calibration or treat the scores as provisional.
- [§4.6.2, Table 12] The 'Previous-Gen Baseline' is unnamed, and no sample size or number of raters is given; the win rates are hard to interpret without this information.
- [§4.5, Figure 5] Speaker-adapted evaluation covers only four anonymized speakers. Please state how these speakers were selected and whether the gains are statistically consistent across a larger pool.
- [§4.6.1] No inter-annotator agreement metric is reported for the dialect ratings. Since each utterance is rated by three native speakers, please report agreement (e.g., Fleiss' κ or Krippendorff's α).
Circularity Check
No significant circularity: the paper's results rest on external benchmarks, independent leaderboard data, and disclosed evaluation protocols rather than on equations or parameters that reduce to their own inputs.
full rationale
The paper is an experimental systems report, not a derivation whose predictions collapse into fitted inputs. The core claims are supported by objective evaluations on external or independently defined benchmarks (SEED-TTS-Eval, CV3-Eval), by an independent third-party leaderboard (Artificial Analysis), and by comparative tables against open-source and API baselines. Self-citations such as CosyVoice2 [10], CosyVoice3 [4], and FlowTTS-GRPO [21] are used as methodological provenance or implementation references; they are not invoked as proof of the present results, and the relevant methods are described with equations and training stages in the paper itself. The self-constructed Qwen-Audio-TTS-Eval is a disclosed benchmark with an LLM judge whose agreement with humans is explicitly calibrated and reported (70.0% single-attribute agreement, McNemar p=0.007); this is a validity/risk concern about evaluation, not a circular reduction. No fitted parameter is later renamed as a prediction, and no 'uniqueness theorem' is imported from the authors' prior work to force a choice. The reader's noted evaluation-bias concern is legitimate but does not constitute circularity under the definition used here.
Axiom & Free-Parameter Ledger
free parameters (4)
- Reward weights (lambda_content, lambda_dur, lambda_div, lambda_prosody; lambda1, lambda2, lambda3) =
not disclosed
- Exploration intensity a (sigma_t) =
not disclosed
- Codebook size / FSQ configuration =
59,049 (3^10) at 12.5 Hz
- High-quality annealing subset composition =
not specified
axioms (4)
- domain assumption ASR CER/WER and DNSMOS are valid proxies for content consistency and perceptual quality
- domain assumption Gemini-2.5-Pro judgements are a faithful proxy for human listening in instruction-following
- domain assumption The 12.5 Hz supervised tokenizer retains enough content and speaker information for high-quality synthesis
- domain assumption Leaderboard first-by-point-estimate is a meaningful 'rank first' claim despite overlapping confidence intervals
read the original abstract
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~Hz low-frame-rate speech tokenizer for reduced inference latency with a five-stage progressive training paradigm for coordinated language model (LM) and flow-matching model (FM) optimization. The model provides production-level control through free-style natural-language instructions and fine-grained inline tags, while supporting 16 languages, 20 Chinese dialect regions, one-pass long-form synthesis up to 3 minutes, and robust generation from noisy, reverberant, or unclear reference speech. Across SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness evaluations, Qwen-Audio-3.0-TTS achieves state-of-the-art performance on many reported dimensions or the strongest aggregate results. It also ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard. These results establish Qwen-Audio-3.0-TTS as a strong foundation for production-level speech synthesis.
Figures
Reference graph
Works this paper leans on
-
[1]
Neural codec language models are zero-shot text to speech synthesizers.CoRR, abs/2301.02111, 2023
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei. Neural codec language models are zero-shot text to speech synthesizers.CoRR, abs/2301.02111, 2023
Pith/arXiv arXiv 2023
-
[2]
Seed-tts: A family of high-quality versatile speech generation models.CoRR, abs/2406.02430, 2024
Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, Mingqing Gong, Peisong Huang, Qingqing Huang, Zhiying Huang, Yuanyuan Huo, Dongya Jia, Chumin Li, Feiya Li, Hui Li, Jiaxin Li, Xiaoyang Li, Xingxing Li, Lin Liu, Shouda Liu, Sichao Liu, Xudong Liu, Yuchen Liu, Zhengxi Liu, Lu Lu, J...
Pith/arXiv arXiv 2024
-
[3]
V oicebox: Text- guided multilingual universal speech generation at scale
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, and Wei-Ning Hsu. V oicebox: Text- guided multilingual universal speech generation at scale. InNeurIPS, 2023
2023
-
[4]
Cosyvoice 3: Towards in-the-wild speech generation via scaling-up and post-training, 2025
Zhihao Du, Changfeng Gao, Yuxuan Wang, Fan Yu, Tianyu Zhao, Hao Wang, Xiang Lv, Hui Wang, Chongjia Ni, Xian Shi, Keyu An, Guanrou Yang, Yabin Li, Yanni Chen, Zhifu Gao, Qian Chen, Yue Gu, Mengzhe Chen, Yafeng Chen, Shiliang Zhang, Wen Wang, and Jieping Ye. Cosyvoice 3: Towards in-the-wild speech generation via scaling-up and post-training, 2025
2025
-
[5]
Xinsheng Wang, Mingqi Jiang, Ziyang Ma, Ziyu Zhang, Songxiang Liu, Linqin Li, Zheng Liang, Qixi Zheng, Rui Wang, Xiaoqin Feng, et al. Spark-tts: An efficient llm-based text-to- speech model with single-stream decoupled speech tokens.arXiv preprint arXiv:2503.01710, 2025
Pith/arXiv arXiv 2025
-
[6]
Qwen3-tts technical report, 2026
Hangrui Hu, Xinfa Zhu, Ting He, Dake Guo, Bin Zhang, Xiong Wang, Zhifang Guo, Ziyue Jiang, Hongkun Hao, Zishan Guo, Xinyu Zhang, Pei Zhang, Baosong Yang, Jin Xu, Jingren Zhou, and Junyang Lin. Qwen3-tts technical report, 2026
2026
-
[7]
E2 TTS: embarrassingly easy fully non-autoregressive zero-shot TTS.CoRR, abs/2406.18009, 2024
Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker, Canrun Li, Chung-Hsien Tsai, Zhen Xiao, Hemin Yang, Zirun Zhu, Min Tang, Xu Tan, Yanqing Liu, Sheng Zhao, and Naoyuki Kanda. E2 TTS: embarrassingly easy fully non-autoregressive zero-shot TTS.CoRR, abs/2406.18009, 2024
Pith/arXiv arXiv 2024
-
[8]
F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching
Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, and Xie Chen. F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching. CoRR, abs/2410.06885, 2024
Pith/arXiv arXiv 2024
-
[9]
Zhihao Du, Qian Chen, Shiliang Zhang, Kai Hu, Heng Lu, Yexin Yang, Hangrui Hu, Siqi Zheng, Yue Gu, Ziyang Ma, Zhifu Gao, and Zhijie Yan. Cosyvoice: A scalable multi- lingual zero-shot text-to-speech synthesizer based on supervised semantic tokens.CoRR, abs/2407.05407, 2024
Pith/arXiv arXiv 2024
-
[10]
Zhihao Du, Yuxuan Wang, Qian Chen, Xian Shi, Xiang Lv, Tianyu Zhao, Zhifu Gao, Yexin Yang, Changfeng Gao, Hui Wang, et al. Cosyvoice 2: Scalable streaming speech synthesis with large language models.arXiv preprint arXiv:2412.10117, 2024
Pith/arXiv arXiv 2024
-
[11]
Dongya Jia, Zhuo Chen, Jiawei Chen, Chenpeng Du, Jian Wu, Jian Cong, Xiaobin Zhuang, Chumin Li, Zhen Wei, Yuping Wang, et al. Ditar: Diffusion transformer autoregressive mod- eling for speech generation.arXiv preprint arXiv:2502.03930, 2025
arXiv 2025
-
[12]
dots.tts technical report, 2026
Shi Lian, Changtao Li, Bohan Li, Hankun Wang, Da Zheng, Junfeng Tian, Yufeng Ma, Colin Zhang, and Kai Yu. dots.tts technical report, 2026
2026
-
[13]
V oxcpm2 tech- nical report, 2026
Yixuan Zhou, Guoyang Zeng, Xin Liu, Xiang Li, Renjie Yu, Jiancheng Gui, Jiaheng Wu, Ziyang Wang, Xudong Shen, Runchuan Ye, Zhisheng Zhang, Jiuyang Zhou, Bingsong Bai, Weiyue Sun, Mengyuan Deng, Qundong Shi, Zhiyong Wu, and Zhiyuan Liu. V oxcpm2 tech- nical report, 2026
2026
-
[14]
Bigvgan: A universal neural vocoder with large-scale training
Sang gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon. Bigvgan: A universal neural vocoder with large-scale training. InThe Eleventh International Conference on Learning Representations, 2023. 17
2023
-
[15]
Funaudiollm: V oice understanding and generation foundation models for natural interaction between humans and llms, 2024
Keyu An, Qian Chen, Chong Deng, Zhihao Du, Changfeng Gao, Zhifu Gao, Yue Gu, Ting He, Hangrui Hu, Kai Hu, Shengpeng Ji, Yabin Li, Zerui Li, Heng Lu, Haoneng Luo, Xiang Lv, Bin Ma, Ziyang Ma, Chongjia Ni, Changhe Song, Jiaqi Shi, Xian Shi, Hao Wang, Wen Wang, Yuxuan Wang, Zhangyu Xiao, Zhijie Yan, Yexin Yang, Bin Zhang, Qinglin Zhang, Shiliang Zhang, Nan Z...
2024
-
[16]
Finite scalar quantization: VQ-V AE made simple
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. Finite scalar quantization: VQ-V AE made simple. InICLR. OpenReview.net, 2024
2024
-
[17]
Qian Chen, Yafeng Chen, Yanni Chen, Mengzhe Chen, Yingda Chen, Chong Deng, Zhihao Du, Ruize Gao, Changfeng Gao, Zhifu Gao, et al. MinMo: A multimodal large language model for seamless voice interaction.arXiv preprint arXiv:2501.06282, 2025
Pith/arXiv arXiv 2025
-
[18]
Joyvoice: Long-context conditioning for anthropomorphic multi-speaker conversational synthesis, 2025
Fan Yu, Tao Wang, You Wu, Lin Zhu, Wei Deng, Weisheng Han, Wenchao Wang, Lin Hu, Xiangyu Liang, Xiaodong He, Yankun Huang, Yu Gu, Yuan Liu, Yuxuan Wang, Zhangyu Xiao, Ziteng Wang, Boya Dong, Feng Dang, Jinming Chen, Jingdong Li, Jun Wang, Yechen Jin, Yuan Zhang, Zhengyan Sheng, and Xin Wang. Joyvoice: Long-context conditioning for anthropomorphic multi-sp...
2025
-
[19]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y . K. Li, Y . Wu, and Daya Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
2024
-
[20]
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. InInternational Conference on Learning Representations, 2017
2017
-
[21]
Flowtts-grpo: On- line reinforcement learning with multi-objective reward optimization for flow-matching based text-to-speech, 2026
Haoxu Wang, Biao Tian, Weiqin Li, Xiang Lv, Han Zhao, and Xiangang Li. Flowtts-grpo: On- line reinforcement learning with multi-objective reward optimization for flow-matching based text-to-speech, 2026
2026
-
[22]
Flowse-grpo: Training flow matching speech enhancement via online reinforce- ment learning
Haoxu Wang, Biao Tian, Yiheng Jiang, Zexu Pan, Shengkui Zhao, Bin Ma, Daren Chen, and Xiangang Li. Flowse-grpo: Training flow matching speech enhancement via online reinforce- ment learning. InICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 16182–16186, 2026
2026
-
[23]
Flow-grpo: Training flow matching models via online rl.arXiv preprint arXiv:2505.05470, 2025
Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, and Wanli Ouyang. Flow-grpo: Training flow matching models via online rl.arXiv preprint arXiv:2505.05470, 2025
Pith/arXiv arXiv 2025
-
[24]
Classifier-free diffusion guidance.CoRR, abs/2207.12598, 2022
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance.CoRR, abs/2207.12598, 2022
Pith/arXiv arXiv 2022
-
[25]
Jianlin Su, Murtadha H. M. Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Ro- former: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
2024
-
[26]
Qwen2.5: A party of foundation models, September 2024
Qwen Team. Qwen2.5: A party of foundation models, September 2024
2024
-
[27]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. InInter- national Conference on Learning Representations, 2022
2022
-
[28]
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors,International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Hon- olulu, Hawa...
2023
-
[29]
Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition
Zhifu Gao, Shiliang Zhang, Ian McLoughlin, and Zhijie Yan. Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition. InInterspeech, pages 2063–2067. ISCA, 2022
2063
-
[30]
Funasr: A fundamental end-to-end speech recognition toolkit.arXiv preprint arXiv:2305.11013, 2023
Zhifu Gao, Zerui Li, Jiaming Wang, Haoneng Luo, Xian Shi, Mengzhe Chen, Yabin Li, Lingyun Zuo, Zhihao Du, Zhangyu Xiao, et al. Funasr: A fundamental end-to-end speech recognition toolkit.arXiv preprint arXiv:2305.11013, 2023. 18
Pith/arXiv arXiv 2023
-
[31]
Yafeng Chen, Siqi Zheng, Hui Wang, Luyao Cheng, Qian Chen, and Jiajun Qi. An en- hanced res2net with local and global feature fusion for speaker verification.arXiv preprint arXiv:2305.12838, 2023
Pith/arXiv arXiv 2023
-
[32]
Chandan K. A. Reddy, Vishak Gopal, and Ross Cutler. Dnsmos P.835: A non-intrusive percep- tual objective speech quality metric to evaluate noise suppressors. InICASSP, pages 886–890. IEEE, 2022
2022
-
[33]
Large-scale self-supervised speech representation learning for automatic speaker verification
Zhengyang Chen, Sanyuan Chen, Yu Wu, Yao Qian, Chengyi Wang, Shujie Liu, Yanmin Qian, and Michael Zeng. Large-scale self-supervised speech representation learning for automatic speaker verification. InICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6147–6151. IEEE, 2022
2022
-
[34]
Xiaohui Sun, Ruitong Xiao, Jianye Mo, Bowen Wu, Qun Yu, and Baoxun Wang. F5r-tts: Improving flow matching based text-to-speech with group relative policy optimization.arXiv preprint arXiv:2504.02407, 2025
Pith/arXiv arXiv 2025
-
[35]
Longcat-audiodit: High-fidelity diffusion text-to-speech in the waveform latent space, 2026
Detai Xin, Shujie Hu, Chengzuo Yang, Chen Huang, Guoqiao Yu, Guanglu Wan, and Xunliang Cai. Longcat-audiodit: High-fidelity diffusion text-to-speech in the waveform latent space, 2026
2026
-
[36]
Fireredtts-2: Towards long conversational speech generation for podcast and chatbot, 2025
Kun Xie, Feiyu Shen, Junjie Li, Fenglong Xie, Xu Tang, and Yao Hu. Fireredtts-2: Towards long conversational speech generation for podcast and chatbot, 2025
2025
-
[37]
Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech, 2025
Siyi Zhou, Yiquan Zhou, Yi He, Xun Zhou, Jinchao Wang, Wei Deng, and Jingchen Shu. Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech, 2025
2025
-
[38]
Qwen2.5-omni technical report.arXiv preprint arXiv:2503.20215, 2025
Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, et al. Qwen2.5-omni technical report.arXiv preprint arXiv:2503.20215, 2025
Pith/arXiv arXiv 2025
-
[39]
Qwen3.5-omni technical report, 2026
Qwen Team. Qwen3.5-omni technical report, 2026
2026
-
[40]
Minimax-speech: Intrinsic zero-shot text-to-speech with a learnable speaker encoder, 2025
Bowen Zhang, Congchao Guo, Geng Yang, Hang Yu, Haozhe Zhang, Heidi Lei, Jialong Mai, Junjie Yan, Kaiyue Yang, Mingqi Yang, Peikai Huang, Ruiyang Jin, Sitan Jiang, Wei- hua Cheng, Yawei Li, Yichen Xiao, Yiying Zhou, Yongmao Zhang, Yuan Lu, and Yucen He. Minimax-speech: Intrinsic zero-shot text-to-speech with a learnable speaker encoder, 2025
2025
-
[41]
Se Jin Park, Julian Salazar, Aren Jansen, Keisuke Kinoshita, Yong Man Ro, and R. J. Skerry- Ryan. Long-form speech generation with spoken language models.CoRR, abs/2412.18603, 2024
Pith/arXiv arXiv 2024
-
[42]
Common voice: A massively-multilingual speech corpus.arXiv preprint arXiv:1912.06670, 2019
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. Common voice: A massively-multilingual speech corpus.arXiv preprint arXiv:1912.06670, 2019
Pith/arXiv arXiv 1912
-
[43]
Fleurs: Few-shot learning evaluation of universal representations of speech
Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna. Fleurs: Few-shot learning evaluation of universal representations of speech. In2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805. IEEE, 2023
2023
-
[44]
Gemini 2.5: Pushing the frontier with advanced reasoning, multi- modality, long context, and next generation agentic capabilities, 2025
Gheorghe Comanici et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multi- modality, long context, and next generation agentic capabilities, 2025. 19
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.