REVIEW 3 major objections 1 cited by
Synthetic speech can replace most real training audio for ASR once you know where the model separates the two and how to blur that gap.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 11:10 UTC pith:4RH6VKKN
load-bearing objection Useful LLM-backbone probe plus a practical RIR+LWP recipe; the 25% real “match” is a single-seed 0.02 WER point and should not be the headline. the 3 major comments →
How to Leverage Synthetic Speech for LLM-Based ASR Systems?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Inside a SLAM-style ASR stack (frozen speech encoder, projector, LoRA-adapted LLM), real-versus-synthetic discrimination concentrates in the early-to-middle LLM layers and is most disrupted by temporal and prosodic perturbations; that separability does not directly forecast ASR gains. Room-impulse-response convolution narrows the gap by reproducing the acoustic irregularities of real recordings rather than raising perceptual quality. Combining that augmentation with per-token layer-wise weighted pooling over the decoder layers matches a fully real baseline using only 25 percent real speech (13.6 hours) and surpasses it at every higher real fraction.
What carries the argument
Layer-wise weighted pooling (LWP): a softmax-weighted mix of all LLM hidden states per token, optionally with RIR-convolved synthetic audio, so the model can draw on depths that matter for decoding after the acoustic domain gap has been reduced.
Load-bearing premise
The location of the synthetic-real gap, the filter effects, and the RIR-plus-layer-selection gains seen on this one TTS system, telephone banking English, and one speech-LLM stack will transfer to other synthesizers, languages, channels, and model families.
What would settle it
Train the same LWP-plus-RIR recipe with a different open TTS engine and a non-telephone or non-English test set; if matching the all-real baseline still requires far more than 25 percent real speech, or if early-to-middle layers no longer carry the discrimination, the central transfer claim fails.
If this is right
- Regulated ASR teams can cut real annotated speech to roughly a quarter while holding word error rate with RIR-augmented synthetic data and layer selection.
- Adding more RIR-processed synthetic audio on top of a full real corpus continues to improve accuracy rather than plateauing.
- Pitch and related prosodic filters remain useful co-treatments for synthetic audio even when pure representation overlap does not predict gains.
- Final LLM layers stay the preferred readout for transcription while early layers remain the place to diagnose domain gap.
Where Pith is reading between the lines
- The same early-layer domain signal may limit other speech-LLM tasks (diarization, emotion, deepfake detection) that currently treat synthetic data as a black-box mix-in.
- If room acoustics are the dominant bridge, cheap measured or simulated impulse responses may matter more than ever-higher TTS naturalness scores for privacy-safe training.
- Policy-gradient fine-tuning on large pools of RIR-synthetic utterances could push residual error lower without collecting more customer audio.
- Multilingual and non-telephone channels are the natural next stress tests before claiming general substitution ratios.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies how synthetic speech can replace or augment real speech for LLM-based ASR (SLAM-ASR: frozen WavLM-Large + LoRA-adapted Llama-3.2-3B-Instruct) under privacy constraints. It probes the LLM backbone with silhouette, Wasserstein-on-PC1, and PCA-KDE overlap metrics, localising real/synthetic discrimination in early-to-middle layers (roughly 0–14), where temporal/prosodic filters (time stretch, pitch shift) disrupt it most. Representation-level overlap is shown not to map one-to-one onto WER. RIR convolution on Qwen3-TTS VoiceDesign outputs narrows the gap by injecting real-like acoustic irregularity (despite worse UTMOS), not by improving naturalness. Combining a per-token layer-wise weighted pooling (LWP) module over decoder layers with RIR-augmented synthetic data is reported to match a 100% real baseline (8.68% WER) at 25% real (13.6 h → 8.70%) and to beat it at higher real fractions, with best results when full real data is kept plus RIR-synth.
Significance. If the directional findings hold under broader validation, the work is practically useful for regulated domains (banking/healthcare) where real speech is scarce or restricted, and methodologically useful as one of the first layer-wise analyses of the synthetic/real gap inside a speech-LLM decoder rather than only the encoder. Strengths include a clear experimental pipeline (held-out real test WER, frozen encoder, independent overlap metrics), explicit audio-quality measurements (UTMOS/PESQ) that support the non-obvious claim that RIR helps by making synthetic audio messier, and a simple, zero-init LWP mechanism with residual ablation. The substitution and augmentation tables (IV–VIII) give actionable mix ratios. The main limitation on significance is that all load-bearing WER claims rest on a single seed, one TTS, one telephone banking corpus, and one SLAM-ASR stack, so transfer remains unproven.
major comments (3)
- Abstract and §V-D / Table VII: the headline claim that LWP+RIR “matches” the fully real baseline with only 25% real speech rests on 8.70% vs 8.68% WER under a single seed (42), with no multi-seed averages, error bars, or significance test. That 0.02-point gap is smaller than typical ASR run-to-run variation; without LWP the same 25%+RIR point is 9.20%, so the match is carried almost entirely by one LWP draw. Either report multi-seed mean±std (or bootstrap) for the critical rows of Tables IV–VIII, or soften the abstract/§V-D language to “approaches / is competitive with” rather than “matches … and surpasses it at all higher proportions.”
- §IV and the bridge to §V-C: the paper correctly notes that representation-level separability does not directly predict ASR gains, yet filter selection for Table VI is still justified primarily by early/mid-layer overlap (Time Stretch, Pitch Shift, Band Pass; High Pass deprioritised). Several combinations then behave non-monotonically with RIRs (e.g. Low Pass + Time Stretch best without RIR, High Pass + Pitch Shift best with RIR). Clarify the decision rule: which overlap metric and which layers were used as the selection criterion, and whether any filter was chosen post-hoc after seeing WER. Without that, the “interpretability-guided” claim is only partially supported.
- §III-B / §VI and all main tables: every reported WER uses one TTS (Qwen3-TTS VoiceDesign), one domain (DefinedAI banking telephone English), one encoder–LLM pair, and one held-out test set. The weakest assumption for the practical claim is transfer. At minimum, add one alternative TTS or a second acoustic condition (or a non-banking partition) for the critical 25% real + LWP+RIR and 100% real + 100% RIR-synth settings; otherwise state the scope limitation more prominently in the abstract and conclusions rather than only as future work.
Circularity Check
No circularity: empirical WER and layer-overlap measurements on held-out real speech; nothing reduces by construction to fitted inputs or self-citation.
full rationale
The paper is a standard empirical ASR study. Real/synthetic discrimination is measured with independent overlap metrics (silhouette, Wasserstein-1 on PC1, PCA-2D KDE) on frozen Llama layers; downstream claims are %WER on a held-out real banking test set after LoRA fine-tuning. The headline result (LWP + RIR matching the 100% real baseline at 25% real) is a reported table entry, not a quantity derived from a parameter fitted to that same quantity. RIR convolution and layer-wise weighted pooling are known techniques applied and ablated; their benefit is measured, not assumed by definition. Self-citations ([30], [31] DefinedAI corpus; SLAM-ASR framework) supply data and architecture, not a uniqueness theorem or load-bearing identity that forces the WER numbers. Statistical fragility of a 0.02-point single-seed difference is a reliability concern, not circularity. No self-definitional step, fitted-input-as-prediction, or ansatz smuggled via citation appears in the derivation chain.
Axiom & Free-Parameter Ledger
free parameters (5)
- LoRA rank and alpha (r=16, α=32)
- Learning rate 1e-4, warmup 1000, 5 epochs, seed 42
- Signal-filter hyperparameters (SNR=10 dB, +2 semitones, 1.2× stretch, cutoffs)
- Layer score vector s (zero-init, learned per LWP run)
- Real/synthetic mix fractions and RIR application policy
axioms (5)
- domain assumption A frozen WavLM-Large + single projector + LoRA on Llama-3.2-3B is a representative speech-LLM ASR stack for studying the synthetic/real gap.
- domain assumption Qwen3-TTS VoiceDesign outputs, conditioned on persona prompts, stand in for modern multi-speaker TTS usable in regulated domains.
- domain assumption Convolution with BUT Speech@FIT RIRs injects the acoustic irregularities of real telephone recordings sufficiently to close the domain gap for ASR.
- ad hoc to paper Silhouette, Wasserstein-on-PC1, and PCA-KDE overlap are valid proxies for where the model encodes real vs synthetic.
- standard math Standard optimization and LoRA fine-tuning math (AdamW, residual transformers) behave as usual.
invented entities (1)
-
Per-token layer-wise weighted pooling (LWP) over Llama decoder layers with optional speech residual
no independent evidence
read the original abstract
In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic speech recognition (ASR) without exposing sensitive customer recordings. Yet a persistent distributional gap between synthetic and real data limits how far it can replace genuine recordings. Prior work largely treats this gap as a black box to be engineered around, but in our work, we instead examine its origin directly by probing a SLAM-ASR architecture. Then, we localise where its LLM backbone separates real from synthetic speech and find the discriminative signal concentrated in the early-to-middle layers, where temporal and prosodic perturbations disrupt it most. We further show that representation-level separability, help, but does not directly predict downstream ASR gains. On the other hand, convolving synthetic audio with room impulse responses (RIRs) narrows the gap not by making synthetic speech sound cleaner or more natural, but by reproducing the acoustic irregularities of real recordings. Translating these findings into the training procedure, by adding a layer-selection module combined with RIR augmentation matches a fully real-data baseline using only 25% of the real speech (13.6h) and surpasses it at all higher proportions.
Figures
Forward citations
Cited by 1 Pith paper
-
When Synthetic Speech Is All You Have: Better Call GRPO
On synthetic banking speech alone, GRPO cuts ASR WER 40% relative to SFT (36.71%→22.09%) by improving stopping calibration and attention anchoring to audio.
Reference graph
Works this paper leans on
-
[1]
Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence (arti- ficial intelligence act),
European Parliament and Council of the European Union, “Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence (arti- ficial intelligence act),” Official Journal of the European Union, OJ L, 2024/1689, 12.7.2024, 2024, http://data.europa.eu/eli/reg/2024/1689/oj
2024
-
[2]
Towards explicit acoustic evidence perception in audio llms for speech deepfake detection,
X. Guo, Y . Xie, H. Cheng, J. Zhou, J. Liu, H. Huang, L. Ye, and Q. Zhang, “Towards explicit acoustic evidence perception in audio llms for speech deepfake detection,” 2026. [Online]. Available: https://arxiv.org/abs/2601.23066
arXiv 2026
-
[3]
SALMONN: Towards generic hearing abilities for large language models,
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. MA, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” inThe Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=14rn7HpKVk
2024
-
[4]
Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,
Y . Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou, “Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,” 2023. [Online]. Available: https://arxiv.org/abs/2311.07919
Pith/arXiv arXiv 2023
-
[5]
Audiopalm: A large language model that can speak and listen,
P. K. Rubenstein, C. Asawaroengchai, D. D. Nguyen, A. Bapna, Z. Borsos, F. de Chaumont Quitry, P. Chen, D. E. Badawy, W. Han, E. Kharitonov, H. Muckenhirn, D. Padfield, J. Qin, D. Rozenberg, T. Sainath, J. Schalkwyk, M. Sharifi, M. T. Ramanovich, M. Tagliasacchi, A. Tudor, M. Velimirovi´c, D. Vincent, J. Yu, Y . Wang, V . Zayats, N. Zeghidour, Y . Zhang, ...
Pith/arXiv arXiv 2023
-
[6]
Speech recognition meets large language model: benchmarking, models, and exploration,
Z. Ma, G. Yang, Y . Yang, Z. Gao, J. Wang, Z. Du, F. Yu, Q. Chen, S. Zheng, S. Zhang, and X. Chen, “Speech recognition meets large language model: benchmarking, models, and exploration,” inProceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence and Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence and Fifte...
doi:10.1609/aaai 2025
-
[7]
Wavlm: Large-scale self-supervised pre- training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “Wavlm: Large-scale self-supervised pre- training for full stack speech processing,”IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022
2022
-
[8]
Robust speech recognition via large-scale weak supervi- sion,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervi- sion,” inProceedings of the 40th International Conference on Machine Learning, ser. ICML’23. JMLR.org, 2023
2023
-
[9]
How auditory knowledge in llm backbones shapes audio language models: A holistic evaluation,
K.-H. Lu, S.-W. Fu, C.-H. H. Yang, Z. Chen, S.-F. Huang, C.-K. Yang, Y .-C. Lin, C.-Y . Hsiao, W. Ren, E.-P. Hu, Y .-H. Huang, A.-Y . Cheng, C.-H. Chiang, Y . Tsao, Y .-C. F. Wang, and H. yi Lee, “How auditory knowledge in llm backbones shapes audio language models: A holistic evaluation,” 2026. [Online]. Available: https://arxiv.org/abs/2603.19195
arXiv 2026
-
[10]
Generating synthetic audio data for attention-based speech recognition systems,
N. Rossenbach, A. Zeyer, R. Schl ¨uter, and H. Ney, “Generating synthetic audio data for attention-based speech recognition systems,” inICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 7069–7073. 1Repository withheld for blind review
2020
-
[11]
Speech recognition with augmented synthesized speech,
A. Rosenberg, Y . Zhang, B. Ramabhadran, Y . Jia, P. Moreno, Y . Wu, and Z. Wu, “Speech recognition with augmented synthesized speech,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2019, pp. 996–1002
2019
-
[12]
Large- Scale Self- and Semi-Supervised Learning for Speech Translation,
C. Wang, A. Wu, J. Pino, A. Baevski, M. Auli, and A. Conneau, “Large- Scale Self- and Semi-Supervised Learning for Speech Translation,” in Interspeech 2021, 2021, pp. 2242–2246
2021
-
[13]
Enhancing low-resource asr through versatile tts: Bridging the data gap,
G. Yang, F. Yu, Z. Ma, Z. Du, Z. Gao, S. Zhang, and X. Chen, “Enhancing low-resource asr through versatile tts: Bridging the data gap,” 2024. [Online]. Available: https://arxiv.org/abs/2410.16726
Pith/arXiv arXiv 2024
-
[14]
Towards improved speech recognition through optimized synthetic data generation,
Y . Perrin and G. Boulianne, “Towards improved speech recognition through optimized synthetic data generation,” 2025. [Online]. Available: https://arxiv.org/abs/2508.21631
Pith/arXiv arXiv 2025
-
[15]
The State Of TTS: A Case Study with Human Fooling Rates,
P. Srinivasa Varadhan, S. Thomas, S. Teja M S, S. Bhooshan, and M. M. Khapra, “The State Of TTS: A Case Study with Human Fooling Rates,” inInterspeech 2025, 2025, pp. 2285–2289
2025
-
[16]
Task arithmetic can mitigate synthetic-to-real gap in automatic speech recognition,
H. Su, H. Farn, F.-Y . Sun, S.-T. Chen, and H.-y. Lee, “Task arithmetic can mitigate synthetic-to-real gap in automatic speech recognition,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp...
2024
-
[17]
A self-refining framework for enhancing asr using tts-synthesized data,
C.-K. Chou, C.-J. Hsu, H.-L. Chung, L.-H. Tseng, H.-C. Cheng, Y .-K. Fu, K. P. Huang, and H.-Y . Lee, “A self-refining framework for enhancing asr using tts-synthesized data,” 2025. [Online]. Available: https://arxiv.org/abs/2506.11130
Pith/arXiv arXiv 2025
-
[18]
Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,
X. Tan, J. Chen, H. Liu, J. Cong, C. Zhang, Y . Liu, X. Wang, Y . Leng, Y . Yi, L. He, S. Zhao, T. Qin, F. Soong, and T.-Y . Liu, “Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 6, pp. 4234–4245, 2024
2024
-
[19]
J. Mishra, M. Chhibber, H. jin Shim, and T. H. Kinnunen, “Towards explainable spoofed speech attribution and detection:a probabilistic approach for characterizing speech synthesizer components,” 2025. [Online]. Available: https://arxiv.org/abs/2502.04049
Pith/arXiv arXiv 2025
-
[20]
Asvspoof 2019: A large-scale public database of synthesized, converted and replayed speech,
X. Wang, J. Yamagishi, M. Todisco, H. Delgado, A. Nautsch, N. Evans, M. Sahidullah, V . Vestman, T. Kinnunen, K. A. Lee, L. Juvela, P. Alku, Y .-H. Peng, H.-T. Hwang, Y . Tsao, H.-M. Wang, S. L. Maguer, M. Becker, F. Henderson, R. Clark, Y . Zhang, Q. Wang, Y . Jia, K. Onuma, K. Mushika, T. Kaneda, Y . Jiang, L.-J. Liu, Y .-C. Wu, W.-C. Huang, T. Toda, K....
2019
-
[21]
Towards robust speech deepfake detection via human-inspired reasoning,
A. Dvirniak, E. Kushnir, D. Tarasov, A. Iudin, O. Kiriukhin, M. Pautov, D. Korzh, and O. Y . Rogov, “Towards robust speech deepfake detection via human-inspired reasoning,” 2026. [Online]. Available: https://arxiv.org/abs/2603.10725
arXiv 2026
-
[22]
Specializing Self-Supervised Speech Representations for Speaker Segmentation,
S. Baroudi, T. Pellegrini, and H. Bredin, “Specializing Self-Supervised Speech Representations for Speaker Segmentation,” inInterspeech 2024, 2024, pp. 3769–3773
2024
-
[23]
Speech Self-Supervised Representation Benchmarking: Are We Doing it Right?
S. Zaiem, Y . Kemiche, T. Parcollet, S. Essid, and M. Ravanelli, “Speech Self-Supervised Representation Benchmarking: Are We Doing it Right?” inInterspeech 2023, 2023, pp. 2873–2877
2023
-
[24]
On the use of self- supervised representation learning for speaker diarization and separa- tion,
S. Baroudi, H. Bredin, J. Razik, and R. Marxer, “On the use of self- supervised representation learning for speaker diarization and separa- tion,” in2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025, pp. 1–7
2025
-
[25]
Layer-wise analysis of a self-supervised speech representation model,
A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” in2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2021, pp. 914–921
2021
-
[26]
SUPERB: Speech Processing Universal PERformance Benchmark,
S. wen Yang, P.-H. Chi, Y .-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y . Y . Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee, “SUPERB: Speech Processing Universal PERformance Benchmark,” inInterspeech 2021, 2021, pp. 1194–1198
2021
-
[27]
Anatomy of the modality gap: Dissecting the internal states of end-to-end speech llms,
M.-H. Hsu, X. Zhang, X. Tian, J. Zhang, and Z. Wu, “Anatomy of the modality gap: Dissecting the internal states of end-to-end speech llms,”
-
[28]
Available: https://arxiv.org/abs/2603.01502
[Online]. Available: https://arxiv.org/abs/2603.01502
-
[29]
A. Grattafioriet al., “The llama 3 herd of models,” 2024. [Online]. Available: https://arxiv.org/abs/2407.21783
Pith/arXiv arXiv 2024
-
[30]
LoRA: Low-rank adaptation of large language models,
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” inInternational Conference on Learning Representations, 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9
2022
-
[31]
Distilling conversations: Abstract compression of conversational audio context for llm-based asr,
S. Kumar, E. Villatoro-Tello, S. Burdisso, K. Hacioglu, T. Ba ˜neras- Roux, H. Watawana, D. Sanchez-Cortes, S. Madikeri, P. Motlicek, and A. Stolcke, “Distilling conversations: Abstract compression of conversational audio context for llm-based asr,” 2026. [Online]. Available: https://arxiv.org/abs/2603.26246
arXiv 2026
-
[32]
Text-only adaptation in llm-based asr through text denoising,
A. Carofilis, S. Burdisso, E. Villatoro-Tello, S. Kumar, K. Hacioglu, S. Madikeri, P. Rangappa, M. K. E, P. Motlicek, S. Venkatesan, and A. Stolcke, “Text-only adaptation in llm-based asr through text denoising,” 2026. [Online]. Available: https://arxiv.org/abs/2601.20900
arXiv 2026
-
[33]
H. Hu, X. Zhu, T. He, D. Guo, B. Zhang, X. Wang, Z. Guo, Z. Jiang, H. Hao, Z. Guo, X. Zhang, P. Zhang, B. Yang, J. Xu, J. Zhou, and J. Lin, “Qwen3-tts technical report,” 2026. [Online]. Available: https://arxiv.org/abs/2601.15621
Pith/arXiv arXiv 2026
-
[34]
Z. Du, Q. Chen, S. Zhang, K. Hu, H. Lu, Y . Yang, H. Hu, S. Zheng, Y . Gu, Z. Ma, Z. Gao, and Z. Yan, “Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens,” 2024. [Online]. Available: https://arxiv.org/abs/2407.05407
Pith/arXiv arXiv 2024
-
[35]
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,
E. Casanova, K. Davis, E. G ¨olge, G. G ¨oknar, I. Gulea, L. Hart, A. Alja- fari, J. Meyer, R. Morais, S. Olayemi, and J. Weber, “XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,” inInterspeech 2024, 2024, pp. 4978–4982
2024
-
[36]
Parler-tts,
Y . Lacombe, V . Srivastav, and S. Gandhi, “Parler-tts,” https://github.com/ huggingface/parler-tts, 2024
2024
-
[37]
S. Zhou, Y . Zhou, Y . He, X. Zhou, J. Wang, W. Deng, and J. Shu, “Indextts2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,” 2025. [Online]. Available: https://arxiv.org/abs/2506.21619
Pith/arXiv arXiv 2025
-
[38]
Omnivoice: Towards omnilingual zero-shot text- to-speech with diffusion language models,
H. Zhu, L. Ye, W. Kang, Z. Yao, L. Guo, F. Kuang, Z. Han, W. Zhuang, L. Lin, and D. Povey, “Omnivoice: Towards omnilingual zero-shot text- to-speech with diffusion language models,” 2026. [Online]. Available: https://arxiv.org/abs/2604.00688
Pith/arXiv arXiv 2026
-
[39]
Chatterbox-TTS,
Resemble AI, “Chatterbox-TTS,” https://github.com/resemble-ai/ chatterbox, 2025, gitHub repository
2025
-
[40]
A study on data augmentation of reverberant speech for robust speech recognition,
T. Ko, V . Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 5220–5224
2017
-
[41]
Building and evaluation of a real room impulse response dataset,
I. Sz ¨oke, M. Sk ´acel, L. Mo ˇsner, J. Paliesek, and J. ˇCernock´y, “Building and evaluation of a real room impulse response dataset,”IEEE Journal of Selected Topics in Signal Processing, vol. 13, no. 4, pp. 863–876, 2019
2019
-
[42]
pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,
H. Bredin, “pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,” inInterspeech 2023, 2023, pp. 1983–1987
2023
-
[43]
UTMOS: UTokyo-SaruLab System for V oiceMOS Chal- lenge 2022,
T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-SaruLab System for V oiceMOS Chal- lenge 2022,” inInterspeech 2022, 2022, pp. 4521–4525
2022
-
[44]
Deepseekmath: Pushing the limits of mathematical reasoning in open language models,
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y . K. Li, Y . Wu, and D. Guo, “Deepseekmath: Pushing the limits of mathematical reasoning in open language models,” 2024. [Online]. Available: https://arxiv.org/abs/2402.03300
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.