Pith. sign in

REVIEW 3 major objections 1 cited by

Synthetic speech can replace most real training audio for ASR once you know where the model separates the two and how to blur that gap.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 11:10 UTC pith:4RH6VKKN

load-bearing objection Useful LLM-backbone probe plus a practical RIR+LWP recipe; the 25% real “match” is a single-seed 0.02 WER point and should not be the headline. the 3 major comments →

arxiv 2606.29031 v2 pith:4RH6VKKN submitted 2026-06-27 cs.CL cs.AI

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

classification cs.CL cs.AI
keywords Synthetic DataSpeech LLMASRRoom Impulse ResponseText-to-SpeechLayer-wise AnalysisInterpretabilitySLAM-ASR
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Privacy rules make real customer speech hard to collect for speech recognition training in banking and healthcare, so synthetic speech from text-to-speech is attractive if it can stand in for real recordings. This paper opens the model rather than treating the synthetic-real gap as a black box: it finds that the LLM backbone of a speech ASR system mostly encodes the difference in early-to-middle layers, where temporal and pitch changes disrupt it most. Representation overlap alone does not reliably predict word-error gains. What does work is convolving synthetic audio with room impulse responses so it carries the acoustic messiness of real telephone calls, not cleaner sound. Adding a learnable layer-selection module on top of that recipe matches a full real-data baseline with only about a quarter of the real speech and beats it when more real data is available.

Core claim

Inside a SLAM-style ASR stack (frozen speech encoder, projector, LoRA-adapted LLM), real-versus-synthetic discrimination concentrates in the early-to-middle LLM layers and is most disrupted by temporal and prosodic perturbations; that separability does not directly forecast ASR gains. Room-impulse-response convolution narrows the gap by reproducing the acoustic irregularities of real recordings rather than raising perceptual quality. Combining that augmentation with per-token layer-wise weighted pooling over the decoder layers matches a fully real baseline using only 25 percent real speech (13.6 hours) and surpasses it at every higher real fraction.

What carries the argument

Layer-wise weighted pooling (LWP): a softmax-weighted mix of all LLM hidden states per token, optionally with RIR-convolved synthetic audio, so the model can draw on depths that matter for decoding after the acoustic domain gap has been reduced.

Load-bearing premise

The location of the synthetic-real gap, the filter effects, and the RIR-plus-layer-selection gains seen on this one TTS system, telephone banking English, and one speech-LLM stack will transfer to other synthesizers, languages, channels, and model families.

What would settle it

Train the same LWP-plus-RIR recipe with a different open TTS engine and a non-telephone or non-English test set; if matching the all-real baseline still requires far more than 25 percent real speech, or if early-to-middle layers no longer carry the discrimination, the central transfer claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Regulated ASR teams can cut real annotated speech to roughly a quarter while holding word error rate with RIR-augmented synthetic data and layer selection.
  • Adding more RIR-processed synthetic audio on top of a full real corpus continues to improve accuracy rather than plateauing.
  • Pitch and related prosodic filters remain useful co-treatments for synthetic audio even when pure representation overlap does not predict gains.
  • Final LLM layers stay the preferred readout for transcription while early layers remain the place to diagnose domain gap.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same early-layer domain signal may limit other speech-LLM tasks (diarization, emotion, deepfake detection) that currently treat synthetic data as a black-box mix-in.
  • If room acoustics are the dominant bridge, cheap measured or simulated impulse responses may matter more than ever-higher TTS naturalness scores for privacy-safe training.
  • Policy-gradient fine-tuning on large pools of RIR-synthetic utterances could push residual error lower without collecting more customer audio.
  • Multilingual and non-telephone channels are the natural next stress tests before claiming general substitution ratios.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper studies how synthetic speech can replace or augment real speech for LLM-based ASR (SLAM-ASR: frozen WavLM-Large + LoRA-adapted Llama-3.2-3B-Instruct) under privacy constraints. It probes the LLM backbone with silhouette, Wasserstein-on-PC1, and PCA-KDE overlap metrics, localising real/synthetic discrimination in early-to-middle layers (roughly 0–14), where temporal/prosodic filters (time stretch, pitch shift) disrupt it most. Representation-level overlap is shown not to map one-to-one onto WER. RIR convolution on Qwen3-TTS VoiceDesign outputs narrows the gap by injecting real-like acoustic irregularity (despite worse UTMOS), not by improving naturalness. Combining a per-token layer-wise weighted pooling (LWP) module over decoder layers with RIR-augmented synthetic data is reported to match a 100% real baseline (8.68% WER) at 25% real (13.6 h → 8.70%) and to beat it at higher real fractions, with best results when full real data is kept plus RIR-synth.

Significance. If the directional findings hold under broader validation, the work is practically useful for regulated domains (banking/healthcare) where real speech is scarce or restricted, and methodologically useful as one of the first layer-wise analyses of the synthetic/real gap inside a speech-LLM decoder rather than only the encoder. Strengths include a clear experimental pipeline (held-out real test WER, frozen encoder, independent overlap metrics), explicit audio-quality measurements (UTMOS/PESQ) that support the non-obvious claim that RIR helps by making synthetic audio messier, and a simple, zero-init LWP mechanism with residual ablation. The substitution and augmentation tables (IV–VIII) give actionable mix ratios. The main limitation on significance is that all load-bearing WER claims rest on a single seed, one TTS, one telephone banking corpus, and one SLAM-ASR stack, so transfer remains unproven.

major comments (3)
  1. Abstract and §V-D / Table VII: the headline claim that LWP+RIR “matches” the fully real baseline with only 25% real speech rests on 8.70% vs 8.68% WER under a single seed (42), with no multi-seed averages, error bars, or significance test. That 0.02-point gap is smaller than typical ASR run-to-run variation; without LWP the same 25%+RIR point is 9.20%, so the match is carried almost entirely by one LWP draw. Either report multi-seed mean±std (or bootstrap) for the critical rows of Tables IV–VIII, or soften the abstract/§V-D language to “approaches / is competitive with” rather than “matches … and surpasses it at all higher proportions.”
  2. §IV and the bridge to §V-C: the paper correctly notes that representation-level separability does not directly predict ASR gains, yet filter selection for Table VI is still justified primarily by early/mid-layer overlap (Time Stretch, Pitch Shift, Band Pass; High Pass deprioritised). Several combinations then behave non-monotonically with RIRs (e.g. Low Pass + Time Stretch best without RIR, High Pass + Pitch Shift best with RIR). Clarify the decision rule: which overlap metric and which layers were used as the selection criterion, and whether any filter was chosen post-hoc after seeing WER. Without that, the “interpretability-guided” claim is only partially supported.
  3. §III-B / §VI and all main tables: every reported WER uses one TTS (Qwen3-TTS VoiceDesign), one domain (DefinedAI banking telephone English), one encoder–LLM pair, and one held-out test set. The weakest assumption for the practical claim is transfer. At minimum, add one alternative TTS or a second acoustic condition (or a non-banking partition) for the critical 25% real + LWP+RIR and 100% real + 100% RIR-synth settings; otherwise state the scope limitation more prominently in the abstract and conclusions rather than only as future work.

Circularity Check

0 steps flagged

No circularity: empirical WER and layer-overlap measurements on held-out real speech; nothing reduces by construction to fitted inputs or self-citation.

full rationale

The paper is a standard empirical ASR study. Real/synthetic discrimination is measured with independent overlap metrics (silhouette, Wasserstein-1 on PC1, PCA-2D KDE) on frozen Llama layers; downstream claims are %WER on a held-out real banking test set after LoRA fine-tuning. The headline result (LWP + RIR matching the 100% real baseline at 25% real) is a reported table entry, not a quantity derived from a parameter fitted to that same quantity. RIR convolution and layer-wise weighted pooling are known techniques applied and ablated; their benefit is measured, not assumed by definition. Self-citations ([30], [31] DefinedAI corpus; SLAM-ASR framework) supply data and architecture, not a uniqueness theorem or load-bearing identity that forces the WER numbers. Statistical fragility of a 0.02-point single-seed difference is a reliability concern, not circularity. No self-definitional step, fitted-input-as-prediction, or ansatz smuggled via citation appears in the derivation chain.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 1 invented entities

The work is empirical. Load-bearing choices are architectural (SLAM-ASR + LoRA), data (DefinedAI + Qwen3-TTS + BUT RIRs), and a small set of hand-set filter/training hyperparameters. No new physical entities; the main engineered object is the per-token layer weight vector. Claims rest on the assumption that this stack and telephone banking domain are representative enough for the stated substitution result.

free parameters (5)
  • LoRA rank and alpha (r=16, α=32)
    Adapter capacity chosen by authors; all domain adaptation flows through these adapters.
  • Learning rate 1e-4, warmup 1000, 5 epochs, seed 42
    Training schedule fixed without reported sensitivity; single seed underlies every WER.
  • Signal-filter hyperparameters (SNR=10 dB, +2 semitones, 1.2× stretch, cutoffs)
    Hand-chosen perturbations used both for probing and ASR ablations.
  • Layer score vector s (zero-init, learned per LWP run)
    Only new trainable parameters for LWP; final weights concentrate on layer 28.
  • Real/synthetic mix fractions and RIR application policy
    Experimental design knobs that define the substitution and augmentation tables.
axioms (5)
  • domain assumption A frozen WavLM-Large + single projector + LoRA on Llama-3.2-3B is a representative speech-LLM ASR stack for studying the synthetic/real gap.
    All probing and WER results are on this SLAM-ASR configuration (§III-A).
  • domain assumption Qwen3-TTS VoiceDesign outputs, conditioned on persona prompts, stand in for modern multi-speaker TTS usable in regulated domains.
    Synthetic corpus generation (§III-B); TTS chosen by author listening, not a public benchmark suite.
  • domain assumption Convolution with BUT Speech@FIT RIRs injects the acoustic irregularities of real telephone recordings sufficiently to close the domain gap for ASR.
    Central mechanism claim in abstract and §V; UTMOS falls while WER improves (Table I).
  • ad hoc to paper Silhouette, Wasserstein-on-PC1, and PCA-KDE overlap are valid proxies for where the model encodes real vs synthetic.
    §IV-A metrics; authors later note they do not directly predict ASR gains.
  • standard math Standard optimization and LoRA fine-tuning math (AdamW, residual transformers) behave as usual.
    Background training assumptions, not re-derived.
invented entities (1)
  • Per-token layer-wise weighted pooling (LWP) over Llama decoder layers with optional speech residual no independent evidence
    purpose: Learn which LLM depths to use for transcription when training mixes real and synthetic speech.
    Eq. (1) and Fig. 1; related to SUPERB-style sums but applied here to the LLM backbone for synthetic-data ASR. Validated only in this paper’s tables.

pith-pipeline@v1.1.0-grok45 · 18067 in / 3283 out tokens · 39863 ms · 2026-07-12T11:10:21.391216+00:00 · methodology

0 comments
read the original abstract

In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic speech recognition (ASR) without exposing sensitive customer recordings. Yet a persistent distributional gap between synthetic and real data limits how far it can replace genuine recordings. Prior work largely treats this gap as a black box to be engineered around, but in our work, we instead examine its origin directly by probing a SLAM-ASR architecture. Then, we localise where its LLM backbone separates real from synthetic speech and find the discriminative signal concentrated in the early-to-middle layers, where temporal and prosodic perturbations disrupt it most. We further show that representation-level separability, help, but does not directly predict downstream ASR gains. On the other hand, convolving synthetic audio with room impulse responses (RIRs) narrows the gap not by making synthetic speech sound cleaner or more natural, but by reproducing the acoustic irregularities of real recordings. Translating these findings into the training procedure, by adding a layer-selection module combined with RIR augmentation matches a fully real-data baseline using only 25% of the real speech (13.6h) and surpasses it at all higher proportions.

Figures

Figures reproduced from arXiv: 2606.29031 by Andreas Stolcke, Dairazalia Sanchez-Cortes, Esa\'u Villatoro-Tello, Kadri Hacio\u{g}lu, Manjunath K E, Old\v{r}ich Plchot, Petr Motlicek, Sergio Burdisso, S\'everin Baroudi, Shashi Kumar, Srikanth Madikeri, Yanis Labrak.

Figure 1
Figure 1. Figure 1: Layer-wise Weighted Pooling inside of Llama architecture. All LLM [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Within-corpus speaker diversity (pairwise cosine distance between [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Layer-wise overlap metrics across all ablation conditions for 28 Llama layers. Lower values indicate greater real/synthetic overlap. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Synthetic Speech Is All You Have: Better Call GRPO

    cs.CL 2026-07 conditional novelty 6.0

    On synthetic banking speech alone, GRPO cuts ASR WER 40% relative to SFT (36.71%→22.09%) by improving stopping calibration and attention anchoring to audio.

Reference graph

Works this paper leans on

44 extracted references · 12 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence (arti- ficial intelligence act),

    European Parliament and Council of the European Union, “Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence (arti- ficial intelligence act),” Official Journal of the European Union, OJ L, 2024/1689, 12.7.2024, 2024, http://data.europa.eu/eli/reg/2024/1689/oj

  2. [2]

    Towards explicit acoustic evidence perception in audio llms for speech deepfake detection,

    X. Guo, Y . Xie, H. Cheng, J. Zhou, J. Liu, H. Huang, L. Ye, and Q. Zhang, “Towards explicit acoustic evidence perception in audio llms for speech deepfake detection,” 2026. [Online]. Available: https://arxiv.org/abs/2601.23066

  3. [3]

    SALMONN: Towards generic hearing abilities for large language models,

    C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. MA, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” inThe Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=14rn7HpKVk

  4. [4]

    Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,

    Y . Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou, “Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,” 2023. [Online]. Available: https://arxiv.org/abs/2311.07919

  5. [5]

    Audiopalm: A large language model that can speak and listen,

    P. K. Rubenstein, C. Asawaroengchai, D. D. Nguyen, A. Bapna, Z. Borsos, F. de Chaumont Quitry, P. Chen, D. E. Badawy, W. Han, E. Kharitonov, H. Muckenhirn, D. Padfield, J. Qin, D. Rozenberg, T. Sainath, J. Schalkwyk, M. Sharifi, M. T. Ramanovich, M. Tagliasacchi, A. Tudor, M. Velimirovi´c, D. Vincent, J. Yu, Y . Wang, V . Zayats, N. Zeghidour, Y . Zhang, ...

  6. [6]

    Speech recognition meets large language model: benchmarking, models, and exploration,

    Z. Ma, G. Yang, Y . Yang, Z. Gao, J. Wang, Z. Du, F. Yu, Q. Chen, S. Zheng, S. Zhang, and X. Chen, “Speech recognition meets large language model: benchmarking, models, and exploration,” inProceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence and Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence and Fifte...

  7. [7]

    Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “Wavlm: Large-scale self-supervised pre- training for full stack speech processing,”IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022

  8. [8]

    Robust speech recognition via large-scale weak supervi- sion,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervi- sion,” inProceedings of the 40th International Conference on Machine Learning, ser. ICML’23. JMLR.org, 2023

  9. [9]

    How auditory knowledge in llm backbones shapes audio language models: A holistic evaluation,

    K.-H. Lu, S.-W. Fu, C.-H. H. Yang, Z. Chen, S.-F. Huang, C.-K. Yang, Y .-C. Lin, C.-Y . Hsiao, W. Ren, E.-P. Hu, Y .-H. Huang, A.-Y . Cheng, C.-H. Chiang, Y . Tsao, Y .-C. F. Wang, and H. yi Lee, “How auditory knowledge in llm backbones shapes audio language models: A holistic evaluation,” 2026. [Online]. Available: https://arxiv.org/abs/2603.19195

  10. [10]

    Generating synthetic audio data for attention-based speech recognition systems,

    N. Rossenbach, A. Zeyer, R. Schl ¨uter, and H. Ney, “Generating synthetic audio data for attention-based speech recognition systems,” inICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 7069–7073. 1Repository withheld for blind review

  11. [11]

    Speech recognition with augmented synthesized speech,

    A. Rosenberg, Y . Zhang, B. Ramabhadran, Y . Jia, P. Moreno, Y . Wu, and Z. Wu, “Speech recognition with augmented synthesized speech,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2019, pp. 996–1002

  12. [12]

    Large- Scale Self- and Semi-Supervised Learning for Speech Translation,

    C. Wang, A. Wu, J. Pino, A. Baevski, M. Auli, and A. Conneau, “Large- Scale Self- and Semi-Supervised Learning for Speech Translation,” in Interspeech 2021, 2021, pp. 2242–2246

  13. [13]

    Enhancing low-resource asr through versatile tts: Bridging the data gap,

    G. Yang, F. Yu, Z. Ma, Z. Du, Z. Gao, S. Zhang, and X. Chen, “Enhancing low-resource asr through versatile tts: Bridging the data gap,” 2024. [Online]. Available: https://arxiv.org/abs/2410.16726

  14. [14]

    Towards improved speech recognition through optimized synthetic data generation,

    Y . Perrin and G. Boulianne, “Towards improved speech recognition through optimized synthetic data generation,” 2025. [Online]. Available: https://arxiv.org/abs/2508.21631

  15. [15]

    The State Of TTS: A Case Study with Human Fooling Rates,

    P. Srinivasa Varadhan, S. Thomas, S. Teja M S, S. Bhooshan, and M. M. Khapra, “The State Of TTS: A Case Study with Human Fooling Rates,” inInterspeech 2025, 2025, pp. 2285–2289

  16. [16]

    Task arithmetic can mitigate synthetic-to-real gap in automatic speech recognition,

    H. Su, H. Farn, F.-Y . Sun, S.-T. Chen, and H.-y. Lee, “Task arithmetic can mitigate synthetic-to-real gap in automatic speech recognition,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp...

  17. [17]

    A self-refining framework for enhancing asr using tts-synthesized data,

    C.-K. Chou, C.-J. Hsu, H.-L. Chung, L.-H. Tseng, H.-C. Cheng, Y .-K. Fu, K. P. Huang, and H.-Y . Lee, “A self-refining framework for enhancing asr using tts-synthesized data,” 2025. [Online]. Available: https://arxiv.org/abs/2506.11130

  18. [18]

    Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,

    X. Tan, J. Chen, H. Liu, J. Cong, C. Zhang, Y . Liu, X. Wang, Y . Leng, Y . Yi, L. He, S. Zhao, T. Qin, F. Soong, and T.-Y . Liu, “Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 6, pp. 4234–4245, 2024

  19. [19]

    Towards explainable spoofed speech attribution and detection:a probabilistic approach for characterizing speech synthesizer components,

    J. Mishra, M. Chhibber, H. jin Shim, and T. H. Kinnunen, “Towards explainable spoofed speech attribution and detection:a probabilistic approach for characterizing speech synthesizer components,” 2025. [Online]. Available: https://arxiv.org/abs/2502.04049

  20. [20]

    Asvspoof 2019: A large-scale public database of synthesized, converted and replayed speech,

    X. Wang, J. Yamagishi, M. Todisco, H. Delgado, A. Nautsch, N. Evans, M. Sahidullah, V . Vestman, T. Kinnunen, K. A. Lee, L. Juvela, P. Alku, Y .-H. Peng, H.-T. Hwang, Y . Tsao, H.-M. Wang, S. L. Maguer, M. Becker, F. Henderson, R. Clark, Y . Zhang, Q. Wang, Y . Jia, K. Onuma, K. Mushika, T. Kaneda, Y . Jiang, L.-J. Liu, Y .-C. Wu, W.-C. Huang, T. Toda, K....

  21. [21]

    Towards robust speech deepfake detection via human-inspired reasoning,

    A. Dvirniak, E. Kushnir, D. Tarasov, A. Iudin, O. Kiriukhin, M. Pautov, D. Korzh, and O. Y . Rogov, “Towards robust speech deepfake detection via human-inspired reasoning,” 2026. [Online]. Available: https://arxiv.org/abs/2603.10725

  22. [22]

    Specializing Self-Supervised Speech Representations for Speaker Segmentation,

    S. Baroudi, T. Pellegrini, and H. Bredin, “Specializing Self-Supervised Speech Representations for Speaker Segmentation,” inInterspeech 2024, 2024, pp. 3769–3773

  23. [23]

    Speech Self-Supervised Representation Benchmarking: Are We Doing it Right?

    S. Zaiem, Y . Kemiche, T. Parcollet, S. Essid, and M. Ravanelli, “Speech Self-Supervised Representation Benchmarking: Are We Doing it Right?” inInterspeech 2023, 2023, pp. 2873–2877

  24. [24]

    On the use of self- supervised representation learning for speaker diarization and separa- tion,

    S. Baroudi, H. Bredin, J. Razik, and R. Marxer, “On the use of self- supervised representation learning for speaker diarization and separa- tion,” in2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025, pp. 1–7

  25. [25]

    Layer-wise analysis of a self-supervised speech representation model,

    A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” in2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2021, pp. 914–921

  26. [26]

    SUPERB: Speech Processing Universal PERformance Benchmark,

    S. wen Yang, P.-H. Chi, Y .-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y . Y . Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee, “SUPERB: Speech Processing Universal PERformance Benchmark,” inInterspeech 2021, 2021, pp. 1194–1198

  27. [27]

    Anatomy of the modality gap: Dissecting the internal states of end-to-end speech llms,

    M.-H. Hsu, X. Zhang, X. Tian, J. Zhang, and Z. Wu, “Anatomy of the modality gap: Dissecting the internal states of end-to-end speech llms,”

  28. [28]

    Available: https://arxiv.org/abs/2603.01502

    [Online]. Available: https://arxiv.org/abs/2603.01502

  29. [29]

    The llama 3 herd of models,

    A. Grattafioriet al., “The llama 3 herd of models,” 2024. [Online]. Available: https://arxiv.org/abs/2407.21783

  30. [30]

    LoRA: Low-rank adaptation of large language models,

    E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” inInternational Conference on Learning Representations, 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9

  31. [31]

    Distilling conversations: Abstract compression of conversational audio context for llm-based asr,

    S. Kumar, E. Villatoro-Tello, S. Burdisso, K. Hacioglu, T. Ba ˜neras- Roux, H. Watawana, D. Sanchez-Cortes, S. Madikeri, P. Motlicek, and A. Stolcke, “Distilling conversations: Abstract compression of conversational audio context for llm-based asr,” 2026. [Online]. Available: https://arxiv.org/abs/2603.26246

  32. [32]

    Text-only adaptation in llm-based asr through text denoising,

    A. Carofilis, S. Burdisso, E. Villatoro-Tello, S. Kumar, K. Hacioglu, S. Madikeri, P. Rangappa, M. K. E, P. Motlicek, S. Venkatesan, and A. Stolcke, “Text-only adaptation in llm-based asr through text denoising,” 2026. [Online]. Available: https://arxiv.org/abs/2601.20900

  33. [33]

    Qwen3-tts technical report,

    H. Hu, X. Zhu, T. He, D. Guo, B. Zhang, X. Wang, Z. Guo, Z. Jiang, H. Hao, Z. Guo, X. Zhang, P. Zhang, B. Yang, J. Xu, J. Zhou, and J. Lin, “Qwen3-tts technical report,” 2026. [Online]. Available: https://arxiv.org/abs/2601.15621

  34. [34]

    Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens,

    Z. Du, Q. Chen, S. Zhang, K. Hu, H. Lu, Y . Yang, H. Hu, S. Zheng, Y . Gu, Z. Ma, Z. Gao, and Z. Yan, “Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens,” 2024. [Online]. Available: https://arxiv.org/abs/2407.05407

  35. [35]

    XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,

    E. Casanova, K. Davis, E. G ¨olge, G. G ¨oknar, I. Gulea, L. Hart, A. Alja- fari, J. Meyer, R. Morais, S. Olayemi, and J. Weber, “XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,” inInterspeech 2024, 2024, pp. 4978–4982

  36. [36]

    Parler-tts,

    Y . Lacombe, V . Srivastav, and S. Gandhi, “Parler-tts,” https://github.com/ huggingface/parler-tts, 2024

  37. [37]

    Indextts2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,

    S. Zhou, Y . Zhou, Y . He, X. Zhou, J. Wang, W. Deng, and J. Shu, “Indextts2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,” 2025. [Online]. Available: https://arxiv.org/abs/2506.21619

  38. [38]

    Omnivoice: Towards omnilingual zero-shot text- to-speech with diffusion language models,

    H. Zhu, L. Ye, W. Kang, Z. Yao, L. Guo, F. Kuang, Z. Han, W. Zhuang, L. Lin, and D. Povey, “Omnivoice: Towards omnilingual zero-shot text- to-speech with diffusion language models,” 2026. [Online]. Available: https://arxiv.org/abs/2604.00688

  39. [39]

    Chatterbox-TTS,

    Resemble AI, “Chatterbox-TTS,” https://github.com/resemble-ai/ chatterbox, 2025, gitHub repository

  40. [40]

    A study on data augmentation of reverberant speech for robust speech recognition,

    T. Ko, V . Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 5220–5224

  41. [41]

    Building and evaluation of a real room impulse response dataset,

    I. Sz ¨oke, M. Sk ´acel, L. Mo ˇsner, J. Paliesek, and J. ˇCernock´y, “Building and evaluation of a real room impulse response dataset,”IEEE Journal of Selected Topics in Signal Processing, vol. 13, no. 4, pp. 863–876, 2019

  42. [42]

    pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,

    H. Bredin, “pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,” inInterspeech 2023, 2023, pp. 1983–1987

  43. [43]

    UTMOS: UTokyo-SaruLab System for V oiceMOS Chal- lenge 2022,

    T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-SaruLab System for V oiceMOS Chal- lenge 2022,” inInterspeech 2022, 2022, pp. 4521–4525

  44. [44]

    Deepseekmath: Pushing the limits of mathematical reasoning in open language models,

    Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y . K. Li, Y . Wu, and D. Guo, “Deepseekmath: Pushing the limits of mathematical reasoning in open language models,” 2024. [Online]. Available: https://arxiv.org/abs/2402.03300