REVIEW 2 major objections 4 minor 41 references
Fine-tuning on just 50 Singlish speakers teaches zero-shot TTS the accent, not just the voices.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 03:48 UTC pith:QXPFTMQ7
load-bearing objection First systematic Singlish TTS benchmark with a sound generalization split, but the accent-similarity metric is doing more work than it can support. the 2 major comments →
Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the discovery is that a modest fine-tuning corpus — about an hour per speaker for 50 speakers — is enough to shift the output distribution of two zero-shot TTS systems measurably toward real Singlish. The accent-similarity score against matched ground-truth recordings rises from 0.511 to 0.638 for the weaker baseline and from 0.577 to 0.604 for the one that starts closer, and the gain does not disappear when the models are asked to clone 42 voices they never heard during training. The authors take this persistence as evidence that the adaptation captures the accent itself, not just the training voices. They also find that conditioning on a learned per-speaker index
What carries the argument
The load-bearing mechanism is the adaptation-versus-consistency split. Fifty speakers are seen during fine-tuning and 42 are held out, a per-speaker split ensures no speaker overlap, and each utterance is synthesised from a single reference prompt whose content is never the synthesis target. The headline metric is accent similarity: the cosine distance between generated and matched real speech in an accent-classifier embedding space, treated as an accent-fidelity score because that classifier separates accents at high accuracy. On top of this, fine-tuning is deliberately limited to each system's text-to-token pathway — the autoregressive module in one model, the language-model and flow-match
Load-bearing premise
The headline metric assumes that cosine similarity in an accent-classifier embedding space measures whether a listener would judge the speech as Singlish; if that space is dominated by voice or recording-quality differences, the measured accent gains would not establish real accent transfer.
What would settle it
Present fine-tuned and off-the-shelf outputs to Singapore English listeners in an accent-rating or forced-choice test; if perceived Singlish-ness does not track the ACC-SIM gains, the metric is measuring the wrong thing. A second check: regress speaker embeddings out of the accent embeddings — if the 'accent' gains collapse, they are confounded with voice similarity.
If this is right
- A few dozen speakers — not hundreds of hours — can shift a zero-shot TTS system toward an under-resourced accent, lowering the data bar for accent adaptation.
- Because gains persist on held-out speakers, the fine-tuning procedure can be reused to build accent-specific TTS without voice-memorisation artefacts.
- The two backbones trade off differently: one sacrifices over-clean intelligibility for authentic accent, while the other's fine-tuning also repairs content hallucinations.
- A speaker-index conditioning mode can exceed audio-prompt accent fidelity on seen speakers, at the cost of prompt faithfulness and open-set use.
- The same protocol — matched utterance evaluation plus in-/out-of-domain split — can benchmark accent fidelity for other regional English varieties.
Where Pith is reading between the lines
- Editorial inference: the headline ACC-SIM metric is only as valid as its embedding space; if that space mostly encodes timbre or recording conditions, the reported gains would not prove perceived Singlish transfer. A native-listener accent test would settle this.
- Editorial inference: the observed inverse relationship — weaker initial accent fidelity leads to larger fine-tuning gains — suggests the gap is largely a coverage problem, so scaling pretraining with accent-diverse data might reduce or eliminate the need for fine-tuning.
- Editorial inference: the success of speaker-index conditioning hints at a controllable middle ground, e.g., interpolating between prompt and index conditioning, which could give users an accent-strength dial while retaining open-set cloning.
- Editorial inference: the dataset-filtering pipeline (WER thresholding, quality screening, gender balancing) could be transferred to other low-resource accents, but the paper does not test whether its 50-speaker corpus size is minimal; an ablation on speaker count would be needed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether fine-tuning two zero-shot TTS systems, Chatterbox and CosyVoice 3, on 50 Singapore English (Singlish) speakers from the IMDA National Speech Corpus closes the accent gap between off-the-shelf and real Singlish speech. It evaluates three distributions (real, off-the-shelf, fine-tuned) with the CodecMOS-Accent protocol (UT-MOS, WER, SPK-SIM, ACC-SIM), separates in-domain (seen) from out-of-domain (unseen) speakers, and reports that fine-tuning raises accent similarity on both, concluding that models learn the accent rather than memorize voices.
Significance. If the central claim holds, the work is a useful low-resource accent-adaptation study: it demonstrates that modest fine-tuning of two modern ZS-TTS backbones can shift output toward a regional English variety and that the effect persists on held-out speakers. The in-domain/out-of-domain design is a principled way to separate voice memorization from accent generalization, and the use of external evaluation models and a public corpus supports reproducibility. However, the headline ACC-SIM metric is an unvalidated construct, and the absence of statistical inference makes the magnitude of the reported gains uncertain.
major comments (2)
- [Section V, Table V] All reported gains are single measurements on a fixed split. There are no confidence intervals, significance tests, or per-speaker variances. The headline improvements (+0.126, +0.092, +0.027, +0.028 in ACC-SIM match) could be within run-to-run or speaker-derived noise, especially on the smaller out-of-domain set (42 speakers). Please provide bootstrap CIs or per-speaker standard errors, and ideally multiple fine-tuning runs. This is standard for TTS evaluation and is necessary to support the conclusion that the gains are reliable.
- [Section V, Table V] The observation that F-Chatterbox's prompt ACC-SIM (0.6748) exceeds the Ground Truth ceiling (0.6277) is internally inconsistent with treating ACC-SIM as a faithfulness score: generated speech should not be 'more accent-faithful' than a real same-speaker recording. The paper interprets this as drift toward a 'generalized Singlish voice,' but this also indicates that the prompt variant of ACC-SIM is not a bounded or well-calibrated measure. This anomaly should be explained or the prompt ACC-SIM should be de-emphasized; otherwise it weakens the construct validity of the headline metric.
minor comments (4)
- [Section III] The WER filtering threshold of 50% is described as based on manual inspection but no sensitivity analysis is given. A sentence reporting the retained utterance count and its dependence on the threshold (e.g., 80% or 30%) would help the reader judge robustness.
- [Section IV-D] Chatterbox hyperparameters (Table IV) are said to be 'selected on the in-domain validation split' but the selection criterion is not described. Also, the text says training loss decreases from 6.0 to below 10^-3, which is a very wide range; please clarify the loss metric and whether this is per-token cross-entropy.
- [Section V] The sentence 'F-CosyVoice is flat (0.6200 vs 0.6277, 0.5820 vs 0.5802)' for prompt ACC-SIM is unclear: the two comparisons need labels (in-domain/out-of-domain) to be interpretable.
- [References] Some references contain typographical artifacts (e.g., 'V oice', 'V ALL-E', 'Y . Wu'). Please proofread the bibliography.
Circularity Check
No significant circularity: the study is an empirical fine-tuning evaluation using external corpora and external pretrained metrics; the only self-citation is ancillary and not load-bearing.
full rationale
The paper's central claim is an empirical measurement, not a derived equation. Fine-tuning is performed on an external corpus (IMDA NSC), and all four evaluation metrics (UT-MOS, WER, SPK-SIM, ACC-SIM) come from external pretrained models or protocols (VoiceMOS, Singlish Whisper, ECAPA-TDNN, CommonAccent, CodecMOS-Accent). The in-domain/out-of-domain split directly addresses the memorize-vs-generalize question, and the generalization claim rests on held-out speakers rather than on the fine-tuning set. The only self-citation involving the present authors is reference [19], the RADAR 2026 challenge, which is cited only as a community-contribution statement and does not support any technical inference. The ACC-SIM construct-validity concern (whether the embedding primarily captures accent rather than timbre or channel) is a measurement-assumption risk, not a circularity: the metric was proposed and validated in external work [17], and the paper does not define Singlish-accent fidelity in terms of the metric it later 'predicts.' No equation is shown to reduce to its own inputs, and no fitted parameter is renamed as a prediction. Accordingly, the circularity burden is very low; the appropriate score is 1, reflecting only the presence of a minor, non-load-bearing self-citation.
Axiom & Free-Parameter Ledger
free parameters (3)
- Chatterbox inference hyperparameters (temperature=0.75, cfg_weight=0.50, repetition_penalty=1.35, exaggeration=0.40) =
temperature 0.75; cfg weight 0.50; repetition penalty 1.35; exaggeration 0.40
- CosyVoice fine-tuning configuration (LR=5e-7, epochs=200, LLM+Flow Matching trainable, vocoder frozen) =
LR 5e-7, 200 epochs
- WER filtering threshold (50%) =
50% WER
axioms (5)
- domain assumption Cosine similarity in CommonAccent ECAPA-TDNN embedding space is a valid proxy for perceived Singlish accent fidelity.
- domain assumption Singlish-fine-tuned Whisper ASR WER is a valid intelligibility proxy for synthetic Singlish.
- domain assumption IMDA NSC Part 3 speech, as filtered by the 50% WER threshold, represents colloquial Singlish of the target accent.
- domain assumption A 50-speaker, 55.5-hour in-domain pool is sufficient to learn accent rather than memorize speakers.
- domain assumption CodecMOS-Accent objective protocol validated on other accented English transfers to Singlish.
read the original abstract
Zero-shot text-to-speech (ZS-TTS) achieves near-human quality for standard English, but it copies regional accents poorly. Prompted with a short Singlish utterance, state-of-the-art systems reproduce a speaker's timbre while flattening the accent toward generic English. We investigate whether targeted fine-tuning off-the-shelf ZS-TTS can close the gap for Singapore English (Singlish). We fine-tune two cutting-edge ZS-TTS models, Chatterbox and CosyVoice 3, on 50 Singlish speakers from the IMDA National Speech Corpus. Three speech distributions are evaluated: real recordings against off-the-shelf and fine-tuned generation driven by the same Singlish audio prompts. The evaluation covers four dimensions: naturalness, intelligibility, speaker similarity, and accent similarity. We separate adaptation (in-domain speakers seen during fine-tuning) from consistency (held-out speakers) to test whether accent transfer generalises beyond the training data. Fine-tuning raises accent similarity on in-domain and out-of-domain speakers for both Chatterbox and CosyVoice 3. It moves the generated distribution measurably toward real Singlish, with the gain persisting on held-out speakers. To our knowledge, this is the first systematic study of Singlish-accented TTS.
Figures
Reference graph
Works this paper leans on
-
[1]
Neural codec language models are zero- shot text to speech synthesizers,
S. Chen, C. Wang, Y . Wuet al., “Neural codec language models are zero- shot text to speech synthesizers,”IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 705–718, 2025
2025
-
[2]
V ALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers
S. Chen, S. Liu, L. Zhouet al., “V ALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers.”CoRR, vol. abs/2406.05370, 2024
Pith/arXiv arXiv 2024
-
[3]
Z. Du, Q. Chen, S. Zhanget al., “CosyV oice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic to- kens,”ArXiv, vol. abs/2407.05407, 2024
Pith/arXiv arXiv 2024
-
[4]
CosyV oice 2: Scalable streaming speech synthesis with large language models,
Z. Du, Y . Wang, Q. Chenet al., “CosyV oice 2: Scalable streaming speech synthesis with large language models,”ArXiv, vol. abs/2412.10117, 2024
Pith/arXiv arXiv 2024
-
[5]
CosyV oice 3: Towards in-the-wild speech generation via scaling-up and post-training,
Z. Du, C. Gao, Y . Wanget al., “CosyV oice 3: Towards in-the-wild speech generation via scaling-up and post-training,”ArXiv, vol. abs/2505.17589, 2025
Pith/arXiv arXiv 2025
-
[6]
F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,
Y . Chen, Z. Niu, Z. Maet al., “F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: ACL, Jul. 2025, pp. 6255–6271
2025
-
[7]
NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Z. Ju, Y . Wang, K. Shenet al., “NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.” inICML, ser. Proceedings of Machine Learning Research, vol. 235. PMLR / OpenReview.net, 2024, pp. 22 605–22 623
2024
-
[8]
Chatterbox-TTS,
Resemble AI, “Chatterbox-TTS,” https://github.com/resemble-ai/ chatterbox, 2025, gitHub repository
2025
-
[9]
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation,
J. Zhong, K. Richmond, Z. Suet al., “AccentBox: Towards High-Fidelity Zero-Shot Accent Generation,” inICASSP. IEEE, Apr. 2025, p. 1–5
2025
-
[10]
C. M. H. Poon, P. C. Ng, X. Miaoet al., “Clarity: Contextual linguistic adaptation and accent retrieval for dual-bias mitigation in text-to-speech generation,”arXiv preprint arXiv:2511.11104, 2025
arXiv 2025
-
[11]
Deterding,Singapore English
D. Deterding,Singapore English. Edinburgh University Press, 2007
2007
-
[12]
The Homogenization of Ethnic Differences in Singapore English? A Consonantal Production Study,
K. Kalaivanan, F. Sumartono, and Y .-Y . Tan, “The Homogenization of Ethnic Differences in Singapore English? A Consonantal Production Study,”Language and Speech, vol. 64, no. 1, p. 123–140, 2020
2020
-
[13]
Building the Singapore English national speech corpus,
J. X. Koh, A. Mislan, K. Khooet al., “Building the Singapore English national speech corpus,” inInterspeech, Graz, Austria, 2019
2019
-
[14]
MNSC: Advancing Singlish Speech Understanding with Carefully Curated Corpora,
B. Wang, X. Zou, S. Sunet al., “MNSC: Advancing Singlish Speech Understanding with Carefully Curated Corpora,” inASRU, 2025, pp. 1–8
2025
-
[15]
MERaLiON-AudioLLM: Advancing Speech and Language Understanding for Singapore,
Y . He, Z. Liu, G. Linet al., “MERaLiON-AudioLLM: Advancing Speech and Language Understanding for Singapore,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguis- tics (Volume 3: System Demonstrations). Association for Computational Linguistics, Jul. 2025, pp. 22–30
2025
-
[16]
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
M. Huzaifah, G. Lin, T. Liuet al., “MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond.” arXiv:2412.11538, 2024
Pith/arXiv arXiv 2024
-
[17]
W.-C. Huang, N. Sanders, and E. Cooper, “CodecMOS-Accent: A MOS benchmark of resynthesized and TTS speech from neural codecs across English accents,” vol. arXiv:2603.14328, 2026
arXiv 2026
-
[18]
whisper-large-v3-singlish: A Singlish-fine- tuned Whisper-large-v3 model,
M. J. Wong, “whisper-large-v3-singlish: A Singlish-fine- tuned Whisper-large-v3 model,” https://huggingface.co/mjwong/ whisper-large-v3-singlish, 2024
2024
-
[19]
Radar challenge 2026: Robust audio deepfake recognition under media transformations,
H.-T. Luong, X. Liu, I. Kukanovet al., “Radar challenge 2026: Robust audio deepfake recognition under media transformations,”arXiv preprint arXiv:2605.09568, 2026
Pith/arXiv arXiv 2026
-
[20]
Accented text-to-speech synthesis with limited data,
X. Zhou, M. Zhang, Y . Zhouet al., “Accented text-to-speech synthesis with limited data,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 1699–1711, 2024
2024
-
[21]
Low-resource multilingual and zero- shot multispeaker tts,
F. Lux, J. Koch, and N. T. Vu, “Low-resource multilingual and zero- shot multispeaker tts,” inProceedings of the 2nd Conference of the Asia- Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2022, pp. 741–751
2022
-
[22]
Scalable controllable accented tts,
H. L. Xinyuan, Z. Cai, A. Garget al., “Scalable controllable accented tts,”arXiv preprint arXiv:2508.07426, 2025
Pith/arXiv arXiv 2025
-
[23]
Parameter-efficient learning for text-to-speech accent adaptation,
L.-J. Yang, C.-H. H. Yang, and J.-T. Chien, “Parameter-efficient learning for text-to-speech accent adaptation,” inProc. Interspeech. ISCA, 2023, pp. 4354–4358
2023
-
[24]
Accent vector: Controllable accent manipulation for multilingual tts without accented data,
T. Lertpetchpun, T. Trachu, J. Leeet al., “Accent vector: Controllable accent manipulation for multilingual tts without accented data,”arXiv preprint arXiv:2603.07534, 2026
arXiv 2026
-
[25]
Xtts: a massively multilingual zero-shot text-to-speech model,
E. Casanova, K. Davis, E. G ¨olgeet al., “Xtts: a massively multilingual zero-shot text-to-speech model,” inProc. Interspeech 2024, 2024, pp. 4978–4982
2024
-
[26]
Learning-free l2-accented speech generation using phonological rules,
T. Lertpetchpun, Y . Lee, J. Leeet al., “Learning-free l2-accented speech generation using phonological rules,”arXiv preprint arXiv:2603.07550, 2026
arXiv 2026
-
[27]
Macst: Multi-accent speech synthesis via text transliteration for accent conversion,
S. Inoue, S. Wang, W. Wanget al., “Macst: Multi-accent speech synthesis via text transliteration for accent conversion,” inICASSP 2025- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5
2025
-
[28]
The anatomy of singlish: globalisation, multiculturalism and the construction of the ‘local’in singapore,
R. B. Goh, “The anatomy of singlish: globalisation, multiculturalism and the construction of the ‘local’in singapore,”Journal of Multilingual and Multicultural Development, vol. 37, no. 8, pp. 748–758, 2016
2016
-
[29]
The roles of singapore standard english and singlish,
S. Harada, “The roles of singapore standard english and singlish,”Joho Kenkyu, vol. 40, pp. 69–81, 2009
2009
-
[30]
Exploring the unique morphological and syntactic features of singlish (singapore english),
N. S. Ningsih and F. Rahman, “Exploring the unique morphological and syntactic features of singlish (singapore english),”Journal of English in Academic and Professional Communication, vol. 9, no. 2, pp. 72–80, 2023
2023
-
[31]
Singlish: a controversial yet unique creole of singaporean,
A. M. Kareba, S. Aminahet al., “Singlish: a controversial yet unique creole of singaporean,” inInternational Seminar Commemorating the 100th Annniversary of Tamansiswa, vol. 1, no. 1, 2022, pp. 39–45
2022
-
[32]
Singlish as a dialect in singapore,
C. F. Peng and S. D. Madawan, “Singlish as a dialect in singapore,” International Journal of Physical and Social Sciences, vol. 3, no. 4, pp. 50–70, 2013
2013
-
[33]
The Homogenization of Ethnic Differences in Singapore English? A Consonantal Production Study,
K. Kalaivanan, F. Sumartono, and Y .-Y . Tan, “The Homogenization of Ethnic Differences in Singapore English? A Consonantal Production Study,”Language and Speech, vol. 64, no. 1, pp. 123–140, 2021
2021
-
[34]
ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,” inInterspeech. ISCA, 2020, pp. 3830– 3834
2020
-
[35]
CommonAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common V oice,
J. Zuluaga-Gomez, S. Ahmed, D. Visockaset al., “CommonAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common V oice,” inINTERSPEECH. ISCA, Aug. 2023, p. 5291–5295
2023
-
[36]
Utmos: Utokyo-sarulab system for voicemos challenge 2022,
T. Saeki, D. Xin, W. Nakataet al., “Utmos: Utokyo-sarulab system for voicemos challenge 2022,”Interspeech 2022, 2022
2022
-
[37]
TTSDS - Text-to-Speech Distribution Score,
C. Minixhofer, O. Klejch, and P. Bell, “TTSDS - Text-to-Speech Distribution Score,” inSLT. IEEE, Dec. 2024, p. 766–773
2024
-
[38]
TTSDS2: Robust Objective Evaluation for Human-Quality Synthetic Speech,
C. Minixhofer, O. Klejch, and P. Bell, “TTSDS2: Robust Objective Evaluation for Human-Quality Synthetic Speech,” in13th edition of the Speech Synthesis Workshop. ISCA, Aug. 2025, p. 68–75
2025
-
[39]
Cosyvoice,
FunAudioLLM, “Cosyvoice,” https:// github.com/FunAudioLLM/CosyV oice/tree/ 4d7295a9a7076b7b656f63c35f4dad199f3af33e, 2025, gitHub repository, commit 4d7295a9a7076b7b656f63c35f4dad199f3af33e. Accessed: 2026-06-29
2025
-
[40]
Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS,
D. Seo, G. Park, and K. Nam, “Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS,” 2026. [Online]. Available: https://arxiv.org/abs/2605.30748
Pith/arXiv arXiv 2026
-
[41]
The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains,
E. Cooper, W.-C. Huang, Y . Tsaoet al., “The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023, pp. 1–7
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.