A Taiwan-specific tokenizer, language model, and bridge to a reused acoustic stack cut code-switching TTS CER from 11.45% to 4.81%, with 65.6% listener preference.
Language models are unsupervised multitask learners,
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
Hallucinations are mapped as outputs of self-attention entity confusion, MLE lack of factual constraint, and autoregressive error cascade, using an existing taxonomy.
citing papers explorer
-
BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech
A Taiwan-specific tokenizer, language model, and bridge to a reused acoustic stack cut code-switching TTS CER from 11.45% to 4.81%, with 65.6% listener preference.
-
From Architecture to Output: Structural Origins of Hallucination in Large Language Models and the Amplifying Role of Data
Hallucinations are mapped as outputs of self-attention entity confusion, MLE lack of factual constraint, and autoregressive error cascade, using an existing taxonomy.