ReNikud improves Hebrew G2P by combining ASR pseudo-labeling from unlabeled audio with character-level IPA prediction, outperforming prior methods on benchmarks including a new spoken Hebrew test set.
Powsm: A phonetic open whisper-style speech foundation model
3 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 3representative citing papers
Whisper-style speech encoders show semantic cross-lingual alignment beyond phonetics in final layers, and early-exiting boosts ASR performance on low-resource languages.
A 33M-parameter raw-audio CTC model with 19-block RoPE E-Branchformer achieves 9.19% whitespace-insensitive IPA CER on a 16,660-utterance 41-language test set, outperforming a 575M-parameter PhoneticXEUS baseline at 9.78% under matched normalization.
citing papers explorer
-
ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion
ReNikud improves Hebrew G2P by combining ASR pseudo-labeling from unlabeled audio with character-level IPA prediction, outperforming prior methods on benchmarks including a new spoken Hebrew test set.
-
Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically
Whisper-style speech encoders show semantic cross-lingual alignment beyond phonetics in final layers, and early-exiting boosts ASR performance on low-resource languages.
-
BranchShine: Compact Raw-Audio-to-IPA Transcription with a RoPE E-Branchformer Encoder
A 33M-parameter raw-audio CTC model with 19-block RoPE E-Branchformer achieves 9.19% whitespace-insensitive IPA CER on a 16,660-utterance 41-language test set, outperforming a 575M-parameter PhoneticXEUS baseline at 9.78% under matched normalization.