Subword tokenization impairs phonological knowledge encoding in LMs, but an IPA-based fine-tuning method restores it with minimal impact on other capabilities.
PhonologyBench: Evaluating phonological skills of large language models,
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
Larger and Japanese-specialized LLMs outperform conventional morphological analyzers on Japanese G2P conversion, with parse-mode prompting yielding the lowest character error rates below 0.52%.
citing papers explorer
-
How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them
Subword tokenization impairs phonological knowledge encoding in LMs, but an IPA-based fine-tuning method restores it with minimal impact on other capabilities.
-
Benchmarking Large Language Models for Grapheme-to-Phoneme Conversion: A Japanese Case Study
Larger and Japanese-specialized LLMs outperform conventional morphological analyzers on Japanese G2P conversion, with parse-mode prompting yielding the lowest character error rates below 0.52%.