REVIEW 5 cited by
Language Model Tokenizers Introduce Unfairness Between Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent language models have shown impressive multilingual performance, even when not explicitly trained for it. Despite this, there are concerns about the quality of their outputs across different languages. In this paper, we show how disparity in the treatment of different languages arises at the tokenization stage, well before a model is even invoked. The same text translated into different languages can have drastically different tokenization lengths, with differences up to 15 times in some cases. These disparities persist even for tokenizers that are intentionally trained for multilingual support. Character-level and byte-level models also exhibit over 4 times the difference in the encoding length for some language pairs. This induces unfair treatment for some language communities in regard to the cost of accessing commercial language services, the processing time and latency, as well as the amount of content that can be provided as context to the models. Therefore, we make the case that we should train future language models using multilingually fair subword tokenizers.
Forward citations
Cited by 5 Pith papers
-
Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
Multilingual LLM evaluators systematically inflate scores for lower-resource languages, and the standard pairwise-accuracy metric cannot detect the resulting safety-threshold disparities.
-
Causal Estimation of Tokenisation Bias
Using regression discontinuity, the paper shows that adding a subword to a tokenizer's vocabulary can raise the model's probability for that string by up to about 17 times in small models.
-
Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models
DZEN, a parallel Dzongkha-English benchmark of 5,161 school science exam questions, shows large LLM accuracy gaps in Dzongkha; adding English translations narrows the gap.
-
Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts
A layer-wise expert allocation algorithm based on hidden-state similarity, plus a routing classifier, improves parameter efficiency and reduces forgetting when expanding LLMs to new languages.
-
BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis
A 32K-token SentencePiece tokenizer trained on a balanced seven-language corpus cuts Indic text token counts by ~25% vs mBART-50 and ~90% vs GPT-2 on IKS sentences, with some term-level gains guaranteed by reserved vo...
Discussion (0). Continue with ORCID to comment.