Pith. sign in

REVIEW 5 cited by

Language Model Tokenizers Introduce Unfairness Between Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.15425 v2 pith:36CDGDZM submitted 2023-05-17 cs.CL cs.LG

classification cs.CLcs.LG
keywords languagedifferentlanguagesmodelsevensometokenizersmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent language models have shown impressive multilingual performance, even when not explicitly trained for it. Despite this, there are concerns about the quality of their outputs across different languages. In this paper, we show how disparity in the treatment of different languages arises at the tokenization stage, well before a model is even invoked. The same text translated into different languages can have drastically different tokenization lengths, with differences up to 15 times in some cases. These disparities persist even for tokenizers that are intentionally trained for multilingual support. Character-level and byte-level models also exhibit over 4 times the difference in the encoding length for some language pairs. This induces unfair treatment for some language communities in regard to the cost of accessing commercial language services, the processing time and latency, as well as the amount of content that can be provided as context to the models. Therefore, we make the case that we should train future language models using multilingually fair subword tokenizers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 29 citations worldwide. Full citation record

  1. Lower-Resource, Higher Scores: Language Bias in LLM Evaluators

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Multilingual LLM evaluators systematically inflate scores for lower-resource languages, and the standard pairwise-accuracy metric cannot detect the resulting safety-threshold disparities.

  2. Causal Estimation of Tokenisation Bias

    cs.CL 2025-06 conditional novelty 7.0 of 10

    Using regression discontinuity, the paper shows that adding a subword to a tokenizer's vocabulary can raise the model's probability for that string by up to about 17 times in small models.

  3. Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    DZEN, a parallel Dzongkha-English benchmark of 5,161 school science exam questions, shows large LLM accuracy gaps in Dzongkha; adding English translations narrows the gap.

  4. Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A layer-wise expert allocation algorithm based on hidden-state similarity, plus a routing classifier, improves parameter efficiency and reduces forgetting when expanding LLMs to new languages.

  5. BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

    cs.CL 2026-07 conditional novelty 4.0 of 10

    A 32K-token SentencePiece tokenizer trained on a balanced seven-language corpus cuts Indic text token counts by ~25% vs mBART-50 and ~90% vs GPT-2 on IKS sentences, with some term-level gains guaranteed by reserved vo...

Pith tools