Pith. sign in

Language Model Tokenizers Introduce Unfairness Between Languages

7 Pith papers cite this work, alongside 29 external citations. Polarity classification is still indexing.

7 Pith papers citing it
29 external citations · external index

fields

cs.CL 6 cs.HC 1

years

2026 7

verdicts

UNVERDICTED 7

representative citing papers

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base

cs.CL · 2026-05-28 · unverdicted · novelty 7.0

BrahmicTokenizer-131K is a 131K-vocab tokenizer constructed via script-prune crop and linear-programming retrofit to o200k_base, achieving 26.7% fewer tokens on Indic text while matching o200k_base on English fertility and outperforming alternatives on code/math benchmarks.

Tokenization with Split Trees

cs.CL · 2026-05-21 · unverdicted · novelty 7.0

ToaST uses vocabulary-independent split trees and integer programming to produce tokenizers with over 11% fewer tokens than BPE, WordPiece, and UnigramLM while improving 1.5B-parameter LM scores on CORE.

citing papers explorer

Showing 7 of 7 citing papers.