Pith. sign in

REVIEW 1 major objections 6 minor 67 references

Explicit Boundary Markers for Subword Vocabularies

T0 review · 1 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Common words are stored twice by subword tokenizers, and explicit boundary markers remove the duplication and improve language modeling without improving compression.

desk verdict A well-argued tokenization proposal with a genuinely new marker scheme, but the headline LM gain may rest on an unspecified training budget. read the letter →

arxiv 2608.08847 v2 pith:55RQXMGN submitted 2026-08-09 cs.CL

classification cs.CL
keywords subwordtokenizationwordboundarymarkersleading-spaceduplicationcapitalizationcodespretokenizationlanguagemodelingbitsperbytevocabularycompression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Subword tokenizers store many common words twice in space-using writing systems, once with a leading space and once without, so the same word occupies separate embedding rows that are trained independently. This paper proposes replacing the leading-space convention with an explicit word-boundary marker, encoding spaces as pairs of markers and title or upper case with shift codes, before any vocabulary is learned. The scheme eliminates space-driven duplicate entries, leaves compression within roughly one percent of the baseline, and every marker variant tested achieves lower bits per byte in downstream language modeling. The paper concludes that duplication carries a cost that compression alone does not measure, making explicit boundary markers a drop-in pretokenization change worth adopting when a canonical word form is wanted.

What carries the argument

The central mechanism is an atomic boundary marker token (written ¦) placed on both sides of word, punctuation, or digit spans during pretokenization; a single space between two marked spans is removed, so a pair of adjacent markers becomes the encoding of a space, and decoding reverses this by replacing each marker pair with a space and deleting the remaining markers. Two additional atomic shift codes, ↑ for title case and ⇑ for upper case, are attached before a span's opening marker while the span is lowercased, letting case variants share the same entry. Invertibility rests on the rule that word spans are always marked on both sides and never lie adjacent without a real space, which is guaranteed by merging word runs across script changes and by marking only a fixed list of twenty space-using scripts.

What would settle it

Run the proposed encoding on a corpus containing a script outside the fixed list of twenty space-separating scripts and round-trip decode; if a single spurious space appears, the invertibility guarantee fails. Alternatively, reproduce the English bits-per-byte comparison at a larger model scale: if marker schemes no longer beat the baseline, the claimed uncaptured duplication cost is not a general effect.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that attaching a leading space to a word token creates systematic duplication, with a single word appearing in up to six forms once capitalization is counted, and that this duplication hurts language modeling even when tokenization compression is essentially unchanged. The proposed encoding marks word spans with an atomic boundary token, encodes inter-word spaces as pairs of those tokens, and uses title-case and upper-case codes so lower-case, title-case and upper-case forms can share one internal representation. Across six languages, two vocabulary-learning algorithms and three marker configurations, compression stays within one percent of the baseline for the punctuation- and digit-marking schemes, while all marker schemes tested improve bits per byte over the leading-space baseline at matched vocabulary size. The paper therefore claims that the benefit of removing duplication is real but invisible to compression metrics.

Load-bearing premise

The encoding is invertible only if two words never appear side by side without a real space between them; the method enforces this with a fixed list of twenty space-using scripts and by treating script changes within a word as one marking unit, so an unlisted or edge-case script could make decoding insert a spurious space.

Editorial extensions

If this is right

  • At matched vocabulary size, switching to explicit boundary markers lowers language-modeling bits per byte relative to the leading-space baseline, for both vocabulary-learning algorithms and for every marker configuration tested.
  • Because compression is essentially unchanged for the punctuation- and digit-marking schemes, the downstream gain is a separate effect of removing duplicated entries, not a side effect of better compression.
  • Space-driven duplicate vocabulary entries fall to zero by construction, and case-driven duplication drops as well when case codes are enabled.
  • The change is confined to pretokenization and leaves the vocabulary-learning algorithm untouched, so it can be applied to byte-pair-encoding-style, Unigram-style, and other tokenizers as a drop-in modification.
  • Words are emitted as single tokens far more often under the marker schemes than under the baseline, so morphological-alignment scores for whole words shift accordingly.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the proposed mechanism would count effective embedding updates per word form before and after marking: the duplication-cost account predicts that the rare form of a common word receives many more updates under the marker scheme.
  • The digit-marking scheme is the least settled part of the design; splitting digit runs rather than marking whole runs could remove the four-form duplication of numbers and may reclaim the small compression losses seen in the current results.
  • If the bits-per-byte gain persists at larger scale and in non-English languages, the leading-space convention itself becomes the suspect design choice, and tokenizer evaluations would need to include downstream modeling quality, not just compression.
  • The case-code idea suggests a natural extension to mixed-case spans, though invertibility would likely require a more expressive code than a single prefix marker.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 6 minor

Summary. The paper proposes an alternative to the standard leading-space convention in subword tokenization. Instead of attaching a space to the following word, it introduces an explicit boundary marker (and optional case-shift codes) emitted during pretokenization, so that words have a canonical form regardless of preceding whitespace and capitalization. The encoding is invertible. The authors evaluate five marking schemes on six languages with two vocabulary-learning algorithms (BPE and MinGram), measuring compression, morphological alignment (MorphScore), and downstream language-modeling bits per byte on a depth-12 nanochat model. They find that the best marker schemes are approximately compression-neutral, but that all schemes tested downstream achieve lower bits per byte than the leading-space baseline, which they interpret as evidence that vocabulary duplication carries a cost not captured by compression.

Significance. If the downstream result survives a matched-compute comparison, the paper offers a genuinely simple and practical pretokenization change: a drop-in modification applicable to any vocabulary-learning algorithm, with no change to the training objective or architecture, that reduces embedding duplication and improves language-modeling efficiency at a near-matched vocabulary size. The compression and duplication analyses over six languages and two tokenizers are careful, and the invertibility argument is well specified. The reported code availability and the use of standard deviations over three seeds are also strengths. However, the central headlining claim currently rests on a downstream comparison whose training budget is unspecified, so the significance is conditional.

major comments (1)
  1. [Section 5, Table 3] The sentence 'Loss is normalized by the text's true UTF-8 length, so schemes that emit different numbers of tokens stay comparable' only makes the evaluation metric comparable, not the training budget. The number of optimizer steps, training tokens, or training bytes used for the nanochat runs is not reported. Since boundary[w] has about 9% worse compression than plain (Table 2), a fixed-token budget yields about 9% fewer bytes of training text for boundary[w], whereas a fixed-byte budget yields more optimizer updates per byte; in neither case is the comparison controlled for compute. The reported 0.5–1% bits-per-byte advantage of the marker schemes could therefore reflect a training-compute difference rather than the representation itself. Please report the exact training budget and include a matched-compute comparison (matched bytes or matched FLOPs) to support the claim that duplication carries a cost that compression does not capture.
minor comments (6)
  1. [Section 1] The phrase 'Incl100k' is missing a space and should read 'In cl100k'.
  2. [Section 3.1] The example sentence 'the in "the is marked"¦the¦' is difficult to read; rephrase as 'the word "the" in "the is marked" is encoded as ¦the¦, giving the same pretoken as the span in "the cat".'
  3. [Appendix B.4, Table 6] In the encoding of 'the, cat', the printed string '¦the¦ ,¦ ¦cat¦' shows a space between the comma's marker and the following marker; since that space is removed by the scheme, please adjust the formatting or add a note so that the marker adjacency is not misinterpreted.
  4. [Section 4, Tokenizer training] The sentence 'Every tokenizer learns 32,768 tokens beyond its atomic alphabet, which differs by at most three entries between schemes' means that the total vocabulary sizes are not exactly equal; please state the total vocabulary size for each scheme explicitly, or use the phrase 'matched learned-vocabulary size' to be precise.
  5. [Section 5, MorphScore discussion] The reversal of the ordering when whole gold words are excluded deserves more analysis than the current one-sentence explanation; readers may otherwise interpret the high credit-mode scores as evidence of better morphological alignment, when the scores largely track the share of words left whole. Consider reporting the exclude-mode result as the primary metric or adding a brief analysis of why split words are aligned more poorly.
  6. [Section 5, Downstream evaluation] The claim that every scheme beats plain at p<0.01 is based on a two-sided paired t-test with only three shared seeds; please also report effect sizes and consider a permutation test, since the t-test's validity at this sample size is sensitive to distributional assumptions.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the marker-scheme comparison is controlled within the same SCRIPT framework, and no fitted parameter is renamed as a prediction; self-citations are motivational or implementational, not load-bearing.

full rationale

The paper's central empirical claims are measurements rather than derivations. All schemes share the same SCRIPT base encoding and differ only in which spans receive markers, so the plain baseline is a within-framework control rather than an input that forces the marker results. The downstream bits-per-byte comparison is normalized by true UTF-8 length and is not derived from any fitted parameter; the finding that boundary[w] is best despite worse compression is an empirical outcome, not a prediction from compression. The MorphScore discussion is transparent: the credit setting rewards whole gold words, and the paper explicitly reports that excluding whole words reverses the ordering, so no metric result is hidden as a consequence of the method's definition. Self-citations (Land 2026b for inspiration, Land and Arnett 2025 for SCRIPT, Land 2026a and Land and Pinter 2026 for MinGram, and Land and Bartolo 2024 for under-trained tokens) are motivational, implementational, or external to the central comparison; none is invoked as a uniqueness theorem or as proof of the marker schemes' effectiveness. The invertibility argument is a proof from the stated span definitions, not a circular premise. The paper's own limitations admit the English-only, single-scale downstream evaluation and the unsettled digit-marking design, which are scope caveats rather than circularity. The skeptical concern about matched token vs. byte training budgets is a correctness/experimental-control risk, not a circularity of the derivation chain.

Assumptions & free parameters 0 free parameters · 3 assumptions · 3 invented entities

The method and its evaluation rest on Unicode category and script assignments, the space-using script list, and the invertibility of case folding for the tested languages. The marker tokens are design artifacts introduced by the paper rather than externally validated entities. No numerical parameters are fitted to data.

assumptions (3)
  • domain assumption Unicode general category and script properties, plus the fixed list of 20 space-using scripts, are accurate and stable.
    Marker placement in Section B.1/B.3 depends on these assignments; an error could leave scripts unmarked or break invertibility.
  • domain assumption Two word spans are never adjacent without an intervening space after merging word runs across script changes.
    This property underlies the invertibility proof in Section 3.1; without it, adjacent boundary markers would be misread as an elided space.
  • domain assumption For the tested languages, lowercasing a span and re-applying the detected case pattern reconstructs the original span.
    Case codes in Section 3.2 rely on invertible case mapping; the paper notes the Turkish dotted-I exception, which is outside the test set.
invented entities (3)
  • Boundary marker token '¦'
    purpose: Delimits word, punctuation, and digit spans; a pair encodes an elided single space.
    Design artifact introduced and evaluated only in this paper; no external validation outside its experiments.
  • Title case code token '↑'
    purpose: Marks a span entirely in title case so the lowercased word can be reused.
    Same: internal device, tested in one downstream configuration.
  • Upper case code token '⇑'
    purpose: Marks a span entirely in upper case so the lowercased word can be reused.
    Same: internal device, tested in one downstream configuration.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Explicit Boundary Markers for Subword Vocabularies." pith.science (2026). https://pith.science/paper/55RQXMGN

@misc{pith2026260808847,
  author       = {Pith},
  title        = {Pith review of: Explicit Boundary Markers for Subword Vocabularies},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/55RQXMGN}},
  note         = {Machine review of arXiv:2608.08847}
}
read the original abstract

Subword tokenizers represent many common words twice in space-using writing systems, once with a leading space and once without. The two entries have separate embeddings in models, so occurrences of one word are divided across rows that are trained independently, and the two forms need not even segment the string the same way: " together" may be a single entry while the same word without a preceding space is tokenized as "to|gether". Capitalization divides a word further, into as many as six forms. We introduce an alternative to standard whitespace conventions using an explicit word boundary marker, which prevents such duplication. Words are delimited by the boundary markers, and spaces between words are represented as pairs of such markers. Two shift codes do the same for title case and upper case, allowing one internal representation of a word to be re-used across different settings. Switching to this convention mitigates the duplicate-entry issue, but does not improve tokenization compression: for both vocabulary-learning algorithms, the best marker scheme stays within one percent of the baseline in characters per token, averaged across six languages. It does result in better language modeling performance. Every marker scheme tested downstream reaches lower bits per byte than the baseline, suggesting that duplication carries a cost that compression does not capture.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

67 extracted references · 24 canonical work pages

  1. [1]

    2025 , booktitle=

    Arnett, Catherine and Hudspeth, Marisa and O'Connor, Brendan , title =. 2025 , booktitle=

  2. [2]

    2019 , eprint=

    Novel Applications of Factored Neural Machine Translation , author=. 2019 , eprint=

  3. [3]

    Say Anything but This: When Tokenizer Betrays Reasoning in

    Navid Ayoobi and Marcus I Armstrong and Arjun Mukherjee , year=. Say Anything but This: When Tokenizer Betrays Reasoning in. 2601.14658 , archivePrefix=

  4. [4]

    Improving Tokenisation by Alternative Treatment of Spaces

    Gow-Smith, Edward and Tayyar Madabushi, Harish and Scarton, Carolina and Villavicencio, Aline. Improving Tokenisation by Alternative Treatment of Spaces. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. doi:10.18653/v1/2022.emnlp-main.786

  5. [5]

    Tokenization Falling Short: On Subword Robustness in Large Language Models

    Chai, Yekun and Fang, Yewei and Peng, Qiwei and Li, Xuhong. Tokenization Falling Short: On Subword Robustness in Large Language Models. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024. doi:10.18653/v1/2024.findings-emnlp.86

  6. [6]
  7. [7]

    Lindsey, Jack and Gurnee, Wes and Ameisen, Emmanuel and Chen, Brian and Pearce, Adam and Turner, Nicholas L. and Citro, Craig and Abrahams, David and Carter, Shan and Hosmer, Basil and Marcus, Jonathan and Sklar, Michael and Templeton, Adly and Bricken, Trenton and McDougall, Callum and Cunningham, Hoagy and Henighan, Thomas and Jermyn, Adam and Jones, An...

  8. [8]

    doi:10.5281/zenodo.7199437 , url =

    Robyn Speer , title =. doi:10.5281/zenodo.7199437 , url =

Show all 67 references
  1. [9]

    2026 , month = jul, day =

    Land, Sander , title =. 2026 , month = jul, day =

  2. [10]

    2022 , url =

    OpenAI , title =. 2022 , url =

  3. [11]

    Guilherme Penedo and Hynek Kydl. The. 2024 , eprint =

  4. [12]

    Rosetta Code , howpublished =

  5. [13]

    and Reddy, Varshini and Zhang, Haoran and Alameddine, Alec and Uzan, Omri and Pinter, Yuval and Tanner, Chris

    Schmidt, Craig W. and Reddy, Varshini and Zhang, Haoran and Alameddine, Alec and Uzan, Omri and Pinter, Yuval and Tanner, Chris. Tokenization Is More Than Compression. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1...

  6. [14]

    and Tanner, Chris and Pinter, Yuval

    Uzan, Omri and Schmidt, Craig W. and Tanner, Chris and Pinter, Yuval. Greed is All You Need: An Evaluation of Tokenizer Inference Methods. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2024. doi:10.18653/v1/20...

  7. [15]

    Tokenization and the Noiseless Channel

    Zouhar, Vil \'e m and Meister, Clara and Gastaldi, Juan and Du, Li and Sachan, Mrinmaya and Cotterell, Ryan. Tokenization and the Noiseless Channel. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18...

  8. [16]

    Two Counterexamples to Tokenization and the Noiseless Channel

    Cognetta, Marco and Zouhar, Vil \'e m and Moon, Sangwhan and Okazaki, Naoaki. Two Counterexamples to Tokenization and the Noiseless Channel. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). 2024

  9. [17]

    NLLB Team and Costa-juss \`a , Marta R. and Cross, James and C elebi, Onur and Elbayad, Maha and Heafield, Kenneth and Heffernan, Kevin and Kalbassi, Elahe and Lam, Janice and Licht, Daniel and Maillard, Jean and Sun, Anna and Wang, Skyler and Wenzek, Guillaume and Youngblood,...

  10. [18]

    Byte Pair Encoding is Suboptimal for Language Model Pretraining

    Bostrom, Kaj and Durrett, Greg. Byte Pair Encoding is Suboptimal for Language Model Pretraining. Findings of the Association for Computational Linguistics: EMNLP 2020. 2020. doi:10.18653/v1/2020.findings-emnlp.414

  11. [19]

    Incorporating Context into Subword Vocabularies

    Yehezkel, Shaked and Pinter, Yuval. Incorporating Context into Subword Vocabularies. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. doi:10.18653/v1/2023.eacl-main.45

  12. [20]

    Rethinking Tokenization for Rich Morphology: The Dominance of U nigram over BPE and Morphological Alignment

    Vemula, Saketh Reddy and Dandapat, Sandipan and Sharma, Dipti and Krishnamurthy, Parameswari. Rethinking Tokenization for Rich Morphology: The Dominance of U nigram over BPE and Morphological Alignment. The 14th International Joint Conference on Natural Language Processing and...

  13. [21]

    2502.00894 , archivePrefix=

    Ehsaneddin Asgari and Yassine El Kheir and Mohammad Ali Sadraei Javaheri , year=. 2502.00894 , archivePrefix=

  14. [22]

    Which Pieces Does U nigram Tokenization Really Need?

    Land, Sander and Pinter, Yuval. Which Pieces Does U nigram Tokenization Really Need?. Findings of the A ssociation for C omputational L inguistics: ACL 2026. 2026. doi:10.18653/v1/2026.findings-acl.316

  15. [23]

    2025 , publisher =

    Guilherme Penedo , title =. 2025 , publisher =

  16. [24]

    2020 , eprint=

    BPE-Dropout: Simple and Effective Subword Regularization , author=. 2020 , eprint=

  17. [25]

    Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features

    Stephen, Abishek and Libovick \'y , Jind r ich. Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features. Findings of the A ssociation for C omputational L inguistics: EACL 2026. 2026. doi:10.18653/v1/2026.findings-eacl.196

  18. [26]

    2025 , eprint=

    Evaluating Morphological Alignment of Tokenizers in 70 Languages , author=. 2025 , eprint=

  19. [27]

    A. P. Dempster and N. M. Laird and D. B. Rubin , journal =. Maximum Likelihood from Incomplete Data via the

  20. [28]

    SentencePiece : A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

    Kudo, Taku and Richardson, John. SentencePiece : A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 2018. doi:10.18653/v1...

  21. [29]

    Combinatorial Pattern Matching , pages=

    Linear-Time Longest-Common-Prefix Computation in Suffix Arrays and Its Applications , author=. Combinatorial Pattern Matching , pages=. 2001 , organization=

  22. [30]

    ICML 2025 Tokenization Workshop , url=

    Sander Land and Catherine Arnett , year=. ICML 2025 Tokenization Workshop , url=

  23. [31]

    BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training

    Chizhov, Pavel and Arnett, Catherine and Korotkova, Elizaveta and Yamshchikov, Ivan P. BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024....

  24. [32]

    and Bergen, Benjamin

    Arnett, Catherine and Chang, Tyler A. and Bergen, Benjamin. Proceedings of the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages @ LREC-COLING 2024. 2024

  25. [33]

    Neural Machine Translation of Rare Words with Subword Units

    Sennrich, Rico and Haddow, Barry and Birch, Alexandra. Neural Machine Translation of Rare Words with Subword Units. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. doi:10.18653/v1/P16-1162

  26. [34]

    Proceedings of the 31st International Conference on Computational Linguistics

    Velayuthan, Menan and Sarveswaran, Kengatharaiyer. Proceedings of the 31st International Conference on Computational Linguistics. 2025

  27. [35]

    Pre-tokenization on punctuation in

    Sander Land , year=. Pre-tokenization on punctuation in

  28. [36]

    Catherine Arnett , year=

  29. [37]

    A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages

    Ortiz Suarez, Pedro Javier and Romary, Laurent and Sagot, Benoit. A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020

  30. [38]

    Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures , series =

    Pedro Javier. Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures , series =. 2019 , language =. doi:10.14618/ids-pub-9021 , url =

  31. [39]

    Rossi and Thien Huu Nguyen , year=

    Thuat Nguyen and Chien Van Nguyen and Viet Dac Lai and Hieu Man and Nghia Trung Ngo and Franck Dernoncourt and Ryan A. Rossi and Thien Huu Nguyen , year=. 2309.09400 , archivePrefix=

  32. [40]

    Jamo-Level Subword Tokenization in Low-Resource K orean Machine Translation

    Lee, Junyoung and Cognetta, Marco and Moon, Sangwhan and Okazaki, Naoaki. Jamo-Level Subword Tokenization in Low-Resource K orean Machine Translation. Proceedings of the Eighth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2025). 2025

  33. [41]

    Fishing for M agikarp: Automatically Detecting Under-trained Tokens in Large Language Models

    Land, Sander and Bartolo, Max. Fishing for M agikarp: Automatically Detecting Under-trained Tokens in Large Language Models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.649

  34. [42]

    Rush , year=

    Thomas Wolf and Lysandre Debut and Victor Sanh and Julien Chaumond and Clement Delangue and Anthony Moi and Pierric Cistac and Tim Rault and Rémi Louf and Morgan Funtowicz and Joe Davison and Sam Shleifer and Patrick von Platen and Clara Ma and Yacine Jernite and Julien Plu an...

  35. [43]

    2017 , eprint=

    Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling , author=. 2017 , eprint=

  36. [44]

    Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates

    Kudo, Taku. Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. doi:10.18653/v1/P18-1007

  37. [45]

    BPE -Dropout: Simple and Effective Subword Regularization

    Provilkov, Ivan and Emelianenko, Dmitrii and Voita, Elena. BPE -Dropout: Simple and Effective Subword Regularization. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.170

  38. [46]

    Mielke and Zaid Alyafeai and Elizabeth Salesky and Colin Raffel and Manan Dey and Matthias Gallé and Arun Raja and Chenglei Si and Wilson Y

    Sabrina J. Mielke and Zaid Alyafeai and Elizabeth Salesky and Colin Raffel and Manan Dey and Matthias Gallé and Arun Raja and Chenglei Si and Wilson Y. Lee and Benoît Sagot and Samson Tan , year=. Between words and characters: A Brief History of Open-Vocabulary Modeling and To...

  39. [47]

    Too Much in Common: Shifting of Embeddings in Transformer Language Models and its Implications

    Bi \'s , Daniel and Podkorytov, Maksim and Liu, Xiuwen. Too Much in Common: Shifting of Embeddings in Transformer Language Models and its Implications. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lang...

  40. [48]

    2025 , url=

    Alisa Liu and Jonathan Hayase and Valentin Hofmann and Sewoong Oh and Noah A Smith and Yejin Choi , booktitle=. 2025 , url=

  41. [49]

    2025 , eprint=

    Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization , author=. 2025 , eprint=

  42. [50]

    2026 , url =

    Proceedings of the Fourteenth International Conference on Learning Representations , author =. 2026 , url =

  43. [51]

    Language Model Tokenizers Introduce Unfairness Between Languages , url =

    Petrov, Aleksandar and La Malfa, Emanuele and Torr, Philip and Bibi, Adel , booktitle =. Language Model Tokenizers Introduce Unfairness Between Languages , url =

  44. [52]

    Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models

    Ahia, Orevaoghene and Kumar, Sachin and Gonen, Hila and Kasai, Jungo and Mortensen, David and Smith, Noah and Tsvetkov, Yulia. Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natu...

  45. [53]

    Tokenizer Choice For LLM Training: Negligible or Crucial?

    Ali, Mehdi and Fromm, Michael and Thellmann, Klaudia and Rutmann, Richard and L \"u bbering, Max and Leveling, Johannes and Klug, Katrin and Ebert, Jan and Doll, Niclas and Buschhoff, Jasper and Jain, Charvi and Weber, Alexander and Jurkschat, Lena and Abdelwahab, Hammam and J...

  46. [54]

    and Lopes, Ant \'o nio V

    Lotz, Jonas F. and Lopes, Ant \'o nio V. and Peitz, Stephan and Setiawan, Hendra and Emili, Leonardo. Beyond Text Compression: Evaluating Tokenizers Across Scales. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). ...

  47. [55]

    Why do language models perform worse for morphologically complex languages?

    Arnett, Catherine and Bergen, Benjamin. Why do language models perform worse for morphologically complex languages?. Proceedings of the 31st International Conference on Computational Linguistics. 2025

  48. [56]

    How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models

    Rust, Phillip and Pfeiffer, Jonas and Vuli \'c , Ivan and Ruder, Sebastian and Gurevych, Iryna. How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics a...

  49. [57]

    Superbizarre Is Not Superb: Derivational Morphology Improves BERT ' s Interpretation of Complex Words

    Hofmann, Valentin and Pierrehumbert, Janet and Sch. Superbizarre Is Not Superb: Derivational Morphology Improves BERT ' s Interpretation of Complex Words. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint C...

  50. [58]

    2026 , eprint=

    Tokenisation via Convex Relaxations , author=. 2026 , eprint=

  51. [59]

    Proceedings of the 15th Language Resources and Evaluation Conference (LREC) , year=

    Goldfish: Monolingual Language Models for 350 Languages , author=. Proceedings of the 15th Language Resources and Evaluation Conference (LREC) , year=

  52. [60]

    2025 , isbn =

    Lian, Haoran and Xiong, Yizhe and Niu, Jianwei and Mo, Shasha and Su, Zhenpeng and Lin, Zijia and Chen, Hui and Han, Jungong and Ding, Guiguang , title =. 2025 , isbn =. doi:10.1609/aaai.v39i23.34633 , booktitle =

  53. [61]

    The U niversity of E dinburgh ' s Neural MT Systems for WMT 17

    Sennrich, Rico and Birch, Alexandra and Currey, Anna and Germann, Ulrich and Haddow, Barry and Heafield, Kenneth and Miceli Barone, Antonio Valerio and Williams, Philip. The U niversity of E dinburgh ' s Neural MT Systems for WMT 17. Proceedings of the Second Conference on Mac...

  54. [62]

    Second Conference on Language Modeling , year=

    Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier , author=. Second Conference on Language Modeling , year=

  55. [63]

    Investigating the Effectiveness of BPE : The Power of Shorter Sequences

    Gall \'e , Matthias. Investigating the Effectiveness of BPE : The Power of Shorter Sequences. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. ...

  56. [64]

    2026 , eprint=

    Tokenization with Split Trees , author=. 2026 , eprint=

  57. [65]

    Nemotron-

    Shizhe Diao and Yu Yang and Yonggan Fu and Xin Dong and Dan Su and Markus Kliegl and Zijia Chen and Peter Belcak and Yoshi Suhara and Hongxu Yin and Mostofa Patwary and Yingyan Lin and Jan Kautz and Pavlo Molchanov , year=. Nemotron-. 2504.13161 , archivePrefix=

  58. [66]

    2025 , publisher =

    Andrej Karpathy , title =. 2025 , publisher =

  59. [67]

    and Carmon, Yair and Dave, Achal and Schmidt, Ludwig and Shankar, Vaishaal , booktitle =

    Li, Jeffrey and Fang, Alex and Smyrnis, Georgios and Ivgi, Maor and Jordan, Matt and Gadre, Samir and Bansal, Hritik and Guha, Etash and Keh, Sedrick and Arora, Kushal and Garg, Saurabh and Xin, Rui and Muennighoff, Niklas and Heckel, Reinhard and Mercat, Jean and Chen, Mayee ...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.