Enforcing canonical BPE tokenizations through conditioning or architectural constraints improves held-out likelihood for GPT-2 and Llama models.
■ 26More generally, Lemma 3 holds for any tokenization model with a bigram-based canonicality test (i.e., δ ∈ D ⇐ ⇒BIGRAMS (δ) ⊆ B
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Language Models over Canonical Byte-Pair Encodings
Enforcing canonical BPE tokenizations through conditioning or architectural constraints improves held-out likelihood for GPT-2 and Llama models.