Enforcing canonical BPE tokenizations through conditioning or architectural constraints improves held-out likelihood for GPT-2 and Llama models.
Consider the following subcases characterizing the possible positions for this merge: (a) The merge is in a(t)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Language Models over Canonical Byte-Pair Encodings
Enforcing canonical BPE tokenizations through conditioning or architectural constraints improves held-out likelihood for GPT-2 and Llama models.