REVIEW 4 major objections 4 minor 119 references
Beyond Text Compression: Evaluating Tokenizers Across Scales
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Small-model tokenizer rankings predict large-model performance on machine translation, where a 350M model with the right tokenizer matches a 2.7B model.
desk verdict A useful MT-specific empirical result, but the parameter/compute mismatch and overbroad claims need fixing before I'd trust it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the scale-consistency conjecture: if a tokenizer significantly affects model quality, its impact will manifest consistently across model scales. On the measurement side, the paper introduces four intrinsic metrics computed from the token-frequency distribution on task data, namely CARDINALITY (number of unique tokens), AUC (area under the log-log rank-frequency curve), SLOPE (the fitted power-law exponent), and POWER LAW (mean absolute deviation from a fitted power law), plus the standard COMPRESSION baseline. These feed a two-stage predictor: logistic regression or a linear-kernel SVM on pairwise metric differences, followed by a Bradley-Terry model that converts pairwise win probabilities into a global tokenizer ranking. The key identity at work is that a tokenizer whose token frequencies approximate a Zipfian power law is better aligned with natural-language statistics, and this alignment is what the metrics operationalize.
What would settle it
Retrain the same six tokenizers at a third scale, say 7B parameters, and recompute the tokenizer ranking on the same WMT21 translation pairs. If the 350M-to-2.7B ranking (Kendall's tau 0.87) does not persist at the larger jump, or if an English multiple-choice benchmark starts to show scale-consistent tokenizer effects beyond 2.7B, the scale-consistency conjecture and the task-specificity conclusion would need revision.
Extended reading notes
Core claim
The central claim is that tokenizer quality can be evaluated cheaply and reliably before large-scale training: rankings of tokenizers measured on 350M-parameter models predict rankings at 2.7B parameters with Kendall's tau of 0.87 for machine translation, at roughly 15% of the pretraining cost. The paper further claims that this scale-consistency does not hold for English-centric tasks, with multiple-choice benchmarks giving tau 0.33 and summarization giving -0.07, so the proxy method is task-specific. On the intrinsic side, the paper claims that deviation from a Zipfian power-law token-frequency distribution (POWER LAW) is the single most informative predictor of multilingual performance, that combining it with vocabulary cardinality and the slope of the rank-frequency curve in an SVM improves pairwise predictions, and that a Bradley-Terry model over those pairwise comparisons recovers the ground-truth tokenizer ranking across Czech, German, Russian, and Chinese. A direct corollary, stated in Section 4.3, is that a 350M model with the AYA23 tokenizer performs comparably to GPT-NEOX at 2.7B and surpasses GPT-2 at 2.7B, showing tokenizer choice can offset a roughly fivefold increase in parameters for translation.
Load-bearing premise
The entire proxy method rests on the assumption that tokenizer effects rank the same way at 350M and 2.7B parameters, an assumption the paper's own numbers confirm for machine translation but violate for English benchmarks and summarization.
Editorial extensions
If this is right
- Pretraining a 350M-parameter proxy instead of a 2.7B model cuts the compute of an extrinsic tokenizer evaluation by roughly 85%, making systematic tokenizer comparison feasible before large runs.
- For English-only modeling, tokenizer choice among the six evaluated has negligible downstream impact, so selection can focus on compression and inference speed rather than quality.
- For multilingual generation, tokenizer selection is a first-order decision: the AYA23 tokenizer at 350M matches or beats English-centric tokenizers at 2.7B on translation, but larger vocabularies slow inference.
- Intrinsic metrics based on Zipf's-law alignment, especially POWER LAW, predict multilingual performance better than the compression baseline, and combining CARDINALITY, POWER LAW, and SLOPE in a pairwise-then-Bradley-Terry framework yields reliable tokenizer rankings without training a model.
- Tokenizers that overrepresent high-frequency tokens or rely on a long tail of rare units are penalized by the Zipfian metrics, giving a concrete design target for new multilingual tokenizers.
Reading between the lines
- If scale-consistency holds for translation but not for English benchmarks, a plausible generalization is that tokenizer impact is scale-consistent exactly when the task stresses vocabulary coverage for the target language; one testable extension is to apply the 350M-proxy protocol to code generation or biomedical text, where the authors expect the metric mix to change.
- The Zipfian metrics are cheap enough to compute on unlabeled target-language text before any training, so a practical recipe suggested but not validated by the paper is to screen candidate tokenizers on a small sample of the target corpus and only pretrain proxies for the top few.
- The paper's single-seed, single-pretraining-pipeline design leaves open how much of the measured ranking is tokenizer effect versus optimization noise; a direct follow-up would re-run the AYA23-versus-GPT-2 comparison with several seeds to bound the ranking's stability.
- Because the AYA23 tokenizer's advantage shows up mainly on Chinese and other non-Latin scripts, the Zipfian metrics may be most informative for scripts where English-centric tokenizers fragment text heavily; testing on more languages beyond five would reveal whether the tau equals 0.87 result generalizes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper trains decoder-only transformer models at nominal 350M and 2.7B scales using six published tokenizers (four English-centric, two multilingual) on the FineWeb corpus, evaluates them on multiple-choice benchmarks, English summarization (X-SUM), and multilingual machine translation (four language pairs with English), and proposes four intrinsic metrics (CARDINALITY, AUC, SLOPE, POWER LAW) derived from Zipfian rank-frequency plots. It reports that tokenizer choice has little impact on English tasks, that a model using the AYA23 tokenizer outperforms larger English-tokenizer models in translation, and that the proposed metrics, alone or combined, can predict multilingual performance better than text compression. The paper closes with a two-stage framework (pairwise logistic/SVM plus Bradley-Terry aggregation) for intrinsic tokenizer ranking.
Significance. If the central claims were supported, the paper would offer a practical, low-cost way to select tokenizers for multilingual model development and new intrinsic diagnostics. The study is usefully broad: it spans six tokenizers, two scales, three task families, and four non-English languages across three scripts, with transparent pretraining and tuning details. However, the current evidence does not support the abstract's general claim of cross-scale predictability: scale consistency holds only for translation, the machine-translation results are confounded by vocabulary-size-induced changes in parameter count and compute, and the reported significance levels in the correlation analysis are not attainable with six tokenizers. The paper's strengths are its detailed reporting and its clear documentation of where tokenizer effects do and do not appear; these make the manuscript a useful starting point after substantial revision.
major comments (4)
- [Table 3] Table 3 reports Spearman correlations with n=6 tokenizers and marks several coefficients as significant at p<0.01 (e.g., COMPRESSION -0.59 for multiple-choice, CARDINALITY -0.79 for machine translation). With six data points, the exact two-tailed p-value for rho=-0.59 is approximately 0.22, and even rho=0.77 does not reach p<0.01; the critical rho for p<0.01 with n=6 is 1.0. The reported significance stars are therefore incorrect and the claim that the proposed metrics "correlate more strongly than text compression" lacks the stated statistical support. Please remove the asterisks, provide exact p-values (e.g., from permutation tests), or rephrase the correlation claims as descriptive effect sizes.
- [Section 4 and Tables 1/7] The claim in Section 4 that "the models vary only in their choice of tokenizer" is contradicted by Table 1 (and Table 7). The AYA23 350M model has 566M trainable parameters, 490 H100-hours, and 8.17 TFLOPs, whereas GPT-2 at the same nominal scale has 356M parameters, 210 H100-hours, and 5.58 TFLOPs. Thus the headline result that a "350M" AYA23 model surpasses a 2.7B GPT-2 model in translation conflates tokenizer choice with vocabulary size, embedding/softmax capacity, and training compute. The cross-scale tau=0.87 for machine translation in Table 3 may simply reflect that tokenizers with larger vocabularies are ordered the same way at both scales. To support the central tokenizer-selection claim, the study would need to control for parameter count or compute (e.g., by adjusting hidden size or training budget), or the claims must be substantially narrowed to "a model with a larger-vocabulary tokenizer and higher compute can outperform a smaller-vocabulary model with more parameters."
- [Section 3.1 and Table 3] The scale-consistency conjecture in Section 3.1 ("if a tokenizer significantly affects model quality, its impact will manifest consistently across different model scales") is not supported by the data: Table 3 shows across-scale Kendall's tau of 0.33 for multiple-choice benchmarks and -0.07 for summarization, with only machine translation reaching tau=0.87. The abstract's opening claim that "smaller models can accurately predict significant differences in tokenizer impact on larger models" is therefore overbroad. Either the paper should restrict its predictive claim to the translation scenario where consistency is actually observed, or it should provide a theoretical or empirical account of why consistency should hold selectively.
- [Abstract and Section 4.3] The abstract and Section 4.3 use the word "significant" without any variance estimate; each condition is trained with a single seed, so there is no confidence interval for the differences between tokenizers. The Limitations acknowledge this, but the abstract's "significant differences" and the comparison between AYA23 350M and GPT-2 2.7B are stated as if they were established effects. Please add uncertainty quantification or soften the wording to "reported differences."
minor comments (4)
- [Section 3.4] Section 3.4 states "we propose four new metrics"; CARDINALITY (number of unique tokens) is a standard measure in tokenizer analysis and is not new. Please rephrase to distinguish the genuinely new metrics from the existing one.
- [Figures 1 and 2] The legend in Figures 1 and 2 uses only color to distinguish the six tokenizers, making the curves difficult to separate in grayscale print. Consider adding markers or distinct line styles.
- [Table 6] Table 6 should state explicitly in the caption that Kendall's tau compares the Bradley-Terry predicted ranking with the ground-truth (MetricX-based) ranking, and should define the significance thresholds used for the asterisks.
- [Section 5.1] In Section 5.1, the F1 scores are reported as point estimates without confidence intervals or a chance-level baseline; given the small number of tokenizers and the correlated pairwise outcomes, please include a null baseline or a permutation interval to aid interpretation.
Circularity Check
No significant circularity: scale-consistency and intrinsic-metric claims rest on independent empirical evaluations.
full rationale
The paper's chain of reasoning is empirical and self-contained rather than definitional. The central scale-consistency claim (Section 3.1 conjecture; Table 3 'Across scales 0.87' for MT) is a post-hoc rank correlation between independently trained 350M and 2.7B models; no parameter is fitted to the 2.7B performances and then reported as a prediction. The intrinsic metrics (CARDINALITY, AUC, SLOPE, POWER LAW) are defined a priori from Zipf's law and token-frequency statistics on task training data (Section 3.4, Figure 1), not from downstream MetricX or chrF scores, so their correlation with translation quality is an independent test of the stated hypothesis. The pairwise and Bradley-Terry frameworks (Sections 5.1-5.2) use leave-one-tokenizer-out and leave-one-language-out splits, which are legitimate supervised evaluations. The only self-citation (Rust et al. 2023 in Related Work, co-authored by J.F. Lotz) is descriptive and not load-bearing. The paper's Limitations section candidly states that larger-scale trends are unverified and that seed/hyperparameter sensitivity was not explored. A separate correctness concern, not a circularity, is the claim in Section 4 that 'the models vary only in their choice of tokenizer,' since Table 1/7 show AYA23 at the 350M tier uses 566M parameters, 490 H100 hours, and 8.17 TFLOPs versus GPT-2's 356M/210h/5.58 TFLOPs; this confounds tokenizer identity with vocabulary-driven capacity/compute, but it does not make any derived quantity equal to its input by construction.
Assumptions & free parameters
free parameters (3)
- Upper rank cutoff for Zipf metrics =
log(rank) <= 6
- Intrinsic metric combination C+P+S =
CARDINALITY, POWER LAW, SLOPE
- Model hyperparameters for SVM/logistic regression =
cross-validated, not reported
assumptions (3)
- domain assumption Tokenizer effects are scale-consistent.
- domain assumption Zipf's law holds for natural language and token distributions should align with it.
- standard math Power-law behavior only applies above a minimum rank.
Cite this review
Pith. "Pith review of Beyond Text Compression: Evaluating Tokenizers Across Scales." pith.science (2026). https://pith.science/paper/EB5FDIVG
@misc{pith2026250603101,
author = {Pith},
title = {Pith review of: Beyond Text Compression: Evaluating Tokenizers Across Scales},
year = {2026},
howpublished = {\url{https://pith.science/paper/EB5FDIVG}},
note = {Machine review of arXiv:2506.03101}
}
read the original abstract
The choice of tokenizer can profoundly impact language model performance, yet accessible and reliable evaluations of tokenizer quality remain an open challenge. Inspired by scaling consistency, we show that smaller models can accurately predict significant differences in tokenizer impact on larger models at a fraction of the compute cost. By systematically evaluating both English-centric and multilingual tokenizers, we find that tokenizer choice has negligible effects on tasks in English but results in consistent performance differences in multilingual settings. We propose new intrinsic tokenizer metrics inspired by Zipf's law that correlate more strongly with downstream performance than text compression when modeling unseen languages. By combining several metrics to capture multiple aspects of tokenizer behavior, we develop a reliable framework for intrinsic tokenizer evaluations. Our work offers a more efficient path to informed tokenizer selection in future language model development.
Figures
Reference graph
Works this paper leans on
-
[1]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, and 110 others. 2024. https://doi.org/10.48550/arXiv.2404.14219 Phi-3 technical rep...
-
[2]
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, and Yulia Tsvetkov. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.614 Do all languages cost the same? tokenization in the era of commercial language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9904--9923, ...
-
[3]
Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ond r ej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina Espa \ n a-Bonet, Angela Fan, Christian Federmann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter, Kenneth Heafield, Christopher Homan, Matthias Huck, Kwabena Amponsah-Kaakyire, ...
2021
-
[4]
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max L \"u bbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Buschhoff, Charvi Jain, Alexander Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, and 2 others. 2024. https://doi.org/10.18653/v1/2024.fin...
-
[5]
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. https://doi.org/10.48550/arXiv.2311.16867 The falcon series of open language models . arXiv preprint
-
[6]
Norah Alzahrani, Hisham Alyahya, Yazeed Alnumay, Sultan AlRashed, Shaykhah Alsubaie, Yousef Almushayqih, Faisal Mirza, Nouf Alotaibi, Nora Al-Twairesh, Areeb Alowisheq, M Saiful Bari, and Haidar Khan. 2024. https://doi.org/10.18653/v1/2024.acl-long.744 When benchmarks are targets: Revealing the sensitivity of large language model leaderboards . In Proceed...
-
[7]
Reinald Kim Amplayo, Peter J Liu, Yao Zhao, and Shashi Narayan. 2023. https://openreview.net/forum?id=OIe3kpwl40D SMART : Sentences as basic units for text evaluation . In The Eleventh International Conference on Learning Representations
2023
-
[8]
Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Kelly Marchisio, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Phil Blunsom, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker. 2024. https://doi.org/10.48550/arXiv.2404.14619 Aya 23: Open weight releases to further multilingua...
Show all 119 references
- [9]
- [10]
-
[11]
Lo \"i c Barrault, Magdalena Biesialska, Ond r ej Bojar, Marta R. Costa-juss \`a , Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljube s i \'c , Christof Monz, Makoto Morishita, Ma...
2020
-
[12]
Lo \"i c Barrault, Ond r ej Bojar, Marta R. Costa-juss \`a , Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias M \"u ller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. https://doi.org/10.1...
2019 doi
-
[13]
Lisa Beinborn and Yuval Pinter. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.272 Analyzing cognitive plausibility of subword tokenization . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4478--4486, Singapore. Association ...
2023 doi
- [14]
-
[15]
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, and 1 others. 2023. Pythia: A suite for analyzing large language models across training and scali...
2023
-
[16]
Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, and 374 others
BigScience Workshop , Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi,...
-
[17]
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, and 1 others. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432--7439
2020
-
[18]
Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. https://doi...
2022 doi
-
[19]
Ond r ej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz. 2018. https://doi.org/10.18653/v1/W18-6401 Findings of the 2018 conference on machine translation ( WMT 18) . In Proceedings of the Third Conference ...
2018 doi
-
[20]
Kaj Bostrom and Greg Durrett. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.414 Byte pair encoding is suboptimal for language model pretraining . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4617--4624, Online. Association for Computa...
2020 doi
-
[21]
Ralph Allan Bradley and Milton E Terry. 1952. Rank analysis of incomplete block designs i: The method of paired comparisons. Biometrika, 39(3/4):324--345
1952
-
[22]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...
2020
-
[23]
Yekun Chai, Yewei Fang, Qiwei Peng, and Xuhong Li. 2024 a . https://doi.org/10.18653/v1/2024.findings-emnlp.86 Tokenization falling short: On subword robustness in large language models . In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 1582--159...
2024 doi
-
[24]
Yekun Chai, Qingyi Liu, Jingwu Xiao, Shuohuan Wang, Yu Sun, and Hua Wu. 2024 b . https://doi.org/10.18653/v1/2024.emnlp-main.182 Autoregressive pre-training on pixels and texts . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 3...
2024 doi
- [25]
- [26]
-
[27]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vino...
2023
-
[28]
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30
2017
- [29]
-
[30]
Aaron Clauset, Cosma Rohilla Shalizi, and M. E. J. Newman. 2009. https://doi.org/10.1137/070710111 Power-law distributions in empirical data . SIAM Review, 51(4):661--703
2009 doi
-
[31]
Marco Cognetta, Vil \'e m Zouhar, Sangwhan Moon, and Naoaki Okazaki. 2024. https://aclanthology.org/2024.lrec-main.1469/ Two counterexamples to tokenization and the noiseless channel . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Lang...
2024
-
[32]
Gautier Dagan, Gabriel Synnaeve, and Baptiste Rozi\` e re. 2024. Getting the most out of your tokenizer for pre-training and domain adaptation. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org
2024
-
[33]
Miguel Domingo, Mercedes Garc \' a-Mart \' nez, Alexandre Helle, Francisco Casacuberta, and Manuel Herranz. 2019. https://link.springer.com/chapter/10.1007/978-3-031-24337-0_38 How much does tokenization affect neural machine translation? In International Conference on Computa...
2019 doi
- [34]
-
[35]
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023. https://doi.org/10.18653/v1/2023.wmt-1.51 Results of ...
2023 doi
-
[36]
Matthias Gall \'e . 2019. https://doi.org/10.18653/v1/D19-1141 Investigating the effectiveness of BPE : The power of shorter sequences . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natu...
2019 doi
-
[37]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...
2024 doi
-
[38]
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein. 2024. https://openreview.net/forum?id=Y5inHAjMu0 Coercing LLM s to do and reveal (almost) anything . In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models
2024
-
[39]
Daniela Gerz, Ivan Vuli \'c , Edoardo Maria Ponti, Roi Reichart, and Anna Korhonen. 2018. https://doi.org/10.18653/v1/D18-1029 On the relation between linguistic typology and (limitations of) multilingual language modeling . In Proceedings of the 2018 Conference on Empirical M...
2018 doi
-
[40]
Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao, Idan Szpektor, and Reut Tsarfaty. 2024. https://doi.org/10.18653/v1/2024.findings-acl.134 Unpacking tokenization: Evaluating text compression and its correlation with model performance . In Findings of the Association for Comp...
2024 doi
-
[41]
Thamme Gowda and Jonathan May. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.352 Finding the optimal vocabulary size for neural machine translation . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3955--3964, Online. Association for Com...
2020 doi
-
[42]
Gregory Grefenstette. 1999. https://doi.org/10.1007/978-94-015-9273-4_9 Tokenization , pages 117--133. Springer Netherlands, Dordrecht
1999 doi
-
[43]
Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, and 2...
2024 doi
- [44]
-
[45]
Jonathan Hayase, Alisa Liu, Yejin Choi, Sewoong Oh, and Noah A. Smith. 2024. https://openreview.net/forum?id=0SRg6Cwx3h Data mixture inference attack: BPE tokenizers reveal training data compositions . In ICML 2024 Workshop on Foundation Models in the Wild
2024
-
[46]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. https://openreview.net/forum?id=d7KBjmI3GmQ Measuring massive multitask language understanding . In International Conference on Learning Representations
2021
-
[47]
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Thomas Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Sim...
2022
-
[48]
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Sch \"u tze. 2021. https://doi.org/10.18653/v1/2021.acl-long.279 Superbizarre is not superb: Derivational morphology improves BERT `s interpretation of complex words . In Proceedings of the 59th Annual Meeting of the Associati...
2021 doi
-
[49]
Valentin Hofmann, Hinrich Schuetze, and Janet Pierrehumbert. 2022. https://doi.org/10.18653/v1/2022.acl-short.43 An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers . In Proceedings of the 60th Annual Meeting of the Associ...
2022 doi
- [50]
-
[51]
Juraj Juraska, Mara Finkelstein, Daniel Deutsch, Aditya Siddhant, Mehdi Mirzazadeh, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.wmt-1.63 M etric X -23: The G oogle submission to the WMT 2023 metrics shared task . In Proceedings of the Eighth Conference on Machin...
2023 doi
-
[52]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. https://doi.org/10.48550/arXiv.2001.08361 Scaling laws for neural language models . arXiv preprint
-
[53]
Stav Klein and Reut Tsarfaty. 2020. https://doi.org/10.18653/v1/2020.sigmorphon-1.24 Getting the \# \# life out of living: How adequate are word-pieces for modelling complex morphology? In Proceedings of the 17th SIGMORPHON Workshop on Computational Research in Phonetics, Phon...
2020 doi
-
[54]
Tom Kocmi, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Nov \'a k, Ma...
2022
-
[55]
Taku Kudo. 2018. https://doi.org/10.18653/v1/P18-1007 Subword regularization: Improving neural network translation models with multiple subword candidates . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), page...
2018 doi
-
[56]
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. https://doi.org/10.18653/v1/D17-1082 RACE : Large-scale R e A ding comprehension dataset from examinations . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages...
2017 doi
-
[57]
Sander Land and Max Bartolo. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.649 Fishing for magikarp: Automatically detecting under-trained tokens in large language models . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 116...
2024 doi
-
[58]
Teven Le Scao, Thomas Wang, Daniel Hesslow, Stas Bekman, M Saiful Bari, Stella Biderman, Hady Elsahar, Niklas Muennighoff, Jason Phang, Ofir Press, Colin Raffel, Victor Sanh, Sheng Shen, Lintang Sutawika, Jaesung Tae, Zheng Xin Yong, Julien Launay, and Iz Beltagy. 2022. https:...
2022 doi
-
[59]
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, and 40 o...
2024
- [60]
-
[61]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. https://doi.org/10.18653/v1/2022.acl-long.229 T ruthful QA : Measuring how models mimic human falsehoods . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pa...
2022 doi
- [62]
-
[63]
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, Richard Fan, Yi Gu, Victor Miller, Yonghao Zhuang, Guowei He, Haonan Li, Fajri Koto, Liping Tang, Nikhil Ranjan, and 8 others. 2024. https://ope...
2024
-
[64]
Ilya Loshchilov and Frank Hutter. 2017. https://openreview.net/forum?id=Skq89Scxx SGDR : Stochastic gradient descent with warm restarts . In International Conference on Learning Representations
2017
-
[65]
Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled weight decay regularization . In International Conference on Learning Representations
2019
-
[66]
Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes
Lovish Madaan, Aaditya K. Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes. 2024. https://doi.org/10.48550/arXiv.2406.10229 Quantifying variance in evaluation benchmarks . arXiv preprint
-
[67]
Guerreiro, Ricardo Rei, Duarte M
Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, and André F. T. Martins. 2024. https://d...
-
[68]
Sachin Mehta, Mohammad Sekhavat, Qingqing Cao, Max Horton, Yanzi Jin, Frank Sun, Iman Mirzadeh, Mahyar Najibikohnehshahri, Dmitry Belenko, Peter Zatloukal, and Mohammad Rastegari. 2024. https://arxiv.org/abs/2404.14619 Openelm: An efficient language model family with open trai...
2024 arXiv
-
[69]
Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y
Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y. Lee, Benoît Sagot, and Samson Tan. 2021. https://doi.org/10.48550/arXiv.2112.10508 Between words and characters: A brief history of open-vocabulary m...
-
[70]
Isabel Moreno-Sánchez, Francesc Font-Clos, and Álvaro Corral. 2016. https://doi.org/10.1371/journal.pone.0147073 Large-scale analysis of zipf’s law in english texts . PLOS ONE, 11(1):1--19
2016 doi
-
[71]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.154 C row S -pairs: A challenge dataset for measuring social biases in masked language models . In Proceedings of the 2020 Conference on Empirical Methods in Na...
2020 doi
-
[72]
Cohen, and Mirella Lapata
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. https://doi.org/10.18653/v1/D18-1206 Don`t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization . In Proceedings of the 2018 Conference on Empirical Methods in Natura...
2018 doi
-
[73]
MEJ Newman. 2005. https://doi.org/10.1080/00107510500052444 Power laws, pareto distributions and zipf's law . Contemporary Physics, 46(5):323--351
2005 doi
-
[74]
OpenAI . 2023. Gpt-4 technical report
2023
-
[75]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, a...
2022
-
[76]
Guilherme Penedo, Hynek Kydl\' c ek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. https://proceedings.neurips.cc/paper_files/paper/2024/file/370df50ccfdf8bde18f8f9c2d9151bda-Paper-Datasets_and_Benchmarks_Track.pdf ...
2024
-
[77]
Aleksandar Petrov, Emanuele La Malfa, Philip Torr, and Adel Bibi. 2023. https://openreview.net/forum?id=78yDLKi95p Language model tokenizers introduce unfairness between languages . In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[78]
Piantadosi
Steven T. Piantadosi. 2014. https://doi.org/10.3758/s13423-014-0585-6 Zipf's word frequency law in natural language: A critical review and future directions . Psychonomic Bulletin & Review, 21(5):1112--1130
2014 doi
-
[79]
John C. Platt. 1999. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. In Advances in Large Margin Classifiers, pages 61--74. MIT Press
1999
-
[80]
Maja Popovi \'c . 2015. https://doi.org/10.18653/v1/W15-3049 chr F : character n-gram F -score for automatic MT evaluation . In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392--395, Lisbon, Portugal. Association for Computational Linguistics
2015 doi
-
[81]
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020. https://doi.org/10.18653/v1/2020.acl-main.170 BPE -dropout: Simple and effective subword regularization . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1882--1892, O...
2020 doi
-
[82]
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf Language models are unsupervised multitask learners
2019
-
[83]
Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott
Phillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott. 2023. https://openreview.net/forum?id=FkSp8VW8RjH Language modelling with pixels . In The Eleventh International Conference on Learning Representations
2023
-
[84]
Phillip Rust, Jonas Pfeiffer, Ivan Vuli \'c , Sebastian Ruder, and Iryna Gurevych. 2021. https://doi.org/10.18653/v1/2021.acl-long.243 How good is your tokenizer? on the monolingual performance of multilingual language models . In Proceedings of the 59th Annual Meeting of the ...
2021 doi
-
[85]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99--106
2021
-
[86]
Elizabeth Salesky, David Etter, and Matt Post. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.576 Robust open-vocabulary translation from visual text representations . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7235--725...
2021 doi
-
[87]
Jonne Saleva and Constantine Lignos. 2023. https://doi.org/10.18653/v1/2023.insights-1.7 What changes when you randomly choose BPE merge operations? not much. In Proceedings of the Fourth Workshop on Insights from Negative Results in NLP, pages 59--66, Dubrovnik, Croatia. Asso...
2023 doi
-
[88]
Craig W Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.40 Tokenization is more than compression . In Proceedings of the 2024 Conference on Empirical Methods in Natural Languag...
2024 doi
-
[89]
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. https://doi.org/10.18653/v1/2020.acl-main.704 BLEURT : Learning robust metrics for text generation . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881--7892, Online. Ass...
2020 doi
-
[90]
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. https://doi.org/10.18653/v1/P16-1162 Neural machine translation of rare words with subword units . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages ...
2016 doi
-
[91]
Claude E. Shannon. 1948. https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=6773024 A mathematical theory of communication . The Bell System Technical Journal, 27(3):379--423
1948
- [92]
- [93]
-
[94]
Jimin Sun, Patrick Fernandes, Xinyi Wang, and Graham Neubig. 2023. https://doi.org/10.18653/v1/2023.findings-eacl.128 A multi-dimensional evaluation of tokenizer-free multilingual pretrained models . In Findings of the Association for Computational Linguistics: EACL 2023, page...
2023 doi
-
[95]
Yintao Tai, Xiyang Liao, Alessandro Suglia, and Antonio Vergari. 2024. https://doi.org/10.18653/v1/2024.findings-acl.874 PIXAR : Auto-regressive language modeling in pixel space . In Findings of the Association for Computational Linguistics: ACL 2024, pages 14673--14695, Bangk...
2024 doi
- [96]
-
[97]
Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, and Ngai Wong. 2024. https://openreview.net/forum?id=sKCKPr8cRL Scaling laws with vocabulary: Larger models deserve larger vocabularies . In The Thirty-eighth Annual Conference on Neural In...
2024
-
[98]
Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler
Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler. 2023. https://openreview.net/forum?id=6ruVLB727MC UL 2: Unifying language learning paradigms . ...
2023
-
[99]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 a . https://arxiv.org/abs/2302.13971 LLa...
2023 arXiv
- [100]
-
[101]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html Attention is all you need . In Advances in Neural I...
2017
-
[102]
Shibo Wang and Pankaj Kanwar. 2019. https://cloud.google.com/blog/products/ai-machine-learning/bfloat16-the-secret-to-high-performance-on-cloud-tpus Bfloat16: The secret to high performance on cloud tpus . Blog Post
2019
-
[103]
Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.72 Frequency effects on syntactic rule learning in transformers . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 932--948...
2021 doi
-
[104]
Jason Wei, Najoung Kim, Yi Tay, and Quoc Le. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.963 Inverse scaling can become U -shaped . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 15580--15591, Singapore. Association for C...
2023 doi
-
[105]
Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. https://openreview.net/forum?id=yzkSU5zd...
2022
-
[106]
Haoran Xu, Young Jin Kim, Amr Sharaf, and Hany Hassan Awadalla. 2024 a . https://openreview.net/forum?id=farT6XXntP A paradigm shift in machine translation: Boosting translation performance of large language models . In The Twelfth International Conference on Learning Representations
2024
-
[107]
Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, and Young Jin Kim. 2024 b . https://openreview.net/forum?id=51iwkioZpn Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation . In F...
2024
-
[108]
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022. https://doi.org/10.1162/tacl_a_00461 B y T 5: Towards a token-free future with pre-trained byte-to-byte models . Transactions of the Association for Computa...
2022 doi
-
[109]
Shaked Yehezkel and Yuval Pinter. 2023. https://doi.org/10.18653/v1/2023.eacl-main.45 Incorporating context into subword vocabularies . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 623--635, Dubrovnik, Cr...
2023 doi
-
[110]
Zheng Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, Genta Winata, Stella Biderman, Edward Raff, Dragomir Radev, and Vassilina Nikoulina. 2023. https://doi....
2023 doi
-
[111]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://doi.org/10.18653/v1/P19-1472 H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4...
2019 doi
- [112]
- [113]
-
[114]
George K. Zipf. 1949. Human Behaviour and the Principle of Least Effort. Addison-Wesley
1949
-
[115]
George Kingsley Zipf. 1935. The Psychobiology of Language. Houghton-Mifflin, New York, NY, USA
1935
- [116]
-
[117]
Vil \'e m Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. 2023. https://doi.org/10.18653/v1/2023.acl-long.284 Tokenization and the noiseless channel . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (...
2023 doi
-
[118]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[119]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.