Pith. sign in

REVIEW 2 cited by

Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.16607 v4 pith:UJXXSEMU submitted 2024-07-23 cs.CL cs.LG

classification cs.CLcs.LG
keywords datatokenizerstrainingmixturetokenizerinferenceinformationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The pretraining data of today's strongest language models is opaque; in particular, little is known about the proportions of various domains or languages represented. In this work, we tackle a task which we call data mixture inference, which aims to uncover the distributional make-up of training data. We introduce a novel attack based on a previously overlooked source of information: byte-pair encoding (BPE) tokenizers, used by the vast majority of modern language models. Our key insight is that the ordered list of merge rules learned by a BPE tokenizer naturally reveals information about the token frequencies in its training data. Given a tokenizer's merge list along with example data for each category of interest, we formulate a linear program that solves for the proportion of each category in the tokenizer's training set. In controlled experiments, we show that our attack recovers mixture ratios with high precision for tokenizers trained on known mixtures of natural languages, programming languages, and data sources. We then apply our approach to off-the-shelf tokenizers released with recent LMs. We confirm much publicly disclosed information about these models, and also make several new inferences: GPT-4o and Mistral NeMo's tokenizers are much more multilingual than their predecessors, training on 39% and 47% non-English language data, respectively; Llama 3 extends GPT-3.5's tokenizer primarily for multilingual (48%) use; GPT-3.5's and Claude's tokenizers are trained on predominantly code (~60%). We hope our work sheds light on current design practices for pretraining data, and inspires continued research into data mixture inference for LMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Emojis in harmful prompts bypass LLM safety more effectively than plain text, across 7 models and 5 languages, through a heterogeneous tokenization channel.

  2. TokAlign: Efficient Vocabulary Adaptation via Token Alignment

    cs.CL 2025-06 conditional novelty 6.0 of 10

    TokAlign aligns source and target BPE token vocabularies using GloVe co-occurrence embeddings and re-initializes LLM embeddings, recovering within 5k steps and enabling token-level distillation.

Pith tools