Pith. sign in

REVIEW 1 cited by

An Analysis of BPE Vocabulary Trimming in Neural Machine Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00397 v1 pith:YCMVUKRJ submitted 2024-03-30 cs.CL

classification cs.CL
keywords subwordstrimmingvocabularymachinemodelperformanceraretokenization
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We explore threshold vocabulary trimming in Byte-Pair Encoding subword tokenization, a postprocessing step that replaces rare subwords with their component subwords. The technique is available in popular tokenization libraries but has not been subjected to rigorous scientific scrutiny. While the removal of rare subwords is suggested as best practice in machine translation implementations, both as a means to reduce model size and for improving model performance through robustness, our experiments indicate that, across a large space of hyperparameter settings, vocabulary trimming fails to improve performance, and is even prone to incurring heavy degradation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pruned BPE: Post-training Visibility Pruning and Token Reallocation for Byte Pair Encoding

    cs.CL 2026-08 conditional novelty 5.0 of 10

    Pruned BPE hides low-final-exposure BPE tokens as internal merge nodes and reallocates their visible vocabulary slots to better-exposed candidates from resumed training, reducing encoded length by about 0.27-0.36% at ...

Pith tools