Pith. sign in

REVIEW 1 cited by

Unigram-Normalized Perplexity as a Language Model Performance Measure with Different Vocabulary Sizes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.13220 v1 pith:WUDNHOJB submitted 2020-11-26 cs.CL cs.LG

classification cs.CLcs.LG
keywords performancelanguagemodelperplexityvocabularycorpusdifferentmetric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although Perplexity is a widely used performance metric for language models, the values are highly dependent upon the number of words in the corpus and is useful to compare performance of the same corpus only. In this paper, we propose a new metric that can be used to evaluate language model performance with different vocabulary sizes. The proposed unigram-normalized Perplexity actually presents the performance improvement of the language models from that of simple unigram model, and is robust on the vocabulary size. Both theoretical analysis and computational experiments are reported.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dictionaries to the Rescue: Cross-Lingual Vocabulary Transfer for Low-Resource Languages Using Bilingual Dictionaries

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A dictionary-only vocabulary transfer method, built on iterative BPE subword removal, beats the FOCUS baseline on several low-resource languages, with the largest gains for Manchu.

Pith tools