Pith. sign in

REVIEW 1 cited by

Ranking LLMs by compression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14171 v1 pith:J3SM4MTU submitted 2024-06-20 cs.AI cs.CL

classification cs.AIcs.CL
keywords compressionlanguagelargemodelmodelscodinglengthllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We conceptualize the process of understanding as information compression, and propose a method for ranking large language models (LLMs) based on lossless data compression. We demonstrate the equivalence of compression length under arithmetic coding with cumulative negative log probabilities when using a large language model as a prior, that is, the pre-training phase of the model is essentially the process of learning the optimal coding length. At the same time, the evaluation metric compression ratio can be obtained without actual compression, which greatly saves overhead. In this paper, we use five large language models as priors for compression, then compare their performance on challenging natural language processing tasks, including sentence completion, question answering, and coreference resolution. Experimental results show that compression ratio and model performance are positively correlated, so it can be used as a general metric to evaluate large language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GPT as a Monte Carlo Language Tree: A Probabilistic Perspective

    cs.CL 2025-01 conditional novelty 5.0 of 10

    GPT models trained on a corpus are shown to approximate a tree of empirical next-token probabilities derived from the same corpus, with alignment increasing with model size.

Pith tools