Pith. sign in

REVIEW 5 cited by

LLMZip: Lossless Text Compression using Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.04050 v2 pith:JOPDBC2U submitted 2023-06-06 cs.IT cs.CLcs.LGmath.IT

classification cs.ITcs.CLcs.LGmath.IT
keywords compressionlanguagelargelosslesstextciteenglishestimates
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We provide new estimates of an asymptotic upper bound on the entropy of English using the large language model LLaMA-7B as a predictor for the next token given a window of past tokens. This estimate is significantly smaller than currently available estimates in \cite{cover1978convergent}, \cite{lutati2023focus}. A natural byproduct is an algorithm for lossless compression of English text which combines the prediction from the large language model with a lossless compression scheme. Preliminary results from limited experiments suggest that our scheme outperforms state-of-the-art text compression schemes such as BSC, ZPAQ, and paq8h.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-based Source Code Compression via Thresholded Symbol Ranking

    cs.IT 2026-07 conditional novelty 6.0 of 10

    Bounding LLM next-token ranks to T=63 (or T=1) with escaped exceptions beats unbounded LLM ranking by up to 37% ratio and 40% speed on source code, and beats general-purpose compressors by up to 82% ratio.

  2. StateSMix: Online Lossless Compression via Mamba State Space Models and Sparse N-gram Context Mixing

    cs.LG 2026-04 conditional novelty 6.0 of 10

    An online-trained Mamba SSM mixed with sparse n-gram logit bias beats xz on enwik8 up to 10MB without pre-trained weights or a GPU.

  3. PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database

    cs.LG 2025-07 conditional novelty 5.0 of 10

    PMKLC is a parallel multi-knowledge learning-based genomic compressor that achieves the best average compression ratio and throughput among tested baselines, with gains over the strongest baseline under 1%.

  4. Channel-Agnostic Semantic Compression for Bandwidth-Limited Visual Communication

    cs.IT 2026-08 conditional novelty 4.0 of 10

    RQ-NAC, a residual-quantized image codec with n-gram arithmetic coding, reports 671 times compression over uncompressed dashcam frames at modest reconstruction quality (SSIM 0.74, PSNR 23.6 dB).

  5. Joint Lossless Compression and Steganography for Medical Images via Large Language Models

    eess.IV 2025-08 reject novelty 4.0 of 10

    A joint lossless compression and steganography framework for medical images that splits bit planes into a VAE-compressed global part and an LLM-compressed local part, embedding secret messages in the local part.

Pith tools