Pith. sign in

REVIEW 3 cited by

AlphaZip: Neural Network-Enhanced Lossless Text Compression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.15046 v1 pith:XRDZEKMQ submitted 2024-09-23 cs.IT cs.AIcs.LGmath.IT

classification cs.ITcs.AIcs.LGmath.IT
keywords compressionneuraltextcompressinglosslessadaptivealgorithmsalphazip
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Data compression continues to evolve, with traditional information theory methods being widely used for compressing text, images, and videos. Recently, there has been growing interest in leveraging Generative AI for predictive compression techniques. This paper introduces a lossless text compression approach using a Large Language Model (LLM). The method involves two key steps: first, prediction using a dense neural network architecture, such as a transformer block; second, compressing the predicted ranks with standard compression algorithms like Adaptive Huffman, LZ77, or Gzip. Extensive analysis and benchmarking against conventional information-theoretic baselines demonstrate that neural compression offers improved performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-based Source Code Compression via Thresholded Symbol Ranking

    cs.IT 2026-07 conditional novelty 6.0 of 10

    Bounding LLM next-token ranks to T=63 (or T=1) with escaped exceptions beats unbounded LLM ranking by up to 37% ratio and 40% speed on source code, and beats general-purpose compressors by up to 82% ratio.

  2. Separate Source Channel Coding Is Still What You Need: An LLM-based Rethinking

    cs.IT 2025-01 conditional novelty 6.0 of 10

    LLM-based arithmetic coding plus ECCT-enhanced LDPC decoding makes separate source and channel coding competitive with, and in these tests superior to, joint source-channel coding for text under a total-energy comparison.

  3. GPT as a Monte Carlo Language Tree: A Probabilistic Perspective

    cs.CL 2025-01 conditional novelty 5.0 of 10

    GPT models trained on a corpus are shown to approximate a tree of empirical next-token probabilities derived from the same corpus, with alignment increasing with model size.

Pith tools