REVIEW 4 cited by
DeepZip: Lossless Data Compression using Recurrent Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sequential data is being generated at an unprecedented pace in various forms, including text and genomic data. This creates the need for efficient compression mechanisms to enable better storage, transmission and processing of such data. To solve this problem, many of the existing compressors attempt to learn models for the data and perform prediction-based compression. Since neural networks are known as universal function approximators with the capability to learn arbitrarily complex mappings, and in practice show excellent performance in prediction tasks, we explore and devise methods to compress sequential data using neural network predictors. We combine recurrent neural network predictors with an arithmetic coder and losslessly compress a variety of synthetic, text and genomic datasets. The proposed compressor outperforms Gzip on the real datasets and achieves near-optimal compression for the synthetic datasets. The results also help understand why and where neural networks are good alternatives for traditional finite context models
Forward citations
Cited by 4 Pith papers
-
LLM-based Source Code Compression via Thresholded Symbol Ranking
Bounding LLM next-token ranks to T=63 (or T=1) with escaped exceptions beats unbounded LLM ranking by up to 37% ratio and 40% speed on source code, and beats general-purpose compressors by up to 82% ratio.
-
The 2026 Algorithmic Information Theory Data Compression Challenge
The paper introduces and reports results from a new benchmark challenge for general-purpose lossless data compressors using public training and hidden test sets.
-
PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database
PMKLC is a parallel multi-knowledge learning-based genomic compressor that achieves the best average compression ratio and throughput among tested baselines, with gains over the strongest baseline under 1%.
-
TextEconomizer: Enhancing Lossy Text Compression with Denoising Transformers and Entropy Coding
Presents TextEconomizer, a transformer-based encoder-decoder for lossy text compression claiming 5.39x ratio, near-perfect semantic quality via standard metrics, and 153x fewer parameters than comparables.
Discussion (0). Continue with ORCID to comment.