Pith. sign in

REVIEW 4 cited by

DeepZip: Lossless Data Compression using Recurrent Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1811.08162 v1 pith:6KBSGSKA submitted 2018-11-20 cs.CL eess.SPq-bio.GN

classification cs.CLeess.SPq-bio.GN
keywords dataneuralcompressiondatasetsnetworkscompressgenomiclearn
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sequential data is being generated at an unprecedented pace in various forms, including text and genomic data. This creates the need for efficient compression mechanisms to enable better storage, transmission and processing of such data. To solve this problem, many of the existing compressors attempt to learn models for the data and perform prediction-based compression. Since neural networks are known as universal function approximators with the capability to learn arbitrarily complex mappings, and in practice show excellent performance in prediction tasks, we explore and devise methods to compress sequential data using neural network predictors. We combine recurrent neural network predictors with an arithmetic coder and losslessly compress a variety of synthetic, text and genomic datasets. The proposed compressor outperforms Gzip on the real datasets and achieves near-optimal compression for the synthetic datasets. The results also help understand why and where neural networks are good alternatives for traditional finite context models

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-based Source Code Compression via Thresholded Symbol Ranking

    cs.IT 2026-07 conditional novelty 6.0 of 10

    Bounding LLM next-token ranks to T=63 (or T=1) with escaped exceptions beats unbounded LLM ranking by up to 37% ratio and 40% speed on source code, and beats general-purpose compressors by up to 82% ratio.

  2. The 2026 Algorithmic Information Theory Data Compression Challenge

    cs.IT 2026-06 unverdicted novelty 6.0 of 10

    The paper introduces and reports results from a new benchmark challenge for general-purpose lossless data compressors using public training and hidden test sets.

  3. PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database

    cs.LG 2025-07 conditional novelty 5.0 of 10

    PMKLC is a parallel multi-knowledge learning-based genomic compressor that achieves the best average compression ratio and throughput among tested baselines, with gains over the strongest baseline under 1%.

  4. TextEconomizer: Enhancing Lossy Text Compression with Denoising Transformers and Entropy Coding

    cs.CL 2026-06 unverdicted novelty 3.0 of 10

    Presents TextEconomizer, a transformer-based encoder-decoder for lossy text compression claiming 5.39x ratio, near-perfect semantic quality via standard metrics, and 153x fewer parameters than comparables.

Pith tools