Pith. sign in

REVIEW 4 cited by

Towards Image Understanding from Deep Compression without Decoding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1803.06131 v1 pith:RDWHSXAC submitted 2018-03-16 cs.CV

classification cs.CV
keywords compressedcompressionimagenetworksrepresentationsclassificationimagesmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Motivated by recent work on deep neural network (DNN)-based image compression methods showing potential improvements in image quality, savings in storage, and bandwidth reduction, we propose to perform image understanding tasks such as classification and segmentation directly on the compressed representations produced by these compression methods. Since the encoders and decoders in DNN-based compression methods are neural networks with feature-maps as internal representations of the images, we directly integrate these with architectures for image understanding. This bypasses decoding of the compressed representation into RGB space and reduces computational cost. Our study shows that accuracies comparable to networks that operate on compressed RGB images can be achieved while reducing the computational complexity up to $2\times$. Furthermore, we show that synergies are obtained by jointly training compression networks with classification networks on the compressed representations, improving image quality, classification accuracy, and segmentation performance. We find that inference from compressed representations is particularly advantageous compared to inference from compressed RGB images for aggressive compression rates.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision

    cs.CV 2025-01 conditional novelty 7.0 of 10

    UG-ICM trains a single learned image codec with global, local, and instance-level CLIP supervision and a preference-conditioned decoder to support both human viewing and unseen machine analytics tasks from one bitstream.

  2. From Raw Data to Structural Semantics: Trade-offs among Distortion, Rate, and Inference Accuracy

    cs.IT 2024-12 reject novelty 6.0 of 10

    Using persistence diagrams as transmitted semantic summaries can drastically cut bit rate and improve error robustness for topology-based classification, but the reported rate advantage rests on a self-defined metric.

  3. Faster and Accurate Classification for JPEG2000 Compressed Images in Networked Applications

    cs.CV 2019-09 conditional novelty 6.0 of 10

    A CNN can classify JPEG2000 images directly from their CDF 9/7 DWT coefficients, saving most of the decoding time while matching or slightly beating RGB-domain accuracy.

  4. Versatile Volumetric Medical Image Coding for Human-Machine Vision

    eess.IV 2024-12 conditional novelty 5.0 of 10

    A learned volumetric medical image codec transfers inter-slice latent features across slices so a single compressed stream supports both image reconstruction and direct organ segmentation.

Pith tools