REVIEW 3 cited by
Sandwiched Compression: Repurposing Standard Codecs with Neural Network Wrappers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Sandwiched Compression: Repurposing Standard Codecs with Neural Network Wrappers
read the original abstract
We propose sandwiching standard image and video codecs between pre- and post-processing neural networks. The networks are jointly trained through a differentiable codec proxy to minimize a given rate-distortion loss. This sandwich architecture not only improves the standard codec's performance on its intended content, but more importantly, adapts the codec to other types of image/video content and to other distortion measures. The sandwich learns to transmit ``neural code images'' that optimize and improve overall rate-distortion performance, with the improvements becoming significant especially when the overall problem is well outside of the scope of the codec's design. We apply the sandwich architecture to standard codecs with mismatched sources transporting different numbers of channels, higher resolution, higher dynamic range, computer graphics, and with perceptual distortion measures. The results demonstrate substantial improvements (up to 9 dB gains or up to 30\% bitrate reductions) compared to alternative adaptations. We establish optimality properties for sandwiched compression and design differentiable codec proxies approximating current standard codecs. We further analyze model complexity, visual quality under perceptual metrics, as well as sandwich configurations that offer interesting potentials in video compression and streaming.
Forward citations
Cited by 3 Pith papers
-
A Projection-Based Surrogate Gradient Interpretation for Neural Codec Wrappers
SCALED surrogate gradient is reinterpreted as a projection-based first-order local approximation of non-differentiable video codecs, enabling effective training of full neural wrappers with BD-Rate gains up to 23.59% on x264.
-
Differentiable Proxy Learning for Adaptive Quantization Control in H.264 Video Coding
A soft-indexed differentiable proxy of H.264 enables adaptive global and spatial QP control that yields BD-rate savings up to 17% for segmentation and 15% for MS-SSIM versus fixed-QP baselines.
-
SEAOTTER: Sensor Embedded Autoencoding with One-Time Transcode for Efficient Reconstruction
SEAOTTER pairs a frozen sensor autoencoder with a learnable JPEG color/quantization transcode to deliver 200:1 compression, 7x faster encoding and 3.5x faster decoding than AVIF while improving ImageNet accuracy and r...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.