Pith. sign in

REVIEW 71 cited by

Variational image compression with a scale hyperprior

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.01436 v2 pith:NPABLLZE submitted 2018-02-01 eess.IV cs.ITmath.IT

classification eess.IVcs.ITmath.IT
keywords compressionimagemodelhyperpriorautoencodermethodsvariationalwhen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We describe an end-to-end trainable model for image compression based on variational autoencoders. The model incorporates a hyperprior to effectively capture spatial dependencies in the latent representation. This hyperprior relates to side information, a concept universal to virtually all modern image codecs, but largely unexplored in image compression using artificial neural networks (ANNs). Unlike existing autoencoder compression methods, our model trains a complex prior jointly with the underlying autoencoder. We demonstrate that this model leads to state-of-the-art image compression when measuring visual quality using the popular MS-SSIM index, and yields rate-distortion performance surpassing published ANN-based methods when evaluated using a more traditional metric based on squared error (PSNR). Furthermore, we provide a qualitative comparison of models trained for different distortion metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 71 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. See all 71 Pith citations

  1. Dual-Constrained Diffusion Image Compression for Operational Rate-Distortion-Perception Optimization

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    DCIC uses dual constraints on a diffusion decoder to realize adjustable RDP operating points in neural image compression without extra rate cost.

  2. GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    GVCC achieves the lowest LPIPS on UVG at bitrates down to 0.003 bpp by encoding stochastic innovations in a marginal-preserving stochastic process derived from a pretrained rectified-flow video model, with 65% LPIPS r...

  3. Compress-Align-Detect: onboard change detection from unregistered images

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A single neural network performs compression, co-registration, and change detection onboard a satellite, achieving F1 up to about 70% at low bitrates on simulated unregistered image pairs.

  4. Compressed Feature Quality Assessment: Dataset and Baselines

    cs.CV 2025-06 conditional novelty 7.0 of 10

    The first compressed feature quality assessment benchmark is released, and three standard similarity metrics are shown to correlate inconsistently with task-level semantic distortion.

  5. Deep Convolutional Compression for Massive MIMO CSI Feedback

    cs.IT 2019-07 unverdicted novelty 7.0 of 10

    DeepCMC is a convolutional autoencoder architecture that compresses CSI matrices while jointly optimizing compression rate and reconstruction quality, outperforming prior schemes at equivalent bit rates.

  6. FARI: Robust One-Step Inversion for Watermarking in Diffusion Models

    cs.CR 2026-07 accept novelty 6.0 of 10

    One-step adversarially LoRA-tuned inversion exploits low-curvature reverse trajectories to beat 50-step DDIM on watermark robustness after ~20 minutes of fine-tuning.

  7. ECoNGS: Efficient Compressive Neural Gaussian Splats for Volume Visualization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    ECoNGS compresses volume-visualization scenes into entropy-coded neural Gaussian splats that are up to 6x smaller, train up to 6x faster, and render more accurately than the prior iVR-GS method.

  8. Locality-Aware Density Control for Efficient Gaussian-based Image Representation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A locality-aware density-control framework for 2D Gaussian image representation that densifies coherent high-error regions and merges redundant similar Gaussians, improving PSNR at fixed budgets.

  9. DCVC-MB: Neural B-Frame Video Compression using State Space Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    DCVC-MB, a neural B-frame video codec using Mamba state-space fusion, reports BD-rate savings up to 8.98% over prior neural codecs and up to 30.45% over VTM-19.0-LDP.

  10. MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    MambaRaw uses SSM-based context modeling with TileMambaBlock and EAR modules for efficient JPEG-guided 4K raw reconstruction, reporting 1.2-1.4 dB PSNR gains and 9% lower latency over baselines on Sony, Olympus, and S...

  11. MoECodec: Image Compression for joint human and machine perception via Mixture-of-Experts

    eess.IV 2026-06 unverdicted novelty 6.0 of 10

    MoECodec replaces FFN layers with token-wise MoE plus stable routing and GShMLP experts to support multiple downstream tasks in a single image compression model.

  12. Benchmarking Neural Speech Compression from a Rate-Distortion Perspective

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    ECC integrates hyperprior side information, channel-wise context, latent residual prediction, temporal modeling, and entropy skip into a learned entropy model, yielding 39.9% and 76.3% average BD-rate reductions on Vi...

  13. Few-step Generative Models as Lossy Compression

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Few-step generative models can be reformulated as lossy codecs in the reverse channel coding framework without retraining, yielding faster encoding/decoding on low-resolution image benchmarks.

  14. Implicit Structural Modeling via Generative Diffusion Frameworks

    physics.geo-ph 2026-06 unverdicted novelty 6.0 of 10

    Diffusion model trained on synthetic geological data and conditioned via a dedicated encoder generalizes from stylized normal faults to strike-slip, flower, and thrust nappe structures.

  15. A Geometric Lens on Physics-Aligned Data Compression

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Develops a local tangent-space rate-distortion theory and eigenspace-overlap diagnostic showing when physics-aligned compression necessarily degrades standard fidelity due to misaligned sensitivity directions.

  16. Motion-Compensated Weight Compression

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    MCWC aligns permutation-symmetric blocks across layers to enable sequential prediction and residual entropy coding, improving rate-accuracy tradeoffs versus quantization and prior codecs on language and vision models.

  17. Benchmarking and Enhancing VLM for Compressed Image Understanding

    cs.CV 2025-12 unverdicted novelty 6.0 of 10

    Introduces a benchmark for VLMs on compressed images and a universal adaptor to improve performance across codecs and bitrates.

  18. HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A hyperprior predicts a Gaussian in codebook space and converts it to index probabilities, enabling content-adaptive entropy coding for VQ image compression.

  19. SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding

    eess.SP 2025-09 conditional novelty 6.0 of 10

    A semantic communication framework uses MLLM-derived scenario-aware importance labels to allocate coding resources, improving PSNR of important image regions at comparable or lower bandwidth than prior JSCC systems.

  20. JPEG Processing Neural Operator for Backward-Compatible Coding

    eess.IV 2025-07 conditional novelty 6.0 of 10

    JPNeO improves JPEG compression with neural operators at encoding and decoding, without changing the JPEG bitstream format.

  21. LINR-PCGC: Lossless Implicit Neural Representations for Point Cloud Geometry Compression

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A lossless point cloud geometry codec overfits a small sparse-convolution network per group of frames and entropy-codes octree occupancy using the network's predicted probabilities.

  22. CosmoFlow: Scale-Aware Representation Learning for Cosmology with Flow Matching

    astro-ph.CO 2025-07 conditional novelty 6.0 of 10

    Flow matching on CDM fields yields an 8 number latent that reconstructs fields and estimates Omega_m and sigma_8 nearly as well as a raw-field network, with channels tied to spatial scales.

  23. Generative Latent Diffusion for Efficient Spatiotemporal Data Reduction

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A latent diffusion model conditioned on keyframe latents reconstructs non-key frames, giving higher compression ratios than prior scientific data compressors.

  24. SIEDD: Shared-Implicit Encoder with Discrete Decoders

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A shared encoder trained on a few video frames, followed by frozen-encoder parallel decoder training, cuts neural video encoding time by 20 to 30 times at similar quality.

  25. StableCodec: Taming One-Step Diffusion for Extreme Image Compression

    eess.IV 2025-06 conditional novelty 6.0 of 10

    A one-step diffusion codec that compresses noisy latents at 64x and decodes with a single denoising step, setting state-of-the-art FID, KID, and DISTS at ultra-low bitrates.

  26. GaussMarker: Robust Dual-Domain Watermark for Diffusion Models

    cs.CR 2025-06 conditional novelty 6.0 of 10

    GaussMarker embeds watermarks in both the spatial and frequency domains of initial diffusion noise and adds a learned restorer, reporting near-perfect detection across eight distortions and four attacks on three Stabl...

  27. Generalized Gaussian Entropy Model for Point Cloud Attribute Compression with Dynamic Likelihood Intervals

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A generalized Gaussian entropy model with dynamically adjusted likelihood intervals reduces bitrate by 6 to 11 percent across three point-cloud attribute compression baselines.

  28. Distributed Image Semantic Communication via Nonlinear Transform Coding

    cs.IT 2025-06 conditional novelty 6.0 of 10

    D-NTSC and D-NTSCC, built on nonlinear transform coding with joint entropy modeling and spatial alignment, outperform existing distributed image transmission baselines on KITTI and Cityscapes.

  29. Flexible Mixed Precision Quantization for Learned Image Compression

    eess.IV 2025-06 conditional novelty 6.0 of 10

    A rate-distortion sensitivity criterion assigns per-layer bit-widths, yielding about 1 to 2 percent BD-Rate improvement over 8-bit fixed-precision quantization at matched model size for learned image compression.

  30. Generative Image Compression by Estimating Gradients of the Rate-variable Feature Distribution

    eess.IV 2025-05 conditional novelty 6.0 of 10

    A single rate-variable generative compression model treats quantization as a forward corruption and reverses it with a two-step denoiser, outperforming prior generative codecs on perceptual quality benchmarks.

  31. S2CFormer: Revisiting the RD-Latency Trade-off in Transformer-based Learned Image Compression

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Channel aggregation by feed-forward networks, not spatial attention, drives rate-distortion performance in transformer-based learned image compression, and simplified models achieve state-of-the-art results with over ...

  32. Rate-Aware Learned Speech Compression

    eess.AS 2025-01 conditional novelty 6.0 of 10

    A rate-distortion trained speech codec using a channel-wise entropy model and CNN-RWKV blocks reports 53.51% average BD-rate savings over four baselines.

  33. Towards Loss-Resilient Image Coding for Unstable Satellite Networks

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A loss-resilient learned image codec using spatial-channel rearrangement, mask-conditional decoding, and Gilbert-Elliot loss simulation beats prior progressive codecs under satellite packet loss.

  34. SNeRV: Spectra-preserving Neural Representation for Video

    eess.IV 2025-01 conditional novelty 6.0 of 10

    SNeRV decomposes frames with wavelet transforms, embeds only low-frequency content, and regenerates high-frequency details, outperforming prior NeRV models on reconstruction and interpolation.

  35. Exploiting Latent Properties to Optimize Neural Codecs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Using uniform lattice grids for quantization and entropy-gradient latent shifting improves rate-distortion performance of off-the-shelf neural codecs by 1-3% without retraining.

  36. AsymLLIC: Asymmetric Lightweight Learned Image Compression

    eess.IV 2024-12 conditional novelty 6.0 of 10

    AsymLLIC uses a two-stage training scheme to replace complex decoder modules with simpler ones, cutting decoder MACs to 51.47 GMACs while keeping RD performance close to VVC.

  37. Point Cloud-Assisted Neural Image Compression

    eess.IV 2024-12 conditional novelty 6.0 of 10

    Point cloud depth projected onto the image, fused through a new attention module, improves learned image compression on KITTI by 54.5% BD-rate over Cheng2020.

  38. Unicorn: Unified Neural Image Compression with One Number Reconstruction

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An image set is compressed by training one conditional diffusion model to memorize the set, after which each image is transmitted as only a log2(M)-bit index while the decoder cost is amortized.

  39. Vision Transformer-based Semantic Communications With Importance-Aware Quantization

    eess.SP 2024-12 conditional novelty 6.0 of 10

    A pretrained ViT's attention scores allocate quantization bits to image patches, improving classification accuracy per transmitted bit over uniform or top-k patch quantization.

  40. Feature Coding in the Era of Large Models: Dataset, Test Conditions, and Benchmark

    cs.MM 2024-12 conditional novelty 6.0 of 10

    A public benchmark and unified test conditions for compressing intermediate features of large models, with two image-codec baselines evaluated.

  41. LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LL-ICM jointly trains a neural image codec and a vision-language-guided diffusion restorer, claiming large bit-rate savings for low-level machine vision tasks.

  42. DeepFGS: Fine-Grained Scalable Coding for Learned Image Compression

    eess.IV 2024-11 conditional novelty 6.0 of 10

    DeepFGS is a learned image codec that produces a single fine-grained scalable bitstream, outperforming prior scalable codecs on Kodak in PSNR and MS-SSIM.

  43. Generalized Gaussian Model for Learned Image Compression

    eess.IV 2024-11 conditional novelty 6.0 of 10

    A generalized Gaussian entropy model with a learned shape parameter and two training fixes improves rate-distortion performance of learned image codecs compared to Gaussian and mixture models.

  44. Robust Deep Joint Source-Channel Coding Enabled Distributed Image Transmission with Imperfect Channel State Information

    eess.SP 2024-11 unverdicted novelty 6.0 of 10

    RDJSCC introduces CVIE and CCF mechanisms to leverage source correlation for better image reconstruction in distributed transmission over severe fading channels with imperfect CSI.

  45. Watermarking Visual Concepts for Diffusion Models

    cs.CR 2024-11 conditional novelty 6.0 of 10

    ConceptWM binds a watermark to a specific visual concept in diffusion model outputs and adds adversarial noise that degrades models fine-tuned on those watermarked images.

  46. An End-to-End Real-World Camera Imaging Pipeline

    eess.IV 2024-11 conditional novelty 6.0 of 10

    A single end-to-end neural network performs RAW-to-RGB conversion and image compression jointly, reporting rate-distortion gains over separate ISP-plus-codec baselines.

  47. Efficient Progressive Image Compression with Variance-aware Masking

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A progressive image codec that ranks residual latent elements by predicted standard deviation and transmits them from most to least important, matching state-of-the-art RD with far lower decode cost.

  48. Rethink Before You Execute: Adaptive Execution for World Action Models

    cs.RO 2026-08 reject novelty 5.0 of 10

    TempoWAM adapts the replanning frequency of world action models based on an online estimate of task progress, reducing inference calls on easy tasks and improving success on hard tasks.

  49. FrequencyFormer: A Co-Designed Sensor-to-Processor Pipeline for Frequency-Domain Vision Transformer Inference

    eess.IV 2026-06 unverdicted novelty 5.0 of 10

    FrequencyFormer co-designs a multi-scale DCT tokenizer, LUT-based near-sensor hardware, and modified MIPI communication to enable frequency-domain ViT inference with up to 128x data reduction and 230x lower communicat...

  50. t-gems: text-guided exit modules for decreasing clip image encoder

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Proposes T-GEMs plus a rate-based regularizer for early exits in CLIP encoders guided by text semantics to lower encoder usage costs.

  51. Spectral and Spatial Graph Learning for Multispectral Solar Image Compression

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A graph-based learned codec modeling wavelength-to-wavelength relationships plus windowed spatial attention reports modest PSNR/MS-SSIM gains and 20.15% lower MSID on six-channel solar images than two self-defined baselines.

  52. SDGIC: A Semantic Disambiguation-Guided Generative Image Compression Method for Ultra-Low Bitrates

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A diffusion-based image codec guided by text, a highly compressed image, and CLIP-derived semantic pseudo-words improves semantic consistency at bitrates below 0.05 bpp.

  53. RAVQ-HoloNet: Rate-Adaptive Vector-Quantized Hologram Compression

    cs.LG 2025-11 conditional novelty 5.0 of 10

    RAVQ-HoloNet compresses phase-only holograms with a hierarchical VQ-VAE whose codebook size is shrunk by an LSTM module, enabling multiple bitrates from one model and claiming BD-Rate -33.91% vs DPRC.

  54. Efficient Learned Image Compression Through Knowledge Distillation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Knowledge-distilled students with 64 or more channels match the rate-distortion performance of a 128-channel teacher while cutting memory by 68% and energy by 34%.

  55. PVINet: Point-Voxel Interlaced Network for Point Cloud Compression

    cs.CV 2025-09 conditional novelty 5.0 of 10

    An interlaced point-voxel compression network with routing-weight-guided sparse convolutions reports 15.5% and 8% BD-rate savings over PCGCv2 and DeepPCC, but 22.9% worse than SparsePCGC.

  56. DiSC-Med: Diffusion-based Semantic Communications for Robust Medical Image Transmission

    cs.LG 2025-07 reject novelty 5.0 of 10

    A diffusion-based semantic communication framework for CT images that sends segmentation and edge maps as conditions and regenerates images at the receiver, with a channel-aware denoising module for noise robustness.

  57. Fast Training-free Perceptual Image Compression

    eess.IV 2025-06 conditional novelty 5.0 of 10

    A noise-then-denoise decoder with a pre-trained diffusion model turns any existing codec into a fast, training-free perceptual codec with a KL-divergence guarantee and 0.1-10s decoding.

  58. Compress image to patches for Vision Transformer

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Using a frozen learned-compression encoder as the ViT patch embedder yields a 4x token reduction and 63% FLOP savings, with accuracy gains shown only on one small dataset.

  59. CMamba: Learned Image Compression with State Space Models

    eess.IV 2025-02 conditional novelty 5.0 of 10

    A hybrid CNN and Mamba (state space model) image compression codec reports BD-Rate savings of 14.95% to 18.83% over VVC with fewer parameters, FLOPs, and lower decoding time than the prior best learned method.

  60. Versatile Volumetric Medical Image Coding for Human-Machine Vision

    eess.IV 2024-12 conditional novelty 5.0 of 10

    A learned volumetric medical image codec transfers inter-slice latent features across slices so a single compressed stream supports both image reconstruction and direct organ segmentation.

See all 71 Pith citations

Pith tools