REVIEW 71 cited by
Variational image compression with a scale hyperprior
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We describe an end-to-end trainable model for image compression based on variational autoencoders. The model incorporates a hyperprior to effectively capture spatial dependencies in the latent representation. This hyperprior relates to side information, a concept universal to virtually all modern image codecs, but largely unexplored in image compression using artificial neural networks (ANNs). Unlike existing autoencoder compression methods, our model trains a complex prior jointly with the underlying autoencoder. We demonstrate that this model leads to state-of-the-art image compression when measuring visual quality using the popular MS-SSIM index, and yields rate-distortion performance surpassing published ANN-based methods when evaluated using a more traditional metric based on squared error (PSNR). Furthermore, we provide a qualitative comparison of models trained for different distortion metrics.
Forward citations
Showing 60 of 71 Pith papers that cite this
-
Dual-Constrained Diffusion Image Compression for Operational Rate-Distortion-Perception Optimization
DCIC uses dual constraints on a diffusion decoder to realize adjustable RDP operating points in neural image compression without extra rate cost.
-
GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow
GVCC achieves the lowest LPIPS on UVG at bitrates down to 0.003 bpp by encoding stochastic innovations in a marginal-preserving stochastic process derived from a pretrained rectified-flow video model, with 65% LPIPS r...
-
Compress-Align-Detect: onboard change detection from unregistered images
A single neural network performs compression, co-registration, and change detection onboard a satellite, achieving F1 up to about 70% at low bitrates on simulated unregistered image pairs.
-
Compressed Feature Quality Assessment: Dataset and Baselines
The first compressed feature quality assessment benchmark is released, and three standard similarity metrics are shown to correlate inconsistently with task-level semantic distortion.
-
Deep Convolutional Compression for Massive MIMO CSI Feedback
DeepCMC is a convolutional autoencoder architecture that compresses CSI matrices while jointly optimizing compression rate and reconstruction quality, outperforming prior schemes at equivalent bit rates.
-
FARI: Robust One-Step Inversion for Watermarking in Diffusion Models
One-step adversarially LoRA-tuned inversion exploits low-curvature reverse trajectories to beat 50-step DDIM on watermark robustness after ~20 minutes of fine-tuning.
-
ECoNGS: Efficient Compressive Neural Gaussian Splats for Volume Visualization
ECoNGS compresses volume-visualization scenes into entropy-coded neural Gaussian splats that are up to 6x smaller, train up to 6x faster, and render more accurately than the prior iVR-GS method.
-
Locality-Aware Density Control for Efficient Gaussian-based Image Representation
A locality-aware density-control framework for 2D Gaussian image representation that densifies coherent high-error regions and merges redundant similar Gaussians, improving PSNR at fixed budgets.
-
DCVC-MB: Neural B-Frame Video Compression using State Space Models
DCVC-MB, a neural B-frame video codec using Mamba state-space fusion, reports BD-rate savings up to 8.98% over prior neural codecs and up to 30.45% over VTM-19.0-LDP.
-
MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction
MambaRaw uses SSM-based context modeling with TileMambaBlock and EAR modules for efficient JPEG-guided 4K raw reconstruction, reporting 1.2-1.4 dB PSNR gains and 9% lower latency over baselines on Sony, Olympus, and S...
-
MoECodec: Image Compression for joint human and machine perception via Mixture-of-Experts
MoECodec replaces FFN layers with token-wise MoE plus stable routing and GShMLP experts to support multiple downstream tasks in a single image compression model.
-
Benchmarking Neural Speech Compression from a Rate-Distortion Perspective
ECC integrates hyperprior side information, channel-wise context, latent residual prediction, temporal modeling, and entropy skip into a learned entropy model, yielding 39.9% and 76.3% average BD-rate reductions on Vi...
-
Few-step Generative Models as Lossy Compression
Few-step generative models can be reformulated as lossy codecs in the reverse channel coding framework without retraining, yielding faster encoding/decoding on low-resolution image benchmarks.
-
Implicit Structural Modeling via Generative Diffusion Frameworks
Diffusion model trained on synthetic geological data and conditioned via a dedicated encoder generalizes from stylized normal faults to strike-slip, flower, and thrust nappe structures.
-
A Geometric Lens on Physics-Aligned Data Compression
Develops a local tangent-space rate-distortion theory and eigenspace-overlap diagnostic showing when physics-aligned compression necessarily degrades standard fidelity due to misaligned sensitivity directions.
-
Motion-Compensated Weight Compression
MCWC aligns permutation-symmetric blocks across layers to enable sequential prediction and residual entropy coding, improving rate-accuracy tradeoffs versus quantization and prior codecs on language and vision models.
-
Benchmarking and Enhancing VLM for Compressed Image Understanding
Introduces a benchmark for VLMs on compressed images and a universal adaptor to improve performance across codecs and bitrates.
-
HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression
A hyperprior predicts a Gaussian in codebook space and converts it to index probabilities, enabling content-adaptive entropy coding for VQ image compression.
-
SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
A semantic communication framework uses MLLM-derived scenario-aware importance labels to allocate coding resources, improving PSNR of important image regions at comparable or lower bandwidth than prior JSCC systems.
-
JPEG Processing Neural Operator for Backward-Compatible Coding
JPNeO improves JPEG compression with neural operators at encoding and decoding, without changing the JPEG bitstream format.
-
LINR-PCGC: Lossless Implicit Neural Representations for Point Cloud Geometry Compression
A lossless point cloud geometry codec overfits a small sparse-convolution network per group of frames and entropy-codes octree occupancy using the network's predicted probabilities.
-
CosmoFlow: Scale-Aware Representation Learning for Cosmology with Flow Matching
Flow matching on CDM fields yields an 8 number latent that reconstructs fields and estimates Omega_m and sigma_8 nearly as well as a raw-field network, with channels tied to spatial scales.
-
Generative Latent Diffusion for Efficient Spatiotemporal Data Reduction
A latent diffusion model conditioned on keyframe latents reconstructs non-key frames, giving higher compression ratios than prior scientific data compressors.
-
SIEDD: Shared-Implicit Encoder with Discrete Decoders
A shared encoder trained on a few video frames, followed by frozen-encoder parallel decoder training, cuts neural video encoding time by 20 to 30 times at similar quality.
-
StableCodec: Taming One-Step Diffusion for Extreme Image Compression
A one-step diffusion codec that compresses noisy latents at 64x and decodes with a single denoising step, setting state-of-the-art FID, KID, and DISTS at ultra-low bitrates.
-
GaussMarker: Robust Dual-Domain Watermark for Diffusion Models
GaussMarker embeds watermarks in both the spatial and frequency domains of initial diffusion noise and adds a learned restorer, reporting near-perfect detection across eight distortions and four attacks on three Stabl...
-
Generalized Gaussian Entropy Model for Point Cloud Attribute Compression with Dynamic Likelihood Intervals
A generalized Gaussian entropy model with dynamically adjusted likelihood intervals reduces bitrate by 6 to 11 percent across three point-cloud attribute compression baselines.
-
Distributed Image Semantic Communication via Nonlinear Transform Coding
D-NTSC and D-NTSCC, built on nonlinear transform coding with joint entropy modeling and spatial alignment, outperform existing distributed image transmission baselines on KITTI and Cityscapes.
-
Flexible Mixed Precision Quantization for Learned Image Compression
A rate-distortion sensitivity criterion assigns per-layer bit-widths, yielding about 1 to 2 percent BD-Rate improvement over 8-bit fixed-precision quantization at matched model size for learned image compression.
-
Generative Image Compression by Estimating Gradients of the Rate-variable Feature Distribution
A single rate-variable generative compression model treats quantization as a forward corruption and reverses it with a two-step denoiser, outperforming prior generative codecs on perceptual quality benchmarks.
-
S2CFormer: Revisiting the RD-Latency Trade-off in Transformer-based Learned Image Compression
Channel aggregation by feed-forward networks, not spatial attention, drives rate-distortion performance in transformer-based learned image compression, and simplified models achieve state-of-the-art results with over ...
-
Rate-Aware Learned Speech Compression
A rate-distortion trained speech codec using a channel-wise entropy model and CNN-RWKV blocks reports 53.51% average BD-rate savings over four baselines.
-
Towards Loss-Resilient Image Coding for Unstable Satellite Networks
A loss-resilient learned image codec using spatial-channel rearrangement, mask-conditional decoding, and Gilbert-Elliot loss simulation beats prior progressive codecs under satellite packet loss.
-
SNeRV: Spectra-preserving Neural Representation for Video
SNeRV decomposes frames with wavelet transforms, embeds only low-frequency content, and regenerates high-frequency details, outperforming prior NeRV models on reconstruction and interpolation.
-
Exploiting Latent Properties to Optimize Neural Codecs
Using uniform lattice grids for quantization and entropy-gradient latent shifting improves rate-distortion performance of off-the-shelf neural codecs by 1-3% without retraining.
-
AsymLLIC: Asymmetric Lightweight Learned Image Compression
AsymLLIC uses a two-stage training scheme to replace complex decoder modules with simpler ones, cutting decoder MACs to 51.47 GMACs while keeping RD performance close to VVC.
-
Point Cloud-Assisted Neural Image Compression
Point cloud depth projected onto the image, fused through a new attention module, improves learned image compression on KITTI by 54.5% BD-rate over Cheng2020.
-
Unicorn: Unified Neural Image Compression with One Number Reconstruction
An image set is compressed by training one conditional diffusion model to memorize the set, after which each image is transmitted as only a log2(M)-bit index while the decoder cost is amortized.
-
Vision Transformer-based Semantic Communications With Importance-Aware Quantization
A pretrained ViT's attention scores allocate quantization bits to image patches, improving classification accuracy per transmitted bit over uniform or top-k patch quantization.
-
Feature Coding in the Era of Large Models: Dataset, Test Conditions, and Benchmark
A public benchmark and unified test conditions for compressing intermediate features of large models, with two image-codec baselines evaluated.
-
LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model
LL-ICM jointly trains a neural image codec and a vision-language-guided diffusion restorer, claiming large bit-rate savings for low-level machine vision tasks.
-
DeepFGS: Fine-Grained Scalable Coding for Learned Image Compression
DeepFGS is a learned image codec that produces a single fine-grained scalable bitstream, outperforming prior scalable codecs on Kodak in PSNR and MS-SSIM.
-
Generalized Gaussian Model for Learned Image Compression
A generalized Gaussian entropy model with a learned shape parameter and two training fixes improves rate-distortion performance of learned image codecs compared to Gaussian and mixture models.
-
Robust Deep Joint Source-Channel Coding Enabled Distributed Image Transmission with Imperfect Channel State Information
RDJSCC introduces CVIE and CCF mechanisms to leverage source correlation for better image reconstruction in distributed transmission over severe fading channels with imperfect CSI.
-
Watermarking Visual Concepts for Diffusion Models
ConceptWM binds a watermark to a specific visual concept in diffusion model outputs and adds adversarial noise that degrades models fine-tuned on those watermarked images.
-
An End-to-End Real-World Camera Imaging Pipeline
A single end-to-end neural network performs RAW-to-RGB conversion and image compression jointly, reporting rate-distortion gains over separate ISP-plus-codec baselines.
-
Efficient Progressive Image Compression with Variance-aware Masking
A progressive image codec that ranks residual latent elements by predicted standard deviation and transmits them from most to least important, matching state-of-the-art RD with far lower decode cost.
-
Rethink Before You Execute: Adaptive Execution for World Action Models
TempoWAM adapts the replanning frequency of world action models based on an online estimate of task progress, reducing inference calls on easy tasks and improving success on hard tasks.
-
FrequencyFormer: A Co-Designed Sensor-to-Processor Pipeline for Frequency-Domain Vision Transformer Inference
FrequencyFormer co-designs a multi-scale DCT tokenizer, LUT-based near-sensor hardware, and modified MIPI communication to enable frequency-domain ViT inference with up to 128x data reduction and 230x lower communicat...
-
t-gems: text-guided exit modules for decreasing clip image encoder
Proposes T-GEMs plus a rate-based regularizer for early exits in CLIP encoders guided by text semantics to lower encoder usage costs.
-
Spectral and Spatial Graph Learning for Multispectral Solar Image Compression
A graph-based learned codec modeling wavelength-to-wavelength relationships plus windowed spatial attention reports modest PSNR/MS-SSIM gains and 20.15% lower MSID on six-channel solar images than two self-defined baselines.
-
SDGIC: A Semantic Disambiguation-Guided Generative Image Compression Method for Ultra-Low Bitrates
A diffusion-based image codec guided by text, a highly compressed image, and CLIP-derived semantic pseudo-words improves semantic consistency at bitrates below 0.05 bpp.
-
RAVQ-HoloNet: Rate-Adaptive Vector-Quantized Hologram Compression
RAVQ-HoloNet compresses phase-only holograms with a hierarchical VQ-VAE whose codebook size is shrunk by an LSTM module, enabling multiple bitrates from one model and claiming BD-Rate -33.91% vs DPRC.
-
Efficient Learned Image Compression Through Knowledge Distillation
Knowledge-distilled students with 64 or more channels match the rate-distortion performance of a 128-channel teacher while cutting memory by 68% and energy by 34%.
-
PVINet: Point-Voxel Interlaced Network for Point Cloud Compression
An interlaced point-voxel compression network with routing-weight-guided sparse convolutions reports 15.5% and 8% BD-rate savings over PCGCv2 and DeepPCC, but 22.9% worse than SparsePCGC.
-
DiSC-Med: Diffusion-based Semantic Communications for Robust Medical Image Transmission
A diffusion-based semantic communication framework for CT images that sends segmentation and edge maps as conditions and regenerates images at the receiver, with a channel-aware denoising module for noise robustness.
-
Fast Training-free Perceptual Image Compression
A noise-then-denoise decoder with a pre-trained diffusion model turns any existing codec into a fast, training-free perceptual codec with a KL-divergence guarantee and 0.1-10s decoding.
-
Compress image to patches for Vision Transformer
Using a frozen learned-compression encoder as the ViT patch embedder yields a 4x token reduction and 63% FLOP savings, with accuracy gains shown only on one small dataset.
-
CMamba: Learned Image Compression with State Space Models
A hybrid CNN and Mamba (state space model) image compression codec reports BD-Rate savings of 14.95% to 18.83% over VVC with fewer parameters, FLOPs, and lower decoding time than the prior best learned method.
-
Versatile Volumetric Medical Image Coding for Human-Machine Vision
A learned volumetric medical image codec transfers inter-slice latent features across slices so a single compressed stream supports both image reconstruction and direct organ segmentation.
Discussion (0). Continue with ORCID to comment.