REVIEW 2 cited by
PNVC: Towards Practical INR-based Video Compression
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Neural video compression has recently demonstrated significant potential to compete with conventional video codecs in terms of rate-quality performance. These learned video codecs are however associated with various issues related to decoding complexity (for autoencoder-based methods) and/or system delays (for implicit neural representation (INR) based models), which currently prevent them from being deployed in practical applications. In this paper, targeting a practical neural video codec, we propose a novel INR-based coding framework, PNVC, which innovatively combines autoencoder-based and overfitted solutions. Our approach benefits from several design innovations, including a new structural reparameterization-based architecture, hierarchical quality control, modulation-based entropy modeling, and scale-aware positional embedding. Supporting both low delay (LD) and random access (RA) configurations, PNVC outperforms existing INR-based codecs, achieving nearly 35%+ BD-rate savings against HEVC HM 18.0 (LD) - almost 10% more compared to one of the state-of-the-art INR-based codecs, HiNeRV and 5% more over VTM 20.0 (LD), while maintaining 20+ FPS decoding speeds for 1080p content. This represents an important step forward for INR-based video coding, moving it towards practical deployment. The source code will be available for public evaluation.
Forward citations
Cited by 2 Pith papers
-
Good, Cheap, and Fast: Overfitted Image Compression with Wasserstein Distortion
Optimizing the overfitted C3 codec with Wasserstein Distortion plus shared decoder noise yields generative-level perceptual quality at a fraction of the decoding cost.
-
BVI-CR: A Multi-View Human Dataset for Volumetric Video Compression
The BVI-CR dataset contributes 18 multi-view RGB-D human captures with textured meshes, plus a benchmark where INR codecs beat the MPEG TMIV anchor by up to 38.5% BD-rate.
Discussion (0). Continue with ORCID to comment.