Pith. sign in

REVIEW 2 cited by

NVC-1B: A Large Neural Video Coding Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.19402 v1 pith:NBVSX4Y2 submitted 2024-07-28 cs.CV eess.IV

classification cs.CVeess.IV
keywords modelvideocodinglargeneuralcompressionarchitecturesmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emerging large models have achieved notable progress in the fields of natural language processing and computer vision. However, large models for neural video coding are still unexplored. In this paper, we try to explore how to build a large neural video coding model. Based on a small baseline model, we gradually scale up the model sizes of its different coding parts, including the motion encoder-decoder, motion entropy model, contextual encoder-decoder, contextual entropy model, and temporal context mining module, and analyze the influence of model sizes on video compression performance. Then, we explore to use different architectures, including CNN, mixed CNN-Transformer, and Transformer architectures, to implement the neural video coding model and analyze the influence of model architectures on video compression performance. Based on our exploration results, we design the first neural video coding model with more than 1 billion parameters -- NVC-1B. Experimental results show that our proposed large model achieves a significant video compression performance improvement over the small baseline model, and represents the state-of-the-art compression efficiency. We anticipate large models may bring up the video coding technologies to the next level.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Conditional Residual Coding with Explicit-Implicit Temporal Buffering for Learned Video Compression

    eess.IV 2025-08 conditional novelty 6.0 of 10

    A hybrid explicit-implicit temporal buffer, one decoded frame plus a 3-channel learned feature, achieves most of the coding gain of large feature buffers in conditional residual video coding.

  2. Neural Video Compression with Context Modulation

    eess.IV 2025-05 conditional novelty 6.0 of 10

    DCMVC modulates the propagated temporal context with an additional oriented context from the reference frame, reporting 10.1 percent bitrate savings over DCVC-FM and 22.7 percent over VVC on standard test sets.

Pith tools