Pith. sign in

REVIEW 2 major objections 5 minor 65 references

Augmented Deep Contexts for Spatially Embedded Video Coding

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read SEVC claims that embedding a 4×-downsampled spatial reference lets neural video codecs beat temporal-only prediction, cutting bitrate by 11.9% at equal quality.

desk verdict A credible spatially scalable neural video codec with honest 11.9% bitrate savings; the main open question is the undisclosed fine-tuning subset selection. read the letter →

arxiv 2505.05309 v1 pith:IXLOX62Y submitted 2025-05-08 eess.IV cs.CV

classification eess.IVcs.CV
keywords neuralvideocodingspatialscalabilityconditionaldeepcontextslatentpriorrate-distortionoptimizationlargemotionemergingobjects
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that a neural video codec no longer has to rely only on past frames: by first compressing a 4× downsampled version of the current frame, the codec can obtain spatial references that survive large motions and bring back objects that temporal references cannot describe. SEVC embeds a base low-resolution codec, then augments the low-resolution motion vectors, spatial feature, and latent representation with temporal information, and trains the whole system with a joint spatial-temporal loss. The authors report that this reduces bitrate by 11.9% more than the previous best learned codec at the same quality, while also emitting a separately decodable low-resolution bitstream. If true, this gives learned video coding a route past the temporal-reference ceiling without abandoning the conditional-coding paradigm.

What carries the argument

The load-bearing mechanism is the Motion and Feature Co-Augmentation (MFCA) module: a multi-scale, multi-stage loop in which each Augment Stage first sharpens the low-resolution motion vectors by adding a residual predicted from temporal features and the current spatial feature, then warps the temporal feature with the sharpened motion and uses it to refine the spatial feature. This progressive refinement generates the hybrid spatial-temporal contexts \($C_t^{1}$, $C_t^{2}$, $C_t^{3}$\). The second load-bearing device is the spatial-guided latent prior: the spatial latent \(\hat{y}_t^b\), upsampled to full resolution, serves as the query that aligns \(\hat{y}_{t-1}, \hat{y}_{t-2}, \hat{y}_{t-3}\) through Transformer blocks, replacing the single misaligned prior \(\hat{y}_{t-1}\). A joint spatial-temporal optimization with a small reconstruction constraint on the base layer lets the network learn how many bits the low-resolution stream should spend, rather than fixing that by hand.

What would settle it

Encode a fast-motion or object-appearance sequence with SEVC and with a version of SEVC in which the low-resolution spatial branch is replaced by a constant input (the average low-resolution color) while the base bits are reallocated to the full-resolution layer; if BD-rate over DCVC-DC does not degrade, the spatial reference is not the source of the gain. A second check is to blur the low-resolution input before encoding and see whether the reported 11.9% saving shrinks.

Watch

Extended reading notes

Core claim

The central discovery is that a lossy low-resolution reconstruction of the current frame can be converted into a high-value prediction asset rather than a separate bitstream burden. SEVC's MFCA module co-augments base motion vectors and a spatial feature by alternating residual prediction with temporal-feature alignment across three scales, yielding hybrid contexts that describe regions where motion estimation or temporal reference is unreliable. In parallel, the low-resolution latent, upsampled and used as a transformer query, aligns and fuses the previous three latent representations into a spatial-guided prior for the entropy model. The paper reports average BD-rate savings of 11.9% over DCVC-FM at an intra period of -1 and 8.3% over DCVC-DC at an intra period of 32, with the largest margins on sequences containing large motions or emerging objects.

Load-bearing premise

The design banks on the idea that a 4×-smaller, compressed copy of the current frame retains enough real spatial detail to guide full-resolution prediction; if that copy is too blurry or too lossy, the motion-and-feature augmentation and the latent prior query have nothing useful to add.

Editorial extensions

If this is right

  • On sequences with large motion or newly appearing objects, SEVC should show its largest BD-rate advantage over temporal-only codecs, since that is the regime the spatial branch is designed to repair.
  • A SEVC bitstream can be partially decoded into a low-resolution video, so fast preview and skimming become possible from the same compressed representation.
  • The spatial-guided prior should make rate estimation more accurate for frames that differ strongly from the previous frame, lowering bitrate at equal quality.
  • Base-layer bit allocation can be learned end-to-end, removing the need to hand-tune quality ratios between the low- and full-resolution layers.
  • The spatial-embedding strategy is an augmentation over a base codec, so improved future temporal-only codecs could inherit the same gain by being wrapped in the same structure.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A stress test on scene cuts and content swaps would likely amplify the reported gains, because temporal references become nearly useless there and the spatial branch carries almost the entire prediction load; the paper's tested sets contain such moments but do not isolate them.
  • The 4× downsampling factor is chosen, not proven optimal; a smaller factor would send more spatial detail at higher base cost, so the best trade-off may shift with resolution, bitrate, and content, and locating the optimum would require a sweep the paper does not report.
  • Because the base layer is a low-resolution decodable stream, SEVC has a natural fit for adaptive streaming use cases where a client first requests the low-resolution layer and later upgrades; the paper does not explore that deployment path.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. SEVC embeds a low-resolution base codec (DCVC-DC style) into a full-resolution learned video codec: the input frame is 4x downsampled and compressed, yielding base motion vectors, a spatial feature, and a spatial latent representation. These spatial references are then augmented with temporal references in two ways: the Motion and Feature Co-Augmentation (MFCA) module progressively refines base MVs and the spatial feature to produce hybrid spatial-temporal contexts, and a spatial-guided Transformer-based latent prior aligns and fuses multiple past latents. A joint spatial-temporal optimization loss adjusts base-layer bit allocation. Experiments report BD-Rate savings over VTM-13.2 and several learned codecs, including an 11.9% additional saving over DCVC-FM at IP=-1, together with targeted evidence on large-motion and emerging-object sequences. Code and trained models are released.

Significance. If the empirical claims hold, SEVC is a meaningful contribution: it transfers the spatial-reference idea from scalable and super-resolution coding into conditional NVC feature-space context generation, and it includes the base-layer bits in the total rate, so the reported savings are not obtained by hiding the extra low-resolution bitstream. The ablation structure is sensible and supports the individual design choices, and the public code and models strengthen reproducibility. The main caveat is that the headline result is produced by a fine-tuning stage whose training subset is not specified, which makes the central number hard to verify; the large-motion/emerging-object claim also rests on a very small set of sequences. These issues are fixable with disclosure or additional experiments, so the contribution is potentially publishable in a major revision.

major comments (2)
  1. [4.1 (Training Setup)] The joint optimization is conducted on "a selected subset of 9000 sequences from the original videos of the Vimeo-90k dataset," but the selection criterion is never given. Because the headline 11.9% BD-Rate improvement and the large-motion/emerging-object results are outcomes of this fine-tuning stage, the subset choice is load-bearing: if it was chosen to emphasize large-motion or emerging-object content, or to resemble the test sequences in Table 3, the reported gains could be inflated. Please state the exact selection criterion, specify whether the selection was random, and release the list of sequence indices; if the selection was content-based, repeat the fine-tuning on a random subset and report both sets of results.
  2. [4.2 and Table 3] The claim that SEVC "effectively alleviates the limitations in handling large motions or emerging objects" is supported by only three named sequences (USTC BicycleDriving, videoSRC21, BasketballDrive) with no a priori protocol for choosing them and no error bars. One of the three comes from USTC-TD, a dataset introduced by the same group. To make the claim convincing, the paper should either provide a systematic evaluation over a larger set of sequences with a defined threshold or annotation for "large motion" and "emerging object," or substantially soften the claim to a qualitative observation. Without this, the strong wording in the abstract and conclusion exceeds what the evidence supports.
minor comments (5)
  1. [Various] There are several typos: Section 2.1 has "the the superior potential," Section 3.1 has "persepecitive," and the supplementary material has "Euqation" and "Architechture." These should be corrected.
  2. [3.3] The statement that the hyperprior is discarded "due to similar characteristics of hyper encoder/decoder and our base codec" is asserted without an ablation. Please either add an experiment keeping the hyperprior or soften the claim.
  3. [4.3 (Table 5)] The bullet notation in Table 5 is ambiguous: it is not immediately clear which components are present in M1 and M4, especially because the text says M4 discards the spatial latent. Please spell out each baseline in words or use explicit check marks with a legend.
  4. [Equations (4) and (5)] Define R_t explicitly as the total bitrate including the base-layer bits. The text implies this, but a formal definition would prevent readers from misinterpreting the BD-Rate comparisons as excluding the low-resolution bitstream.
  5. [Tables 1-6] No error bars, confidence intervals, or repeated-seed results are reported. Some differences are small (for example, 1.5% between M1 and M2 in Table 5), so a statement about variance or a release of per-sequence numbers would strengthen the empirical claims.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation found; empirical benchmark claims stand on held-out tests, with a training-subset disclosure caveat.

full rationale

I walked the derivation chain. Equation (1), H(X) = H(X^b) + H(X|X^b), is a standard information-theoretic identity used only as motivation, and the supplementary derivation is correct; it does not define SEVC's performance in terms of itself. The reported BD-Rate gains (Tables 1-3) are measured on separate test sets (HEVC B-E, MCL-JCV, UVG, USTC-TD) against VTM, with the base-layer bits included in the total rate, so the 11.9% claim is not a fitted input renamed as a prediction. Hyperparameters such as lambda and w_l are fixed regularization/loss weights selected by ablations, not inverse-fitted to test RD values. Self-citations (e.g., [6] LSSVC, [32] USTC-TD, [51] Sheng-2024) are design or test-set references and are not used as a uniqueness theorem or to forbid alternatives; the base codec DCVC-DC [28] is external prior work. The one flagged issue is not circular: Section 4.1 says joint optimization is 'conducted on a selected subset of 9000 sequences from the original videos of the Vimeo-90k dataset' without specifying the selection criterion. This is a training-data selection and reproducibility concern that could bias the claimed gains, but it does not make any equation reduce to its own input or any fitted parameter masquerade as a prediction. Hence the low circularity score.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central RD gains rely on standard hyperparameters and architectural assumptions; no new physical entities are introduced. The free parameters are design choices rather than quantities fitted to the test data.

free parameters (5)
  • lambda (RD tradeoff) = {50, 95, 200, 400}
    Standard rate-distortion Lagrange multipliers used for training; chosen by hand and not derived.
  • wl (base reconstruction weight) = 0.05 in ablations; final value not stated
    Joint spatial-temporal loss weight; the paper reports best at 0.05 in ablation but does not clearly state which value is used for the final model.
  • number of augment stages = 2
    Chosen as a compromise between performance and complexity based on Table 4.
  • number of temporal latent representations = 3
    Queue of y_{t-1}, y_{t-2}, y_{t-3}; chosen as a design, complexity grows linearly.
  • downsampling factor = 4x
    Spatial reference resolution; chosen without a sensitivity study.
assumptions (5)
  • standard math H(X)=H(Xb)+H(X|Xb) for deterministic downsampling
    Supplementary Section C derives the identity given p(xb|x)=1; it serves as motivation for two-stage coding, not as part of the learned optimization.
  • domain assumption DCVC-DC reconstruction provides reliable spatial references
    Section 3.1 assumes the LR base codec's outputs are informative; the paper does not analyze LR reconstruction error propagation.
  • ad hoc to paper Transformer with spatial latent as query aligns temporal latents
    Section 3.3 borrows this from VSR and validates indirectly through ablation M2 vs M3; there is no theoretical guarantee of alignment for entropy priors.
  • domain assumption Removing the hyperprior does not degrade RD performance
    Section 3.3 discards the hyperprior citing posterior collapse; the paper does not ablate whether retaining it would improve results.
  • domain assumption Bicubic downsampling models the spatial reference degradation
    Section 4.1 generates LR frames by bicubic downsampling; real scalable coding may encounter other scaling kernels.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Augmented Deep Contexts for Spatially Embedded Video Coding." pith.science (2026). https://pith.science/paper/IXLOX62Y

@misc{pith2026250505309,
  author       = {Pith},
  title        = {Pith review of: Augmented Deep Contexts for Spatially Embedded Video Coding},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IXLOX62Y}},
  note         = {Machine review of arXiv:2505.05309}
}
read the original abstract

Most Neural Video Codecs (NVCs) only employ temporal references to generate temporal-only contexts and latent prior. These temporal-only NVCs fail to handle large motions or emerging objects due to limited contexts and misaligned latent prior. To relieve the limitations, we propose a Spatially Embedded Video Codec (SEVC), in which the low-resolution video is compressed for spatial references. Firstly, our SEVC leverages both spatial and temporal references to generate augmented motion vectors and hybrid spatial-temporal contexts. Secondly, to address the misalignment issue in latent prior and enrich the prior information, we introduce a spatial-guided latent prior augmented by multiple temporal latent representations. At last, we design a joint spatial-temporal optimization to learn quality-adaptive bit allocation for spatial references, further boosting rate-distortion performance. Experimental results show that our SEVC effectively alleviates the limitations in handling large motions or emerging objects, and also reduces 11.9% more bitrate than the previous state-of-the-art NVC while providing an additional low-resolution bitstream. Our code and model are available at https://github.com/EsakaK/SEVC.

Figures

Figures reproduced from arXiv: 2505.05309 by the authors.

Figure 1
Figure 1. Compared with previous Neural Video Codecs (NVCs), [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Comparison of our Spatially Embedded Video Codec (SEVC) and previous NVCs. (a) compresses the residual between the [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Overview of our SEVC. Our SEVC embeds a base [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Diagrams of the Motion and Feature Co-Augmentation module. “UP” represents upsampling using subpixel convolution [ [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Visualization of base MVs and the spatial feature within [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 7
Figure 7. Figure 7: Illustration of latent residuals for different priors. The [PITH_FULL_IMAGE:figures/full_fig_p005_7.png]
Figure 8
Figure 8. Figure 8: Rate and distortion curves on four 1080p datasets. The Intra Period is 32 with 96 frames. [PITH_FULL_IMAGE:figures/full_fig_p006_8.png]
Figure 9
Figure 9. Figure 9: Illustration of base codec bits proportion. The joint [PITH_FULL_IMAGE:figures/full_fig_p006_9.png]
Figure 10
Figure 10. Figure 10: Rate and distortion curves on four 1080p datasets. The Intra Period is –1 with 96 frames. [PITH_FULL_IMAGE:figures/full_fig_p007_10.png]
Figure 11
Figure 11. Figure 11: Visualization of the MVs and contexts in DCVC-DC and our SEVC. [PITH_FULL_IMAGE:figures/full_fig_p008_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

65 extracted references · 61 canonical work pages

  1. [1]

    https://vcgit.hhi.fraunhofer.de/ jvet/VVCSoftware_VTM

    VTM-17.0. https://vcgit.hhi.fraunhofer.de/ jvet/VVCSoftware_VTM . Accessed July 28, 2024. 6, 12

  2. [2]

    https://github.com/sanghyun- son/bicubic_pytorch

    bicubic-pytorch. https://github.com/sanghyun- son/bicubic_pytorch. Accessed July 28, 2024. 6

  3. [3]

    Scale-space flow for end-to-end optimized video compression

    Eirikur Agustsson, David Minnen, Nick Johnston, Johannes Balle, Sung Jin Hwang, and George Toderici. Scale-space flow for end-to-end optimized video compression. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8503–8512, 2020. 2

  4. [4]

    Hierarchical B-frame Video Coding Using Two- Layer CANF without Motion Coding

    David Alexandre, Hsueh-Ming Hang, and Wen-Hsiao Peng. Hierarchical B-frame Video Coding Using Two- Layer CANF without Motion Coding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10249–10258, 2023. 3

  5. [5]

    Variational image compres- sion with a scale hyperprior

    Johannes Ball ´e, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Variational image compres- sion with a scale hyperprior. In International Conference on Learning Representations (ICLR) , pages 4961–5007, 2018. 4

  6. [6]

    LSSVC: A Learned Spatially Scalable Video Coding Scheme

    Yifan Bian, Xihua Sheng, Li Li, and Dong Liu. LSSVC: A Learned Spatially Scalable Video Coding Scheme. IEEE Transactions on Image Processing: a publication of the IEEE Signal Processing Society, 33:3314–3327, 2024. 2

  7. [7]

    Common test conditions and software reference configurations

    Frank Bossen et al. Common test conditions and software reference configurations. JCTVC-L1100, 12(7):1, 2013. 6, 7

  8. [8]

    Overview of SHVC: Scalable Extensions of the High Efficiency Video Coding Standard

    Jill M Boyce, Yan Ye, Jianle Chen, and Adarsh K Ramasub- ramonian. Overview of SHVC: Scalable Extensions of the High Efficiency Video Coding Standard. IEEE Transactions on Circuits and Systems for Video Technology, 26(1):20–34,

Show all 65 references
  1. [9]

    Overview of the Versatile Video Coding (VVC) Standard and its Applica- tions

    Benjamin Bross, Ye-Kui Wang, Yan Ye, Shan Liu, Jianle Chen, Gary J Sullivan, and Jens-Rainer Ohm. Overview of the Versatile Video Coding (VVC) Standard and its Applica- tions. IEEE Transactions on Circuits and Systems for Video Technology, 31(10):3736–3764, 2021. 1

  2. [10]

    BasicVSR: The Search for Essential Components in Video Super-Resolution and Beyond

    Kelvin CK Chan, Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. BasicVSR: The Search for Essential Components in Video Super-Resolution and Beyond. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 4947–4956, 2021. 1, 3

  3. [11]

    BasicVSR++: Improving Video Super- Resolution with Enhanced Propagation and Alignment

    Kelvin CK Chan, Shangchen Zhou, Xiangyu Xu, and Chen Change Loy. BasicVSR++: Improving Video Super- Resolution with Enhanced Propagation and Alignment. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR) , pages 5972–5981,

  4. [12]

    NeRV: Neural Representations for Videos

    Hao Chen, Bo He, Hanyu Wang, Yixuan Ren, Ser Nam Lim, and Abhinav Shrivastava. NeRV: Neural Representations for Videos. Advances in Neural Information Processing Systems (NeurIPS), 34:21557–21568, 2021. 2

  5. [13]

    HNeRV: A Hybrid Neural Representation for Videos

    Hao Chen, Matthew Gwilliam, Ser-Nam Lim, and Abhinav Shrivastava. HNeRV: A Hybrid Neural Representation for Videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 10270–10279, 2023. 2

  6. [14]

    Elements of information theory

    Thomas M Cover. Elements of information theory . John Wiley & Sons, 1999. 3, 13

  7. [15]

    Neural Inter-Frame Com- pression for Video Coding

    Abdelaziz Djelouah, Joaquim Campos, Simone Schaub- Meyer, and Christopher Schroers. Neural Inter-Frame Com- pression for Video Coding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 6421–6429, 2019. 2

  8. [16]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, G Heigold, S Gelly, et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on...

  9. [17]

    Asymmetric numeral systems: entropy coding combining speed of Huffman coding with compression rate of arithmetic coding

    Jarek Duda. Asymmetric numeral systems: entropy coding combining speed of Huffman coding with compression rate of arithmetic coding. arXiv preprint arXiv:1311.2540, 2013. 4

  10. [18]

    Comparison of the H.263 and H.261 video compression stan- dards

    Bernd Girod, Eckehard G Steinbach, and Niko Faerber. Comparison of the H.263 and H.261 video compression stan- dards. In Standards and Common Interfaces for Video Infor- mation Systems: A Critical Review , pages 230–248. SPIE,

  11. [19]

    CANF-VC: Conditional Augmented Normalizing Flows for Video Compression

    Yung-Han Ho, Chih-Peng Chang, Peng-Yu Chen, Alessan- dro Gnutti, and Wen-Hsiao Peng. CANF-VC: Conditional Augmented Normalizing Flows for Video Compression. In European Conference on Computer Vision (ECCV) , pages 207–223. Springer, 2022. 2, 5

  12. [20]

    FVC: A New Framework towards Deep Video Compression in Feature Space

    Zhihao Hu, Guo Lu, and Dong Xu. FVC: A New Framework towards Deep Video Compression in Feature Space. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 1502–1511, 2021. 1, 2, 3

  13. [21]

    Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction

    Zhihao Hu, Guo Lu, Jinyang Guo, Shan Liu, Wei Jiang, and Dong Xu. Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5921–5930, 2022. 2, 3, 5

  14. [22]

    Sparse Point Clouds Assisted Learned Image Compression

    Yiheng Jiang, Haotian Zhang, Li Li, Dong Liu, and Zhu Li. Sparse Point Clouds Assisted Learned Image Compression. arXiv preprint arXiv:2412.15752, 2024. 2

  15. [23]

    Video Super-Resolution With Con- volutional Neural Networks

    Kappeler, Armin and Yoo, Seunghwan and Dai, Qiqin and Katsaggelos, Aggelos K. Video Super-Resolution With Con- volutional Neural Networks. IEEE Transactions on Compu- tational Imaging, 2(2):109–122, 2016. 1, 3

  16. [24]

    Optical Flow and Mode Se- lection for Learning-based Video Coding

    Th ´eo Ladune, Pierrick Philippe, Wassim Hamidouche, Lu Zhang, and Olivier D ´eforges. Optical Flow and Mode Se- lection for Learning-based Video Coding. In 2020 IEEE 22nd International Workshop on Multimedia Signal Process- ing (MMSP), pages 1–6. IEEE, 2020. 2

  17. [25]

    Conditional Coding and Vari- able Bitrate for Practical Learned Video Coding

    Th ´eo Ladune, Pierrick Philippe, Wassim Hamidouche, Lu Zhang, and Olivier D´eforges. Conditional Coding and Vari- able Bitrate for Practical Learned Video Coding. arXiv preprint arXiv:2104.09103, 2021

  18. [26]

    Deep Contextual Video Com- pression

    Jiahao Li, Bin Li, and Yan Lu. Deep Contextual Video Com- pression. Advances in Neural Information Processing Sys- tems (NeurIPS), 34:18114–18125, 2021. 2, 3, 4, 5, 6 9

  19. [27]

    Hybrid Spatial-Temporal En- tropy Modelling for Neural Video Compression

    Jiahao Li, Bin Li, and Yan Lu. Hybrid Spatial-Temporal En- tropy Modelling for Neural Video Compression. InProceed- ings of the 30th ACM International Conference on Multime- dia (ACM MM), pages 1503–1511, 2022. 1, 2, 3, 4, 6, 7, 8, 12, 14

  20. [28]

    Neural Video Compression with Diverse Contexts

    Jiahao Li, Bin Li, and Yan Lu. Neural Video Compression with Diverse Contexts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22616–22626, 2023. 1, 2, 3, 4, 5, 6, 7, 12, 13, 14

  21. [29]

    Neural Video Compression with Feature Modulation

    Jiahao Li, Bin Li, and Yan Lu. Neural Video Compression with Feature Modulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26099–26108, 2024. 1, 2, 3, 4, 6, 7, 13, 14

  22. [30]

    E-NeRV: Expedite Neural Video Representation with Disentangled Spatial-Temporal Con- text

    Zizhang Li, Mengmeng Wang, Huaijin Pi, Kechun Xu, Jian- biao Mei, and Yong Liu. E-NeRV: Expedite Neural Video Representation with Disentangled Spatial-Temporal Con- text. In European Conference on Computer Vision (ECCV), pages 267–284. Springer, 2022. 2

  23. [31]

    Uniformly Accelerated Motion Model for Inter Prediction

    Zhuoyuan Li, Yao Li, Chuanbo Tang, Li Li, Dong Liu, and Feng Wu. Uniformly Accelerated Motion Model for Inter Prediction. arXiv preprint arXiv:2407.11541, 2024. 1

  24. [32]

    USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020s

    Zhuoyuan Li, Junqi Liao, Chuanbo Tang, Haotian Zhang, Yuqi Li, Yifan Bian, Xihua Sheng, Xinmin Feng, Yao Li, Changsheng Gao, et al. USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020s. arXiv preprint arXiv:2409.08481, 2024. 6, 7, 15

  25. [33]

    Object Segmentation-Assisted Inter Predic- tion for Versatile Video Coding

    Zhuoyuan Li, Zikun Yuan, Li Li, Dong Liu, Xiaohu Tang, and Feng Wu. Object Segmentation-Assisted Inter Predic- tion for Versatile Video Coding. IEEE Transactions on Broadcasting, 70(4):1236–1253, 2024. 1

  26. [34]

    SwinIR: Image Restoration Using Swin Transformer

    Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. SwinIR: Image Restoration Using Swin Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1833–1844, 2021. 5, 12, 13

  27. [35]

    M- LVC: Multiple Frames Prediction for Learned Video Com- pression

    Jianping Lin, Dong Liu, Houqiang Li, and Feng Wu. M- LVC: Multiple Frames Prediction for Learned Video Com- pression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 3546–3554, 2020. 3

  28. [36]

    Neural Video Coding using Mul- tiscale Motion Compensation and Spatiotemporal Context Model

    Haojie Liu, Ming Lu, Zhan Ma, Fan Wang, Zhihuang Xie, Xun Cao, and Yao Wang. Neural Video Coding using Mul- tiscale Motion Compensation and Spatiotemporal Context Model. IEEE Transactions on Circuits and Systems for Video Technology, 31(8):3182–3196, 2020. 2

  29. [37]

    Swin Transformer: Hierarchical Vision Transformer using Shifted Windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 10012–10022, 2021. 2, 5

  30. [38]

    DVC: An End-to-end Deep Video Compression Framework

    Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, and Zhiyong Gao. DVC: An End-to-end Deep Video Compression Framework. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11006–11015, 2019. 2, 3, 5

  31. [39]

    An End-to-End Learning Framework for Video Compression

    Guo Lu, Xiaoyun Zhang, Wanli Ouyang, Li Chen, Zhiyong Gao, and Dong Xu. An End-to-End Learning Framework for Video Compression. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(10):3292–3308, 2020. 2, 3, 5

  32. [40]

    Don’t Blame the Elbo! A Linear Vae Perspec- tive on Posterior Collapse

    James Lucas, George Tucker, Roger B Grosse, and Moham- mad Norouzi. Don’t Blame the Elbo! A Linear Vae Perspec- tive on Posterior Collapse. Advances in Neural Information Processing Systems (NeurIPS), 32, 2019. 5

  33. [41]

    Uncertainty-Aware Deep Video Compression With Ensembles

    Wufei Ma, Jiahao Li, Bin Li, and Yan Lu. Uncertainty-Aware Deep Video Compression With Ensembles. IEEE Transac- tions on Multimedia, 26:7863–7872, 2024. 1, 3

  34. [42]

    Spatial scalability with VVC: coding performance and complexity

    Gwenaelle Marquant, Charles Salmon-Legagneur, Fabrice Urban, and Philippe de Lagrange. Spatial scalability with VVC: coding performance and complexity. In Applications of Digital Image Processing XLV, pages 10–15. SPIE, 2022. 5

  35. [43]

    VCT: A Video Compression Transformer

    Fabian Mentzer, George Toderici, David Minnen, Sung-Jin Hwang, Sergi Caelles, Mario Lucic, and Eirikur Agustsson. VCT: A Video Compression Transformer. arXiv preprint arXiv:2206.07307, 2022. 3

  36. [44]

    UVG Dataset: 50/120fps 4K Sequences for Video Codec Analysis and Development

    Alexandre Mercat, Marko Viitanen, and Jarno Vanne. UVG Dataset: 50/120fps 4K Sequences for Video Codec Analysis and Development. In Proceedings of the 11th ACM Multi- media Systems Conference (ACM MMSys) , pages 297–302,

  37. [45]

    Optical Flow Estima- tion using A Spatial Pyramid Network

    Anurag Ranjan and Michael J Black. Optical Flow Estima- tion using A Spatial Pyramid Network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4161–4170, 2017. 7

  38. [46]

    H.263: Video coding for low-bit-rate commu- nication

    Karel Rijkse. H.263: Video coding for low-bit-rate commu- nication. IEEE Communications magazine , 34(12):42–45,

  39. [47]

    Learned Video Compression

    Oren Rippel, Sanjay Nair, Carissa Lew, Steve Branson, Alexander G Anderson, and Lubomir Bourdev. Learned Video Compression. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV) , pages 3454–3463, 2019. 2

  40. [48]

    ELF-VC: Effi- cient Learned Flexible-Rate Video Coding

    Oren Rippel, Alexander G Anderson, Kedar Tatwawadi, San- jay Nair, Craig Lytle, and Lubomir Bourdev. ELF-VC: Effi- cient Learned Flexible-Rate Video Coding. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 14479–14488, 2021. 2

  41. [49]

    Overview of the Scalable Video Coding Extension of the H.264/A VC Standard

    Heiko Schwarz, Detlev Marpe, and Thomas Wiegand. Overview of the Scalable Video Coding Extension of the H.264/A VC Standard. IEEE Transactions on Circuits and Systems for Video Technology, 17(9):1103–1120, 2007. 5

  42. [50]

    Temporal Context Mining for Learned Video Compression

    Xihua Sheng, Jiahao Li, Bin Li, Li Li, Dong Liu, and Yan Lu. Temporal Context Mining for Learned Video Compression. IEEE Transactions on Multimedia, 25:7311–7322, 2022. 1, 2, 3, 4, 5, 6, 12

  43. [51]

    Spatial De- composition and Temporal Fusion Based Inter Prediction for Learned Video Compression

    Xihua Sheng, Li Li, Dong Liu, and Houqiang Li. Spatial De- composition and Temporal Fusion Based Inter Prediction for Learned Video Compression. IEEE Transactions on Circuits and Systems for Video Technology, 34(7):6460–6473, 2024. 1, 2, 3, 4, 6, 8

  44. [52]

    Rethinking Alignment in Video Super-Resolution Transformers

    Shuwei Shi, Jinjin Gu, Liangbin Xie, Xintao Wang, Yu- jiu Yang, and Chao Dong. Rethinking Alignment in Video Super-Resolution Transformers. Advances in Neural In- formation Processing Systems (NeurIPS) , 35:36081–36093,

  45. [53]

    Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network

    Wenzhe Shi, Jose Caballero, Ferenc Husz ´ar, Johannes Totz, Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network. In Proceedings of the IEEE/CVF Conference on C...

  46. [54]

    Overview of the High Efficiency Video Coding (HEVC) Standard

    Gary J Sullivan, Jens-Rainer Ohm, Woo-Jin Han, and Thomas Wiegand. Overview of the High Efficiency Video Coding (HEVC) Standard. IEEE Transactions on Circuits and Systems for Video Technology, 22(12):1649–1668, 2012. 1

  47. [55]

    Standardized Exten- sions of High Efficiency Video Coding (HEVC).IEEE Jour- nal of selected topics in Signal Processing, 7(6):1001–1016,

    Gary J Sullivan, Jill M Boyce, Ying Chen, Jens-Rainer Ohm, C Andrew Segall, and Anthony Vetro. Standardized Exten- sions of High Efficiency Video Coding (HEVC).IEEE Jour- nal of selected topics in Signal Processing, 7(6):1001–1016,

  48. [56]

    Offline and Online Optical Flow En- hancement for Deep Video Compression

    Chuanbo Tang, Xihua Sheng, Zhuoyuan Li, Haotian Zhang, Li Li, and Dong Liu. Offline and Online Optical Flow En- hancement for Deep Video Compression. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 5118– 5126, 2024. 1

  49. [57]

    RAFT: Recurrent All-Pairs Field Transforms for Optical Flow

    Zachary Teed and Jia Deng. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. In European Conference on Computer Vision (ECCV) , pages 402–419. Springer, 2020. 13

  50. [58]

    Attention is All You Need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is All You Need. Advances in Neural Information Processing Systems (NeurIPS), 30, 2017. 3, 5

  51. [59]

    MCL-JCV: a JND-based H.264/A VC video quality assessment dataset

    Haiqiang Wang, Weihao Gan, Sudeng Hu, Joe Yuchieh Lin, Lina Jin, Longguang Song, Ping Wang, Ioannis Katsavouni- dis, Anne Aaron, and C-C Jay Kuo. MCL-JCV: a JND-based H.264/A VC video quality assessment dataset. In2016 IEEE International Conference on Image Processing (ICIP), ...

  52. [60]

    Yao Wang and O. Lee. Use of two-dimensional deformable mesh structures for video coding .I. The synthesis problem: mesh-based function approximation and mapping. IEEE Transactions on Circuits and Systems for Video Technology, 6(6):636–646, 1996. 1

  53. [61]

    Overview of the H.264/A VC Video Coding Standard

    Thomas Wiegand, Gary J Sullivan, Gisle Bjontegaard, and Ajay Luthra. Overview of the H.264/A VC Video Coding Standard. IEEE Transactions on Circuits and Systems for Video Technology, 13(7):560–576, 2003. 1

  54. [62]

    Affine Multipicture Motion-compensated Prediction

    Thomas Wiegand, Eckehard Steinbach, and Bernd Girod. Affine Multipicture Motion-compensated Prediction. IEEE Transactions on Circuits and Systems for Video Technology, 15(2):197–209, 2005. 1

  55. [63]

    Video Enhancement with Task-oriented Flow

    Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T Freeman. Video Enhancement with Task-oriented Flow. International Journal of Computer Vision, 127:1106– 1125, 2019. 6

  56. [64]

    Learned Low Bitrate Video Compres- sion with Space-time Super-resolution

    Jiayu Yang, Chunhui Yang, Fei Xiong, Feng Wang, and Ronggang Wang. Learned Low Bitrate Video Compres- sion with Space-time Super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1786–1790, 2022. 3

  57. [65]

    DNeRV: Model- ing Inherent Dynamics via Difference Neural Representation for Videos

    Qi Zhao, M Salman Asif, and Zhan Ma. DNeRV: Model- ing Inherent Dynamics via Difference Neural Representation for Videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2031–2040, 2023. 2 11 Supplementary Material This supple...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.