Pith. sign in

REVIEW 2 cited by

BVI-AOM: A New Training Dataset for Deep Video Compression Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.03265 v3 pith:AOXVTEHM submitted 2024-08-06 eess.IV

classification eess.IV
keywords datasettrainingbvi-aomcodingdeepvideoperformanceterms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning is now playing an important role in enhancing the performance of conventional hybrid video codecs. These learning-based methods typically require diverse and representative training material for optimization in order to achieve model generalization and optimal coding performance. However, existing datasets either offer limited content variability or come with restricted licensing terms constraining their use to research purposes only. To address these issues, we propose a new training dataset, named BVI-AOM, which contains 956 uncompressed sequences at various resolutions from 270p to 2160p, covering a wide range of content and texture types. The dataset comes with more flexible licensing terms and offers competitive performance when used as a training set for optimizing deep video coding tools. The experimental results demonstrate that when used as a training set to optimize two popular network architectures for two different coding tools, the proposed dataset leads to additional bitrate savings of up to 0.29 and 2.98 percentage points in terms of PSNR-Y and VMAF, respectively, compared to an existing training dataset, BVI-DVC, which has been widely used for deep video coding. The BVI-AOM dataset is available at https://github.com/fan-aaron-zhang/bvi-aom

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RTSR: A Real-Time Super-Resolution Model for AV1 Compressed Content

    eess.IV 2024-11 conditional novelty 4.0 of 10

    RTSR is a low-complexity CNN super-resolution model for AV1 compressed video that reported the best complexity-performance trade-off in the AIM 2024 Efficient Real-Time Video Super-Resolution competition.

  2. Compressed Video Super-Resolution based on Hierarchical Encoding

    eess.IV 2025-06 conditional novelty 2.0 of 10

    VSR-HE, a per-frame transformer trained with perceptual and GAN losses, reports improved 4x super-resolution quality on HEVC-compressed conferencing video versus bicubic, EDSR, CVEGAN, and SwinIR.

Pith tools