Pith. sign in

REVIEW 3 cited by

VideoGigaGAN: Towards Detail-rich Video Super-Resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.12388 v2 pith:ATNSGJ4W submitted 2024-04-18 cs.CV

classification cs.CV
keywords temporalvideogigaganconsistencyvideovideosgenerativeimagesuper-resolution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Video super-resolution (VSR) approaches have shown impressive temporal consistency in upsampled videos. However, these approaches tend to generate blurrier results than their image counterparts as they are limited in their generative capability. This raises a fundamental question: can we extend the success of a generative image upsampler to the VSR task while preserving the temporal consistency? We introduce VideoGigaGAN, a new generative VSR model that can produce videos with high-frequency details and temporal consistency. VideoGigaGAN builds upon a large-scale image upsampler -- GigaGAN. Simply inflating GigaGAN to a video model by adding temporal modules produces severe temporal flickering. We identify several key issues and propose techniques that significantly improve the temporal consistency of upsampled videos. Our experiments show that, unlike previous VSR methods, VideoGigaGAN generates temporally consistent videos with more fine-grained appearance details. We validate the effectiveness of VideoGigaGAN by comparing it with state-of-the-art VSR models on public datasets and showcasing video results with $8\times$ super-resolution.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhance-A-Video: Better Generated Video for Free

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Enhance-A-Video computes the mean off-diagonal temporal attention weight and uses it, scaled by a per-prompt temperature, to boost the attention residual in DiT models during inference.

  2. Sequence Matters: Harnessing Video Models in 3D Super-Resolution

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Ordering low-resolution multi-view images into video-like sequences lets off-the-shelf video super-resolution models outperform existing 3D super-resolution pipelines.

  3. Generative Adversarial Networks Bridging Art and Machine Intelligence

    cs.LG 2025-02 unverdicted novelty 1.0 of 10

    This paper is a textbook-style review of generative adversarial networks, covering theory, classic variants, training methods, and applications; no new architecture, theorem, or experimental result is introduced.

Pith tools