Pith. sign in

REVIEW 1 cited by

Motion-Guided Latent Diffusion for Temporally Consistent Real-world Video Super-resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.00853 v2 pith:IW352RTX submitted 2023-12-01 cs.CV

classification cs.CV
keywords diffusionlatentreal-worldmodelsmotion-guidedqualitytemporalvideo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Real-world low-resolution (LR) videos have diverse and complex degradations, imposing great challenges on video super-resolution (VSR) algorithms to reproduce their high-resolution (HR) counterparts with high quality. Recently, the diffusion models have shown compelling performance in generating realistic details for image restoration tasks. However, the diffusion process has randomness, making it hard to control the contents of restored images. This issue becomes more serious when applying diffusion models to VSR tasks because temporal consistency is crucial to the perceptual quality of videos. In this paper, we propose an effective real-world VSR algorithm by leveraging the strength of pre-trained latent diffusion models. To ensure the content consistency among adjacent frames, we exploit the temporal dynamics in LR videos to guide the diffusion process by optimizing the latent sampling path with a motion-guided loss, ensuring that the generated HR video maintains a coherent and continuous visual flow. To further mitigate the discontinuity of generated details, we insert temporal module to the decoder and fine-tune it with an innovative sequence-oriented loss. The proposed motion-guided latent diffusion (MGLD) based VSR algorithm achieves significantly better perceptual quality than state-of-the-arts on real-world VSR benchmark datasets, validating the effectiveness of the proposed model design and training strategies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SCST reports the best perceptual quality (LPIPS/DISTS) on four synthetic benchmarks and the best no-reference quality scores on the real-world VideoLQ benchmark by adding spatio-temporal Mamba and contrastive ControlN...

Pith tools