Pith. sign in

REVIEW 1 cited by

RealisVSR: Detail-enhanced Diffusion for Real-World 4K Video Super-Resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.19138 v1 pith:MD2PBIWF submitted 2025-07-25 eess.IV cs.CV

classification eess.IVcs.CV
keywords diffusionvideohigh-frequencysuper-resolutionchallengescomplexdetaildetail-enhanced
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video Super-Resolution (VSR) has achieved significant progress through diffusion models, effectively addressing the over-smoothing issues inherent in GAN-based methods. Despite recent advances, three critical challenges persist in VSR community: 1) Inconsistent modeling of temporal dynamics in foundational models; 2) limited high-frequency detail recovery under complex real-world degradations; and 3) insufficient evaluation of detail enhancement and 4K super-resolution, as current methods primarily rely on 720P datasets with inadequate details. To address these challenges, we propose RealisVSR, a high-frequency detail-enhanced video diffusion model with three core innovations: 1) Consistency Preserved ControlNet (CPC) architecture integrated with the Wan2.1 video diffusion to model the smooth and complex motions and suppress artifacts; 2) High-Frequency Rectified Diffusion Loss (HR-Loss) combining wavelet decomposition and HOG feature constraints for texture restoration; 3) RealisVideo-4K, the first public 4K VSR benchmark containing 1,000 high-definition video-text pairs. Leveraging the advanced spatio-temporal guidance of Wan2.1, our method requires only 5-25% of the training data volume compared to existing approaches. Extensive experiments on VSR benchmarks (REDS, SPMCS, UDM10, YouTube-HQ, VideoLQ, RealisVideo-720P) demonstrate our superiority, particularly in ultra-high-resolution scenarios.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MoCRA: Mixture of Compositional Rank-1 Atoms for 4K All-in-One Video Restoration

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A single 3.6M-parameter model, MoCRA, restores haze, rain, noise, and low light at native 4K in 0.48 seconds per frame, besting eleven retrained baselines on the mean PSNR of the authors' new UHV-4K-AIO benchmark.

Pith tools