Presents Decoupled Time Guidance (DTG) for training-free generative video super-resolution by temporally decoupling conditional and unconditional diffusion signals.
Star: Spatial-temporal augmentation with text-to-video models for real-world video super-resolution.arXiv preprint arXiv:2501.02976
8 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 8representative citing papers
ChopGrad truncates backpropagation to local frame windows in video diffusion models, reducing memory from linear in frame count to constant while enabling pixel-wise loss fine-tuning.
LiteVSR performs video super-resolution on a completely frozen Diffusion Transformer via a lightweight State-Aware Adapter that uses dual-stream extraction and time-dependent cross-attention, reaching competitive quality with 11.25% trainable parameters after 12 GPU-hours.
DiffST delivers state-of-the-art real-world space-time video super-resolution with 17x faster inference than prior diffusion methods by using one-step sampling, cross-frame context aggregation, and video representation guidance.
DVFace uses a spatio-temporal dual-codebook and asymmetric fusion in a one-step diffusion model to deliver better video face restoration quality, temporal consistency, and identity preservation than recent methods.
SATB-VR trains few-step video restoration diffusion models via SNR-aware trajectory blending of predictor outputs with ground-truth and a denoiser-driven consistency loss to achieve favorable performance on benchmarks.
SDIR is a dual-path iterative refinement model using scale-adaptive transformers and Fourier operators plus a physically consistent spectral loss to improve both spatial accuracy and turbulence-consistent frequency content in precipitation nowcasting.
OASIS reduces redundancy in diffusion models for real-world video super-resolution via attention specialization routing and progressive training, delivering state-of-the-art quality with 6.2x faster inference than prior one-step baselines.
citing papers explorer
-
DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution
Presents Decoupled Time Guidance (DTG) for training-free generative video super-resolution by temporally decoupling conditional and unconditional diffusion signals.
-
ChopGrad: Pixel-Wise Losses for Latent Video Diffusion via Truncated Backpropagation
ChopGrad truncates backpropagation to local frame windows in video diffusion models, reducing memory from linear in frame count to constant while enabling pixel-wise loss fine-tuning.
-
LiteVSR: Lightweight Adaptation of Frozen Diffusion Transformers for Video Super-Resolution
LiteVSR performs video super-resolution on a completely frozen Diffusion Transformer via a lightweight State-Aware Adapter that uses dual-stream extraction and time-dependent cross-attention, reaching competitive quality with 11.25% trainable parameters after 12 GPU-hours.
-
DiffST: Spatiotemporal-Aware Diffusion for Real-World Space-Time Video Super-Resolution
DiffST delivers state-of-the-art real-world space-time video super-resolution with 17x faster inference than prior diffusion methods by using one-step sampling, cross-frame context aggregation, and video representation guidance.
-
DVFace: Spatio-Temporal Dual-Prior Diffusion for Video Face Restoration
DVFace uses a spatio-temporal dual-codebook and asymmetric fusion in a one-step diffusion model to deliver better video face restoration quality, temporal consistency, and identity preservation than recent methods.
-
SATB-VR: Training Few-Step Video Restoration Diffusion Model using SNR-Aware Trajectory Blending
SATB-VR trains few-step video restoration diffusion models via SNR-aware trajectory blending of predictor outputs with ground-truth and a denoiser-driven consistency loss to achieve favorable performance on benchmarks.
-
Learning to Refine: Spectral-Decoupled Iterative Refinement Framework for Precipitation Nowcasting
SDIR is a dual-path iterative refinement model using scale-adaptive transformers and Fourier operators plus a physically consistent spectral loss to improve both spatial accuracy and turbulence-consistent frequency content in precipitation nowcasting.
-
Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution
OASIS reduces redundancy in diffusion models for real-world video super-resolution via attention specialization routing and progressive training, delivering state-of-the-art quality with 6.2x faster inference than prior one-step baselines.