Pith. sign in

REVIEW 6 cited by

STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.02976 v1 pith:3EUHEESA submitted 2025-01-06 cs.CV

classification cs.CV
keywords modelstextbfreal-worldsuper-resolutionvideotemporalartifactsconsistency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively. Integrating text-to-video (T2V) models into video super-resolution for improved temporal modeling is straightforward. However, two key challenges remain: artifacts introduced by complex degradations in real-world scenarios, and compromised fidelity due to the strong generative capacity of powerful T2V models (\textit{e.g.}, CogVideoX-5B). To enhance the spatio-temporal quality of restored videos, we introduce\textbf{~\name} (\textbf{S}patial-\textbf{T}emporal \textbf{A}ugmentation with T2V models for \textbf{R}eal-world video super-resolution), a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency. Specifically, we introduce a Local Information Enhancement Module (LIEM) before the global attention block to enrich local details and mitigate degradation artifacts. Moreover, we propose a Dynamic Frequency (DF) Loss to reinforce fidelity, guiding the model to focus on different frequency components across diffusion steps. Extensive experiments demonstrate\textbf{~\name}~outperforms state-of-the-art methods on both synthetic and real-world datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A latent-cascaded video generation framework with dual frequency-split experts reports state-of-the-art 2K/4K video generation on VBench, FIDpatch, and human preference.

  2. Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Vivid-VR restores degraded videos by distilling the text-concept understanding of a pretrained text-to-video model into a ControlNet-based pipeline, improving perceptual quality at a large cost in pixel fidelity.

  3. TurboVSR: Fantastic Video Upscalers and Where to Find Them

    cs.CV 2025-06 conditional novelty 6.0 of 10

    TurboVSR uses a high-compression video autoencoder with factorized conditioning and non-uniform shortcut sampling to achieve near-state-of-the-art perceptual video super-resolution at roughly 100x lower compute cost.

  4. UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion Space

    cs.CV 2025-05 conditional novelty 6.0 of 10

    UltraVSR achieves video super-resolution in one diffusion step, using a degradation estimate to set the noise level and lightweight temporal feature shifting to maintain frame consistency.

  5. Persistent Free Volume Governs (Anti)plasticization in Chitosan-Water Mixtures

    cond-mat.soft 2026-04 unverdicted novelty 5.0 of 10

    Dynamically accessible free volume, enabled by connected water-accessible regions, is proposed to govern antiplasticization then plasticization of elastic properties in chitosan–water mixtures.

  6. LiftVSR: Lifting Image Diffusion to Video Super-Resolution via Hybrid Temporal Modeling with Only 4$\times$RTX 4090s

    cs.CV 2025-06 conditional novelty 5.0 of 10

    LiftVSR combines short-segment dynamic temporal attention, a long-term attention memory cache, and Diffusion Forcing style asymmetric sampling to achieve strong perceptual video super-resolution scores with dramatical...

Pith tools