Pith. sign in

REVIEW 15 cited by

VEnhancer: Generative Space-Time Enhancement for Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07667 v1 pith:VNZCLVYU submitted 2024-07-10 cs.CV eess.IV

classification cs.CVeess.IV
keywords videovenhancerspace-timediffusiongeneratedmodelspatialtemporal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present VEnhancer, a generative space-time enhancement framework that improves the existing text-to-video results by adding more details in spatial domain and synthetic detailed motion in temporal domain. Given a generated low-quality video, our approach can increase its spatial and temporal resolution simultaneously with arbitrary up-sampling space and time scales through a unified video diffusion model. Furthermore, VEnhancer effectively removes generated spatial artifacts and temporal flickering of generated videos. To achieve this, basing on a pretrained video diffusion model, we train a video ControlNet and inject it to the diffusion model as a condition on low frame-rate and low-resolution videos. To effectively train this video ControlNet, we design space-time data augmentation as well as video-aware conditioning. Benefiting from the above designs, VEnhancer yields to be stable during training and shares an elegant end-to-end training manner. Extensive experiments show that VEnhancer surpasses existing state-of-the-art video super-resolution and space-time super-resolution methods in enhancing AI-generated videos. Moreover, with VEnhancer, exisiting open-source state-of-the-art text-to-video method, VideoCrafter-2, reaches the top one in video generation benchmark -- VBench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution

    cs.CV 2026-07 conditional novelty 6.0 of 10

    TRaM-VSR routes and merges video tokens by importance during one-step diffusion super-resolution, cutting token cost by roughly a fifth while improving temporal consistency over the DOVE baseline.

  2. LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A latent-cascaded video generation framework with dual frequency-split experts reports state-of-the-art 2K/4K video generation on VBench, FIDpatch, and human preference.

  3. HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly

    cs.CV 2025-07 reject novelty 6.0 of 10

    A dual-branch video classifier uses depth and spatiotemporal features plus rank-weighted losses to categorize human-centric AI forgeries into spatial, appearance, and motion anomaly types on a new auto-labeled benchmark.

  4. Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration

    cs.CV 2025-07 conditional novelty 6.0 of 10

    IP-FVR restores degraded face videos with consistent identity by conditioning a video diffusion model on a reference photo of the same person.

  5. TurboVSR: Fantastic Video Upscalers and Where to Find Them

    cs.CV 2025-06 conditional novelty 6.0 of 10

    TurboVSR uses a high-compression video autoencoder with factorized conditioning and non-uniform shortcut sampling to achieve near-state-of-the-art perceptual video super-resolution at roughly 100x lower compute cost.

  6. Enhance-A-Video: Better Generated Video for Free

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Enhance-A-Video computes the mean off-diagonal temporal attention weight and uses it, scaled by a per-prompt temperature, to boost the attention residual in DiT models during inference.

  7. DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex Degradations

    cs.CV 2025-01 conditional novelty 6.0 of 10

    DiffVSR uses a three-stage progressive training curriculum, plus latent interpolation and noise rescheduling, to make a latent diffusion model robust to severely degraded videos while keeping temporal consistency.

  8. SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A 2.48B-parameter diffusion transformer with shifted-window attention and a causal video autoencoder reports competitive perceptual-quality video restoration across synthetic, real-world, and AI-generated benchmarks.

  9. VideoDPO: Omni-Preference Alignment for Video Diffusion Generation

    cs.CV 2024-12 reject novelty 6.0 of 10

    VideoDPO shows that DPO-style training on automatically selected best and worst video pairs improves overall VBench scores on three open text-to-video models, with some sub-metrics degrading and weak gains on external...

  10. FADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face Restoration

    cs.CV 2026-07 conditional novelty 5.0 of 10

    FADRA restores degraded face videos by adding LQ-guided step-wise residual refinement and a frequency-aware loss to a frozen text-to-video diffusion model, achieving state-of-the-art temporal coherence and spatial fidelity.

  11. GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GigaVideo-1 fine-tunes Wan2.1 on synthetic weakness-targeted prompts with VLM reward reweighting and reports ~4% average VBench-2.0 gains per dimension at 4 GPU-hours each, though joint training gains less.

  12. LiftVSR: Lifting Image Diffusion to Video Super-Resolution via Hybrid Temporal Modeling with Only 4$\times$RTX 4090s

    cs.CV 2025-06 conditional novelty 5.0 of 10

    LiftVSR combines short-segment dynamic temporal attention, a long-term attention memory cache, and Diffusion Forcing style asymmetric sampling to achieve strong perceptual video super-resolution scores with dramatical...

  13. Goku: Flow Based Video Generative Foundation Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A joint image-video generation model family reports state-of-the-art benchmark scores using rectified flow transformers, with all key evidence self-reported and no artifacts released.

  14. STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A video super-resolution method that adapts T2V diffusion models with a local attention module and a dynamic frequency loss to restore real-world degraded videos.

  15. Hunyuan-Game: Industrial-grade Intelligent Game Creation Model

    cs.CV 2025-05 reject novelty 4.0 of 10

    Tencent's Hunyuan-Game applies diffusion transformers to game asset creation across nine image and video generation tasks, with self-reported gains that are partly contradicted by its own evaluation table.

Pith tools