REVIEW 15 cited by
VEnhancer: Generative Space-Time Enhancement for Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present VEnhancer, a generative space-time enhancement framework that improves the existing text-to-video results by adding more details in spatial domain and synthetic detailed motion in temporal domain. Given a generated low-quality video, our approach can increase its spatial and temporal resolution simultaneously with arbitrary up-sampling space and time scales through a unified video diffusion model. Furthermore, VEnhancer effectively removes generated spatial artifacts and temporal flickering of generated videos. To achieve this, basing on a pretrained video diffusion model, we train a video ControlNet and inject it to the diffusion model as a condition on low frame-rate and low-resolution videos. To effectively train this video ControlNet, we design space-time data augmentation as well as video-aware conditioning. Benefiting from the above designs, VEnhancer yields to be stable during training and shares an elegant end-to-end training manner. Extensive experiments show that VEnhancer surpasses existing state-of-the-art video super-resolution and space-time super-resolution methods in enhancing AI-generated videos. Moreover, with VEnhancer, exisiting open-source state-of-the-art text-to-video method, VideoCrafter-2, reaches the top one in video generation benchmark -- VBench.
Forward citations
Cited by 15 Pith papers
-
TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution
TRaM-VSR routes and merges video tokens by importance during one-step diffusion super-resolution, cutting token cost by roughly a fifth while improving temporal consistency over the DOVE baseline.
-
LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts
A latent-cascaded video generation framework with dual frequency-split experts reports state-of-the-art 2K/4K video generation on VBench, FIDpatch, and human preference.
-
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
A dual-branch video classifier uses depth and spatiotemporal features plus rank-weighted losses to categorize human-centric AI forgeries into spatial, appearance, and motion anomaly types on a new auto-labeled benchmark.
-
Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration
IP-FVR restores degraded face videos with consistent identity by conditioning a video diffusion model on a reference photo of the same person.
-
TurboVSR: Fantastic Video Upscalers and Where to Find Them
TurboVSR uses a high-compression video autoencoder with factorized conditioning and non-uniform shortcut sampling to achieve near-state-of-the-art perceptual video super-resolution at roughly 100x lower compute cost.
-
Enhance-A-Video: Better Generated Video for Free
Enhance-A-Video computes the mean off-diagonal temporal attention weight and uses it, scaled by a per-prompt temperature, to boost the attention residual in DiT models during inference.
-
DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex Degradations
DiffVSR uses a three-stage progressive training curriculum, plus latent interpolation and noise rescheduling, to make a latent diffusion model robust to severely degraded videos while keeping temporal consistency.
-
SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
A 2.48B-parameter diffusion transformer with shifted-window attention and a causal video autoencoder reports competitive perceptual-quality video restoration across synthetic, real-world, and AI-generated benchmarks.
-
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
VideoDPO shows that DPO-style training on automatically selected best and worst video pairs improves overall VBench scores on three open text-to-video models, with some sub-metrics degrading and weak gains on external...
-
FADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face Restoration
FADRA restores degraded face videos by adding LQ-guided step-wise residual refinement and a frequency-aware loss to a frozen text-to-video diffusion model, achieving state-of-the-art temporal coherence and spatial fidelity.
-
GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning
GigaVideo-1 fine-tunes Wan2.1 on synthetic weakness-targeted prompts with VLM reward reweighting and reports ~4% average VBench-2.0 gains per dimension at 4 GPU-hours each, though joint training gains less.
-
LiftVSR: Lifting Image Diffusion to Video Super-Resolution via Hybrid Temporal Modeling with Only 4$\times$RTX 4090s
LiftVSR combines short-segment dynamic temporal attention, a long-term attention memory cache, and Diffusion Forcing style asymmetric sampling to achieve strong perceptual video super-resolution scores with dramatical...
-
Goku: Flow Based Video Generative Foundation Models
A joint image-video generation model family reports state-of-the-art benchmark scores using rectified flow transformers, with all key evidence self-reported and no artifacts released.
-
STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
A video super-resolution method that adapts T2V diffusion models with a local attention module and a dynamic frequency loss to restore real-world degraded videos.
-
Hunyuan-Game: Industrial-grade Intelligent Game Creation Model
Tencent's Hunyuan-Game applies diffusion transformers to game asset creation across nine image and video generation tasks, with self-reported gains that are partly contradicted by its own evaluation table.
Discussion (0). Continue with ORCID to comment.