REVIEW 56 cited by
ControlVideo: Training-free Controllable Text-to-Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temporal modeling. Besides the training burden, the generated videos also suffer from appearance inconsistency and structural flickers, especially in long video synthesis. To address these challenges, we design a \emph{training-free} framework called \textbf{ControlVideo} to enable natural and efficient text-to-video generation. ControlVideo, adapted from ControlNet, leverages coarsely structural consistency from input motion sequences, and introduces three modules to improve video generation. Firstly, to ensure appearance coherence between frames, ControlVideo adds fully cross-frame interaction in self-attention modules. Secondly, to mitigate the flicker effect, it introduces an interleaved-frame smoother that employs frame interpolation on alternated frames. Finally, to produce long videos efficiently, it utilizes a hierarchical sampler that separately synthesizes each short clip with holistic coherency. Empowered with these modules, ControlVideo outperforms the state-of-the-arts on extensive motion-prompt pairs quantitatively and qualitatively. Notably, thanks to the efficient designs, it generates both short and long videos within several minutes using one NVIDIA 2080Ti. Code is available at https://github.com/YBYBZhang/ControlVideo.
Forward citations
Cited by 56 Pith papers
-
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
Interactive image editing can be cast as image-to-video generation: initializing from Stable Video Diffusion plus a new matching attention mechanism yields high-quality sketch, drag, and coarse-edit results with far l...
-
Class Balance Matters to Active Class-Incremental Learning
A distribution-matching, class-balanced selection method (CBS) improves incremental learning from unlabeled pools, beating random and standard active learning baselines on five datasets.
-
EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
EmoWorld adds three training-free steering operators to a frozen video diffusion transformer that separately control atmosphere, affect-bearing cues, and temporal emotion transitions in generated videos.
-
Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation
SGC quantifies 3D geometric consistency of generated videos by measuring divergence among local camera poses estimated only on static background sub-regions.
-
ANYPORTAL: Zero-Shot Consistent Video Background Replacement
A training-free video background replacement pipeline that keeps the foreground pixel-consistent by projecting refined latents through a deterministic reparameterization.
-
DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
A diffusion-based disentanglement method that separates object content from domain style achieves state-of-the-art unsupervised cross-domain image retrieval on three benchmarks.
-
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.
-
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
EVS combines a text-to-image and a text-to-video diffusion model in a single denoising pass, improving frame quality and temporal consistency without retraining.
-
HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation
HairShifter transfers a reference hairstyle onto a person throughout a video by animating a high-quality anchor frame and using a gated decoder that preserves non-hair regions.
-
AnyI2V: Animating Any Conditional Image with Motion Control
AnyI2V animates arbitrary conditional images with user-defined trajectories by injecting debiased diffusion features and aligning attention queries across frames, without training.
-
When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators
MotionEcho adaptively re-injects teacher-model guidance into few-step distilled video generators so reference motion can be copied at test time without training.
-
Controllable Coupled Image Generation via Diffusion Models
A cross-attention control method that couples backgrounds across multiple generated images by blending LLM-extracted background and entity prompts with a time-varying weight optimized for background similarity and tex...
-
UNIC: Unified In-Context Video Editing
One diffusion transformer handles ID insert, swap, delete, stylization, propagation, and re-camera control in a single model using in-context token concatenation with task-aware positional encoding and bias.
-
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
By training a semantic expert and a LoRA-based detail expert, DCM reaches nearly teacher-level VBench scores with 4-step video sampling on HunyuanVideo and CogVideoX.
-
Interactive Video Generation via Domain Adaptation
A training-free method combines mask normalization and temporal intrinsic denoising to improve trajectory control and perceptual quality in text-to-video diffusion.
-
EF-VI: Enhancing End-Frame Injection for Video Inbetweening
EF-VI injects temporally expanded end-frame features into a transformer-based image-to-video diffusion model, improving video inbetweening quality over direct fine-tuning and bidirectional sampling baselines.
-
Controllable Weather Synthesis and Removal with Video Diffusion Models
A video diffusion system that adds controllable weather effects to arbitrary videos and removes existing weather, trained on synthetic, generated, and auto-labeled real data.
-
AnimateAnywhere: Rouse the Background in Human Image Animation
AnimateAnywhere learns to animate the background of human videos directly from human pose sequences, using an epipolar-constrained 3D attention mechanism, and reports state-of-the-art results without camera trajectories.
-
Satellite to GroundScape -- Large-scale Consistent Ground View Generation from Satellite Views
Satellite-to-ground generation is extended from single images to consistent multi-view sequences by adding satellite-guided and satellite-temporal conditioning to a frozen latent diffusion model, with a new 100k-pair dataset.
-
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance
AnyCharV is a two-stage diffusion method that places a reference character into a target video scene using pose and mask guidance, with a self-boosting stage that improves identity preservation.
-
Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models
A temporal attention purification and skip-connection rerouting method that separates motion learning from appearance learning in text-to-video diffusion model customization.
-
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
Video Depth Anything adapts Depth Anything V2 to produce temporally consistent depth for arbitrarily long videos using a temporal attention head, an optical-flow-free gradient loss, and key-frame-based stitching.
-
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
A 300M-parameter video generation model reaches about 5-dan level in Go and near-oracle robot control by predicting next frames and compact latent codes, with no search or reward.
-
TDM: Temporally-Consistent Diffusion Model for All-in-One Real-World Video Restoration
TDM restores five kinds of video degradation with one ControlNet-fine-tuned Stable Diffusion model, using task prompts in training and windowed cross-frame attention plus DDIM inversion at inference for temporal consistency.
-
ACE: Anti-Editing Concept Erasure in Text-to-Image Models
ACE trains a LoRA adapter on both conditional and unconditional noise predictions so that erased concepts are suppressed during both generation and text-guided editing.
-
UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity
A multi-granularity mixture-of-experts image restoration model that routes each degraded image to an expert using both degradation and granularity estimates, outperforming all-in-one baselines.
-
GaussianPainter: Painting Point Cloud into 3D Gaussians with Normal Guidance
GaussianPainter produces 3D Gaussians from a point cloud and reference image in one forward pass by constraining Gaussian rotations with predicted surface normals.
-
Memory Efficient Matting with Adaptive Token Routing
Adaptive token routing with a lightweight refinement branch lets a ViT matting model run on full-resolution high-res images at about 12% of the memory of the ViTMatte baseline with only a small accuracy drop on Compos...
-
SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
A cross-view attention module and hybrid training recipe let a frozen text-to-video model generate synchronized multi-camera videos from arbitrary viewpoints given only a text prompt and camera poses.
-
Learning Spatially Decoupled Color Representations for Facial Image Colorization
FCNet decouples facial colorization into per-component color codes, enabling controllable reference-based, automatic, and diverse colorization of face images.
-
MoViE: Mobile Diffusion for Video Editing
MoViE distills a diffusion-based video editor into a single-step mobile model, achieving 12 fps on a Snapdragon 8 Gen 3 phone with modest quality loss.
-
ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model
ROSE uses patch-wise perception in a large multimodal model to predict dense masks and generate open-set category names without predefined prompts.
-
Trajectory Attention for Fine-grained Video Motion Control
An auxiliary trajectory attention branch, added to temporal attention in video diffusion models, improves camera motion control precision while preserving generation quality.
-
PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors
A pipeline that uses multi-view diffusion-generated depth images as shape priors, fused with the partial point cloud via attention and confidence filtering, achieves state-of-the-art completion on custom single-view b...
-
FloAt: Flow Warping of Self-Attention for Clothing Animation Generation
FloAtControlNet animates clothing by flow-warping self-attention maps in a normal-map-conditioned ControlNet, improving temporal coherence and reducing background flicker.
-
AnimateAnything: Consistent and Controllable Animation for Video Generation
A two-stage video generation system that converts camera, drag, and reference-video controls into unified optical flows, and adds a frequency-domain stabilizer to reduce flicker.
-
FlipSketch: Flipping Static Drawings to Text-Guided Sketch Animations
A text-guided system that animates a single raster sketch into a dynamic video by fine-tuning a text-to-video diffusion model and steering its attention maps with the input drawing.
-
Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
Fine-tuning video generators on driving data can improve visual fidelity while degrading how accurately the model predicts the movement of cars and pedestrians.
-
DAPE: Dual-Stage Parameter-Efficient Fine-Tuning for Consistent Video Editing with Diffusion Models
A dual-stage fine-tuning recipe, norm tuning followed by a visual adapter, improves temporal consistency and text alignment in one-shot video editing, with a new 232-video benchmark.
-
Understanding Attention Mechanism in Video Diffusion Models
Attention maps with high entropy correlate with video quality while low entropy maps carry structure, and this paper uses entropy-guided attention replacement to improve video generation and editing.
-
Multi-Modality Driven LoRA for Adverse Condition Depth Estimation
MMD-LoRA uses text-guided LoRA adapters and contrastive learning to adapt a depth estimator to unseen adverse weather conditions, setting new reported numbers on nuScenes and Oxford RobotCar.
-
Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective
A two-stage multi-task training method that trains auxiliary tasks equally in task-specific decoders and weights their shared-encoder gradients by uncertainty and gradient norm improves primary-task performance relati...
-
Re-Attentional Controllable Video Diffusion Editing
ReAtCo improves text-guided video editing by using attention-map gradients to place edited objects in user-specified regions and by re-injecting the original background during diffusion sampling.
-
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.
-
Collaborative Feature-Logits Contrastive Learning for Open-Set Semi-Supervised Object Detection
A detector trained with feature contrastive and uncertainty classification losses improves open-set semi-supervised object detection, labeling unseen objects as unknown.
-
Multi-View Face and Gesture Animation with Dynamic Gaussians
Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.
-
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
LongVie combines unified noise initialization, global control normalization, and multi-modal depth-plus-keypoint guidance to generate temporally consistent controllable videos of up to one minute.
-
EndoControlMag: Robust Endoscopic Vascular Motion Magnification with Periodic Reference Resetting and Hierarchical Tissue-aware Dual-Mask Control
A training-free Lagrangian motion magnification framework with periodic reference resetting and tissue-aware dual-mask control improves vascular pulsation visibility in endoscopic surgery videos.
-
DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing
DFVEdit edits videos by iteratively subtracting a conditional delta flow vector, the difference between the model's predictions under the target and source prompts, from the latent representation of the source video.
-
DiffuEraser: A Diffusion Model for Video Inpainting
DiffuEraser is a stable-diffusion video inpainting model that injects ProPainter priors via DDIM inversion and expands temporal receptive fields for long-sequence consistency.
-
MetricDepth: Enhancing Monocular Depth Estimation with Deep Metric Learning
MetricDepth adds a multi-range contrastive loss to monocular depth estimation, defining positive and negative feature pairs by ground-truth depth differences.
-
Geo-ConvGRU: Geographically Masked Convolutional Gated Recurrent Unit for Bird-Eye View Segmentation
A ConvGRU temporal module with a geographic visibility mask raises BEV segmentation accuracy on nuScenes by about 1.3 to 1.8 IoU points over the Fiery baseline.
-
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
A mask-guided attention modulation framework for tuning-free multi-attribute video editing, built on Stable Diffusion, with keyframe feature propagation.
-
Unsupervised Region-Based Image Editing of Denoising Diffusion Models
A masking and Jacobian projection technique discovers unsupervised semantic directions in diffusion model latent space, enabling region-local editing without fine-tuning.
-
Video Diffusion Transformers are In-Context Learners
Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.
-
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
A diffusion transformer with joint attention over all camera views and frames, a lightweight BEV controller, and a re-weighted object loss reports state-of-the-art multi-view driving video generation (FVD 37.8 on nuScenes).
Discussion (0). Continue with ORCID to comment.