Pith. sign in

REVIEW 56 cited by

ControlVideo: Training-free Controllable Text-to-Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13077 v1 pith:3VI2VK7W submitted 2023-05-22 cs.CV

classification cs.CV
keywords controlvideogenerationlongmodulesvideovideosappearanceefficient
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temporal modeling. Besides the training burden, the generated videos also suffer from appearance inconsistency and structural flickers, especially in long video synthesis. To address these challenges, we design a \emph{training-free} framework called \textbf{ControlVideo} to enable natural and efficient text-to-video generation. ControlVideo, adapted from ControlNet, leverages coarsely structural consistency from input motion sequences, and introduces three modules to improve video generation. Firstly, to ensure appearance coherence between frames, ControlVideo adds fully cross-frame interaction in self-attention modules. Secondly, to mitigate the flicker effect, it introduces an interleaved-frame smoother that employs frame interpolation on alternated frames. Finally, to produce long videos efficiently, it utilizes a hierarchical sampler that separately synthesizes each short clip with holistic coherency. Empowered with these modules, ControlVideo outperforms the state-of-the-arts on extensive motion-prompt pairs quantitatively and qualitatively. Notably, thanks to the efficient designs, it generates both short and long videos within several minutes using one NVIDIA 2080Ti. Code is available at https://github.com/YBYBZhang/ControlVideo.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 56 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors

    cs.CV 2025-01 conditional novelty 7.0 of 10

    Interactive image editing can be cast as image-to-video generation: initializing from Stable Video Diffusion plus a new matching attention mechanism yields high-quality sketch, drag, and coarse-edit results with far l...

  2. Class Balance Matters to Active Class-Incremental Learning

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A distribution-matching, class-balanced selection method (CBS) improves incremental learning from unlabeled pools, beating random and standard active learning baselines on five datasets.

  3. EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    EmoWorld adds three training-free steering operators to a frozen video diffusion transformer that separately control atmosphere, affect-bearing cues, and temporal emotion transitions in generated videos.

  4. Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    SGC quantifies 3D geometric consistency of generated videos by measuring divergence among local camera poses estimated only on static background sub-regions.

  5. ANYPORTAL: Zero-Shot Consistent Video Background Replacement

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A training-free video background replacement pipeline that keeps the foreground pixel-consistent by projecting refined latents through a deterministic reparameterization.

  6. DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A diffusion-based disentanglement method that separates object content from domain style achieves state-of-the-art unsupervised cross-domain image retrieval on three benchmarks.

  7. Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.

  8. Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    EVS combines a text-to-image and a text-to-video diffusion model in a single denoising pass, improving frame quality and temporal consistency without retraining.

  9. HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    HairShifter transfers a reference hairstyle onto a person throughout a video by animating a high-quality anchor frame and using a gated decoder that preserves non-hair regions.

  10. AnyI2V: Animating Any Conditional Image with Motion Control

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AnyI2V animates arbitrary conditional images with user-defined trajectories by injecting debiased diffusion features and aligning attention queries across frames, without training.

  11. When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MotionEcho adaptively re-injects teacher-model guidance into few-step distilled video generators so reference motion can be copied at test time without training.

  12. Controllable Coupled Image Generation via Diffusion Models

    cs.CV 2025-06 reject novelty 6.0 of 10

    A cross-attention control method that couples backgrounds across multiple generated images by blending LLM-extracted background and entity prompts with a time-varying weight optimized for background similarity and tex...

  13. UNIC: Unified In-Context Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    One diffusion transformer handles ID insert, swap, delete, stylization, propagation, and re-camera control in a single model using in-context token concatenation with task-aware positional encoding and bias.

  14. Dual-Expert Consistency Model for Efficient and High-Quality Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    By training a semantic expert and a LoRA-based detail expert, DCM reaches nearly teacher-level VBench scores with 4-step video sampling on HunyuanVideo and CogVideoX.

  15. Interactive Video Generation via Domain Adaptation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A training-free method combines mask normalization and temporal intrinsic denoising to improve trajectory control and perceptual quality in text-to-video diffusion.

  16. EF-VI: Enhancing End-Frame Injection for Video Inbetweening

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EF-VI injects temporally expanded end-frame features into a transformer-based image-to-video diffusion model, improving video inbetweening quality over direct fine-tuning and bidirectional sampling baselines.

  17. Controllable Weather Synthesis and Removal with Video Diffusion Models

    cs.GR 2025-05 conditional novelty 6.0 of 10

    A video diffusion system that adds controllable weather effects to arbitrary videos and removes existing weather, trained on synthetic, generated, and auto-labeled real data.

  18. AnimateAnywhere: Rouse the Background in Human Image Animation

    cs.CV 2025-04 conditional novelty 6.0 of 10

    AnimateAnywhere learns to animate the background of human videos directly from human pose sequences, using an epipolar-constrained 3D attention mechanism, and reports state-of-the-art results without camera trajectories.

  19. Satellite to GroundScape -- Large-scale Consistent Ground View Generation from Satellite Views

    cs.CV 2025-04 conditional novelty 6.0 of 10

    Satellite-to-ground generation is extended from single images to consistent multi-view sequences by adding satellite-guided and satellite-temporal conditioning to a frozen latent diffusion model, with a new 100k-pair dataset.

  20. AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance

    cs.CV 2025-02 conditional novelty 6.0 of 10

    AnyCharV is a two-stage diffusion method that places a reference character into a target video scene using pose and mask guidance, with a self-boosting stage that improves identity preservation.

  21. Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A temporal attention purification and skip-connection rerouting method that separates motion learning from appearance learning in text-to-video diffusion model customization.

  22. Video Depth Anything: Consistent Depth Estimation for Super-Long Videos

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Video Depth Anything adapts Depth Anything V2 to produce temporally consistent depth for arbitrarily long videos using a temporal attention head, an optical-flow-free gradient loss, and key-frame-based stitching.

  23. VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A 300M-parameter video generation model reaches about 5-dan level in Go and near-oracle robot control by predicting next frames and compact latent codes, with no search or reward.

  24. TDM: Temporally-Consistent Diffusion Model for All-in-One Real-World Video Restoration

    cs.CV 2025-01 conditional novelty 6.0 of 10

    TDM restores five kinds of video degradation with one ControlNet-fine-tuned Stable Diffusion model, using task prompts in training and windowed cross-frame attention plus DDIM inversion at inference for temporal consistency.

  25. ACE: Anti-Editing Concept Erasure in Text-to-Image Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    ACE trains a LoRA adapter on both conditional and unconditional noise predictions so that erased concepts are suppressed during both generation and text-guided editing.

  26. UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A multi-granularity mixture-of-experts image restoration model that routes each degraded image to an expert using both degradation and granularity estimates, outperforming all-in-one baselines.

  27. GaussianPainter: Painting Point Cloud into 3D Gaussians with Normal Guidance

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GaussianPainter produces 3D Gaussians from a point cloud and reference image in one forward pass by constraining Gaussian rotations with predicted surface normals.

  28. Memory Efficient Matting with Adaptive Token Routing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Adaptive token routing with a lightweight refinement branch lets a ViT matting model run on full-resolution high-res images at about 12% of the memory of the ViTMatte baseline with only a small accuracy drop on Compos...

  29. SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A cross-view attention module and hybrid training recipe let a frozen text-to-video model generate synchronized multi-camera videos from arbitrary viewpoints given only a text prompt and camera poses.

  30. Learning Spatially Decoupled Color Representations for Facial Image Colorization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FCNet decouples facial colorization into per-component color codes, enabling controllable reference-based, automatic, and diverse colorization of face images.

  31. MoViE: Mobile Diffusion for Video Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MoViE distills a diffusion-based video editor into a single-step mobile model, achieving 12 fps on a Snapdragon 8 Gen 3 phone with modest quality loss.

  32. ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ROSE uses patch-wise perception in a large multimodal model to predict dense masks and generate open-set category names without predefined prompts.

  33. Trajectory Attention for Fine-grained Video Motion Control

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An auxiliary trajectory attention branch, added to temporal attention in video diffusion models, improves camera motion control precision while preserving generation quality.

  34. PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A pipeline that uses multi-view diffusion-generated depth images as shape priors, fused with the partial point cloud via attention and confidence filtering, achieves state-of-the-art completion on custom single-view b...

  35. FloAt: Flow Warping of Self-Attention for Clothing Animation Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FloAtControlNet animates clothing by flow-warping self-attention maps in a normal-map-conditioned ControlNet, improving temporal coherence and reducing background flicker.

  36. AnimateAnything: Consistent and Controllable Animation for Video Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A two-stage video generation system that converts camera, drag, and reference-video controls into unified optical flows, and adds a frequency-domain stabilizer to reduce flicker.

  37. FlipSketch: Flipping Static Drawings to Text-Guided Sketch Animations

    cs.GR 2024-11 conditional novelty 6.0 of 10

    A text-guided system that animates a single raster sketch into a dynamic video by fine-tuning a text-to-video diffusion model and steering its attention maps with the input drawing.

  38. Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Fine-tuning video generators on driving data can improve visual fidelity while degrading how accurately the model predicts the movement of cars and pedestrians.

  39. DAPE: Dual-Stage Parameter-Efficient Fine-Tuning for Consistent Video Editing with Diffusion Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A dual-stage fine-tuning recipe, norm tuning followed by a visual adapter, improves temporal consistency and text alignment in one-shot video editing, with a new 232-video benchmark.

  40. Understanding Attention Mechanism in Video Diffusion Models

    cs.CV 2025-04 conditional novelty 5.0 of 10

    Attention maps with high entropy correlate with video quality while low entropy maps carry structure, and this paper uses entropy-guided attention replacement to improve video generation and editing.

  41. Multi-Modality Driven LoRA for Adverse Condition Depth Estimation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    MMD-LoRA uses text-guided LoRA adapters and contrastive learning to adapt a depth estimator to unseen adverse weather conditions, setting new reported numbers on nuScenes and Oxford RobotCar.

  42. Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A two-stage multi-task training method that trains auxiliary tasks equally in task-specific decoders and weights their shared-encoder gradients by uncertainty and gradient norm improves primary-task performance relati...

  43. Re-Attentional Controllable Video Diffusion Editing

    cs.CV 2024-12 conditional novelty 5.0 of 10

    ReAtCo improves text-guided video editing by using attention-map gradients to place edited objects in user-specified regions and by re-injecting the original background during diffusion sampling.

  44. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.

  45. Collaborative Feature-Logits Contrastive Learning for Open-Set Semi-Supervised Object Detection

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A detector trained with feature contrastive and uncertainty classification losses improves open-set semi-supervised object detection, labeling unseen objects as unknown.

  46. Multi-View Face and Gesture Animation with Dynamic Gaussians

    cs.CV 2026-08 conditional novelty 4.0 of 10

    Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.

  47. LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation

    cs.CV 2025-08 conditional novelty 4.0 of 10

    LongVie combines unified noise initialization, global control normalization, and multi-modal depth-plus-keypoint guidance to generate temporally consistent controllable videos of up to one minute.

  48. EndoControlMag: Robust Endoscopic Vascular Motion Magnification with Periodic Reference Resetting and Hierarchical Tissue-aware Dual-Mask Control

    eess.IV 2025-07 conditional novelty 4.0 of 10

    A training-free Lagrangian motion magnification framework with periodic reference resetting and tissue-aware dual-mask control improves vascular pulsation visibility in endoscopic surgery videos.

  49. DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing

    cs.CV 2025-06 conditional novelty 4.0 of 10

    DFVEdit edits videos by iteratively subtracting a conditional delta flow vector, the difference between the model's predictions under the target and source prompts, from the latent representation of the source video.

  50. DiffuEraser: A Diffusion Model for Video Inpainting

    cs.CV 2025-01 reject novelty 4.0 of 10

    DiffuEraser is a stable-diffusion video inpainting model that injects ProPainter priors via DDIM inversion and expands temporal receptive fields for long-sequence consistency.

  51. MetricDepth: Enhancing Monocular Depth Estimation with Deep Metric Learning

    cs.CV 2024-12 conditional novelty 4.0 of 10

    MetricDepth adds a multi-range contrastive loss to monocular depth estimation, defining positive and negative feature pairs by ground-truth depth differences.

  52. Geo-ConvGRU: Geographically Masked Convolutional Gated Recurrent Unit for Bird-Eye View Segmentation

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A ConvGRU temporal module with a geographic visibility mask raises BEV segmentation accuracy on nuScenes by about 1.3 to 1.8 IoU points over the Fiery baseline.

  53. MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A mask-guided attention modulation framework for tuning-free multi-attribute video editing, built on Stable Diffusion, with keyframe feature propagation.

  54. Unsupervised Region-Based Image Editing of Denoising Diffusion Models

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A masking and Jacobian projection technique discovers unsupervised semantic directions in diffusion model latent space, enabling region-local editing without fine-tuning.

  55. Video Diffusion Transformers are In-Context Learners

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.

  56. Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A diffusion transformer with joint attention over all camera views and frames, a lightweight BEV controller, and a re-weighted object loss reports state-of-the-art multi-view driving video generation (FVD 37.8 on nuScenes).

Pith tools