Pith. sign in

REVIEW 1 cited by

ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07697 v2 pith:5AYBQJ7T submitted 2023-10-11 cs.CV

classification cs.CV
keywords conditionvideogenerationvideoaccuracybi-directionalbranchcondition-guidedconditional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent works have successfully extended large-scale text-to-image models to the video domain, producing promising results but at a high computational cost and requiring a large amount of video data. In this work, we introduce ConditionVideo, a training-free approach to text-to-video generation based on the provided condition, video, and input text, by leveraging the power of off-the-shelf text-to-image generation methods (e.g., Stable Diffusion). ConditionVideo generates realistic dynamic videos from random noise or given scene videos. Our method explicitly disentangles the motion representation into condition-guided and scenery motion components. To this end, the ConditionVideo model is designed with a UNet branch and a control branch. To improve temporal coherence, we introduce sparse bi-directional spatial-temporal attention (sBiST-Attn). The 3D control network extends the conventional 2D controlnet model, aiming to strengthen conditional generation accuracy by additionally leveraging the bi-directional frames in the temporal domain. Our method exhibits superior performance in terms of frame consistency, clip score, and conditional accuracy, outperforming other compared methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MoTrans transfers specific motions from reference videos to new subjects using a two-stage fine-tuning scheme with recaptioned prompts, appearance injection, and a motion-specific verb embedding.

Pith tools