REVIEW 3 cited by
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-guided image-to-video (I2V) generation aims to generate a coherent video that preserves the identity of the input image and semantically aligns with the input prompt. Existing methods typically augment pretrained text-to-video (T2V) models by either concatenating the image with noised video frames channel-wise before being fed into the model or injecting the image embedding produced by pretrained image encoders in cross-attention modules. However, the former approach often necessitates altering the fundamental weights of pretrained T2V models, thus restricting the model's compatibility within the open-source communities and disrupting the model's prior knowledge. Meanwhile, the latter typically fails to preserve the identity of the input image. We present I2V-Adapter to overcome such limitations. I2V-Adapter adeptly propagates the unnoised input image to subsequent noised frames through a cross-frame attention mechanism, maintaining the identity of the input image without any changes to the pretrained T2V model. Notably, I2V-Adapter only introduces a few trainable parameters, significantly alleviating the training cost and also ensures compatibility with existing community-driven personalized models and control tools. Moreover, we propose a novel Frame Similarity Prior to balance the motion amplitude and the stability of generated videos through two adjustable control coefficients. Our experimental results demonstrate that I2V-Adapter is capable of producing high-quality videos. This performance, coupled with its agility and adaptability, represents a substantial advancement in the field of I2V, particularly for personalized and controllable applications.
Forward citations
Cited by 3 Pith papers
-
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
A $500 LoRA fine-tune of Wan2.1-T2V with per-frame random timesteps matches Wan-I2V's benchmark quality and adds zero-shot start-end and video-extension capabilities.
-
MotionPro: A Precise Motion Controller for Image-to-Video Generation
MotionPro uses region-wise trajectories and a motion mask to control object and camera motion in image-to-video generation, reporting improved trajectory alignment over prior methods.
-
Light-A-Video: Training-free Video Relighting via Progressive Light Fusion
Light-A-Video relights videos without training by injecting image relight results into a video diffusion model's denoising loop with cross-frame attention and progressive blending.
Discussion (0). Continue with ORCID to comment.