REVIEW 4 cited by
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion models have obtained substantial progress in image-to-video generation. However, in this paper, we find that these models tend to generate videos with less motion than expected. We attribute this to the issue called conditional image leakage, where the image-to-video diffusion models (I2V-DMs) tend to over-rely on the conditional image at large time steps. We further address this challenge from both inference and training aspects. First, we propose to start the generation process from an earlier time step to avoid the unreliable large-time steps of I2V-DMs, as well as an initial noise distribution with optimal analytic expressions (Analytic-Init) by minimizing the KL divergence between it and the actual marginal distribution to bridge the training-inference gap. Second, we design a time-dependent noise distribution (TimeNoise) for the conditional image during training, applying higher noise levels at larger time steps to disrupt it and reduce the model's dependency on it. We validate these general strategies on various I2V-DMs on our collected open-domain image benchmark and the UCF101 dataset. Extensive results show that our methods outperform baselines by producing higher motion scores with lower errors while maintaining image alignment and temporal consistency, thereby yielding superior overall performance and enabling more accurate motion control. The project page: \url{https://cond-image-leak.github.io/}.
Forward citations
Cited by 4 Pith papers
-
MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation
A single photo is converted into a 3D mesh with PBR textures using a render-enhanced auto-encoder, two data-augmentation schemes, and a multi-view texturing pipeline, with the claimed result being the best quality amo...
-
MotiF: Making Text Count in Image Animation with Motion Focal Loss
A motion-weighted training loss (MotiF) improves text alignment and object motion in text-image-to-video generation, winning 72% of human-preference comparisons against nine baselines on a new benchmark.
-
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...
-
Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion Models
Masking the image-feature dimensions most correlated with the style reference's content text reduces content leakage and improves text fidelity in text-to-image style transfer diffusion models.
Discussion (0). Continue with ORCID to comment.