REVIEW 3 cited by
ViBiDSampler: Enhancing Video Interpolation Using Bidirectional Diffusion Sampler
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent progress in large-scale text-to-video (T2V) and image-to-video (I2V) diffusion models has greatly enhanced video generation, especially in terms of keyframe interpolation. However, current image-to-video diffusion models, while powerful in generating videos from a single conditioning frame, need adaptation for two-frame (start & end) conditioned generation, which is essential for effective bounded interpolation. Unfortunately, existing approaches that fuse temporally forward and backward paths in parallel often suffer from off-manifold issues, leading to artifacts or requiring multiple iterative re-noising steps. In this work, we introduce a novel, bidirectional sampling strategy to address these off-manifold issues without requiring extensive re-noising or fine-tuning. Our method employs sequential sampling along both forward and backward paths, conditioned on the start and end frames, respectively, ensuring more coherent and on-manifold generation of intermediate frames. Additionally, we incorporate advanced guidance techniques, CFG++ and DDS, to further enhance the interpolation process. By integrating these, our method achieves state-of-the-art performance, efficiently generating high-quality, smooth videos between keyframes. On a single 3090 GPU, our method can interpolate 25 frames at 1024 x 576 resolution in just 195 seconds, establishing it as a leading solution for keyframe interpolation.
Forward citations
Cited by 3 Pith papers
-
TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation
TLB-VFI reports state-of-the-art FID, LPIPS, and FloLPIPS on VFI benchmarks using a latent Brownian bridge diffusion model with temporal-aware encoding and wavelet gating.
-
Semantic Frame Interpolation
The authors define Semantic Frame Interpolation, build a 300k-clip dataset and benchmark, and propose a Mixture-of-LoRA adaptation of Wan2.1 that improves temporal smoothness but does not preserve the given start and ...
-
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
DiffuseSlide boosts the frame rate of latent diffusion videos via latent interpolation, noise re-injection, and sliding-window denoising, reporting better FVD, PSNR, and SSIM than several baselines on WebVid-10M.
Discussion (0). Continue with ORCID to comment.