REVIEW 16 cited by
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts. A major challenge in this task is to ensure that all frames in the edited video are visually consistent. Most recent works apply advanced text-to-image diffusion models to this task by inflating 2D spatial attention in the U-Net into spatio-temporal attention. Although temporal context can be added through spatio-temporal attention, it may introduce some irrelevant information for each patch and therefore cause inconsistency in the edited video. In this paper, for the first time, we introduce optical flow into the attention module in the diffusion model's U-Net to address the inconsistency issue for text-to-video editing. Our method, FLATTEN, enforces the patches on the same flow path across different frames to attend to each other in the attention module, thus improving the visual consistency in the edited videos. Additionally, our method is training-free and can be seamlessly integrated into any diffusion-based text-to-video editing methods and improve their visual consistency. Experiment results on existing text-to-video editing benchmarks show that our proposed method achieves the new state-of-the-art performance. In particular, our method excels in maintaining the visual consistency in the edited videos.
Forward citations
Cited by 16 Pith papers
-
OSVE: One Step Video Editing with One Step Diffusion Models
OSVE performs text-guided video editing with one-step diffusion models by training a single-pass inversion encoder and unifying frame latents for cross-frame attention, achieving quality comparable to multi-step metho...
-
ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics
A training-free method preserves proxy-video dynamics via region-wise latent noising and Stochastic Flow Relaxation, outperforming editing and motion-transfer baselines on a custom 76-video set.
-
Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution
SeTe-VSR injects high-level semantic and spatio-temporal guidance into latent diffusion space to improve fidelity and temporal consistency in video super-resolution.
-
TITAN-Guide: Taming Inference-Time AligNment for Guided Text-to-Video Diffusion Models
A training-free guidance method that uses forward gradients instead of backpropagation to steer text-to-video diffusion latents with lower GPU memory.
-
HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation
HairShifter transfers a reference hairstyle onto a person throughout a video by animating a high-quality anchor frame and using a gated decoder that preserves non-hair regions.
-
Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models
A video-diffusion pipeline conditioned on geometry maps, followed by component-wise UV inpainting, produces more coherent and seam-free textures for 3D meshes than Text2Tex, Paint3D, and Meshy in the reported tests.
-
UNIC: Unified In-Context Video Editing
One diffusion transformer handles ID insert, swap, delete, stylization, propagation, and re-camera control in a single model using in-context token concatenation with task-aware positional encoding and bias.
-
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.
-
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
A two-stage video object remover that removes text conditioning and uses minimax adversarial noise to achieve high-quality removal in 6 sampling steps without classifier-free guidance.
-
Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing
A reference-based video editing pipeline that guides cross-image attention with diffusion correspondence, then trains a per-video restoration model to clean up the zero-shot output.
-
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
EPiC trains a 30M-parameter visibility-aware ControlNet on mask-based anchor videos from 5,000 in-the-wild videos and 500 steps, reaching SOTA camera accuracy on RealEstate10K and MiraData.
-
Consistent and Editable: A Balanced Framework for Text-Guided Video Editing
EquiEdit balances temporal consistency and editability in diffusion-based text-guided video editing via a temporal Mamba module and spectral noise injection on initial latents.
-
DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
DreamSwapV performs mask-guided, subject-agnostic subject swapping in videos using multiple conditions, an adaptive mask strategy, and a new benchmark.
-
Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction
A temporal feature warping and attention fusion module for amodal completion improves occlusion handling and temporal stability in monocular HOI videos, and the completed frames support 3D Gaussian Splatting reconstruction.
-
LiftVSR: Lifting Image Diffusion to Video Super-Resolution via Hybrid Temporal Modeling with Only 4$\times$RTX 4090s
LiftVSR combines short-segment dynamic temporal attention, a long-term attention memory cache, and Diffusion Forcing style asymmetric sampling to achieve strong perceptual video super-resolution scores with dramatical...
-
Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
Follow-Your-Creation fine-tunes the Wan2.1 video inpainting model on composite point-cloud and editing masks so a single monocular video can be converted into editable 4D video with new camera motion.
Discussion (0). Sign in to comment.