REVIEW 7 cited by
Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a generative model that, given a coarsely edited image, synthesizes a photorealistic output that follows the prescribed layout. Our method transfers fine details from the original image and preserve the identity of its parts. Yet, it adapts it to the lighting and context defined by the new layout. Our key insight is that videos are a powerful source of supervision for this task: objects and camera motions provide many observations of how the world changes with viewpoint, lighting, and physical interactions. We construct an image dataset in which each sample is a pair of source and target frames extracted from the same video at randomly chosen time intervals. We warp the source frame toward the target using two motion models that mimic the expected test-time user edits. We supervise our model to translate the warped image into the ground truth, starting from a pretrained diffusion model. Our model design explicitly enables fine detail transfer from the source frame to the generated image, while closely following the user-specified layout. We show that by using simple segmentations and coarse 2D manipulations, we can synthesize a photorealistic edit faithful to the user's input while addressing second-order effects like harmonizing the lighting and physical interactions between edited objects.
Forward citations
Cited by 7 Pith papers
-
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
Interactive image editing can be cast as image-to-video generation: initializing from Stable Video Diffusion plus a new matching attention mechanism yields high-quality sketch, drag, and coarse-edit results with far l...
-
Generative Omnimatte: Learning to Decompose Video into Layers
A finetuned video diffusion model (Casper) removes objects along with their shadows and reflections, and a test-time optimization converts the outputs into editable omnimatte layers without assuming a static backgroun...
-
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
A dual-stream diffusion model trained with Blender-render conditioning, source masking, and object jittering performs 3D-grounded multi-object editing and compositing better than existing baselines on three video datasets.
-
Generative Image Layer Decomposition with Visual Effects
A diffusion-based model decomposes an image into a clean background and a transparent foreground layer that retains shadows and reflections, enabling object removal and spatial edits.
-
Pathways on the Image Manifold: Image Editing via Video Generation
Frame2Frame performs text-based image editing by generating a short video transition from the source image and selecting the best resulting frame, achieving competitive or better benchmark scores than single-image dif...
-
ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions
A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.
-
Motion Prompting: Controlling Video Generation with Motion Trajectories
A single-stage ControlNet on the Lumiere video model, conditioned only on dense point tracks, generalizes to sparse and dense trajectory control for object, camera, and transferred motions.
Discussion (0). Continue with ORCID to comment.