REVIEW 6 cited by
CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Incorporating a customized object into image generation presents an attractive feature in text-to-image generation. However, existing optimization-based and encoder-based methods are hindered by drawbacks such as time-consuming optimization, insufficient identity preservation, and a prevalent copy-pasting effect. To overcome these limitations, we introduce CustomNet, a novel object customization approach that explicitly incorporates 3D novel view synthesis capabilities into the object customization process. This integration facilitates the adjustment of spatial position relationships and viewpoints, yielding diverse outputs while effectively preserving object identity. Moreover, we introduce delicate designs to enable location control and flexible background control through textual descriptions or specific user-defined images, overcoming the limitations of existing 3D novel view synthesis methods. We further leverage a dataset construction pipeline that can better handle real-world objects and complex backgrounds. Equipped with these designs, our method facilitates zero-shot object customization without test-time optimization, offering simultaneous control over the viewpoints, location, and background. As a result, our CustomNet ensures enhanced identity preservation and generates diverse, harmonious outputs.
Forward citations
Cited by 6 Pith papers
-
Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation
A single reference image guides a diffusion model to insert coherent objects into camera-plus-lidar driving scenes and to insert mammographic anomalies into new scans.
-
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
A dual-stream diffusion model trained with Blender-render conditioning, source masking, and object jittering performs 3D-grounded multi-object editing and compositing better than existing baselines on three video datasets.
-
Multitwine: Multi-Object Compositing with Text and Layout Control
A single diffusion model simultaneously composites multiple objects into a scene with text and layout control, outperforming sequential insertion on interacting cases.
-
PIXELS: Progressive Image Xemplar-based Editing with Latent Surgery
PIXELS performs exemplar-based image editing at inference time with a progressive latent mixing mask that provides region-wise strength control, needs no training, and accepts any number of exemplars.
-
MObI: Multimodal Object Inpainting Using Diffusion Models
MObI jointly inpaints camera and lidar views of driving scenes, inserting objects from a single reference image at a user-specified 3D bounding box.
-
DIVE: Taming DINO for Subject-Driven Video Editing
DIVE uses DINOv2 feature maps as automatic video correspondences to carry source motion, while LoRA adapters carry the target identity.
Discussion (0). Continue with ORCID to comment.