Pith. sign in

REVIEW 2 cited by

ZeroComp: Zero-shot Object Compositing from Image Intrinsics via Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08168 v2 pith:5LN424NT submitted 2024-10-10 cs.CV

classification cs.CV
keywords zerocompcompositingimagesimagediffusionduringeffectiveintrinsic
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present ZeroComp, an effective zero-shot 3D object compositing approach that does not require paired composite-scene images during training. Our method leverages ControlNet to condition from intrinsic images and combines it with a Stable Diffusion model to utilize its scene priors, together operating as an effective rendering engine. During training, ZeroComp uses intrinsic images based on geometry, albedo, and masked shading, all without the need for paired images of scenes with and without composite objects. Once trained, it seamlessly integrates virtual 3D objects into scenes, adjusting shading to create realistic composites. We developed a high-quality evaluation dataset and demonstrate that ZeroComp outperforms methods using explicit lighting estimations and generative techniques in quantitative and human perception benchmarks. Additionally, ZeroComp extends to real and outdoor image compositing, even when trained solely on synthetic indoor data, showcasing its effectiveness in image compositing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IntrinsicEdit: Precise generative image manipulation in intrinsic space

    cs.GR 2025-05 conditional novelty 6.0 of 10

    Exact diffusion inversion plus prompt tuning lets users edit intrinsic image channels precisely while preserving identity and automatically resolving lighting effects.

  2. LumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relighting

    cs.CV 2024-11 conditional novelty 6.0 of 10

    LumiNet transfers lighting between indoor scenes from images alone by conditioning a diffusion model on latent intrinsics from the source and a lighting code from the target.

Pith tools