Pith. sign in

REVIEW 6 cited by

ControlCom: Controllable Image Composition using Diffusion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.10040 v1 pith:5STD2MRX submitted 2023-08-19 cs.CV

classification cs.CV
keywords imagecompositionforegroundcompositediffusionimagesmodelcontrollable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image composition targets at synthesizing a realistic composite image from a pair of foreground and background images. Recently, generative composition methods are built on large pretrained diffusion models to generate composite images, considering their great potential in image generation. However, they suffer from lack of controllability on foreground attributes and poor preservation of foreground identity. To address these challenges, we propose a controllable image composition method that unifies four tasks in one diffusion model: image blending, image harmonization, view synthesis, and generative composition. Meanwhile, we design a self-supervised training framework coupled with a tailored pipeline of training data preparation. Moreover, we propose a local enhancement module to enhance the foreground details in the diffusion model, improving the foreground fidelity of composite images. The proposed method is evaluated on both public benchmark and real-world data, which demonstrates that our method can generate more faithful and controllable composite images than existing approaches. The code and model will be available at https://github.com/bcmi/ControlCom-Image-Composition.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interact-Custom: Customized Human Object Interaction Image Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Interact-Custom generates customized human-object interaction images by first generating a foreground mask from the prompt and then using that mask to guide identity-preserving diffusion generation.

  2. HOComp: Interaction-Aware Human-Object Composition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A diffusion-transformer method that composes a foreground object into a human image with MLLM-chosen interaction regions, pose keypoint supervision, and appearance/background consistency losses, plus a new paired dataset.

  3. BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A dual-stream diffusion model trained with Blender-render conditioning, source masking, and object jittering performs 3D-grounded multi-object editing and compositing better than existing baselines on three video datasets.

  4. ORIDa: Object-centric Real-world Image Composition Dataset

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ORIDa is a public real-world dataset of 200 objects in 30,000+ images with multiple positions per scene, designed for object compositing training and evaluation.

  5. MV-CoLight: Efficient Object Compositing with Consistent Lighting and Shadow Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A feed-forward two-stage compositing framework that harmonizes inserted objects across views using a Hilbert-ordered Gaussian color mapping, trained and evaluated on a new 480k-scene synthetic dataset.

  6. CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A synthetic data pipeline for digital measurement devices plus a real-image benchmark improves LVLM reading performance from 32.92% to 96.04% ANLS for InternVL.

Pith tools