Pith. sign in

REVIEW 10 cited by

InstanceDiffusion: Instance-level Control for Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03290 v1 pith:BGMHOPKO submitted 2024-02-05 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modelscontrolinstance-levelinstancediffusiontext-to-imageimageinstanceblock
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Text-to-image diffusion models produce high quality images but do not offer control over individual instances in the image. We introduce InstanceDiffusion that adds precise instance-level control to text-to-image diffusion models. InstanceDiffusion supports free-form language conditions per instance and allows flexible ways to specify instance locations such as simple single points, scribbles, bounding boxes or intricate instance segmentation masks, and combinations thereof. We propose three major changes to text-to-image models that enable precise instance-level control. Our UniFusion block enables instance-level conditions for text-to-image models, the ScaleU block improves image fidelity, and our Multi-instance Sampler improves generations for multiple instances. InstanceDiffusion significantly surpasses specialized state-of-the-art models for each location condition. Notably, on the COCO dataset, we outperform previous state-of-the-art by 20.4% AP$_{50}^\text{box}$ for box inputs, and 25.4% IoU for mask inputs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation

    cs.CV 2025-07 conditional novelty 7.0 of 10

    UniMC uses tokenized instance conditions (class, box, keypoints) and a timestep-aware modulator in a DiT backbone to control multi-class human and animal image generation, trained and evaluated on the new HAIG-2.9M dataset.

  2. MARBLE: Material Recomposition and Blending in CLIP-Space

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MARBLE performs material blending and parametric material-attribute control by manipulating CLIP image embeddings and injecting them into a specific U-Net block of a pre-trained diffusion model.

  3. Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    PartCATSeg improves open-vocabulary part segmentation by separating object- and part-level cost volumes, adding a compositional loss, and injecting DINO structural guidance, achieving over 10% h-IoU gains on three benchmarks.

  4. TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward

    cs.AI 2026-05 conditional novelty 5.0 of 10

    A training-free, test-time guidance rule that tilts a diffusion model's samples toward regions where every concept in a prompt is jointly present; it improves several T2ICompBench categories over prior correctors and ...

  5. FLORA: Efficient Synthetic Data Generation for Object Detection in Low-Data Regimes via finetuning Flux LoRA

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A LoRA fine-tuned Flux inpainting pipeline generates synthetic object detection images that outperform ODGEN's 10x larger synthetic set in downstream mAP.

  6. Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A probabilistic overlap measure for object positions yields a human-aligned spatial relationship metric and a training-free generation guidance method for text-to-image models.

  7. CountDiffusion: Text-to-Image Synthesis with Training-Free Counting-Guidance Diffusion

    cs.CV 2025-05 reject novelty 5.0 of 10

    CountDiffusion improves object-count accuracy in text-to-image diffusion by detecting objects in a one-step predicted image and applying attention-map guidance to add or remove instances.

  8. GSEditPro: 3D Gaussian Splatting Editing with Attention-based Progressive Localization

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A text-driven 3D editing framework that tags 3D Gaussian points via cross-attention and uses SDS plus pseudo-GT guidance to edit only the target region.

  9. MixDiffusion: Mixing Diffusion-based Uni-condition Text-to-Image Generation Models for Multi-condition Image Synthesis

    cs.CV 2026-07 conditional novelty 4.0 of 10

    MixDiffusion derives a joint noise prediction as the sum of per-condition noise estimates minus the base model, enabling multi-condition control without training.

  10. SmartSpatial: Enhancing the 3D Spatial Arrangement Capabilities of Stable Diffusion Models and Introducing a Novel 3D Spatial Evaluation Framework

    cs.CV 2025-01 reject novelty 4.0 of 10

    SmartSpatial combines depth injection and attention guidance in Stable Diffusion with a new VLM-based spatial metric, but reported improvements are not statistically significant per the paper's own p-value statement.

Pith tools