Pith. sign in

REVIEW 4 cited by

RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12908 v3 pith:CKLW4OYC submitted 2024-02-20 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modelsdiffusionrealcompotext-to-imagegenerationcompositionalityimagerealism
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion models have achieved remarkable advancements in text-to-image generation. However, existing models still have many difficulties when faced with multiple-object compositional generation. In this paper, we propose RealCompo, a new training-free and transferred-friendly text-to-image generation framework, which aims to leverage the respective advantages of text-to-image models and spatial-aware image diffusion models (e.g., layout, keypoints and segmentation maps) to enhance both realism and compositionality of the generated images. An intuitive and novel balancer is proposed to dynamically balance the strengths of the two models in denoising process, allowing plug-and-play use of any model without extra training. Extensive experiments show that our RealCompo consistently outperforms state-of-the-art text-to-image models and spatial-aware image diffusion models in multiple-object compositional generation while keeping satisfactory realism and compositionality of the generated images. Notably, our RealCompo can be seamlessly extended with a wide range of spatial-aware image diffusion models and stylized diffusion models. Our code is available at: https://github.com/YangLing0818/RealCompo

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LAION-SG: An Enhanced Large-Scale Dataset for Training Complex Image-Text Models with Structural Annotations

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 540,005-image dataset with GPT-4o-produced scene graph annotations improves compositional text-to-image generation when used to fine-tune SDXL-based models.

  2. FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FATE is a monocular full-head avatar system that improves Gaussian efficiency with sampling-based densification, enables UV-space texture editing through neural baking, and completes non-frontal views using SphereHead priors.

  3. Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models

    cs.CV 2025-07 reject novelty 5.0 of 10

    Inversion-DPO uses DDIM inversion to convert winning and losing images into noise trajectories, yielding a simpler DPO loss for diffusion model alignment that trains faster and improves text-to-image and compositional...

  4. MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A training-free multi-agent scene parser plus hierarchical region-aware diffusion improves complex text-to-image generation on T2I-CompBench over several Stable Diffusion baselines.

Pith tools