Pith. sign in

REVIEW 2 cited by

ECNet: Effective Controllable Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.18417 v1 pith:LCHMUZSA submitted 2024-03-27 cs.CV

classification cs.CV
keywords controllablemodelstext-to-imageconditionaldiffusionguidanceinputsupervision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The conditional text-to-image diffusion models have garnered significant attention in recent years. However, the precision of these models is often compromised mainly for two reasons, ambiguous condition input and inadequate condition guidance over single denoising loss. To address the challenges, we introduce two innovative solutions. Firstly, we propose a Spatial Guidance Injector (SGI) which enhances conditional detail by encoding text inputs with precise annotation information. This method directly tackles the issue of ambiguous control inputs by providing clear, annotated guidance to the model. Secondly, to overcome the issue of limited conditional supervision, we introduce Diffusion Consistency Loss (DCL), which applies supervision on the denoised latent code at any given time step. This encourages consistency between the latent code at each time step and the input signal, thereby enhancing the robustness and accuracy of the output. The combination of SGI and DCL results in our Effective Controllable Network (ECNet), which offers a more accurate controllable end-to-end text-to-image generation framework with a more precise conditioning input and stronger controllable supervision. We validate our approach through extensive experiments on generation under various conditions, such as human body skeletons, facial landmarks, and sketches of general objects. The results consistently demonstrate that our method significantly enhances the controllability and robustness of the generated images, outperforming existing state-of-the-art controllable text-to-image models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HPSv3: Towards Wide-Spectrum Human Preference Score

    cs.CV 2025-08 conditional novelty 6.0 of 10

    HPSv3, trained on the new 1.08M-pair HPDv3 dataset, reaches 76.9% pairwise preference accuracy on its own test set and Spearman 0.94 against human model rankings, and is used to iteratively refine generated images (CoHP).

  2. Rethink Sparse Signals for Pose-guided Text-to-image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SP-Ctrl improves pose-guided text-to-image generation with sparse poses by learning keypoint embeddings and supervising keypoint attention maps, nearly matching dense depth-based control.

Pith tools