Pith. sign in

REVIEW 2 cited by

Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17002 v3 pith:GKUP5IKM submitted 2023-11-28 cs.CV

classification cs.CV
keywords panelrannidiffusiongenerationgeneratorinstructionslanguagemiddleware
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Existing text-to-image (T2I) diffusion models usually struggle in interpreting complex prompts, especially those with quantity, object-attribute binding, and multi-subject descriptions. In this work, we introduce a semantic panel as the middleware in decoding texts to images, supporting the generator to better follow instructions. The panel is obtained through arranging the visual concepts parsed from the input text by the aid of large language models, and then injected into the denoising network as a detailed control signal to complement the text condition. To facilitate text-to-panel learning, we come up with a carefully designed semantic formatting protocol, accompanied by a fully-automatic data preparation pipeline. Thanks to such a design, our approach, which we call Ranni, manages to enhance a pre-trained T2I generator regarding its textual controllability. More importantly, the introduction of the generative middleware brings a more convenient form of interaction (i.e., directly adjusting the elements in the panel or using language instructions) and further allows users to finely customize their generation, based on which we develop a practical system and showcase its potential in continuous generation and chatting-based editing. Our project page is at https://ranni-t2i.github.io/Ranni.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers

    cs.CV 2025-05 reject novelty 6.0 of 10

    LayouSyn is a text-to-layout pipeline using a lightweight open-source LLM for object extraction and an aspect-aware diffusion Transformer for bounding-box generation, reporting SOTA on NSR-1K and COCO-GR layout metrics.

  2. Instruction-augmented Multimodal Alignment for Image-Text and Element Matching

    cs.CV 2025-04 conditional novelty 5.0 of 10

    A fine-tuned multimodal score model with soft Q-Align scoring, element-conditioned prompts, and self-training on validation pseudo-labels takes first place in NTIRE 2025 Track 1 image-text alignment.

Pith tools