REVIEW 6 cited by
BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent text-to-image diffusion models have demonstrated an astonishing capacity to generate high-quality images. However, researchers mainly studied the way of synthesizing images with only text prompts. While some works have explored using other modalities as conditions, considerable paired data, e.g., box/mask-image pairs, and fine-tuning time are required for nurturing models. As such paired data is time-consuming and labor-intensive to acquire and restricted to a closed set, this potentially becomes the bottleneck for applications in an open world. This paper focuses on the simplest form of user-provided conditions, e.g., box or scribble. To mitigate the aforementioned problem, we propose a training-free method to control objects and contexts in the synthesized images adhering to the given spatial conditions. Specifically, three spatial constraints, i.e., Inner-Box, Outer-Box, and Corner Constraints, are designed and seamlessly integrated into the denoising step of diffusion models, requiring no additional training and massive annotated layout data. Extensive experimental results demonstrate that the proposed constraints can control what and where to present in the images while retaining the ability of Diffusion models to synthesize with high fidelity and diverse concept coverage. The code is publicly available at https://github.com/showlab/BoxDiff.
Forward citations
Cited by 6 Pith papers
-
Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy
IP-CIR creates imagined proxy images from a query image and caption via LLM-based layout and conditional generation, then blends proxy, query, and text features to improve zero-shot composed image retrieval.
-
TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
A training-free, test-time guidance rule that tilts a diffusion model's samples toward regions where every concept in a prompt is jointly present; it improves several T2ICompBench categories over prior correctors and ...
-
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
A probabilistic overlap measure for object positions yields a human-aligned spatial relationship metric and a training-free generation guidance method for text-to-image models.
-
Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model
A reactive closed-loop driving simulator that uses a diffusion renderer with retrieval from real recordings, plus a nuPlan behavioral controller, to generate sensor images in response to an end-to-end driving model's actions.
-
AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks
AnySynth is a single synthetic-data pipeline that produces layouts, images, and annotations for multiple vision tasks, and its data improves benchmark scores by small but consistent margins.
-
SmartSpatial: Enhancing the 3D Spatial Arrangement Capabilities of Stable Diffusion Models and Introducing a Novel 3D Spatial Evaluation Framework
SmartSpatial combines depth injection and attention guidance in Stable Diffusion with a new VLM-based spatial metric, but reported improvements are not statistically significant per the paper's own p-value statement.
Discussion (0). Continue with ORCID to comment.