REVIEW 2 cited by
Text2Street: Controllable Text-to-image Generation for Street Views
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-to-image generation has made remarkable progress with the emergence of diffusion models. However, it is still a difficult task to generate images for street views based on text, mainly because the road topology of street scenes is complex, the traffic status is diverse and the weather condition is various, which makes conventional text-to-image models difficult to deal with. To address these challenges, we propose a novel controllable text-to-image framework, named \textbf{Text2Street}. In the framework, we first introduce the lane-aware road topology generator, which achieves text-to-map generation with the accurate road structure and lane lines armed with the counting adapter, realizing the controllable road topology generation. Then, the position-based object layout generator is proposed to obtain text-to-layout generation through an object-level bounding box diffusion strategy, realizing the controllable traffic object layout generation. Finally, the multiple control image generator is designed to integrate the road topology, object layout and weather description to realize controllable street-view image generation. Extensive experiments show that the proposed approach achieves controllable street-view text-to-image generation and validates the effectiveness of the Text2Street framework for street views.
Forward citations
Cited by 2 Pith papers
-
Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation
A single reference image guides a diffusion model to insert coherent objects into camera-plus-lidar driving scenes and to insert mammographic anomalies into new scans.
-
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
Driver gaze and accident-reason text are used to train a video diffusion model that can edit and generate egocentric crash videos with the correct causal participants, with a new large gaze dataset for accidents.
Discussion (0). Sign in to comment.