REVIEW 11 cited by
CityCraft: A Real Crafter for 3D City Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
City scene generation has gained significant attention in autonomous driving, smart city development, and traffic simulation. It helps enhance infrastructure planning and monitoring solutions. Existing methods have employed a two-stage process involving city layout generation, typically using Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), or Transformers, followed by neural rendering. These techniques often exhibit limited diversity and noticeable artifacts in the rendered city scenes. The rendered scenes lack variety, resembling the training images, resulting in monotonous styles. Additionally, these methods lack planning capabilities, leading to less realistic generated scenes. In this paper, we introduce CityCraft, an innovative framework designed to enhance both the diversity and quality of urban scene generation. Our approach integrates three key stages: initially, a diffusion transformer (DiT) model is deployed to generate diverse and controllable 2D city layouts. Subsequently, a Large Language Model(LLM) is utilized to strategically make land-use plans within these layouts based on user prompts and language guidelines. Based on the generated layout and city plan, we utilize the asset retrieval module and Blender for precise asset placement and scene construction. Furthermore, we contribute two new datasets to the field: 1)CityCraft-OSM dataset including 2D semantic layouts of urban areas, corresponding satellite images, and detailed annotations. 2) CityCraft-Buildings dataset, featuring thousands of diverse, high-quality 3D building assets. CityCraft achieves state-of-the-art performance in generating realistic 3D cities.
Forward citations
Cited by 11 Pith papers
-
PolyLayout: Hierarchical VLM-Guided Layout Generation Beyond Rectangular Rooms
A hybrid VLM-plus-solver pipeline produces more plausible furniture layouts than existing generators on standard rooms and extends to non-rectangular floor plans with doors and windows.
-
Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery
A satellite-to-3D-city pipeline that generates building entities with OSM geometry priors and MLLM/T2I appearance guidance reports strong gains over existing city-generation baselines.
-
EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
EarthCrafter generates 600-meter-scale 3D Earth scenes using separate latent diffusion models for structure and texture, conditioned on semantics, images, or nothing.
-
Sat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent Diffusion
Sat2City generates explicit 3D city geometry and appearance from a height-map condition using cascaded latent diffusion on sparse voxel grids, beating prior methods on a new synthetic city dataset.
-
Compositional Generative Model of Unbounded 4D Cities
CityDreamer4D is a compositional generative model that creates unbounded, temporally coherent 4D cities by separately generating static scenes, buildings, and vehicles with neural fields.
-
Proc-GS: Procedural Building Generation for City Assembly with 3D Gaussians
Proc-GS constrains 3D Gaussian Splatting with procedural code to extract reusable building assets and assemble new buildings and cities.
-
LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
A multimodal LLM framework generates interactive Unreal-based 3D environments from text and height maps, claiming superior layout accuracy and over 90x faster production than manual methods.
-
3D and 4D World Modeling: A Survey
A survey that defines 3D/4D world modeling, organizes methods into VideoGen, OccGen, and LiDARGen categories, and compiles datasets, metrics, and benchmark numbers.
-
WorldClaw: Agentic 3D Open-World Generation at Scale
WorldClaw generates globally coherent, locally detailed, editable 3D worlds from open-ended text using a coarse-to-fine agentic pipeline.
-
From 2D to 3D Cognition: A Brief Survey of General World Models
A survey proposing a two-pillar, three-capability framework that organizes recent AI world models by their transition from 2D visual prediction to 3D cognition.
-
3D Scene Generation: A Survey
The paper surveys 3D scene generation and organizes methods into four paradigms, with datasets, evaluation metrics, applications, and future directions.
Discussion (0). Continue with ORCID to comment.