SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 3years
2026 3roles
dataset 1polarities
use dataset 1representative citing papers
Proprioception and multi-contact touch, fused with a physics-guided conditional diffusion model over a Structure-VAE SDF latent space, improve metric amodal object reconstruction under severe hand occlusion versus vision-only baselines.
PAD synthesizes 3D geometry in observation space via depth unprojection as anchor to eliminate pose ambiguity in image-to-3D generation.
citing papers explorer
-
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.
-
Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
Proprioception and multi-contact touch, fused with a physics-guided conditional diffusion model over a Structure-VAE SDF latent space, improve metric amodal object reconstruction under severe hand occlusion versus vision-only baselines.
-
Pose-Aware Diffusion for 3D Generation
PAD synthesizes 3D geometry in observation space via depth unprojection as anchor to eliminate pose ambiguity in image-to-3D generation.