GTA generates 3D worlds from single images via a two-stage video diffusion process that prioritizes geometry before appearance to improve structural consistency.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 5roles
method 1polarities
use method 1representative citing papers
Earth-o1 learns continuous atmospheric dynamics from ungridded observations and matches operational IFS forecast skill in hindcasts.
UniE2F conditions a pre-trained video diffusion model on event streams with inter-frame residual guidance to reconstruct, interpolate, and predict frames in a unified zero-shot framework.
GleSAM++ improves SAM robustness on degraded images by using generative enhancement, feature alignment, and adaptive degradation prediction while adding few parameters.
A mask-guided attention-editing framework, CPAM, separates object and background self-attention and uses null-text cross-attention to perform zero-shot, non-rigid edits on real images across SD1.5, SD2.1, and SDXL.
citing papers explorer
-
GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion
GTA generates 3D worlds from single images via a two-stage video diffusion process that prioritizes geometry before appearance to improve structural consistency.
-
Earth-o1: A Grid-free Observation-native Atmospheric World Model
Earth-o1 learns continuous atmospheric dynamics from ungridded observations and matches operational IFS forecast skill in hindcasts.
-
UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models
UniE2F conditions a pre-trained video diffusion model on event streams with inter-frame residual guidance to reconstruct, interpolate, and predict frames in a unified zero-shot framework.
-
Towards Any-Quality Image Segmentation via Generative and Adaptive Latent Space Enhancement
GleSAM++ improves SAM robustness on degraded images by using generative enhancement, feature alignment, and adaptive degradation prediction while adding few parameters.
-
CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing
A mask-guided attention-editing framework, CPAM, separates object and background self-attention and uses null-text cross-attention to perform zero-shot, non-rigid edits on real images across SD1.5, SD2.1, and SDXL.