REVIEW 13 cited by
GenXD: Generating Any 3D and 4D Scenes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent developments in 2D visual generation have been remarkably successful. However, 3D and 4D generation remain challenging in real-world applications due to the lack of large-scale 4D data and effective model design. In this paper, we propose to jointly investigate general 3D and 4D generation by leveraging camera and object movements commonly observed in daily life. Due to the lack of real-world 4D data in the community, we first propose a data curation pipeline to obtain camera poses and object motion strength from videos. Based on this pipeline, we introduce a large-scale real-world 4D scene dataset: CamVid-30K. By leveraging all the 3D and 4D data, we develop our framework, GenXD, which allows us to produce any 3D or 4D scene. We propose multiview-temporal modules, which disentangle camera and object movements, to seamlessly learn from both 3D and 4D data. Additionally, GenXD employs masked latent conditions to support a variety of conditioning views. GenXD can generate videos that follow the camera trajectory as well as consistent 3D views that can be lifted into 3D representations. We perform extensive evaluations across various real-world and synthetic datasets, demonstrating GenXD's effectiveness and versatility compared to previous methods in 3D and 4D generation.
Forward citations
Cited by 13 Pith papers
-
Wonderland: Navigating 3D Scenes from a Single Image
A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.
-
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.
-
4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
A diffusion model trained on 60,000 fitted 4D Gaussian Splatting human clips generates text-prompted, view-consistent dynamic humans directly in 4D, over 10x faster than video-first pipelines.
-
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
A single pixel-space diffusion model jointly performs 3D scene reconstruction and generation by supervising flow matching on rendered multi-view images, matching SOTA reconstruction and outperforming latent-space generation.
-
Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP
Diffusing in a unified 3D representation (geometry plus appearance and semantics) produces more cross-view-consistent 3D Gaussian scenes than 2D latent pipelines.
-
GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation
GeoWorld improves image-to-3D scene generation by conditioning a video-diffusion model on full-frame geometry features extracted by a multi-view geometry model, yielding higher PSNR/SSIM/LPIPS than prior methods.
-
4DNeX: Feed-Forward 4D Generative Modeling Made Easy
4DNeX generates dynamic 3D point clouds and matching RGB video from a single image by fine-tuning a pretrained video diffusion model on a large pseudo-annotated 4D dataset.
-
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...
-
4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion
4Real-Video generates consistent 4D video grids, frames across time and viewpoint, with a parallel two-stream diffusion transformer that synchronizes temporal and viewpoint token streams, achieving faster and higher-q...
-
CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models
CAT4D uses a multi-view video diffusion model to convert monocular video into multi-view video and reconstruct a dynamic 3D Gaussian scene, with competitive results on 4D reconstruction benchmarks.
-
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.
-
Impact-driven Context Filtering For Cross-file Code Completion
The manuscript's abstract claims a new code-completion filtering method, yet the body contains an unrelated 3D animation paper, leaving the claimed work unverifiable.
-
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.
Discussion (0). Continue with ORCID to comment.