Pith. sign in

REVIEW 13 cited by

GenXD: Generating Any 3D and 4D Scenes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.02319 v2 pith:7CBV6XXJ submitted 2024-11-04 cs.CV cs.AI

classification cs.CVcs.AI
keywords datagenxdcameragenerationreal-worldobjectproposelack
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent developments in 2D visual generation have been remarkably successful. However, 3D and 4D generation remain challenging in real-world applications due to the lack of large-scale 4D data and effective model design. In this paper, we propose to jointly investigate general 3D and 4D generation by leveraging camera and object movements commonly observed in daily life. Due to the lack of real-world 4D data in the community, we first propose a data curation pipeline to obtain camera poses and object motion strength from videos. Based on this pipeline, we introduce a large-scale real-world 4D scene dataset: CamVid-30K. By leveraging all the 3D and 4D data, we develop our framework, GenXD, which allows us to produce any 3D or 4D scene. We propose multiview-temporal modules, which disentangle camera and object movements, to seamlessly learn from both 3D and 4D data. Additionally, GenXD employs masked latent conditions to support a variety of conditioning views. GenXD can generate videos that follow the camera trajectory as well as consistent 3D views that can be lifted into 3D representations. We perform extensive evaluations across various real-world and synthetic datasets, demonstrating GenXD's effectiveness and versatility compared to previous methods in 3D and 4D generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Wonderland: Navigating 3D Scenes from a Single Image

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.

  2. UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.

  3. 4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion model trained on 60,000 fitted 4D Gaussian Splatting human clips generates text-prompted, view-consistent dynamic humans directly in 4D, over 10x faster than video-first pipelines.

  4. PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A single pixel-space diffusion model jointly performs 3D scene reconstruction and generation by supervising flow matching on rendered multi-view images, matching SOTA reconstruction and outperforming latent-space generation.

  5. Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Diffusing in a unified 3D representation (geometry plus appearance and semantics) produces more cross-view-consistent 3D Gaussian scenes than 2D latent pipelines.

  6. GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    GeoWorld improves image-to-3D scene generation by conditioning a video-diffusion model on full-frame geometry features extracted by a multi-view geometry model, yielding higher PSNR/SSIM/LPIPS than prior methods.

  7. 4DNeX: Feed-Forward 4D Generative Modeling Made Easy

    cs.CV 2025-08 conditional novelty 6.0 of 10

    4DNeX generates dynamic 3D point clouds and matching RGB video from a single image by fine-tuning a pretrained video diffusion model on a large pseudo-annotated 4D dataset.

  8. You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale

    cs.CV 2024-12 reject novelty 6.0 of 10

    See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...

  9. 4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion

    cs.CV 2024-12 conditional novelty 6.0 of 10

    4Real-Video generates consistent 4D video grids, frames across time and viewpoint, with a parallel two-stream diffusion transformer that synchronizes temporal and viewpoint token streams, achieving faster and higher-q...

  10. CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    CAT4D uses a multi-view video diffusion model to convert monocular video into multi-view video and reconstruct a dynamic 3D Gaussian scene, with competitive results on 4D reconstruction benchmarks.

  11. AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.

  12. Impact-driven Context Filtering For Cross-file Code Completion

    cs.SE 2025-08 unverdicted novelty 5.0 of 10

    The manuscript's abstract claims a new code-completion filtering method, yet the body contains an unrelated 3D animation paper, leaving the claimed work unverifiable.

  13. Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.

Pith tools