Pith. sign in

REVIEW 5 cited by

Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstruction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.06714 v4 pith:QM5LM4EA submitted 2023-04-13 cs.CV

classification cs.CV
keywords diffusiongenerationnerfreconstructionmodelpriorapproachimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D-aware image synthesis encompasses a variety of tasks, such as scene generation and novel view synthesis from images. Despite numerous task-specific methods, developing a comprehensive model remains challenging. In this paper, we present SSDNeRF, a unified approach that employs an expressive diffusion model to learn a generalizable prior of neural radiance fields (NeRF) from multi-view images of diverse objects. Previous studies have used two-stage approaches that rely on pretrained NeRFs as real data to train diffusion models. In contrast, we propose a new single-stage training paradigm with an end-to-end objective that jointly optimizes a NeRF auto-decoder and a latent diffusion model, enabling simultaneous 3D reconstruction and prior learning, even from sparsely available views. At test time, we can directly sample the diffusion prior for unconditional generation, or combine it with arbitrary observations of unseen objects for NeRF reconstruction. SSDNeRF demonstrates robust results comparable to or better than leading task-specific methods in unconditional generation and single/sparse-view 3D reconstruction.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation

    cs.CV 2023-09 unverdicted novelty 7.0 of 10

    DreamGaussian creates high-quality textured 3D meshes from single-view images in 2 minutes via generative Gaussian Splatting with mesh extraction and UV refinement.

  2. BoostDream: Efficient Refining for High-Quality Text-to-3D Generation from Multi-View Diffusion

    cs.CV 2024-01 unverdicted novelty 6.0 of 10

    BoostDream refines coarse feed-forward text-to-3D assets via 3D distillation, multi-view SDS loss from a 2D diffusion model, and prompt-consistent normal maps to produce higher-quality results more efficiently than st...

  3. Predicting 3D structure by latent posterior sampling

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A two-stage latent-variable model uses diffusion-based score matching to sample 3D scenes from posteriors conditioned on varied observations via volumetric rendering likelihoods.

  4. Predicting 3D structure by latent posterior sampling

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A latent-variable approach uses diffusion models on NeRF-encoded scene representations to perform posterior sampling for 3D reconstruction from single-view, multi-view, noisy, sparse-pixel, or sparse-depth inputs.

  5. Predicting 3D structure by latent posterior sampling

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A two-stage method trains NeRF latents then a diffusion prior to sample posteriors for 3D reconstruction from varied observations including single-view, multi-view, noisy, sparse pixels, and sparse depth.

Pith tools