Pith. sign in

REVIEW 9 cited by

Novel View Synthesis with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.04628 v1 pith:SRGGOSVK submitted 2022-10-06 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords viewnovelmodelviewsconditioningdiffusionsingleconsistency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present 3DiM, a diffusion model for 3D novel view synthesis, which is able to translate a single input view into consistent and sharp completions across many views. The core component of 3DiM is a pose-conditional image-to-image diffusion model, which takes a source view and its pose as inputs, and generates a novel view for a target pose as output. 3DiM can generate multiple views that are 3D consistent using a novel technique called stochastic conditioning. The output views are generated autoregressively, and during the generation of each novel view, one selects a random conditioning view from the set of available views at each denoising step. We demonstrate that stochastic conditioning significantly improves the 3D consistency of a naive sampler for an image-to-image diffusion model, which involves conditioning on a single fixed view. We compare 3DiM to prior work on the SRN ShapeNet dataset, demonstrating that 3DiM's generated completions from a single view achieve much higher fidelity, while being approximately 3D consistent. We also introduce a new evaluation methodology, 3D consistency scoring, to measure the 3D consistency of a generated object by training a neural field on the model's output views. 3DiM is geometry free, does not rely on hyper-networks or test-time optimization for novel view synthesis, and allows a single model to easily scale to a large number of scenes.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A single-image feed-forward Gaussian splatting method that samples supports from predicted depth and decodes Gaussian attributes implicitly, improving cross-dataset large-baseline novel view synthesis.

  2. Online Neural Space Time Memory for Dynamic Novel View Synthesis

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Neural Space-Time Memory (NSTM) decouples low-frequency memory updates from per-frame synthesis with cross-view attention, enabling real-time minute-long dynamic novel view synthesis from multi-view streams.

  3. Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    SGC quantifies 3D geometric consistency of generated videos by measuring divergence among local camera poses estimated only on static background sub-regions.

  4. Edit360: 2D Image Edits to 3D Assets from Any Angle

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A training-free method that propagates a 2D edit applied at any chosen viewpoint across a full 360-degree orbit by fusing anchor-view and front-view video-diffusion trajectories.

  5. Bridging Rendering and Generative Modeling with Monte Carlo Transport Scheduling

    cs.CV 2026-02 reject novelty 5.0 of 10

    A common variance-time SDE aligns Monte Carlo rendering noise with diffusion-model denoising, enabling low-spp render refinement and stage-ordered material control.

  6. Perceive-Sample-Compress: Towards Real-Time 3D Gaussian Splatting

    cs.GR 2025-08 conditional novelty 4.0 of 10

    A three-stage perceive-sample-compress framework for 3D Gaussian Splatting improves rendering fidelity and storage efficiency across small and large scenes.

  7. EgoAnimate: Generating Human Animations from Egocentric top-down Views

    cs.CV 2025-07 conditional novelty 4.0 of 10

    EgoAnimate synthesizes a frontal T-pose image from an egocentric top-down photo using a fine-tuned Stable Diffusion model, then animates it with off-the-shelf image-to-motion methods to produce an animatable avatar.

  8. DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multi-view conditioning framework that improves controllable novel view synthesis and 3D reconstruction by injecting fused 3D latents into frozen image and video diffusion models.

  9. DBMovi-GS: Dynamic View Synthesis from Blurry Monocular Video via Sparse-Controlled Gaussian Splatting

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A Gaussian-splatting method densifies sparse points and combines object and camera motion models to produce sharp novel views from blurry monocular video.

Pith tools