Pith. sign in

REVIEW 12 cited by

ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17994 v2 pith:VPHEIWLR submitted 2023-10-27 cs.CV cs.GR

classification cs.CVcs.GR
keywords novelscenesbackgroundsdataproposesynthesisviewzeronvs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a 3D-aware diffusion model, ZeroNVS, for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked backgrounds, we propose new techniques to address challenges introduced by in-the-wild multi-object scenes with complex backgrounds. Specifically, we train a generative prior on a mixture of data sources that capture object-centric, indoor, and outdoor scenes. To address issues from data mixture such as depth-scale ambiguity, we propose a novel camera conditioning parameterization and normalization scheme. Further, we observe that Score Distillation Sampling (SDS) tends to truncate the distribution of complex backgrounds during distillation of 360-degree scenes, and propose "SDS anchoring" to improve the diversity of synthesized novel views. Our model sets a new state-of-the-art result in LPIPS on the DTU dataset in the zero-shot setting, even outperforming methods specifically trained on DTU. We further adapt the challenging Mip-NeRF 360 dataset as a new benchmark for single-image novel view synthesis, and demonstrate strong performance in this setting. Our code and data are at http://kylesargent.github.io/zeronvs/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GPS as a Control Signal for Image Generation

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A diffusion model conditioned on GPS tags and text can generate location-specific images and reconstruct 3D landmarks via score distillation sampling, without explicit pose estimation.

  2. PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A single pixel-space diffusion model jointly performs 3D scene reconstruction and generation by supervising flow matching on rendered multi-view images, matching SOTA reconstruction and outperforming latent-space generation.

  3. LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis

    cs.CV 2026-03 accept novelty 6.0 of 10

    Initializing a highway encoder-decoder NVS network from VGGT 3D-aware features yields 31.4 PSNR on RealEstate10k with real-time decoding and optional unposed inputs.

  4. Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Diffusing in a unified 3D representation (geometry plus appearance and semantics) produces more cross-view-consistent 3D Gaussian scenes than 2D latent pipelines.

  5. CharacterShot: Controllable and Consistent 4D Character Animation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.

  6. Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.

  7. Towards In-the-wild 3D Plane Reconstruction from a Single Image

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ZeroPlane trains a Transformer plane reconstructor on 560K images spanning 10 indoor and outdoor datasets and outperforms prior methods in zero-shot evaluations on NYUv2, 7-Scenes, ParallelDomain, and ApolloScape.

  8. Pippo: High-Resolution Multi-View Humans from a Single Image

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A single-image multi-view diffusion transformer generates 1K-resolution turnaround views of humans, with attention biasing for many views and a new reprojection-error metric.

  9. Matrix3D: Large Photogrammetry Model All-in-One

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A single multi-modal diffusion transformer trained with masked learning performs pose estimation, depth prediction, and novel view synthesis in one model, reporting SOTA pose and NVS numbers.

  10. UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    UniAvatar integrates FLAME-based 3D motion rendering and SH-based illumination rendering into a diffusion talking-head model, enabling separate or combined control of motion and lighting in generated videos.

  11. Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.

  12. Dynamic View Synthesis as an Inverse Problem

    cs.CV 2025-06 reject novelty 3.0 of 10

    Dynamic view synthesis from a monocular video is achieved by redesigning the noise initialization of a pretrained video diffusion model using a recursive interpolation and a stochastic latent modulation.

Pith tools