Pith. sign in

REVIEW 10 cited by

MVDiffusion++: A Dense High-resolution Multi-view Diffusion Model for Single or Sparse-view 3D Object Reconstruction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12712 v3 pith:RZUY4OJ5 submitted 2024-02-20 cs.CV

classification cs.CV
keywords mvdiffusiondensehigh-resolutionobjectreconstructionviewviewsarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper presents a neural architecture MVDiffusion++ for 3D object reconstruction that synthesizes dense and high-resolution views of an object given one or a few images without camera poses. MVDiffusion++ achieves superior flexibility and scalability with two surprisingly simple ideas: 1) A ``pose-free architecture'' where standard self-attention among 2D latent features learns 3D consistency across an arbitrary number of conditional and generation views without explicitly using camera pose information; and 2) A ``view dropout strategy'' that discards a substantial number of output views during training, which reduces the training-time memory footprint and enables dense and high-resolution view synthesis at test time. We use the Objaverse for training and the Google Scanned Objects for evaluation with standard novel view synthesis and 3D reconstruction metrics, where MVDiffusion++ significantly outperforms the current state of the arts. We also demonstrate a text-to-3D application example by combining MVDiffusion++ with a text-to-image generative model. The project page is at https://mvdiffusion-plusplus.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A multi-view diffusion pipeline that segments 3D objects into parts, completes occluded or invisible parts, and reconstructs them into a compositional 3D asset.

  2. MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation

    cs.GR 2025-05 conditional novelty 6.0 of 10

    A single photo is converted into a 3D mesh with PBR textures using a render-enhanced auto-encoder, two data-augmentation schemes, and a multi-view texturing pipeline, with the claimed result being the best quality amo...

  3. LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LiftImage3D generates small-motion video clips from one image, registers them with MASt3R, and fits a distortion-aware 3D Gaussian field whose canonical scene renders new views.

  4. Generalizable Single-view Object Pose Estimation by Two-side Generating and Matching

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Two-side generation and matching of intermediate views with a score distillation loss improves single-reference object pose estimation under large viewpoint changes.

  5. FreeMesh: Boosting Mesh Generation with Coordinates Merging

    cs.GR 2025-05 conditional novelty 5.0 of 10

    FreeMesh shows that rearranging coordinates into same-axis groups before byte-pair encoding reduces per-token entropy and sequence length, improving point-cloud conditioned mesh generation quality.

  6. CheapNVS: Real-Time On-Device Narrow-Baseline Novel View Synthesis

    cs.CV 2025-01 conditional novelty 5.0 of 10

    CheapNVS performs narrow-baseline single-view novel view synthesis on mobile devices by learning warping and inpainting in parallel from a shared latent space.

  7. UnCommon Objects in 3D

    cs.CV 2025-01 conditional novelty 5.0 of 10

    uCO3D is a large, diverse, high-quality real-object video dataset with 3D annotations that improves training of feedforward 3D reconstruction and text-to-3D models.

  8. Pragmatist: Multiview Conditional Diffusion Models for High-Fidelity 3D Reconstruction from Unposed Sparse Views

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Pragmatist turns sparse unposed photos of an object into a high-fidelity 3D mesh by generating consistent canonical views with a diffusion model, reconstructing a triplane mesh, then refining camera poses and texture ...

  9. Make-A-Texture: Fast Shape-Aware Texture Generation in 3 Seconds

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A texture-generation pipeline that produces 1024x1024 textures from text in 3.07 seconds on an H100, with quality comparable to SyncMVD and other prior methods.

  10. MV-Adapter: Multi-view Consistent Image Generation Made Easy

    cs.CV 2024-12 conditional novelty 5.0 of 10

    An adapter bolts multi-view generation onto frozen text-to-image diffusion models, producing consistent views at up to 768 resolution on SDXL.

Pith tools