Pith. sign in

REVIEW 13 cited by

Generic 3D Diffusion Adapter Using Controlled Multi-View Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12032 v2 pith:NL7QRLHV submitted 2024-03-18 cs.CV cs.GR

classification cs.CVcs.GR
keywords diffusionmulti-viewmveditqualitysynthesisviewsachievesadapter
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Open-domain 3D object synthesis has been lagging behind image synthesis due to limited data and higher computational complexity. To bridge this gap, recent works have investigated multi-view diffusion but often fall short in either 3D consistency, visual quality, or efficiency. This paper proposes MVEdit, which functions as a 3D counterpart of SDEdit, employing ancestral sampling to jointly denoise multi-view images and output high-quality textured meshes. Built on off-the-shelf 2D diffusion models, MVEdit achieves 3D consistency through a training-free 3D Adapter, which lifts the 2D views of the last timestep into a coherent 3D representation, then conditions the 2D views of the next timestep using rendered views, without uncompromising visual quality. With an inference time of only 2-5 minutes, this framework achieves better trade-off between quality and speed than score distillation. MVEdit is highly versatile and extendable, with a wide range of applications including text/image-to-3D generation, 3D-to-3D editing, and high-quality texture synthesis. In particular, evaluations demonstrate state-of-the-art performance in both image-to-3D and text-guided texture generation tasks. Additionally, we introduce a method for fine-tuning 2D latent diffusion models on small 3D datasets with limited resources, enabling fast low-resolution text-to-3D initialization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. mrCAD: Multimodal Refinement of Computer-aided Designs

    cs.AI 2025-04 conditional novelty 7.0 of 10

    mrCAD is a large dataset of multimodal human instructions for generating and refining CAD designs, and a benchmark showing that VLMs struggle with refinement instructions while humans excel at them.

  2. SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control

    cs.GR 2026-07 conditional novelty 6.5 of 10

    Automatic view scheduling via a directed generation graph plus object-level identity and adherence conditioning enables high-quality outdoor 3DGS scenes from arbitrary input geometry without user camera paths.

  3. Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A unified 3D multimodal model combines understanding, text-to-3D generation, instruction-guided editing, and part generation in one architecture, trained on an 87M-sample corpus, with claimed state-of-the-art results.

  4. TanGO: Training-Free 3D Editing via Tangent-Space Guidance and Optimization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Per-token tangent-space steering, with strength set by velocity-direction mismatch, improves localized training-free 3D editing over global-scaling baselines.

  5. EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An end-to-end 3D editing framework achieves high-fidelity local edits from coarse bounding boxes and 2D image prompts using region-aware loss reweighting and a large-scale parts-derived training dataset.

  6. Articulate3D: Zero-Shot Text-Driven 3D Object Posing

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A zero-shot pipeline that reposes 3D meshes by generating text-conditioned target images with rewired multi-view attention and aligning mesh keypoints to them.

  7. Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage diffusion model generates realistic depth noise on synthetic CAD data, and pretraining 3D networks on the resulting data improves few-shot real-world 3D tasks.

  8. Pro3D-Editor : A Progressive-Views Perspective for Consistent and Precise 3D Editing

    cs.GR 2025-05 conditional novelty 6.0 of 10

    Pro3D-Editor chooses the most editing-salient view, propagates the edit to other key views with per-view LoRA experts, and refines the 3D scene, improving multi-view consistency.

  9. CMD: Controllable Multiview Diffusion for 3D Editing and Progressive Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CMD performs local 3D editing and progressive 3D generation with a conditional multiview diffusion model, editing from a single image in about 20 seconds while preserving unmodified parts.

  10. ScanEdit: Hierarchically-Guided Functional 3D Scan Editing

    cs.CV 2025-04 conditional novelty 6.0 of 10

    ScanEdit uses hierarchical scene graphs and LLM-based planning, placement, and optimization to rearrange objects in real-world 3D scans from text instructions.

  11. PrEditor3D: Fast and Precise 3D Shape Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PrEditor3D performs fast, precise, training-free 3D shape editing by synchronously editing four rendered views, detecting the edited regions with Grounding DINO and SAM 2, and merging them into the original 3D voxel grid.

  12. Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Text-guided 3D editing is reframed as multiview image inpainting, giving consistent edits in seconds instead of hours.

  13. Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion

    cs.CV 2025-09 conditional novelty 5.0 of 10

    C33D blends a 3D model with an object category by generating a fused front view, then using texture and shape multi-view diffusion plus adaptive inversion to reconstruct a novel, consistent 3D model.

Pith tools