REVIEW 13 cited by
Generic 3D Diffusion Adapter Using Controlled Multi-View Editing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Open-domain 3D object synthesis has been lagging behind image synthesis due to limited data and higher computational complexity. To bridge this gap, recent works have investigated multi-view diffusion but often fall short in either 3D consistency, visual quality, or efficiency. This paper proposes MVEdit, which functions as a 3D counterpart of SDEdit, employing ancestral sampling to jointly denoise multi-view images and output high-quality textured meshes. Built on off-the-shelf 2D diffusion models, MVEdit achieves 3D consistency through a training-free 3D Adapter, which lifts the 2D views of the last timestep into a coherent 3D representation, then conditions the 2D views of the next timestep using rendered views, without uncompromising visual quality. With an inference time of only 2-5 minutes, this framework achieves better trade-off between quality and speed than score distillation. MVEdit is highly versatile and extendable, with a wide range of applications including text/image-to-3D generation, 3D-to-3D editing, and high-quality texture synthesis. In particular, evaluations demonstrate state-of-the-art performance in both image-to-3D and text-guided texture generation tasks. Additionally, we introduce a method for fine-tuning 2D latent diffusion models on small 3D datasets with limited resources, enabling fast low-resolution text-to-3D initialization.
Forward citations
Cited by 13 Pith papers
-
mrCAD: Multimodal Refinement of Computer-aided Designs
mrCAD is a large dataset of multimodal human instructions for generating and refining CAD designs, and a benchmark showing that VLMs struggle with refinement instructions while humans excel at them.
-
SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control
Automatic view scheduling via a directed generation graph plus object-level identity and adherence conditioning enables high-quality outdoor 3DGS scenes from arbitrary input geometry without user camera paths.
-
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
A unified 3D multimodal model combines understanding, text-to-3D generation, instruction-guided editing, and part generation in one architecture, trained on an 87M-sample corpus, with claimed state-of-the-art results.
-
TanGO: Training-Free 3D Editing via Tangent-Space Guidance and Optimization
Per-token tangent-space steering, with strength set by velocity-direction mismatch, improves localized training-free 3D editing over global-scaling baselines.
-
EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning
An end-to-end 3D editing framework achieves high-fidelity local edits from coarse bounding boxes and 2D image prompts using region-aware loss reweighting and a large-scale parts-derived training dataset.
-
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
A zero-shot pipeline that reposes 3D meshes by generating text-conditioned target images with rewired multi-view attention and aligning mesh keypoints to them.
-
Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion
A two-stage diffusion model generates realistic depth noise on synthetic CAD data, and pretraining 3D networks on the resulting data improves few-shot real-world 3D tasks.
-
Pro3D-Editor : A Progressive-Views Perspective for Consistent and Precise 3D Editing
Pro3D-Editor chooses the most editing-salient view, propagates the edit to other key views with per-view LoRA experts, and refines the 3D scene, improving multi-view consistency.
-
CMD: Controllable Multiview Diffusion for 3D Editing and Progressive Generation
CMD performs local 3D editing and progressive 3D generation with a conditional multiview diffusion model, editing from a single image in about 20 seconds while preserving unmodified parts.
-
ScanEdit: Hierarchically-Guided Functional 3D Scan Editing
ScanEdit uses hierarchical scene graphs and LLM-based planning, placement, and optimization to rearrange objects in real-world 3D scans from text instructions.
-
PrEditor3D: Fast and Precise 3D Shape Editing
PrEditor3D performs fast, precise, training-free 3D shape editing by synchronously editing four rendered views, detecting the edited regions with Grounding DINO and SAM 2, and merging them into the original 3D voxel grid.
-
Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects
Text-guided 3D editing is reframed as multiview image inpainting, giving consistent edits in seconds instead of hours.
-
Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion
C33D blends a 3D model with an object category by generating a fused front view, then using texture and shape multi-view diffusion plus adaptive inversion to reconstruct a novel, consistent 3D model.
Discussion (0). Continue with ORCID to comment.