REVIEW 5 cited by
3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multi-view image diffusion models have significantly advanced open-domain 3D object generation. However, most existing models rely on 2D network architectures that lack inherent 3D biases, resulting in compromised geometric consistency. To address this challenge, we introduce 3D-Adapter, a plug-in module designed to infuse 3D geometry awareness into pretrained image diffusion models. Central to our approach is the idea of 3D feedback augmentation: for each denoising step in the sampling loop, 3D-Adapter decodes intermediate multi-view features into a coherent 3D representation, then re-encodes the rendered RGBD views to augment the pretrained base model through feature addition. We study two variants of 3D-Adapter: a fast feed-forward version based on Gaussian splatting and a versatile training-free version utilizing neural fields and meshes. Our extensive experiments demonstrate that 3D-Adapter not only greatly enhances the geometry quality of text-to-multi-view models such as Instant3D and Zero123++, but also enables high-quality 3D generation using the plain text-to-image Stable Diffusion. Furthermore, we showcase the broad application potential of 3D-Adapter by presenting high quality results in text-to-3D, image-to-3D, text-to-texture, and text-to-avatar tasks.
Forward citations
Cited by 5 Pith papers
-
VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image
VecSet-Edit is the first method to perform high-fidelity mesh editing from a single image by analyzing and manipulating spatial token subsets in a pre-trained VecSet LRM.
-
Pro3D-Editor : A Progressive-Views Perspective for Consistent and Precise 3D Editing
Pro3D-Editor chooses the most editing-salient view, propagates the edit to other key views with per-view LoRA experts, and refines the 3D scene, improving multi-view consistency.
-
SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
A test-time-trained feedforward model that propagates 2D edits onto 3D Gaussian attributes at interactive speeds.
-
ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents
ERMV edits 4D multi-view robot videos from one edited frame plus robot states, and VLA policies trained on the edited data show higher success rates in simulation and real robot tests.
-
Advancing high-fidelity 3D and Texture Generation with 2.5D latents
A 2.5D latent combining multiview RGB, normal, and coordinate images, generated by a mixture-of-LoRA fine-tuned Flux model, enables joint 3D geometry and texture generation from text or images.
Discussion (0). Continue with ORCID to comment.