REVIEW 22 cited by
Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this work, we introduce Unique3D, a novel image-to-3D framework for efficiently generating high-quality 3D meshes from single-view images, featuring state-of-the-art generation fidelity and strong generalizability. Previous methods based on Score Distillation Sampling (SDS) can produce diversified 3D results by distilling 3D knowledge from large 2D diffusion models, but they usually suffer from long per-case optimization time with inconsistent issues. Recent works address the problem and generate better 3D results either by finetuning a multi-view diffusion model or training a fast feed-forward model. However, they still lack intricate textures and complex geometries due to inconsistency and limited generated resolution. To simultaneously achieve high fidelity, consistency, and efficiency in single image-to-3D, we propose a novel framework Unique3D that includes a multi-view diffusion model with a corresponding normal diffusion model to generate multi-view images with their normal maps, a multi-level upscale process to progressively improve the resolution of generated orthographic multi-views, as well as an instant and consistent mesh reconstruction algorithm called ISOMER, which fully integrates the color and geometric priors into mesh results. Extensive experiments demonstrate that our Unique3D significantly outperforms other image-to-3D baselines in terms of geometric and textural details.
Forward citations
Cited by 22 Pith papers
-
VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space
A training-free 3D editing method that inverts a source asset into TRELLIS latent space and replaces latents plus attention K/V tokens in unedited regions during re-denosing.
-
MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation
MIDI extends pre-trained image-to-3D object generators to multi-instance diffusion with a multi-instance attention mechanism, producing spatially coherent 3D scenes from a single image in one pass.
-
EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation
Mask-guided differential flow with a soft preservation loss enables training-free local 3D editing that keeps unedited regions close to the source asset.
-
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.
-
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.
-
Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
Ultra3D speeds up sparse-voxel 3D generation by generating a coarse mesh with compact VecSet latents, then refining voxel features with part-localized attention.
-
Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object
Zero-P-to-3 fuses multi-view diffusion, a restoration prior, and a coarse 3D Gaussian rendering in DDIM sampling, then refines with rotated views, and reports improved invisible-region reconstruction from partial-view...
-
Text-based Animatable 3D Avatars with Morphable Model Alignment
The authors introduce a two-stage pipeline, initialization from Portrait3D and dynamic refinement with a normal- and segmentation-conditioned ControlNet, and report better geometric and expression alignment than prior...
-
Consistent Flow Distillation for Text-to-3D Generation
Consistent Flow Distillation (CFD) guides 3D generation by denoising rendered views with a noise field that is consistent across camera views on the object surface.
-
Beyond Words: AuralLLM and SignMST-C for Sign Language Production and Bidirectional Accessibility
Two new Chinese Sign Language datasets and two models are proposed, with a claimed SOTA on PHOENIX2014-T that is unsupported by released artifacts.
-
IDOL: Instant Photorealistic 3D Human Creation from a Single Image
A single-image feed-forward model trains on a 100K-generated-subject multi-view dataset and reconstructs animatable 3D Gaussian human avatars in under one second.
-
HumanRig: Learning Automatic Rigging for Humanoid Character in a Large Scale Dataset
HumanRig provides a large uniform-topology humanoid rigging dataset and an automatic rigging method combining a 2D pose prior with a point transformer and cross-attention, outperforming GNN-based baselines in the repo...
-
Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects
Text-guided 3D editing is reframed as multiview image inpainting, giving consistent edits in seconds instead of hours.
-
DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters
A diffusion-based pipeline that rigs 3D Gaussian characters, including hair and clothing, using a newly curated dataset of 9,420 anime meshes.
-
LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models
Fine-tuning LLaMA-3.1-8B on an OBJ-as-text dataset lets one chat model both answer questions and generate simple 3D meshes, with no vocabulary expansion.
-
SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis
SynthDrive automatically mines images of rare objects, reconstructs them as 3D assets from a single view, and synthesizes driving footage that modestly improves detection of those objects.
-
TwoSquared: 4D Generation from 2D Image Pairs
A method that generates a temporally consistent, textured 4D mesh sequence from only two RGB images showing an object's initial and final poses.
-
MultiFloodSynth: Multi-Annotated Flood Synthetic Dataset Generation
MultiFloodSynth is a parameter-controllable synthetic flood dataset whose addition to real training data improves flood-level detection performance for YOLOv10 models.
-
MultiGO: Towards Multi-level Geometry Learning for Monocular 3D Textured Human Reconstruction
MultiGO combines skeleton, joint, and wrinkle level improvements on a Gaussian-based 3D human reconstruction model and reports SOTA results on CustomHuman and THuman3.0.
-
Gaussian Object Carver: Object-Compositional Gaussian Splatting with surfaces completion
A Gaussian splatting pipeline reconstructs indoor scenes as separable objects and uses a trained completion model to fill in occluded surfaces zero-shot.
-
ARM: Appearance Reconstruction Model for Relightable 3D Generation
ARM is a feed-forward model that reconstructs a 3D mesh and PBR texture maps (albedo, roughness, metalness) from sparse-view images, improving texture sharpness and relighting quality over prior single-image-to-3D methods.
-
3D Object Manipulation in a Single Image using Generative Models
A single-image object manipulation framework that reconstructs an object in 3D, refines its texture with a custom-tuned diffusion model, corrects background lighting, and renders the edited or animated object back int...
Discussion (0). Continue with ORCID to comment.