Pith. sign in

REVIEW 22 cited by

Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20343 v3 pith:Q5Y3PEO6 submitted 2024-05-30 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords diffusionmodelunique3dimage-to-3dmeshmulti-viewresultsfidelity
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this work, we introduce Unique3D, a novel image-to-3D framework for efficiently generating high-quality 3D meshes from single-view images, featuring state-of-the-art generation fidelity and strong generalizability. Previous methods based on Score Distillation Sampling (SDS) can produce diversified 3D results by distilling 3D knowledge from large 2D diffusion models, but they usually suffer from long per-case optimization time with inconsistent issues. Recent works address the problem and generate better 3D results either by finetuning a multi-view diffusion model or training a fast feed-forward model. However, they still lack intricate textures and complex geometries due to inconsistency and limited generated resolution. To simultaneously achieve high fidelity, consistency, and efficiency in single image-to-3D, we propose a novel framework Unique3D that includes a multi-view diffusion model with a corresponding normal diffusion model to generate multi-view images with their normal maps, a multi-level upscale process to progressively improve the resolution of generated orthographic multi-views, as well as an instant and consistent mesh reconstruction algorithm called ISOMER, which fully integrates the color and geometric priors into mesh results. Extensive experiments demonstrate that our Unique3D significantly outperforms other image-to-3D baselines in terms of geometric and textural details.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space

    cs.CV 2025-08 conditional novelty 7.0 of 10

    A training-free 3D editing method that inverts a source asset into TRELLIS latent space and replaces latents plus attention K/V tokens in unedited regions during re-denosing.

  2. MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    MIDI extends pre-trained image-to-3D object generators to multi-instance diffusion with a multi-instance attention mechanism, producing spatially coherent 3D scenes from a single image in one pass.

  3. EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Mask-guided differential flow with a soft preservation loss enables training-free local 3D editing that keeps unedited regions close to the source asset.

  4. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  5. Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.

  6. Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Ultra3D speeds up sparse-voxel 3D generation by generating a coarse mesh with compact VecSet latents, then refining voxel features with part-localized attention.

  7. Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Zero-P-to-3 fuses multi-view diffusion, a restoration prior, and a coarse 3D Gaussian rendering in DDIM sampling, then refines with rotated views, and reports improved invisible-region reconstruction from partial-view...

  8. Text-based Animatable 3D Avatars with Morphable Model Alignment

    cs.CV 2025-04 conditional novelty 6.0 of 10

    The authors introduce a two-stage pipeline, initialization from Portrait3D and dynamic refinement with a normal- and segmentation-conditioned ControlNet, and report better geometric and expression alignment than prior...

  9. Consistent Flow Distillation for Text-to-3D Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Consistent Flow Distillation (CFD) guides 3D generation by denoising rendered views with a noise field that is consistent across camera views on the object surface.

  10. Beyond Words: AuralLLM and SignMST-C for Sign Language Production and Bidirectional Accessibility

    cs.CV 2025-01 reject novelty 6.0 of 10

    Two new Chinese Sign Language datasets and two models are proposed, with a claimed SOTA on PHOENIX2014-T that is unsupported by released artifacts.

  11. IDOL: Instant Photorealistic 3D Human Creation from a Single Image

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single-image feed-forward model trains on a 100K-generated-subject multi-view dataset and reconstructs animatable 3D Gaussian human avatars in under one second.

  12. HumanRig: Learning Automatic Rigging for Humanoid Character in a Large Scale Dataset

    cs.CV 2024-12 conditional novelty 6.0 of 10

    HumanRig provides a large uniform-topology humanoid rigging dataset and an automatic rigging method combining a 2D pose prior with a point transformer and cross-attention, outperforming GNN-based baselines in the repo...

  13. Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Text-guided 3D editing is reframed as multiview image inpainting, giving consistent edits in seconds instead of hours.

  14. DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A diffusion-based pipeline that rigs 3D Gaussian characters, including hair and clothing, using a newly curated dataset of 9,420 anime meshes.

  15. LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Fine-tuning LLaMA-3.1-8B on an OBJ-as-text dataset lets one chat model both answer questions and generate simple 3D meshes, with no vocabulary expansion.

  16. SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis

    cs.CV 2025-09 conditional novelty 5.0 of 10

    SynthDrive automatically mines images of rare objects, reconstructs them as 3D assets from a single view, and synthesizes driving footage that modestly improves detection of those objects.

  17. TwoSquared: 4D Generation from 2D Image Pairs

    cs.CV 2025-04 conditional novelty 5.0 of 10

    A method that generates a temporally consistent, textured 4D mesh sequence from only two RGB images showing an object's initial and final poses.

  18. MultiFloodSynth: Multi-Annotated Flood Synthetic Dataset Generation

    cs.CV 2025-02 conditional novelty 5.0 of 10

    MultiFloodSynth is a parameter-controllable synthetic flood dataset whose addition to real training data improves flood-level detection performance for YOLOv10 models.

  19. MultiGO: Towards Multi-level Geometry Learning for Monocular 3D Textured Human Reconstruction

    cs.CV 2024-12 conditional novelty 5.0 of 10

    MultiGO combines skeleton, joint, and wrinkle level improvements on a Gaussian-based 3D human reconstruction model and reports SOTA results on CustomHuman and THuman3.0.

  20. Gaussian Object Carver: Object-Compositional Gaussian Splatting with surfaces completion

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A Gaussian splatting pipeline reconstructs indoor scenes as separable objects and uses a trained completion model to fill in occluded surfaces zero-shot.

  21. ARM: Appearance Reconstruction Model for Relightable 3D Generation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    ARM is a feed-forward model that reconstructs a 3D mesh and PBR texture maps (albedo, roughness, metalness) from sparse-view images, improving texture sharpness and relighting quality over prior single-image-to-3D methods.

  22. 3D Object Manipulation in a Single Image using Generative Models

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A single-image object manipulation framework that reconstructs an object in 3D, refines its texture with a custom-tuned diffusion model, corrects background lighting, and renders the edited or animated object back int...

Pith tools