Pith. sign in

REVIEW 19 cited by

Wonder3D: Single Image to 3D using Cross-Domain Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15008 v3 pith:GYCG4PRJ submitted 2023-10-23 cs.CV

Wonder3D: Single Image to 3D using Cross-Domain Diffusion

classification cs.CV
keywords cross-domaindiffusionmulti-viewconsistencyefficiencygeometryhigh-qualityimages
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In this work, we introduce Wonder3D, a novel method for efficiently generating high-fidelity textured meshes from single-view images.Recent methods based on Score Distillation Sampling (SDS) have shown the potential to recover 3D geometry from 2D diffusion priors, but they typically suffer from time-consuming per-shape optimization and inconsistent geometry. In contrast, certain works directly produce 3D information via fast network inferences, but their results are often of low quality and lack geometric details. To holistically improve the quality, consistency, and efficiency of image-to-3D tasks, we propose a cross-domain diffusion model that generates multi-view normal maps and the corresponding color images. To ensure consistency, we employ a multi-view cross-domain attention mechanism that facilitates information exchange across views and modalities. Lastly, we introduce a geometry-aware normal fusion algorithm that extracts high-quality surfaces from the multi-view 2D representations. Our extensive evaluations demonstrate that our method achieves high-quality reconstruction results, robust generalization, and reasonably good efficiency compared to prior works.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Mat\'ern Noise for Triangulation-Agnostic Flow Matching on Meshes

    cs.GR 2026-05 unverdicted novelty 7.0

    Proposes discretized Matérn process noise for triangulation-agnostic flow matching on meshes with PoissonNet denoiser, tested on elastic states and humanoid poses for meshes exceeding one million triangles.

  2. Who Generated This 3D Asset? Learning Source Attribution for Generative 3D Models

    cs.CV 2026-05 unverdicted novelty 7.0

    Introduces the first passive source attribution benchmark for 22 generative 3D models and a Transformer achieving 97.22% accuracy under full supervision and 77.17% with 1% training data.

  3. SILICA: Repurposing Diffusion Priors for Joint Glass Segmentation and Depth Estimation

    cs.CV 2026-07 conditional novelty 6.0

    CLIP-conditioned single-step diffusion regression jointly does glass segmentation and affine-invariant depth, then aligns metric depth by masking bad sensor returns, beating prior glass methods on Mirage 18k and public sets.

  4. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  5. SOV-CAD: Stepwise Orthographic Views Guided CAD Modeling Sequence Reconstruction

    cs.CV 2026-07 conditional novelty 6.0

    SOV-CAD recovers CAD modeling sequences from orthographic images via Decision-Transformer offline RL conditioned on stepwise three-views plus sketch canvas and IoU-based rewards, beating holistic baselines with better...

  6. Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting

    cs.CV 2026-06 unverdicted novelty 6.0

    Flux-GS is a mobile-optimized 3D Gaussian Splatting method that compresses specular energy via Monte Carlo aggregation, recovers details with attribute-conditioned SH offsets, and uses multi-view guidance for densific...

  7. 3DMorph: Single-Image-Guided Local 3D Shape Editing and Morphing

    cs.CV 2026-06 unverdicted novelty 6.0

    3DMorph transfers local modifications from a single edited 2D image to the corresponding regions of a 3D mesh without training and supports shape morphing between original and edited versions.

  8. EMA: Effort Metric Attention for Anatomical Effort-Guided Human Motion Diffusion

    cs.CV 2026-05 unverdicted novelty 6.0

    EMA is a new cross-attention module that uses two kinematic metrics to approximate LMA effort factors and enables numerical, region-wise control of motion intensity in human motion diffusion models.

  9. UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

    cs.CV 2026-05 unverdicted novelty 6.0

    UniVidX unifies diverse video generation tasks into one conditional diffusion model using stochastic condition masking, decoupled gated LoRAs, and cross-modal self-attention.

  10. ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment

    cs.CV 2026-04 unverdicted novelty 6.0

    ReplicateAnyScene performs fully automated zero-shot video-to-compositional-3D reconstruction by cascading alignments of generic priors from vision foundation models across textual, visual, and spatial dimensions.

  11. TInR: Exploring Tool-Internalized Reasoning in Large Language Models

    cs.CL 2026-04 unverdicted novelty 6.0

    TInR-U internalizes tool knowledge into LLMs via bidirectional alignment, supervised fine-tuning, and reinforcement learning, outperforming standard tool-integrated reasoning in both in-domain and out-of-domain evaluations.

  12. Art3D: Training-Free 3D Generation from Flat-Colored Illustration

    cs.CV 2025-04 unverdicted novelty 6.0

    Art3D enhances flat-colored 2D illustrations with 3D illusion using pre-trained 2D model features and VLM realism evaluation, then generates 3D, while introducing the Flat-2D benchmark dataset.

  13. InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

    cs.CV 2024-04 unverdicted novelty 6.0

    InstantMesh produces diverse, high-quality 3D meshes from single images in seconds by combining a multi-view diffusion model with a sparse-view large reconstruction model and optimizing directly on meshes.

  14. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 5.0

    Hallo4D mitigates 3D/4D generation hallucinations via LMM-based detection, multi-model voting correction, and motion-aware optimization without retraining base generators.

  15. DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing

    cs.CV 2026-05 unverdicted novelty 5.0

    DreamEdit3D learns separate token embeddings for segmented object components via two-phase multi-view optimization to enable text-guided 3D editing with consistent image generation and mesh reconstruction.

  16. DecoRec: Decomposed 3D Scene Reconstruction from Single-View Images via Object-Level Diffusion

    cs.CV 2026-05 unverdicted novelty 5.0

    DecoRec decomposes single-view 3D scene reconstruction into per-object diffusion reconstructions followed by a differentiable rendering and diffusion-guided merging pipeline.

  17. Few-step Flow for 3D Generation via Marginal-Data Transport Distillation

    cs.CV 2025-09 conditional novelty 5.0

    MDT-dist distills a pretrained 3D flow model into a 1-2 step generator using velocity matching plus velocity distillation, cutting TRELLIS inference from 6.1s to 0.68s while approximately preserving generation quality.

  18. TInR: Exploring Tool-Internalized Reasoning in Large Language Models

    cs.CL 2026-04 unverdicted novelty 4.0

    TInR-U internalizes tool knowledge into an LLM via bidirectional alignment, SFT warm-up, and RL, claiming better in- and out-of-domain tool reasoning without external docs.

  19. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.