Pith. sign in

REVIEW 6 cited by

SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.04689 v1 pith:QPKFRM5K submitted 2025-01-08 cs.CV cs.GR

classification cs.CVcs.GR
keywords spar3dmodelingpointmethodscloudsdirectionsgenerativemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the problem of single-image 3D object reconstruction. Recent works have diverged into two directions: regression-based modeling and generative modeling. Regression methods efficiently infer visible surfaces, but struggle with occluded regions. Generative methods handle uncertain regions better by modeling distributions, but are computationally expensive and the generation is often misaligned with visible surfaces. In this paper, we present SPAR3D, a novel two-stage approach aiming to take the best of both directions. The first stage of SPAR3D generates sparse 3D point clouds using a lightweight point diffusion model, which has a fast sampling speed. The second stage uses both the sampled point cloud and the input image to create highly detailed meshes. Our two-stage design enables probabilistic modeling of the ill-posed single-image 3D task while maintaining high computational efficiency and great output fidelity. Using point clouds as an intermediate representation further allows for interactive user edits. Evaluated on diverse datasets, SPAR3D demonstrates superior performance over previous state-of-the-art methods, at an inference speed of 0.7 seconds. Project page with code and model: https://spar3d.github.io

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    ELSA3D introduces elastic semantic anchoring via sparse anchor tokens and a scale-aware octree tokenizer to unify 3D generation and captioning at reduced computational cost.

  2. Benefit from Reference: Retrieval-Augmented Cross-modal Point Cloud Completion

    cs.CV 2025-07 reject novelty 6.0 of 10

    A retrieval-augmented cross-modal framework with structural shared encoding and gated reference priors is claimed to achieve state-of-the-art point cloud completion, although the evaluation may be tainted by same-obje...

  3. NeuraLeaf: Neural Parametric Leaf Models with Shape and Deformation Disentanglement

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A neural parametric model for leaves that disentangles 2D base shape from 3D deformation, learned from 2D image data plus a new 300-pair 3D scan dataset, and fitted to observations for reconstruction.

  4. T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A zero-training framework that adaptively selects spatial representation extractors per object and per task stage improves real-world robot manipulation success and efficiency over fixed-representation baselines.

  5. AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

    cs.RO 2025-06 conditional novelty 6.0 of 10

    AntiGrounding lifts candidate robot trajectories into the VLM's visual space via multi-view rendering and structured VQA, and reports 57.5% average success across eight manipulation tasks, beating three intermediate-r...

  6. 3D Arena: An Open Platform for Generative 3D Evaluation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A crowdsourced voting platform with 123,000 votes reveals that people judge AI-generated 3D assets mainly by visual appearance rather than technical quality.

Pith tools