Pith. sign in

REVIEW 2 cited by

SPARC: Sparse Render-and-Compare for CAD model alignment in a single RGB image

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.01044 v1 pith:OXAQQJSB submitted 2022-10-03 cs.CV

classification cs.CV
keywords modelimageposealignmentcoordinatesinformationnormalisedobject
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Estimating 3D shapes and poses of static objects from a single image has important applications for robotics, augmented reality and digital content creation. Often this is done through direct mesh predictions which produces unrealistic, overly tessellated shapes or by formulating shape prediction as a retrieval task followed by CAD model alignment. Directly predicting CAD model poses from 2D image features is difficult and inaccurate. Some works, such as ROCA, regress normalised object coordinates and use those for computing poses. While this can produce more accurate pose estimates, predicting normalised object coordinates is susceptible to systematic failure. Leveraging efficient transformer architectures we demonstrate that a sparse, iterative, render-and-compare approach is more accurate and robust than relying on normalised object coordinates. For this we combine 2D image information including sparse depth and surface normal values which we estimate directly from the image with 3D CAD model information in early fusion. In particular, we reproject points sampled from the CAD model in an initial, random pose and compute their depth and surface normal values. This combined information is the input to a pose prediction network, SPARC-Net which we train to predict a 9 DoF CAD model pose update. The CAD model is reprojected again and the next pose update is predicted. Our alignment procedure converges after just 3 iterations, improving the state-of-the-art performance on the challenging real-world dataset ScanNet from 25.0% to 31.8% instance alignment accuracy. Code will be released at https://github.com/florianlanger/SPARC .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hierarchical coarse-to-fine pipeline refines single-image 3D scenes component-by-component using a learned voxel super-resolution model conditioned on coarse voxels.

  2. DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single color photo becomes a 3D room scene by generating each object with a depth-conditioned diffusion model, then optimizing object poses against estimated depth.

Pith tools