Pith. sign in

REVIEW 3 cited by

Zero123-6D: Zero-shot Novel View Synthesis for RGB Category-level 6D Pose Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14279 v2 pith:DZW4RHSC submitted 2024-03-21 cs.CV

classification cs.CV
keywords posecategory-levelestimationobjectssynthesisworkzero-shotdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Estimating the pose of objects through vision is essential to make robotic platforms interact with the environment. Yet, it presents many challenges, often related to the lack of flexibility and generalizability of state-of-the-art solutions. Diffusion models are a cutting-edge neural architecture transforming 2D and 3D computer vision, outlining remarkable performances in zero-shot novel-view synthesis. Such a use case is particularly intriguing for reconstructing 3D objects. However, localizing objects in unstructured environments is rather unexplored. To this end, this work presents Zero123-6D, the first work to demonstrate the utility of Diffusion Model-based novel-view-synthesizers in enhancing RGB 6D pose estimation at category-level, by integrating them with feature extraction techniques. Novel View Synthesis allows to obtain a coarse pose that is refined through an online optimization method introduced in this work to deal with intra-category geometric differences. In such a way, the outlined method shows reduction in data requirements, removal of the necessity of depth information in zero-shot category-level 6D pose estimation task, and increased performance, quantitatively demonstrated through experiments on the CO3D dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Adding a category-level semantic shape prior, reconstructed from partial RGB-D input, improves 6D pose and size estimation for unseen objects on HouseCat6D and NOCS-REAL275.

  2. Segment Anything in Light Fields for Real-Time Applications via Constrained Prompting

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Segment Anything Model 2 is adapted to light fields by disparity-based mask propagation, semantic occlusion filtering, and reprompting, achieving view-consistent masks at real-time speed without retraining.

  3. Diffusion Features for Zero-Shot 6DoF Object Pose Estimation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Zero-shot 6DoF pose estimation using Stable Diffusion features improves average recall by up to 27% over a DINO-based baseline on LMO, YCBV, and TLESS.

Pith tools