Pith. sign in

REVIEW 2 cited by

CPS++: Improving Class-level 6D Pose and Shape Estimation From Monocular Images With Self-Supervised Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.05848 v3 pith:DKZNQZJI submitted 2020-03-12 cs.CV

classification cs.CV
keywords class-levelestimationposemonocularannotationsdatalearningmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contemporary monocular 6D pose estimation methods can only cope with a handful of object instances. This naturally hampers possible applications as, for instance, robots seamlessly integrated in everyday processes necessarily require the ability to work with hundreds of different objects. To tackle this problem of immanent practical relevance, we propose a novel method for class-level monocular 6D pose estimation, coupled with metric shape retrieval. Unfortunately, acquiring adequate annotations is very time-consuming and labor intensive. This is especially true for class-level 6D pose estimation, as one is required to create a highly detailed reconstruction for all objects and then annotate each object and scene using these models. To overcome this shortcoming, we additionally propose the idea of synthetic-to-real domain transfer for class-level 6D poses by means of self-supervised learning, which removes the burden of collecting numerous manual annotations. In essence, after training our proposed method fully supervised with synthetic data, we leverage recent advances in differentiable rendering to self-supervise the model with unannotated real RGB-D data to improve latter inference. We experimentally demonstrate that we can retrieve precise 6D poses and metric shapes from a single RGB image.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Adding a category-level semantic shape prior, reconstructed from partial RGB-D input, improves 6D pose and size estimation for unseen objects on HouseCat6D and NOCS-REAL275.

  2. Glissando-Net: Deep sinGLe vIew category level poSe eStimation ANd 3D recOnstruction

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A jointly trained image-point-cloud network estimates category-level 6D pose and 3D shape from a single RGB image, reporting better results than the closest prior work CPS on most NOCS benchmarks.

Pith tools