Pith. sign in

REVIEW 2 cited by

HouseCat6D -- A Large-Scale Multi-Modal Category Level 6D Object Perception Dataset with Household Objects in Realistic Scenarios

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.10428 v5 pith:YIDCGDL4 submitted 2022-12-20 cs.CV

classification cs.CV
keywords posecategory-leveldatasetannotationsestimationhousecat6dhouseholdlarge-scale
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Estimating 6D object poses is a major challenge in 3D computer vision. Building on successful instance-level approaches, research is shifting towards category-level pose estimation for practical applications. Current category-level datasets, however, fall short in annotation quality and pose variety. Addressing this, we introduce HouseCat6D, a new category-level 6D pose dataset. It features 1) multi-modality with Polarimetric RGB and Depth (RGBD+P), 2) encompasses 194 diverse objects across 10 household categories, including two photometrically challenging ones, and 3) provides high-quality pose annotations with an error range of only 1.35 mm to 1.74 mm. The dataset also includes 4) 41 large-scale scenes with comprehensive viewpoint and occlusion coverage, 5) a checkerboard-free environment, and 6) dense 6D parallel-jaw robotic grasp annotations. Additionally, we present benchmark results for leading category-level pose estimation networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos

    cs.CV 2024-11 conditional novelty 7.0 of 10

    HOT3D releases 833 minutes of hardware-synchronized, egocentric multi-view video from real headsets with motion-capture ground truth for hands and objects, and shows multi-view baselines outperform single-view baselin...

  2. XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity

    cs.CV 2025-05 conditional novelty 6.0 of 10

    XYZ-IBD is a new industrial bin-picking benchmark with 273k pose annotations on 15 reflective, symmetric metal objects, and it demonstrates large performance drops for state-of-the-art pose estimators.

Pith tools