Pith. sign in

REVIEW 2 cited by

Real-time 3D Semantic Scene Perception for Egocentric Robots with Binocular Vision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.11872 v1 pith:MESHDPM6 submitted 2024-02-19 cs.RO

classification cs.RO
keywords pointscenecloudsobjectobjectspipelinerobotsegmentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Perceiving a three-dimensional (3D) scene with multiple objects while moving indoors is essential for vision-based mobile cobots, especially for enhancing their manipulation tasks. In this work, we present an end-to-end pipeline with instance segmentation, feature matching, and point-set registration for egocentric robots with binocular vision, and demonstrate the robot's grasping capability through the proposed pipeline. First, we design an RGB image-based segmentation approach for single-view 3D semantic scene segmentation, leveraging common object classes in 2D datasets to encapsulate 3D points into point clouds of object instances through corresponding depth maps. Next, 3D correspondences of two consecutive segmented point clouds are extracted based on matched keypoints between objects of interest in RGB images from the prior step. In addition, to be aware of spatial changes in 3D feature distribution, we also weigh each 3D point pair based on the estimated distribution using kernel density estimation (KDE), which subsequently gives robustness with less central correspondences while solving for rigid transformations between point clouds. Finally, we test our proposed pipeline on the 7-DOF dual-arm Baxter robot with a mounted Intel RealSense D435i RGB-D camera. The result shows that our robot can segment objects of interest, register multiple views while moving, and grasp the target object. The source code is available at https://github.com/mkhangg/semantic_scene_perception.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distortion-Aware Adversarial Attacks on Bounding Boxes of Object Detectors

    cs.CV 2024-12 conditional novelty 4.0 of 10

    An iterative gradient-based attack, guided by predicted bounding-box masks and controlled by a normalized cross-correlation distortion threshold, causes object detectors to misdetect objects with high reported success.

  2. Volumetric Mapping with Panoptic Refinement via Kernel Density Estimation for Mobile Robots

    cs.RO 2024-12 reject novelty 4.0 of 10

    A KDE-based depth outlier rejection step improves RGB segmentation mask precision and downstream panoptic volumetric mapping quality.

Pith tools