Pith. sign in

REVIEW 4 major objections 5 minor 55 references

Consensus-Driven Uncertainty for Robotic Grasping based on RGB Perception

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Disagreement among pose estimators predicts whether a robotic grasp will succeed, before the robot moves.

desk verdict A useful but currently over-claimed consensus-uncertainty grasp predictor: the cross-object synergy result will not stand until the test-set checkpoint selection is fixed. read the letter →

arxiv 2506.20045 v2 pith:TUDAWXH4 submitted 2025-06-24 cs.RO cs.CV

classification cs.ROcs.CV
keywords 6-DoFobjectposeestimationgraspsuccesspredictionuncertaintyquantificationensembleconsensusRGB-onlyperceptionroboticgraspingMuJoCosimulationmulti-layerperceptron
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that an RGB-only grasping agent can predict, before executing a grasp, whether that grasp will succeed, by looking at how much several off-the-shelf 6-DoF pose estimators disagree with each other. It trains a lightweight MLP on the signed six-dimensional differences between a chosen principal pose estimate and supporting estimates, using simulated grasp outcomes as labels. The resulting predictor outperforms a threshold-based uncertainty baseline, and training one predictor jointly across all objects improves accuracy further. If correct, a robot could use this cheap pre-execution signal to abstain from grasps that would likely fail.

What carries the argument

The load-bearing object is the signed 6D pose-difference vector in $\mathbb{R}^6$: for each pair of principal and supporting estimator, the three translation differences and three Euler-angle rotation differences are concatenated. These difference vectors are computed after transforming all poses into a common reference frame, which preserves each estimator's error relative to ground truth while allowing the simulator to use one handcrafted reference grasp per object. The MLP maps these differences, plus one-hot object or gripper identifiers in the wider training configurations, through five fully connected layers with ReLU activations and a sigmoid output to a predicted grasp-success probability.

What would settle it

Run the same perception-to-grasp pipeline on a physical robot arm with the same objects and grippers, replacing simulated labels with real grasp outcomes; if the MLP's predicted success probabilities do not reliably rank-order or threshold real successes and failures, the central claim that simulated consensus differences predict real grasp success is refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that signed 6D pose differences between a Principal Estimator and Supporting Estimators carry enough information about downstream task outcome that a small multi-layer perceptron can predict grasp success or failure before the grasp is attempted. This consensus-based signal outperforms the ADD-threshold baseline by 4.5% on average per object, and by 7.27% when networks are trained jointly on all objects. The paper further claims that preserving the six translation and rotation components separately, rather than collapsing them into a single ADD value, is what lets the network learn which geometric errors actually cause task failure.

Load-bearing premise

The MuJoCo grasping protocol, with one handcrafted reference grasp per object, approximated masses and friction coefficients, a 5 cm success tolerance, and no scene clutter, faithfully represents real grasping from the estimated poses.

Editorial extensions

If this is right

  • A grasping agent can use the MLP's output as a gating signal, refusing to attempt grasps whose predicted success probability falls below a threshold.
  • Because the predictor operates on pose differences rather than absolute poses, it can be combined with any number of off-the-shelf detectors, pose estimators, and grippers without retraining them.
  • Joint training across objects is beneficial, indicating that diverse objects share enough structure in how pose error translates to grasp failure to support a single predictor.
  • Translation error, especially along the viewing direction, is a stronger driver of grasp failure than rotation error, so uncertainty metrics that preserve this distinction are more useful for downstream tasks.
  • Training a single network across both grippers is less effective, implying that gripper differences are too large to be captured by a one-hot identifier alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors leave implicit that the same consensus signal could serve as a confidence measure for other pose-guided manipulation tasks such as insertion, placement, or tool use, where the cost of failure is also high.
  • A natural testable extension is to replace the simulated grasp labels with real-robot outcomes on a subset of trials; if the MLP's predictions remain well calibrated on real grasps, the simulator proxy is validated, and if not, the difference would quantify the sim-to-real gap.
  • The MLP's independence from the absolute pose suggests a possible transfer path to new objects or unseen viewpoints, provided the distribution of pose differences stays similar; this could be checked by training on one object set and testing on another.
  • The 5 cm success tolerance and handcrafted reference grasps define an implicit task difficulty; varying the tolerance or using optimized grasps would likely shift the learned decision boundary and could be used to probe how conservative the predictor should be.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a method to predict, before execution, whether an open-loop robotic grasp guided by a 6-DoF pose estimate from a single RGB image will succeed. The method designates one of three pose estimators (EPOS, GDRNPP, ZebraPose) as the Principal Estimator, forms signed 6D differences between the principal and the two supporting estimators, and trains small MLPs on these differences with labels generated by MuJoCo grasping trials. Experiments on YCB-V and LM-O with a parallel gripper and an underactuated hand compare per-object MLPs, joint-object (MLP-O), joint-gripper (MLP-G), and joint-object-gripper (MLP-OG) variants against an ADD-threshold baseline derived from Shi et al. The paper reports that MLP-O achieves the highest average prediction success.

Significance. If the quantitative claims hold, this is a practical contribution: it connects pose uncertainty to a downstream manipulation task and shows that consensus among off-the-shelf estimators can predict grasp success. The paper ships code, data, and a reproducible pipeline, and the idea of training jointly across objects is a useful finding. However, the evaluation protocol contains a test-set selection bias that prevents the central quantitative claims from being accepted as reported.

major comments (4)
  1. [Section V.A, Table III] The sentence "We report results for the model checkpoint with best test-set accuracy" means the same 20% test split is used both for model selection and for reporting. The baseline in Section V.B learns its threshold on the training set only. This asymmetry makes the comparisons in Table III unfair: the MLP numbers are optimistically selected, while the baseline is not. The claimed margins, e.g., MLP-O vs. Baseline 0.898 vs. 0.792 for the parallel gripper, could be an artifact of this selection. The evaluation should be redone using a validation set for checkpoint selection (or the final-epoch model), with repeated splits and confidence intervals.
  2. [Section V.A] The authors drop object-image pairs where any of the three estimators fails to make a detection or prediction. This filters out exactly the high-uncertainty cases in which a consensus-based uncertainty estimate is most needed. The reported accuracies are therefore conditional on all estimators succeeding, which may not hold in practice. Please report the number of dropped pairs and discuss or address how the method would handle missing detections at inference time.
  3. [Sections III.B and IV.C] The grasping trials are entirely simulated with a single handcrafted reference grasp per object, approximated mass and friction coefficients, no scene clutter, and a 5 cm success tolerance. No real-robot validation or validation against a real grasping dataset is provided. The abstract's claim that the method predicts whether "a grasp guided by an image-based pose estimate will succeed" is therefore only established in this simplified simulation proxy. This is a load-bearing limitation for the paper's practical significance; the authors should either temper the claim to the simulated setting or add a real-world validation.
  4. [Table III and Section V.C] The central claim that joint training across objects improves prediction (MLP-O vs. MLP) is a key contribution, but it rests on the same test-selected checkpoints. Because all MLP variants are selected using the test split, the ordinal relationships among MLP, MLP-O, MLP-G, and MLP-OG are also potentially biased. After fixing the selection protocol, the synergy claim needs to be re-established with a validation-based comparison. No confidence intervals or significance tests are provided, so it is unclear whether the reported differences are stable across data splits.
minor comments (5)
  1. [Introduction] The phrase "the so-calledsim2real gap" is missing a space; it should read "the so-called sim2real gap."
  2. [Section V.B] The statement that "translation error is a better indicator of grasp failure than rotation error" is presented without supporting quantitative analysis; please provide the evidence or soften the claim.
  3. [Section III.B, Eqs. (1)-(2)] The error functions e_R and e_t are used without being defined; please define these metrics explicitly (e.g., angular error and Euclidean translation error).
  4. [Section V.A] The sentence "Training 90 MLPs (three PEs, two grippers, 15 objects)" appears to count only the per-object, per-gripper MLPs, but the paper also trains MLP-O, MLP-G, and MLP-OG variants; please clarify the total number of trained networks.
  5. [Section III.C] The input representation uses raw Euler angle differences, which are sensitive to the chosen Euler convention and have singularities. Consider using a more canonical pose difference, such as the logarithmic map of the relative SE(3) transformation, to avoid these issues.

Circularity Check

1 steps flagged · score 4.0 of 10

Test-set checkpoint selection partially fits the reported MLP predictions; the supervised training pipeline itself is not derivationally circular.

  1. fitted input called prediction [Section V-A (Training) and Section V-B (Baseline), Table III]
    "We report results for the model checkpoint with best test-set accuracy. ... To ensure fairness, we use the same training set used by the MLPs."

    The MLPs' reported accuracies are produced by selecting, among 3,000 checkpoints, the one with best accuracy on the same 20% test split later used for the reported numbers, while the baseline threshold is fit only on the training split ('the same training set used by the MLPs'). The checkpoint choice is therefore a fit to the test labels, so the Table III comparison of MLP vs baseline and MLP-O vs per-object MLP is not an unbiased held-out comparison; reported advantages can be an artifact of test-set selection. This is evaluation circularity rather than derivation circularity.

full rationale

The derivation chain is a supervised MLP trained on simulator grasp labels from pose differences and evaluated on a disjoint 20% split, so the central prediction is not circular by construction. The pose-difference representation deliberately omits the PE pose itself, and Equations (1)-(2) are a correct equivariance identity, not a tautology. The underactuated hand simulation cites [52]/[53] with author overlap (Long Wang on [52]), but the hand is an imported tool and the central claim does not depend on those citations' theoretical content. The main issue is evaluation, not derivation: Section V.A states checkpoints are selected by best test-set accuracy, while the baseline threshold is fit on the training set only, so the MLP entries in Table III are partially fitted to the same test labels used for reporting. This is a statistical/selection bias that can inflate the reported margins and the MLP-O synergy claim; it is not a formal equivalence of output to input, so it warrants a moderate partial-circularity score rather than a high one.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on simulator-generated labels and on the feature representation (signed 6D PE-SE differences). The handcrafted physical parameters and reference grasps are not derived from data, and the evaluation split does not separate objects. No new physical entities are introduced.

free parameters (3)
  • Object mass and friction coefficients = approximated per object
    BOP metadata lacks these; authors estimate values and assume uniform density; affects simulator grasp outcomes.
  • Grasp success tolerance = 5 cm
    Centroid offset threshold after 15 s lift; chosen threshold determines binary labels for training.
  • Reference grasp poses = one per object-gripper, handcrafted
    Single reference pose and grasp per object; authors note better or worse grasps would shift distributions.
assumptions (4)
  • domain assumption Objects are known, rigid, instance-level, with a detector available; perception is RGB-only.
    Stated assumptions in Section I; they define the scope of the method and exclude depth sensors, category-level pose, and unknown objects.
  • domain assumption MuJoCo simulation with open-loop control, uniform density, and approximated friction faithfully represents real grasping outcomes.
    Underlies all training labels; no real-hardware validation is provided in the paper.
  • domain assumption A single handcrafted reference grasp per object is representative for evaluating pose-error-induced failures.
    Section III-B; the authors acknowledge that grasp optimality is not considered and that different grasps would shift success distributions.
  • domain assumption The 80/20 split by trial, with objects shared across splits, provides a valid evaluation of generalization.
    No object-disjoint split is reported; the model may exploit object-specific patterns that appear in both train and test sets.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Consensus-Driven Uncertainty for Robotic Grasping based on RGB Perception." pith.science (2026). https://pith.science/paper/TUDAWXH4

@misc{pith2026250620045,
  author       = {Pith},
  title        = {Pith review of: Consensus-Driven Uncertainty for Robotic Grasping based on RGB Perception},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TUDAWXH4}},
  note         = {Machine review of arXiv:2506.20045}
}
read the original abstract

Deep object pose estimators are notoriously overconfident. A grasping agent that both estimates the 6-DoF pose of a target object and predicts the uncertainty of its own estimate could avoid task failure by choosing not to act under high uncertainty. Even though object pose estimation improves and uncertainty quantification research continues to make strides, few studies have connected them to the downstream task of robotic grasping. We propose a method for training lightweight, deep networks to predict whether a grasp guided by an image-based pose estimate will succeed before that grasp is attempted. We generate training data for our networks via object pose estimation on real images and simulated grasping. We also find that, despite high object variability in grasping trials, networks benefit from training on all objects jointly, suggesting that a diverse variety of objects can nevertheless contribute to the same goal.

Figures

Figures reproduced from arXiv: 2506.20045 by the authors.

Figure 1
Figure 1. RGB-based pose estimates as green overlays and their corresponding [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Example reference grasps for selected objects in the LM-O dataset (a-h) and YCB-V dataset (i-o). All grasping trials are attempted with both the [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. From RGB images, to object detections, to pose estimates per object, to grasping each object in a physics simulator. We abstract away scene clutter [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: A diagram of our MLP architecture. All MLPs receive differences [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

55 extracted references · 55 canonical work pages

  1. [1]

    Deep Learning on Monocular Object Pose Detection and Tracking: A Comprehensive Overview,

    Z. Fan, Y . Zhu, Y . He, Q. Sun, H. Liu, and J. He, “Deep Learning on Monocular Object Pose Detection and Tracking: A Comprehensive Overview,”ACM Computing Surveys, vol. 55, no. 4, pp. 1–40, 2022

  2. [2]

    A review on object pose recovery: From 3D bounding box detectors to full 6D pose estimators,

    C. Sahin, G. Garcia-Hernando, J. Sock, and T.-K. Kim, “A review on object pose recovery: From 3D bounding box detectors to full 6D pose estimators,”Image and Vision Computing, vol. 96, p. 103898, 2020. 8

  3. [3]

    BOP Challenge 2022 on Detection, Segmen- tation and Pose Estimation of Specific Rigid Objects,

    M. Sundermeyer, T. Hoda ˇn, Y . Labb´e, G. Wang, E. Brachmann, B. Drost, C. Rother, and J. Matas, “BOP Challenge 2022 on Detection, Segmen- tation and Pose Estimation of Specific Rigid Objects,” inCVPR, 2023

  4. [4]

    Challenges for Monocular 6D Object Pose Estimation in Robotics,

    S. Thalhammer, D. Bauer, P. H ¨onig, J.-B. Weibel, J. Garc ´ıa-Rodr´ıguez, M. Vinczeet al., “Challenges for Monocular 6D Object Pose Estimation in Robotics,”IEEE Transactions on Robotics, 2024

  5. [5]

    BOP: Benchmark for 6D Object Pose Estimation,

    T. Hoda ˇn, F. Michel, E. Brachmann, W. Kehl, A. GlentBuch, D. Kraft, B. Drost, J. Vidal, S. Ihrke, X. Zabuliset al., “BOP: Benchmark for 6D Object Pose Estimation,” inECCV, 2018, pp. 19–34

  6. [6]

    DexYCB: A Benchmark for Capturing Hand Grasping of Objects,

    Y .-W. Chao, W. Yang, Y . Xiang, P. Molchanov, A. Handa, J. Tremblay, Y . S. Narang, K. Van Wyk, U. Iqbal, S. Birchfieldet al., “DexYCB: A Benchmark for Capturing Hand Grasping of Objects,” inCVPR, 2021

  7. [7]

    Fast Uncertainty Quantification for Deep Object Pose Estimation,

    G. Shi, Y . Zhu, J. Tremblay, S. Birchfield, F. Ramos, A. Anandkumar, and Y . Zhu, “Fast Uncertainty Quantification for Deep Object Pose Estimation,” inICRA, 2021, pp. 5200–5207

  8. [8]

    FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects,

    B. Wen, W. Yang, J. Kautz, and S. Birchfield, “FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects,” inCVPR, 2024

Show all 55 references
  1. [9]

    SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Sstimation,

    J. Lin, L. Liu, D. Lu, and K. Jia, “SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Sstimation,” inCVPR, 2024, pp. 27 906–27 916

  2. [10]

    GigaPose: Fast and Robust Novel Object Pose Estimation via One Correspondence,

    V . N. Nguyen, T. Groueix, M. Salzmann, and V . Lepetit, “GigaPose: Fast and Robust Novel Object Pose Estimation via One Correspondence,” in CVPR, 2024, pp. 9903–9913

  3. [11]

    GenPose: Generative Category-level Object Pose Estimation via Diffusion Models,

    J. Zhang, M. Wu, and H. Dong, “GenPose: Generative Category-level Object Pose Estimation via Diffusion Models,”NeurIPS, vol. 36, 2023

  4. [12]

    Self-supervised 6D Object Pose Estimation for Robot Manipulation,

    X. Deng, Y . Xiang, A. Mousavian, C. Eppner, T. Bretl, and D. Fox, “Self-supervised 6D Object Pose Estimation for Robot Manipulation,” inICRA, 2020, pp. 3665–3671

  5. [13]

    GraspNet-1Billion: A Large- Scale Benchmark for General Object Grasping,

    H.-S. Fang, C. Wang, M. Gou, and C. Lu, “GraspNet-1Billion: A Large- Scale Benchmark for General Object Grasping,” inCVPR, 2020

  6. [14]

    Contact- GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes,

    M. Sundermeyer, A. Mousavian, R. Triebel, and D. Fox, “Contact- GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes,” in ICRA, 2021, pp. 13 438–13 444

  7. [15]

    Neural Grasp Distance Fields for Robot Manipulation,

    T. Weng, D. Held, F. Meier, and M. Mukadam, “Neural Grasp Distance Fields for Robot Manipulation,” inICRA, 2023, pp. 1814–1821

  8. [16]

    Combined Optimization of Gripper Finger Design and Pose Estimation Processes for Advanced Industrial Assembly,

    F. Hagelskjær, A. Kramberger, A. Wolniakowski, T. R. Savarimuthu, and N. Kr ¨uger, “Combined Optimization of Gripper Finger Design and Pose Estimation Processes for Advanced Industrial Assembly,” inIROS, 2019, pp. 2022–2029

  9. [17]

    SE(3)-DiffusionFields: Learning smooth cost functions for joint grasp and motion optimization through diffusion,

    J. Urain, N. Funk, J. Peters, and G. Chalvatzaki, “SE(3)-DiffusionFields: Learning smooth cost functions for joint grasp and motion optimization through diffusion,” inICRA, 2023, pp. 5923–5930

  10. [18]

    PoseCNN: A Con- volutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes,

    Y . Xiang, T. Schmidt, V . Narayanan, and D. Fox, “PoseCNN: A Con- volutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes,” inRSS, 2018

  11. [19]

    Implicit 3D Orientation Learning for 6D Object Detection from RGB Images,

    M. Sundermeyer, Z.-C. Marton, M. Durner, M. Brucker, and R. Triebel, “Implicit 3D Orientation Learning for 6D Object Detection from RGB Images,” inECCV, 2018, pp. 699–715

  12. [20]

    BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using Depth,

    M. Rad and V . Lepetit, “BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using Depth,” inICCV, 2017, pp. 3828–3836

  13. [21]

    Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects,

    J. Tremblay, T. To, B. Sundaralingam, Y . Xiang, D. Fox, and S. Birch- field, “Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects,” inConference on Robot Learning, 2018

  14. [22]

    Real-Time Seamless Single Shot 6D Object Pose Prediction,

    B. Tekin, S. N. Sinha, and P. Fua, “Real-Time Seamless Single Shot 6D Object Pose Prediction,” inCVPR, 2018, pp. 292–301

  15. [23]

    PVNet: Pixel-wise V oting Network for 6DoF Pose Estimation,

    S. Peng, Y . Liu, Q. Huang, X. Zhou, and H. Bao, “PVNet: Pixel-wise V oting Network for 6DoF Pose Estimation,” inCVPR, 2019

  16. [24]

    Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose Estimation,

    K. Park, T. Patten, and M. Vincze, “Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose Estimation,” inICCV, 2019

  17. [25]

    EPOS: Estimating 6D Pose of Objects with Symmetries,

    T. Hoda ˇn, D. Barath, and J. Matas, “EPOS: Estimating 6D Pose of Objects with Symmetries,” inCVPR, 2020, pp. 11 703–11 712

  18. [26]

    GDR-Net: Geometry- Guided Direct Regression Network for Monocular 6D Object Pose Estimation,

    G. Wang, F. Manhardt, F. Tombari, and X. Ji, “GDR-Net: Geometry- Guided Direct Regression Network for Monocular 6D Object Pose Estimation,” inCVPR, 2021, pp. 16 611–16 621

  19. [27]

    ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose Estimation,

    Y . Su, M. Saleh, T. Fetzer, J. Rambach, N. Navab, B. Busam, D. Stricker, and F. Tombari, “ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose Estimation,” inCVPR, 2022, pp. 6738–6748

  20. [28]

    Neural Correspondence Field for Object Pose Estimation,

    L. Huang, T. Hodan, L. Ma, L. Zhang, L. Tran, C. Twigg, P.-C. Wu, J. Yuan, C. Keskin, and R. Wang, “Neural Correspondence Field for Object Pose Estimation,” inECCV, 2022, pp. 585–603

  21. [29]

    CheckerPose: Progressive Dense Keypoint Lo- calization for Object Pose Estimation with Graph Neural Network,

    R. Lian and H. Ling, “CheckerPose: Progressive Dense Keypoint Lo- calization for Object Pose Estimation with Graph Neural Network,” in ICCV, 2023, pp. 14 022–14 033

  22. [30]

    On the Continuity of Rotation Representations in Neural Networks,

    Y . Zhou, C. Barnes, J. Lu, J. Yang, and H. Li, “On the Continuity of Rotation Representations in Neural Networks,” inCVPR, 2019

  23. [31]

    CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose Estima- tion,

    Z. Li, G. Wang, and X. Ji, “CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose Estima- tion,” inICCV, 2019, pp. 7678–7687

  24. [32]

    SurfEmb: Dense and Continuous Correspondence Distributions for Object Pose Estimation with Learnt Surface Embeddings,

    R. L. Haugaard and A. G. Buch, “SurfEmb: Dense and Continuous Correspondence Distributions for Object Pose Estimation with Learnt Surface Embeddings,” inCVPR, 2022, pp. 6749–6758

  25. [33]

    Uncertainty-Driven 6D Pose Estimation of Objects and Scenes from a Single RGB Image,

    E. Brachmann, F. Michel, A. Krull, M. Y . Yang, S. Gumholdet al., “Uncertainty-Driven 6D Pose Estimation of Objects and Scenes from a Single RGB Image,” inCVPR, 2016, pp. 3364–3372

  26. [34]

    Implicit-PDF: Non-Parametric Representation of Probability Distribu- tions on the Rotation Manifold,

    K. A. Murphy, C. Esteves, V . Jampani, S. Ramalingam, and A. Makadia, “Implicit-PDF: Non-Parametric Representation of Probability Distribu- tions on the Rotation Manifold,” inICML, 2021, pp. 7882–7893

  27. [35]

    HyperPosePDF- Hypernetworks Predicting the Probability Distribution on SO(3),

    T. H ¨ofer, B. Kiefer, M. Messmer, and A. Zell, “HyperPosePDF- Hypernetworks Predicting the Probability Distribution on SO(3),” in WACV, 2023, pp. 2369–2379

  28. [36]

    SpyroPose: SE(3) Pyramids for Object Pose Distribution Estimation,

    R. L. Haugaard, F. Hagelskjær, and T. M. Iversen, “SpyroPose: SE(3) Pyramids for Object Pose Distribution Estimation,” inICCV, 2023

  29. [37]

    Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles,

    B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles,”NeurIPS, vol. 30, 2017

  30. [38]

    Uncertainty Quantification with Deep Ensembles for 6D Object Pose Estimation,

    K. Wursthorn, M. Hillemann, and M. Ulrich, “Uncertainty Quantification with Deep Ensembles for 6D Object Pose Estimation,”ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, 2024

  31. [39]

    Robotic Task Success Evaluation Under Multi-modal Non-Parametric Object Pose Uncertainty,

    L. Naik, T. M. Iversen, A. Kramberger, and N. Kr ¨uger, “Robotic Task Success Evaluation Under Multi-modal Non-Parametric Object Pose Uncertainty,”arXiv preprint arXiv:2403.10874, 2024

  32. [40]

    Grasping Under Uncertainties: Sequential Neural Ratio Estimation for 6-DoF Robotic Grasping,

    N. Marlier, O. Br ¨uls, and G. Louppe, “Grasping Under Uncertainties: Sequential Neural Ratio Estimation for 6-DoF Robotic Grasping,”IEEE Robotics and Automation Letters, 2024

  33. [41]

    The Franka Emika Robot: A Reference Platform for Robotics Research and Educa- tion,

    S. Haddadin, S. Parusel, L. Johannsmeier, S. Golz, S. Gabl, F. Walch, M. Sabaghian, C. J ¨ahne, L. Hausperger, and S. Haddadin, “The Franka Emika Robot: A Reference Platform for Robotics Research and Educa- tion,”IEEE Robotics & Automation Magazine, vol. 29, no. 2, 2022

  34. [42]

    Adaptive synergies for the design and control of the Pisa/IIT SoftHand,

    M. Catalano, G. Grioli, E. Farnioli, A. Serio, C. Piazza, and A. Bicchi, “Adaptive synergies for the design and control of the Pisa/IIT SoftHand,” Int. Journal of Robotics Research, vol. 33, no. 5, pp. 768–782, apr 2014

  35. [43]

    The Velo gripper: A versatile single- actuator design for enveloping, parallel and fingertip grasps,

    M. Ciocarlie, F. M. Hicks, R. Holmberg, J. Hawke, M. Schlicht, J. Gee, S. Stanford, and R. Bahadur, “The Velo gripper: A versatile single- actuator design for enveloping, parallel and fingertip grasps,”Int. Journal of Robotics Research, vol. 33, no. 5, pp. 753–767, 2014

  36. [44]

    The Highly Adaptive SDM Hand: Design and Performance Evaluation,

    A. M. Dollar and R. D. Howe, “The Highly Adaptive SDM Hand: Design and Performance Evaluation,”Int. Journal of Robotics Research, vol. 29, no. 5, pp. 585–597, 2010

  37. [45]

    An Anthropomorphic Under- actuated Robotic Hand with 15 Dofs and a Single Actuator,

    C. Gosselin, F. Pelletier, and T. Laliberte, “An Anthropomorphic Under- actuated Robotic Hand with 15 Dofs and a Single Actuator,” inICRA, 2008, pp. 749–754

  38. [46]

    A compliant, underactuated hand for robust manipulation,

    L. U. Odhner, L. P. Jentoft, M. R. Claffee, N. Corson, Y . Tenzer, R. R. Ma, M. Buehler, R. Kohout, R. D. Howe, and A. M. Dollar, “A compliant, underactuated hand for robust manipulation,”Int. Journal of Robotics Research, vol. 33, no. 5, pp. 736–752, 2014

  39. [47]

    The Ocean One hands: An adaptive design for robust marine manipulation,

    H. Stuart, S. Wang, O. Khatib, and M. R. Cutkosky, “The Ocean One hands: An adaptive design for robust marine manipulation,”Int. Journal of Robotics Research, vol. 36, no. 2, pp. 150–166, 2017

  40. [48]

    A highly-underactuated robotic hand with force and joint angle sensors,

    L. Wang, J. DelPreto, S. Bhattacharyya, J. Weisz, and P. K. Allen, “A highly-underactuated robotic hand with force and joint angle sensors,” inIROS, 2011, pp. 1380–1385

  41. [49]

    Design of the Utah/M.I.T. Dextrous Hand,

    S. Jacobsen, E. Iversen, D. Knutti, R. Johnson, and K. Biggers, “Design of the Utah/M.I.T. Dextrous Hand,” inICRA, vol. 3, 1986

  42. [50]

    Modeling and control of the stanford/JPL hand,

    C. Loucks, V . Johnson, P. Boissiere, G. Starr, and J. Steele, “Modeling and control of the stanford/JPL hand,” inICRA, vol. 4, 1987

  43. [51]

    Mechanisms of the Anatomically Correct Testbed Hand,

    A. D. Deshpande, Z. Xu, M. J. V . Weghe, B. H. Brown, J. Ko, L. Y . Chang, D. D. Wilkinson, S. M. Bidic, and Y . Matsuoka, “Mechanisms of the Anatomically Correct Testbed Hand,”IEEE/ASME Transactions on Mechatronics, vol. 18, no. 1, pp. 238–250, 2013

  44. [52]

    Underactuation Design for Tendon-Driven Hands via Optimization of Mechanically Realizable Manifolds in Posture and Torque Spaces,

    T. Chen, L. Wang, M. Haas-Heger, and M. Ciocarlie, “Underactuation Design for Tendon-Driven Hands via Optimization of Mechanically Realizable Manifolds in Posture and Torque Spaces,”IEEE Transactions on Robotics, vol. 36, no. 3, pp. 708–723, jun 2020

  45. [53]

    Hardware as policy: Mechanical and computational co-optimization using deep reinforcement learning,

    T. Chen, Z. He, and M. Ciocarlie, “Hardware as policy: Mechanical and computational co-optimization using deep reinforcement learning,” inConference on Robot Learning, 2021

  46. [54]

    MuJoCo: A physics engine for model-based control,

    E. Todorov, T. Erez, and Y . Tassa, “MuJoCo: A physics engine for model-based control,” inIROS, 2012, pp. 5026–5033

  47. [55]

    BOP Challenge 2020 on 6D Object Localization,

    T. Hoda ˇn, M. Sundermeyer, B. Drost, Y . Labb ´e, E. Brachmann, F. Michel, C. Rother, and J. Matas, “BOP Challenge 2020 on 6D Object Localization,” inECCV 2020 Workshops. Springer, 2020, pp. 577–594

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.