REVIEW 4 major objections 5 minor 73 references
The next best view for grasping is the one that most reduces the entropy of the grasp-possibility distribution, not the one that covers the most scene.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
ActiveGrasp selects the next camera view that maximizes predicted reduction in grasp-success entropy using a calibrated SE(3) energy-based model, and reports higher grasp success than prior active-grasping methods in simulation and real-robot tests.
T0 review reviewed 2026-08-03 challenge →
load-bearing objection A serious, well-posed active grasping paper with a plausible central claim and real empirical gains, but the theoretical core currently rests on an unverified Hessian approximation, a missing appendix, and a likely sign error in the entropy definition. the 4 major comments →
ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Model
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that information gain for active grasping should be computed directly from the calibrated distribution of grasps, and that this is tractable. The authors define the entropy of the grasp distribution as η(w)=E_{p(g|w)}[h(s(g,w))], where h is the Bernoulli entropy of a single grasp's success probability s. Using a Gaussian Approximation of the Posterior of a 3D Gaussian Splatting scene, the expected entropy reduction from a candidate view becomes half the trace of the curvature of η with respect to scene parameters, times the reduction in scene-parameter covariance. Because the energy model alone gives only relative scores, the paper calibrates its energy to success rate b
What carries the argument
The central object is a calibrated energy-based model of grasp poses on the SE(3) manifold (the six-degree-of-freedom space of positions and orientations). The model outputs separate energy values for success and failure for a scene represented by 3D Gaussian Splatting, and a learnable temperature; it is trained with an average-precision loss and on both successful and failed grasps so that the energy level aligns with success probability. The second piece is the entropy-reduction estimate: the expected information gain of a candidate view is expressed as half the trace of the curvature of the grasp-entropy function with respect to scene parameters, times the reduction in the posterior covar
Load-bearing premise
The view-quality ranking stands or falls on two approximations: the scene posterior being Gaussian with an inverse-Hessian approximated as diagonal, and the curvature of grasp entropy being the gradient outer product plus a diagonal term; if either misestimates the true entropy reduction, the selected views may not actually be the information-optimal ones.
What would settle it
In a simulated scene with a small object and a dense candidate-view set, exhaustively evaluate every pair of additional views by physically executing a large set of grasps after each pair, and measure the true reduction in grasp-success entropy. If the pair ranked highest by ActiveGrasp's predicted entropy reduction does not yield the largest measured reduction within statistical error over repeated trials, the central claim is falsified.
If this is right
- Using two fixed and two actively chosen views, ActiveGrasp reaches 79% grasp success in simulated clutter, against 74.25% for the best baseline planner with the same calibrated grasp model.
- Calibration matters: pairing the same entropy-driven view selection with an uncalibrated energy model yields 73%, and the calibrated model cuts expected calibration error on executed grasps from 0.35 to 0.02.
- Ten random views (eight more than active methods) reach only 76.5%, implying that view placement, not view count, is the dominant factor in grasp success.
- Because information gain is defined directly from the grasp distribution, the approach removes the need for separate affordance predictors, coverage heuristics, or other surrogate rewards.
- In three real-world cluttered scenes, the method succeeds on 9/10, 8/10, and 9/10 grasp attempts, outperforming coverage-, affordance-, and graspness-based planners that share the same calibrated grasp model.
Where Pith is reading between the lines
- A natural extension: the same entropy-reduction objective could choose views for other manipulation skills, such as tool use or assembly, whenever an energy-based model supplies a distribution over the relevant action manifold.
- The rank-one-plus-diagonal curvature approximation is the piece most likely to fail under domain shift or with richer scene representations; a stochastic or block-diagonal Hessian estimator is a direct testable improvement.
- A consequence the authors do not pursue: if view placement matters more than view count, a perception system with a tight view budget could still succeed by spending all of it on entropy-optimal viewpoints, relevant for low-power or time-critical robots.
- Because calibration was performed on simulated data, an online recalibration loop using real executed grasps would test whether the calibration guarantee survives the sim-to-real gap the authors acknowledge.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes ActiveGrasp, a next-best-view planning method for robotic grasping. It defines the information gain of a candidate view as the expected reduction in a grasp-specific entropy, where the grasp distribution is modeled by a calibrated energy-based model on SE(3). The scene is represented by 3D Gaussian Splatting, the posterior covariance is obtained via the Gaussian Approximation of the Posterior, and the Hessian of the grasp entropy is approximated by a rank-one-plus-diagonal form. The energy-based model is calibrated using failure grasps, an average-precision loss, and a learnable temperature. Experiments in a PyBullet benchmark and on a real robot report that ActiveGrasp achieves 79% grasp success with two actively selected views, versus 74.25% for the best baseline with the same calibrated grasp model and view budget. An ablation on ACRONYM and calibration metrics is also presented.
Significance. If the derivations are correct, the paper would make a useful contribution by replacing proxy objectives such as visibility or affordance with a task-specific information-gain estimate computed directly from a calibrated grasp distribution on SE(3). The paper has clear strengths: it addresses a well-motivated problem, provides a reproducible simulated benchmark (code is promised), includes a careful ablation of the calibration components, and the view-selection objective is not circularly fitted to the final success metric. However, the central information-gain derivation is deferred to an appendix that is not present in the submitted text, the key Hessian approximation is ad hoc and unvalidated, and the sign convention of the entropy definition appears internally inconsistent. These issues currently prevent full verification of the paper's central claim.
major comments (4)
- [Sec. 3.2, Eq. (10) and Eq. (15)] The entropy definition in Eq. (10) is h(s)=s ln s+(1-s) ln(1-s), which is the negative of the standard Shannon entropy of a Bernoulli variable. The text states that H[g|w] decreases when more successful grasps are discovered, but with Eq. (10), as s→1, h(s) increases from -ln2 toward 0. Moreover, if the standard entropy were intended (i.e., with a minus sign), then ∇²η(w) would be negative semidefinite, making Eq. (15) negative when evaluated as trace(∇²η (Σ0-Σ1)) with Σ0-Σ1 positive semidefinite. Under the definition as written, the quantity is not entropy and the claimed monotonic behavior in Fig. 3 is reversed. The two equations are mutually inconsistent, and this is load-bearing for the central claim that views are selected by reducing grasp entropy. Please correct the sign convention and re-derive Eqs. (14)-(15) consistently.
- [Sec. 3.2, Eqs. (14), (15), (17)] The proofs of Eqs. (14), (15), and (17) are stated to be left to the appendix, but no appendix is included in the submitted manuscript. These equations are the core of the information-gain formulation, and the sign of Eq. (15), the Gaussian approximation in Eq. (14), and the derivative identity in Eq. (17) cannot be checked without the derivations. This is not a stylistic issue; it blocks verification of the central claim. The final version must include the full derivations or the paper should be reviewed with the appendix attached.
- [Sec. 3.2, Eq. (16)] The approximation ∇²w η(w) ≈ ∇w η(w)∇w η(w)ᵀ + λI is introduced as 'working well in practice' but is not derived, compared with alternative approximations, or validated. The authors themselves list this as a limitation in Sec. 5. Since the entire view-selection signal is the trace of this Hessian against a covariance difference, an inaccurate Hessian can rank candidate views incorrectly. Please provide evidence that the predicted information gain correlates with the actual reduction in grasp entropy (or with grasp success) after acquiring a view, or replace Eq. (16) with a more principled approximation, e.g., a Gauss-Newton or Monte-Carlo estimate of ∇²η.
- [Sec. 4.3, Table 1] The reported headline improvement (79% vs. 74.25%) is not accompanied by any statistical significance measure. With roughly 400 trials, the standard error of a binomial success rate is about 2 percentage points, so the difference is about 1.6 standard errors; the comparison with Random† (76.5%) is even weaker. The claim that ActiveGrasp outperforms the best baseline with the same view budget would be considerably strengthened by confidence intervals, a significance test, or per-scene/per-target variance reporting. As written, the 4.75-point gap may not be reliable.
minor comments (5)
- [Sec. 3.3, Eq. (19)] The '+2' in the denominator of the probability normalization is unexplained. It makes p_S + p_F < 1, leaving residual probability mass. Since calibration is central to the method, please justify this choice or provide an ablation showing its effect.
- [Sec. 4.1 / References [9] and [71]] The implementation details cite 'ACE [9]' for next-best-view selection, but the experiments and Table 1 use ACE with reference [71]. Please correct the citation to avoid ambiguity.
- [Fig. 3] The figure caption states that H is 'the grasp entropy defined in our paper,' but the plotted entropy curve is not labeled with the sign convention. Given the sign issue in Eq. (10), please make the plotted quantity explicit and ensure it matches the corrected definition.
- [Naming] The model is variously called Se3diff, Se3diff Scene, Se3diff Calib, and Se3diff-Calib. Please unify the terminology across Table 1, Table 3, and the text.
- [Availability] The abstract and conclusion state that source code and the simulator benchmark will be released, but no link or release plan is given. For a reproducibility-oriented contribution, please provide a clear availability statement.
Circularity Check
No significant circularity: the active-view objective is computed from a calibrated grasp-energy model and evaluated against external baselines; no equation reduces to its own input.
full rationale
The central derivation is not circular. The information gain I[g,w;y_acq|x_acq,D] (Eq. 15) is defined as the expected reduction of grasp entropy eta(w)=E[h(s(g,w))], where h is the Bernoulli entropy of the calibrated success probability. The posterior covariance comes from the Gaussian Approximation of Posterior (Eq. 1) and a diagonal Gauss-Newton Hessian (Eq. 3), imported from the authors' prior FisherRF [30]. This is a self-citation, but it is not load-bearing in the circularity sense: it supplies an approximation for scene uncertainty, and FisherRF is also included in Table 1 as an external baseline. The novel term is the grasp-entropy objective; it is not fitted to the final success metric. Success is measured by executed grasps in simulation and on a real Kinova arm with YCB objects, and ActiveGrasp is compared against Breyer, ACE, ActiveNGF, and FisherRF using the same calibrated grasp model, so the advantage is not forced by construction. The rank-one-plus-diagonal approximation of Eq. 16 is acknowledged in Section 5 as a limitation; an unvalidated or inaccurate approximation is a correctness/robustness risk, not a definitional circularity. The sign convention in Eq. 10 (h(s)=s ln s+(1-s)ln(1-s)) is the negative of Shannon entropy and, together with the missing appendix proofs for Eqs. 14/15/17, is an internal-consistency and verifiability concern. None of these concerns exhibits a quantity reducing to its own input. Accordingly, no circular step can be quoted, and the score reflects only a minor self-citation and verification gaps, not circularity.
Axiom & Free-Parameter Ledger
free parameters (5)
- lambda in Eq 16 Hessian approximation
- '+2' constant in Eq 19 probability normalization =
2
- learnable temperature T =
learned
- Noise schedule sigma=0.5 and loss weights lambda1=1, lambda2=0.1 =
sigma=0.5; lambda1=1; lambda2=0.1
- Candidate view sampling parameters =
radius 0.3-0.5 m, 128 proposals, 40 candidates
axioms (5)
- domain assumption Posterior over 3D scene parameters is Gaussian: p(w|D) approx N(w*, H''[w|D]^{-1}) (Eq 1).
- domain assumption Posterior Hessian can be approximated as a diagonal Gauss-Newton sum over views (Eq 3).
- ad hoc to paper grad^2_w eta(w) approx grad_w eta(w) grad_w eta(w)^T + lambda I (Eq 16).
- ad hoc to paper The three stated criteria (grasp distribution, SE(3) formulation, calibration) are necessary for unbiased information gain.
- standard math Denoised score matching and annealed Langevin dynamics on SE(3) from SE(3)-DiffusionFields [64] provide a valid generative model of grasp poses.
Cite this review
Pith. "Pith review of ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Model." pith.science (2026). https://pith.science/paper/QGKEOTYL
@misc{pith2026251112795,
author = {Pith},
title = {Pith review of: ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/QGKEOTYL}},
note = {Machine review of arXiv:2511.12795}
}
read the original abstract
Grasping in a densely cluttered environment is a challenging task for robots. Previous methods tried to solve this problem by actively gathering multiple views before grasp pose generation. However, they either overlooked the importance of the grasp distribution for information gain estimation or relied on the projection of the grasp distribution, which ignores the structure of grasp poses on the SE(3) manifold. To tackle these challenges, we propose a calibrated energy-based model for grasp pose generation and an active view selection method that estimates information gain from grasp distribution. Our energy-based model captures the multi-modality nature of grasp distribution on the SE(3) manifold. The energy level is calibrated to the success rate of grasps so that the predicted distribution aligns with the real distribution. The next best view is selected by estimating the information gain for grasp from the calibrated distribution conditioned on the reconstructed environment, which could efficiently drive the robot to explore affordable parts of the target object. Experiments on simulated environments and real robot setups demonstrate that our model could successfully grasp objects in a cluttered environment with limited view budgets compared to previous state-of-the-art models. Our simulated environment can serve as a reproducible platform for future research on active grasping. The source code of our paper will be made public when the paper is released to the public.
Figures
Reference graph
Works this paper leans on
-
[1]
Ac- tive vision for dexterous grasping of novel objects
Ermano Arruda, Jeremy Wyatt, and Marek Kopicki. Ac- tive vision for dexterous grasping of novel objects. In2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2881–2888, 2016. 1, 2
2016
-
[2]
Active perception.Proceedings of the IEEE, 76(8):966–1005, 1988
Ruzena Bajcsy. Active perception.Proceedings of the IEEE, 76(8):966–1005, 1988. 2
1988
-
[3]
Re- visiting active perception.Autonomous Robots, 42:177–196,
Ruzena Bajcsy, Yiannis Aloimonos, and John K Tsotsos. Re- visiting active perception.Autonomous Robots, 42:177–196,
-
[4]
V olumetric grasping network: Real-time 6 dof grasp detection in clutter
Michel Breyer, Jen Jen Chung, Lionel Ott, Roland Siegwart, and Juan Nieto. V olumetric grasping network: Real-time 6 dof grasp detection in clutter. InConference on Robot Learn- ing, pages 1602–1611. PMLR, 2021. 2, 8
2021
-
[5]
Closed-loop next-best-view planning for target- driven grasping
Michel Breyer, Lionel Ott, Roland Siegwart, and Jen Jen Chung. Closed-loop next-best-view planning for target- driven grasping. InIEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2022, Kyoto, Japan, October 23-27, 2022, pages 1411–1416. IEEE, 2022. 1, 2, 6, 7, 8
2022
-
[6]
Real-time collision-free grasp pose detection with geometry-aware refinement using high-resolution volume
Junhao Cai, Jun Cen, Haokun Wang, and Michael Yu Wang. Real-time collision-free grasp pose detection with geometry-aware refinement using high-resolution volume. IEEE Robotics Autom. Lett., 7(2):1888–1895, 2022. 2
2022
-
[7]
Berk Calli, Aaron Walsman, Arjun Singh, Siddhartha Srini- vasa, Pieter Abbeel, and Aaron M. Dollar. Benchmarking in manipulation research: Using the yale-cmu-berkeley object and model set.IEEE Robotics & Automation Magazine, 22 (3):36–52, 2015. 2, 7, 8
2015
-
[8]
Learning to ex- plore using active neural slam
Devendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta, and Ruslan Salakhutdinov. Learning to ex- plore using active neural slam. InInternational Conference on Learning Representations (ICLR), 2020. 2
2020
-
[9]
Jiaqi Chen, Bingqian Lin, Xinmin Liu, Lin Ma, Xiao- dan Liang, and Kwan-Yee K. Wong. Affordances-oriented planning using foundation models for continuous vision- language navigation. InProceedings of the AAAI Conference on Artificial Intelligence, 2025. 7, 8
2025
-
[10]
Transferable active grasp- ing and real embodied dataset
Xiangyu Chen, Zelin Ye, Jiankai Sun, Yuda Fan, Fang Hu, Chenxi Wang, and Cewu Lu. Transferable active grasp- ing and real embodied dataset. In2020 IEEE International Conference on Robotics and Automation, ICRA 2020, Paris, France, May 31 - August 31, 2020, pages 3611–3618. IEEE,
2020
-
[11]
Fu-Jen Chu, Ruinian Xu, and Patricio A. Vela. Real-world multiobject, multigrasp detection.IEEE Robotics Autom. Lett., 3(4):3355–3362, 2018. 2
2018
-
[12]
Pybullet, a python mod- ule for physics simulation for games, robotics and machine learning.http://pybullet.org, 2016–2021
Erwin Coumans and Yunfei Bai. Pybullet, a python mod- ule for physics simulation for games, robotics and machine learning.http://pybullet.org, 2016–2021. 7
2016
-
[13]
Congyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard, Andrea Tagliasacchi, and Leonidas J. Guibas. Vector neu- rons: A general framework for so(3)-equivariant networks. In2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pages 12180–12189. IEEE, 2021. 6
2021
-
[14]
Pred-nbv: Prediction-guided next-best-view planning for 3d object reconstruction
Harnaik Dhami, Vishnu D Sharma, and Pratap Tokekar. Pred-nbv: Prediction-guided next-best-view planning for 3d object reconstruction. In2023 IEEE/RSJ International Con- ference on Intelligent Robots and Systems (IROS), pages 7149–7154. IEEE, 2023. 2
2023
-
[15]
ACRONYM: A large-scale grasp dataset based on simula- tion
Clemens Eppner, Arsalan Mousavian, and Dieter Fox. ACRONYM: A large-scale grasp dataset based on simula- tion. InIEEE International Conference on Robotics and Au- tomation, ICRA 2021, Xi’an, China, May 30 - June 5, 2021, pages 6222–6227. IEEE, 2021. 6, 8
2021
-
[16]
A density-based algorithm for discovering clusters in large spatial databases with noise
Martin Ester, Hans-Peter Kriegel, J ¨org Sander, and Xiaowei Xu. A density-based algorithm for discovering clusters in large spatial databases with noise. InProceedings of the Sec- ond International Conference on Knowledge Discovery and Data Mining (KDD-96), Portland, Oregon, USA, pages 226–
-
[17]
Anygrasp: Robust and efficient grasp perception in spa- tial and temporal domains.IEEE Trans
Haoshu Fang, Chenxi Wang, Hongjie Fang, Minghao Gou, Jirong Liu, Hengxu Yan, Wenhai Liu, Yichen Xie, and Cewu Lu. Anygrasp: Robust and efficient grasp perception in spa- tial and temporal domains.IEEE Trans. Robotics, 39(5): 3929–3945, 2023. 1
2023
-
[18]
Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, Muhammad Shahzad, Wen Yang, Richard Bamler, and Xiaoxiang Zhu
Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna M. Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, Muhammad Shahzad, Wen Yang, Richard Bamler, and Xiaoxiang Zhu. A survey of uncertainty in deep neural networks.Artif. Intell. Rev., 56(S1):1513–1589, 2023. 2
2023
-
[19]
Uncertainty-driven planner for exploration and navigation
Georgios Georgakis, Bernadette Bucher, Anton Arapin, Karl Schmeckpeper, Nikolai Matni, and Kostas Daniilidis. Uncertainty-driven planner for exploration and navigation. InICRA, 2022. 2
2022
-
[20]
Bayes’ Rays: Uncertainty quantifica- tion in neural radiance fields.arXiv, 2023
Lily Goli, Cody Reading, Silvia Sell ´an, Alec Jacobson, and Andrea Tagliasacchi. Bayes’ Rays: Uncertainty quantifica- tion in neural radiance fields.arXiv, 2023. 2
2023
-
[21]
Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, J ¨orn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swer- sky. Your classifier is secretly an energy based model and you should treat it like one. In8th International Confer- ence on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. 3, 5
2020
-
[22]
Viewpoint selection for grasp detection
Marcus Gualtieri and Robert Platt Jr. Viewpoint selection for grasp detection. In2017 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2017, Vancouver, BC, Canada, September 24-28, 2017, pages 258–264. IEEE,
2017
-
[23]
High precision grasp pose detection in dense clutter
Marcus Gualtieri, Andreas ten Pas, Kate Saenko, and Robert Platt Jr. High precision grasp pose detection in dense clutter. In2016 IEEE/RSJ International Conference on In- telligent Robots and Systems, IROS 2016, Daejeon, South Korea, October 9-14, 2016, pages 598–605. IEEE, 2016. 1
2016
-
[24]
Scone: Surface coverage optimization in unknown environ- ments by volumetric integration.Advances in Neural Infor- mation Processing Systems, 35:20731–20743, 2022
Antoine Gu ´edon, Pascal Monasse, and Vincent Lepetit. Scone: Surface coverage optimization in unknown environ- ments by volumetric integration.Advances in Neural Infor- mation Processing Systems, 35:20731–20743, 2022. 2
2022
-
[25]
Macarons: Mapping and coverage anticipa- tion with rgb online self-supervision
Antoine Gu ´edon, Tom Monnier, Pascal Monasse, and Vin- cent Lepetit. Macarons: Mapping and coverage anticipa- tion with rgb online self-supervision. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 940–951, 2023. 2
2023
-
[26]
Weinberger
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural networks. InProceedings of the 34th International Conference on Machine Learn- ing, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, pages 1321–1330. PMLR, 2017. 3
2017
-
[27]
Orbitgrasp: Se(3)-equivariant grasp learning
Boce Hu, Xupeng Zhu, Dian Wang, Zihao Dong, Haojie Huang, Chenghao Wang, Robin Walters, and Robert Platt. Orbitgrasp: Se(3)-equivariant grasp learning. InConference on Robot Learning, 6-9 November 2024, Munich, Germany, pages 2456–2474. PMLR, 2024. 2
2024
-
[28]
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y . Gal- liker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pert...
Pith/arXiv arXiv 2025
-
[29]
Maddox, Polina Kirichenko, Timur Garipov, Dmitry P
Pavel Izmailov, Wesley J. Maddox, Polina Kirichenko, Timur Garipov, Dmitry P. Vetrov, and Andrew Gordon Wil- son. Subspace inference for bayesian deep learning. InPro- ceedings of the Thirty-Fifth Conference on Uncertainty in Artificial Intelligence, UAI 2019, Tel Aviv, Israel, July 22-25, 2019, pages 1169–1179. AUAI Press, 2019. 2
2019
-
[30]
Fisherrf: Ac- tive view selection and mapping with radiance fields using fisher information
Wen Jiang, Boshu Lei, and Kostas Daniilidis. Fisherrf: Ac- tive view selection and mapping with radiance fields using fisher information. InComputer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part XIII, pages 422–440. Springer,
2024
-
[31]
Multimodal llm guided exploration and active map- ping using fisher information
Wen Jiang, Boshu Lei, Katrina Ashton, and Kostas Dani- ilidis. Multimodal llm guided exploration and active map- ping using fisher information. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5392–5404, 2025. 2
2025
-
[32]
Neu-nbv: Next best view planning using uncer- tainty estimation in image-based neural rendering
Liren Jin, Xieyuanli Chen, Julius R ¨uckin, and Marija Popovi´c. Neu-nbv: Next best view planning using uncer- tainty estimation in image-based neural rendering. In2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 11305–11312. IEEE, 2023. 2
2023
-
[33]
Edward Johns, Stefan Leutenegger, and Andrew J. Davison. Deep learning a grasp function for grasping under gripper pose uncertainty. In2016 IEEE/RSJ International Confer- ence on Intelligent Robots and Systems, IROS 2016, Dae- jeon, South Korea, October 9-14, 2016, pages 4461–4468. IEEE, 2016. 2
2016
-
[34]
Bopar- dikar, Julian Ryde, Kenneth Y
Gregory Kahn, Peter Sujan, Sachin Patil, Shaunak D. Bopar- dikar, Julian Ryde, Kenneth Y . Goldberg, and Pieter Abbeel. Active exploration using trajectory optimization for robotic grasping in the presence of occlusions. InIEEE International Conference on Robotics and Automation, ICRA 2015, Seat- tle, WA, USA, 26-30 May, 2015, pages 4783–4790. IEEE,
2015
-
[35]
Spherical fibonacci mapping.ACM Trans
Benjamin Keinert, Matthias Innmann, Michael S ¨anger, and Marc Stamminger. Spherical fibonacci mapping.ACM Trans. Graph., 34(6):193:1–193:7, 2015. 7
2015
-
[36]
3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics, 42 (4), 2023
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics, 42 (4), 2023. 3
2023
-
[37]
Unifying approaches in ac- tive learning and active sampling via fisher information and information-theoretic quantities.Transactions on Machine Learning Research, 2022
Andreas Kirsch and Yarin Gal. Unifying approaches in ac- tive learning and active sampling via fisher information and information-theoretic quantities.Transactions on Machine Learning Research, 2022. Expert Certification. 2, 3
2022
-
[38]
A survey on learning-based robotic grasp- ing.Current Robotics Reports, 1:239–249, 2020
Kilian Kleeberger, Richard Bormann, Werner Kraus, and Marco F Huber. A survey on learning-based robotic grasp- ing.Current Robotics Reports, 1:239–249, 2020. 2
2020
-
[39]
Meelis Kull, Miquel Perell ´o-Nieto, Markus K ¨angsepp, Telmo de Menezes e Silva Filho, Hao Song, and Pe- ter A. Flach. Beyond temperature scaling: Obtaining well- calibrated multiclass probabilities with dirichlet calibration. CoRR, abs/1910.12656, 2019. 3
Pith/arXiv arXiv 1910
-
[40]
A tutorial on energy-based learning.Predicting structured data, 1(0), 2006
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, Fujie Huang, et al. A tutorial on energy-based learning.Predicting structured data, 1(0), 2006. 4
2006
-
[41]
Deep learning for detecting robotic grasps.Int
Ian Lenz, Honglak Lee, and Ashutosh Saxena. Deep learning for detecting robotic grasps.Int. J. Robotics Res., 34(4-5): 705–724, 2015. 2
2015
-
[42]
Eval- uating and calibrating uncertainty prediction in regression tasks.Sensors, 22(15):5540, 2022
Dan Levi, Liran Gispan, Niv Giladi, and Ethan Fetaya. Eval- uating and calibrating uncertainty prediction in regression tasks.Sensors, 22(15):5540, 2022. 3
2022
-
[43]
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data col- lection.Int
Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen. Learning hand-eye coordination for robotic grasping with deep learning and large-scale data col- lection.Int. J. Robotics Res., 37(4-5):421–436, 2018. 1
2018
-
[44]
Byeongdo Lim, Jongmin Kim, Jihwan Kim, Yonghyeon Lee, and Frank C. Park. Equigraspflow: Se(3)-equivariant 6-dof grasp pose generative flows. InConference on Robot Learn- ing, 6-9 November 2024, Munich, Germany, pages 5067–
2024
-
[45]
Active perception for grasp detection via neural graspness field
Haoxiang Ma, Modi Shi, Boyang Gao, and Di Huang. Active perception for grasp detection via neural graspness field. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. 1, 2, 6, 7, 8
2024
-
[46]
Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics
Jeffrey Mahler, Jacky Liang, Sherdil Niyaz, Michael Laskey, Richard Doan, Xinyu Liu, Juan Aparicio Ojea, and Ken Goldberg. Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics. In Robotics: Science and Systems XIII, Massachusetts Institute of Technology, Cambridge, Massachusetts, USA, July 12-16, 2017, 2017. 1, 2
2017
-
[47]
Wells III, Clare M
Alireza Mehrtash, William M. Wells III, Clare M. Tempany, Purang Abolmaesumi, and Tina Kapur. Confidence calibra- tion and predictive uncertainty estimation for deep medical image segmentation.IEEE Trans. Medical Imaging, 39(12): 3868–3878, 2020. 2
2020
-
[48]
Multi- view picking: Next-best-view reaching for improved grasp- ing in clutter
Douglas Morrison, Peter Corke, and J ¨urgen Leitner. Multi- view picking: Next-best-view reaching for improved grasp- ing in clutter. InInternational Conference on Robotics and Automation, ICRA 2019, Montreal, QC, Canada, May 20-24, 2019, pages 8762–8768. IEEE, 2019. 1, 2
2019
-
[49]
Ac- tivenerf: Learning where to see with uncertainty estimation
Xuran Pan, Zihang Lai, Shiji Song, and Gao Huang. Ac- tivenerf: Learning where to see with uncertainty estimation. InECCV, pages 230–246. Springer, 2022. 2
2022
-
[50]
Gabriel Pereyra, George Tucker, Jan Chorowski, Lukasz Kaiser, and Geoffrey E. Hinton. Regularizing neural net- works by penalizing confident output distributions. In5th In- ternational Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Workshop Track Proceedings. OpenReview.net, 2017. 2
2017
-
[51]
Supersizing self- supervision: Learning to grasp from 50k tries and 700 robot hours
Lerrel Pinto and Abhinav Gupta. Supersizing self- supervision: Learning to grasp from 50k tries and 700 robot hours. In2016 IEEE International Conference on Robotics and Automation, ICRA 2016, Stockholm, Sweden, May 16- 21, 2016, pages 3406–3413. IEEE, 2016. 2
2016
-
[52]
Grasp learning: Models, methods, and per- formance.Annual Review of Control, Robotics, and Au- tonomous Systems, 6(1):363–389, 2023
Robert Platt. Grasp learning: Models, methods, and per- formance.Annual Review of Control, Robotics, and Au- tonomous Systems, 6(1):363–389, 2023. 2
2023
-
[53]
Charles Ruizhongtai Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In2017 IEEE Con- ference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pages 77–85. IEEE Computer Society, 2017. 6
2017
-
[54]
Occupancy anticipation for efficient exploration and navigation
Santhosh K Ramakrishnan, Ziad Al-Halah, and Kristen Grauman. Occupancy anticipation for efficient exploration and navigation. InComputer Vision–ECCV 2020: 16th Eu- ropean Conference, Glasgow, UK, August 23–28, 2020, Pro- ceedings, Part V 16, pages 400–418. Springer, 2020. 2
2020
-
[55]
Neurar: Neural uncertainty for autonomous 3d reconstruction with implicit neural representations.IEEE Robotics and Automation Let- ters, 8(2):1125–1132, 2023
Yunlong Ran, Jing Zeng, Shibo He, Jiming Chen, Lincheng Li, Yingfeng Chen, Gimhee Lee, and Qi Ye. Neurar: Neural uncertainty for autonomous 3d reconstruction with implicit neural representations.IEEE Robotics and Automation Let- ters, 8(2):1125–1132, 2023. 2
2023
-
[56]
SAM 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao- Yuan Wu, Ross Girshick, Piotr Dollar, and Christoph Feicht- enhofer. SAM 2: Segment anything in images and videos. In The Thirteenth Inte...
2025
-
[57]
Real-time grasp de- tection using convolutional neural networks
Joseph Redmon and Anelia Angelova. Real-time grasp de- tection using convolutional neural networks. InIEEE In- ternational Conference on Robotics and Automation, ICRA 2015, Seattle, WA, USA, 26-30 May, 2015, pages 1316–1322. IEEE, 2015. 2
2015
-
[58]
Rezende, and C´esar Roberto de Souza
J ´erˆome Revaud, Jon Almaz ´an, Rafael S. Rezende, and C´esar Roberto de Souza. Learning with average preci- sion: Training image retrieval with a listwise loss. In2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 5106–5115. IEEE, 2019. 5
2019
-
[59]
Learn- ing for single-shot confidence calibration in deep neural net- works through stochastic inferences
Seonguk Seo, Paul Hongsuck Seo, and Bohyung Han. Learn- ing for single-shot confidence calibration in deep neural net- works through stochastic inferences. InIEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 9030–9038. Computer Vision Foundation / IEEE, 2019. 2
2019
-
[60]
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. InAd- vances in Neural Information Processing Systems 32: An- nual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 11895–11907, 2019. 4, 5
2019
-
[61]
Contact-graspnet: Efficient 6-dof grasp gen- eration in cluttered scenes
Martin Sundermeyer, Arsalan Mousavian, Rudolph Triebel, and Dieter Fox. Contact-graspnet: Efficient 6-dof grasp gen- eration in cluttered scenes. In2021 IEEE International Con- ference on Robotics and Automation (ICRA), pages 13438– 13444. IEEE, 2021. 6, 7
2021
-
[62]
Rethinking the in- ception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the in- ception architecture for computer vision. In2016 IEEE Con- ference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV , USA, June 27-30, 2016, pages 2818–
2016
-
[63]
Grasp pose detection in point clouds.The International Journal of Robotics Research, 36(13-14):1455–1473, 2017
Andreas Ten Pas, Marcus Gualtieri, Kate Saenko, and Robert Platt. Grasp pose detection in point clouds.The International Journal of Robotics Research, 36(13-14):1455–1473, 2017. 8
2017
-
[64]
Se(3)-diffusionfields: Learning smooth cost func- tions for joint grasp and motion optimization through diffu- sion
Julen Urain, Niklas Funk, Jan Peters, and Georgia Chal- vatzaki. Se(3)-diffusionfields: Learning smooth cost func- tions for joint grasp and motion optimization through diffu- sion. InIEEE International Conference on Robotics and Au- tomation, ICRA 2023, London, UK, May 29 - June 2, 2023, pages 5923–5930. IEEE, 2023. 2, 4, 6, 7, 8
2023
-
[65]
Graspness discovery in clutters for fast and accurate grasp detection.CoRR, abs/2406.11142,
Chenxi Wang, Haoshu Fang, Minghao Gou, Hongjie Fang, Jin Gao, and Cewu Lu. Graspness discovery in clutters for fast and accurate grasp detection.CoRR, abs/2406.11142,
-
[66]
GPR: grasp pose refine- ment network for cluttered scenes
Wei Wei, Yongkang Luo, Fuyu Li, Guangyun Xu, Jun Zhong, Wanyi Li, and Peng Wang. GPR: grasp pose refine- ment network for cluttered scenes. InIEEE International Conference on Robotics and Automation, ICRA 2021, Xi’an, China, May 30 - June 5, 2021, pages 4295–4302. IEEE,
2021
-
[67]
Capgrasp: An $\mathbb{R}ˆ{3}\times\text{SO(2)- Equivariant}$ continuous approach-constrained generative grasp sampler.IEEE Robotics Autom
Zehang Weng, Haofei Lu, Jens Lundell, and Danica Kragic. Capgrasp: An $\mathbb{R}ˆ{3}\times\text{SO(2)- Equivariant}$ continuous approach-constrained generative grasp sampler.IEEE Robotics Autom. Lett., 9(4):3641–3647,
-
[68]
Non-parametric calibration for classification
Jonathan Wenger, Hedvig Kjellstr ¨om, and Rudolph Triebel. Non-parametric calibration for classification. InThe 23rd In- ternational Conference on Artificial Intelligence and Statis- tics, AISTATS 2020, 26-28 August 2020, Online [Palermo, Sicily, Italy], pages 178–190. PMLR, 2020. 3
2020
-
[69]
Active neural mapping
Zike Yan, Haoxiang Yang, and Hongbin Zha. Active neural mapping. InICCV, 2023. 2
2023
-
[70]
Robotic grasping from classical to modern: A survey.arXiv preprint arXiv:2202.03631, 2022
Hanbo Zhang, Jian Tang, Shiguang Sun, and Xuguang Lan. Robotic grasping from classical to modern: A survey.arXiv preprint arXiv:2202.03631, 2022. 2
Pith/arXiv arXiv 2022
-
[71]
Affordance-driven next-best- view planning for robotic grasping
Xuechao Zhang, Dong Wang, Sun Han, Weichuang Li, Bin Zhao, Zhigang Wang, Xiaoming Duan, Chongrong Fang, Xuelong Li, and Jianping He. Affordance-driven next-best- view planning for robotic grasping. InConference on Robot Learning, CoRL 2023, 6-9 November 2023, Atlanta, GA, USA, pages 2849–2862. PMLR, 2023. 1, 2, 6, 7, 8
2023
-
[72]
Open3d: A modern library for 3d data processing.CoRR, abs/1801.09847, 2018
Qian-Yi Zhou, Jaesik Park, and Vladlen Koltun. Open3d: A modern library for 3d data processing.CoRR, abs/1801.09847, 2018. 6
Pith/arXiv arXiv 2018
-
[2826]
IEEE Computer Society, 2016. 2
2016
This paper was first reviewed by deepseek-v4-flash on August 3, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.