Pith. sign in

REVIEW 4 major objections 5 minor 72 references

ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Tactile and joint feedback, applied as springs, rescue visual 6D pose tracking during in-hand manipulation.

desk verdict Useful test-time refinement idea with a clean formulation, but the headline gains rest on an under-described self-collected benchmark. read the letter →

arxiv 2504.13179 v1 pith:VIFM455P submitted 2025-04-17 cs.RO cs.CV

classification cs.ROcs.CV
keywords 6Dposeestimationvisuotactileperceptionzero-shotlearningtest-timeoptimizationspring-massmodeltactilesensingproprioceptionin-handmanipulation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to show that a purely visual object-pose tracker can be made reliable during in-hand manipulation without any visuotactile training data. Its trick is to treat tactile contacts and robot joint readings as physical constraints, check every visual estimate against them, and when a check fails refine the pose by minimizing the energy of a virtual spring–mass system. On a real two-arm robot with five novel objects, this wrapper raises the ADD-S area under the curve by 55% and ADD by 60% relative to FoundationPose, and lowers position error by 80%, with MegaPose improving too. The broader point is that proprioception and touch can serve as zero-shot prior knowledge instead of a second training modality.

What carries the argument

The load-bearing mechanism is the spring–mass test-time optimizer with its feasibility gate. Tactile readings and proprioception are represented as point clouds located through forward kinematics, and the visual pose is accepted only if contacts lie within a threshold of the object, voxelized penetration stays below a bound, and frame-to-frame motion is feasible. When a check fails, the optimizer minimizes the sum of an attractive spring energy $E_a=\frac12 k_a \min_{i,j}\|\Delta(p_i^O)-p_j^S\|^2$ and a repulsive penetration energy $E_r=\frac12 k_r \max(0,\gamma)^2$ over a relative pose $\Delta$ parameterized in angle–axis plus translation. This is what lets the correction inherit the visual model's generalization while adding physical grounding.

What would settle it

Re-run the same grasping, picking, and handover trials while tracking the objects with an independent ground-truth system, such as motion-capture markers read by external high-speed cameras, and recompute ADD-S, ADD, and position error for ViTa-Zero against FoundationPose and MegaPose; if the reported margins shrink substantially or reverse, the central claim fails.

Watch

Extended reading notes

Core claim

The central claim is that tactile and proprioceptive observations, converted into three feasibility checks—contact, penetration, and kinematic feasibility—can tell when a visual pose estimate has gone wrong, and a test-time spring–mass optimization can pull the pose back to a physically plausible state. An attractive spring links active tactile points to the object surface, a repulsive spring penalizes signed penetration into the robot hand, and the resulting pose retracks from the corrected state. No tactile dataset is collected and no visuotactile model is trained; the framework only needs the object mesh, known tactile sensor positions, and robot forward kinematics. The reported experiments show consistent gains over both visual backbones in grasping, object picking, and bimanual handover.

Load-bearing premise

The reported gains rest on a self-collected test dataset whose ground-truth pose annotation procedure is not described; if those ground-truth poses are noisy or were derived from the same robot kinematics and tactile contacts the method uses, the improvements could be inflated.

Editorial extensions

If this is right

  • Any capable visual pose estimator with access to an object mesh and robot kinematics can be wrapped in this constraint-and-refine loop without retraining.
  • State-based manipulation policies should hold up better under occlusion, because tracking errors get corrected instead of accumulated across frames.
  • The 20 Hz online rate reported for the full loop puts the tactile refinement inside a real-time control cycle.
  • Because tactile signals are normalized into point clouds, the same framework should transfer from fingertip taxels to vision-based or force-based tactile sensors.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The reported gains are measured in failure-rich manipulation sequences, so on easy static scenes the improvement should shrink; the headline numbers describe recovery, not baseline accuracy.
  • Systematically sweeping the global thresholds ($\theta_c$, $\theta_p$, $\theta_d$) per object and per grasp could show where the physical constraints bind and whether one constraint alone drives most of the gain.
  • The same feasibility gate could serve as a safety monitor in production manipulation: persistent infeasibility could trigger a stop or regrasp before tracking diverges.
  • Extending the optimization to soft or deformable grippers would require replacing forward-kinematics contact points with estimated contact locations, which is an open problem the paper itself notes.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes ViTa-Zero, a zero-shot visuotactile framework for 6D pose estimation and tracking of in-hand manipulated objects. A visual estimator (FoundationPose or MegaPose) supplies an initial pose, which is then checked against three physical constraints derived from tactile and proprioceptive observations: contact, penetration, and kinematic feasibility. If the visual pose is infeasible, a test-time optimization refines it by minimizing a spring-mass energy consisting of an attractive term pulling the object toward activated tactile sensors and a repulsive term preventing penetration with the robot model. The authors evaluate the framework on a real UR5e/PSYONIC Ability hand platform with five novel objects across grasping, picking, and bimanual handover scenarios, reporting average improvements of 55% in ADD-S AUC, 60% in ADD AUC, and 80% lower position error relative to FoundationPose. The paper also includes ablations on loss terms, refinement algorithm, and optimization initialization, as well as runtime measurements.

Significance. If the quantitative results hold, ViTa-Zero is a practically valuable contribution: it removes the need for tactile training data, represents tactile signals in a sensor-agnostic way, and can wrap arbitrary visual pose estimators. The physical-constraint formulation and the spring-mass refinement are clean and well motivated, and the ablations provide some evidence that each loss term contributes. However, the central quantitative claims currently rest on a self-collected dataset whose ground-truth annotation procedure is never described, and the evaluation lacks statistical support. The method is also not yet reproducible without code or data release. With the evaluation gaps addressed, the framework would be a solid contribution to visuotactile pose estimation; at present, the headline numbers are not fully verifiable.

major comments (4)
  1. [Section V-B, Table I] The manuscript never describes how ground-truth poses for the self-collected test dataset were obtained. The only statement is that the authors 'collect a dataset for testing purposes.' This is load-bearing because ViTa-Zero's tactile point cloud is computed from forward kinematics, and the same robot model is used for penetration constraints. If the ground-truth poses were derived from the robot's kinematic chain or a known object-in-gripper transform, the comparison against vision-only FoundationPose and MegaPose would be biased in the method's favor. Please provide a complete annotation protocol, state whether the ground truth is independent of the kinematic model (e.g., motion capture, fiducial markers, or manually verified poses), and report the number of frames, sequences, and trials per object and per scenario, together with per-object results. This is necessary to support the headline 55%/60%/80% improvement claims.
  2. [Section IV-C, Eqs. (4)-(5)] The optimization objective is not written as a coherent function of a single relative pose TΔ. In the attractive energy, TΔ is applied to the object point cloud (TΔ(p_i^O)), while in the repulsive energy, TΔ is applied to the robot point cloud (TΔ(p_i^R)) when computing γ in Eq. (4). If TΔ is meant to be the relative pose update applied to the object pose, then the repulsive term should compare the transformed object against a fixed robot model; as written, the two energy terms apply TΔ to different bodies, so it is unclear what candidate pose the objective actually evaluates. Please clarify the notation and, if the implementation matches the equations, explain how T* = TΔ ∘ T is consistent with both terms.
  3. [Section V-B, Tables I-IV] No error bars, standard deviations, or numbers of frames are reported for any metric. The aggregate numbers in Table I could be dominated by a small number of severe failure frames, and the paper does not report how often feasibility checking rejects the visual estimate and triggers refinement. Since the method's value proposition is recovering from visual tracking failures, the frequency of refinement triggering and the distribution of improvements across sequences are essential. Please add per-sequence and per-object metrics, repeated-trial statistics, and the fraction of frames in which the visual estimate was deemed infeasible.
  4. [Section V, hyperparameters] The thresholds and loss weights θc = 0.05, θp = 0.008, θd = 0.03, k_a = 1, k_r = 1000, λ = 1000, learning rate 10^-3, and ten optimization iterations are fixed without sensitivity analysis. These parameters define the entire refinement behavior, so the reader cannot tell whether the reported gains are robust or the result of tuning on this five-object test set. Please add a sensitivity study over at least the main loss weights and the contact/penetration thresholds, and report how the metrics vary.
minor comments (5)
  1. [Section V] There is a typo in 'dataesets' (should be 'datasets'), and the author listing contains an unusual spacing artifact in 'Tas ¸kın' that should be fixed.
  2. [Section IV-C] The expression '1/2kx2' should be typeset as (1/2)kx^2, and the notation should be kept consistent with the later use of Ea and Er.
  3. [Section IV-B, Eq. (3)] The term P_n is described as the contact patch on the object model for the n-th frame, but it is never formally defined. Please specify how P_n is computed from the visual pose and the tactile signal.
  4. [Section V] The threshold values θc = 0.05 and θd = 0.03 are given without units. If they are in meters, please state this explicitly.
  5. [Section V-B, Table V] The component runtimes in Table V sum to about 15 Hz if run serially (1/(0.011+0.006+0.051) ≈ 14.7 Hz), so the statement that the framework 'could reach an online refresh rate of around 20 Hz' needs clarification, e.g., which components run in parallel or whether the reported pose tracking time already includes feasibility checking.

Circularity Check

0 steps flagged · score 0.0 of 10

No load-bearing circularity: ViTa-Zero is a physical test-time optimization, and the cited self-works are not load-bearing.

full rationale

ViTa-Zero does not derive its output from its inputs by construction. The three feasibility constraints (Eqs. 1-3) and the spring-mass objective (Eqs. 4-5) encode physical assumptions: contact, non-penetration, and smooth motion. The reported gains over FoundationPose and MegaPose are empirical measurements, not quantities implied by the equations. The hyperparameters (theta_c, theta_p, theta_d, ka, kr, lambda) are hand-set, but there is no claim that they were fitted to reproduce the reported ADD-S, ADD, or PE values. The authors' prior works (ViHOPE [24], HyperTaxel [31], V-HOP [32]) appear only in the related-work survey and are not load-bearing in the method derivation. The main validity risk is the evaluation ground truth: Section V-B says a dataset was collected 'for testing purposes' but does not describe the annotation procedure. If the ground-truth poses were generated from the same robot forward kinematics that the method consumes, the comparison could be biased; however, the paper provides no text or equation exhibiting such a reduction, so under the evidence standard this is a correctness and verifiability concern, not an established circularity. The Limitations section also explicitly states reliance on forward kinematics, which is a stated assumption rather than a concealed input. No circular step can be quoted from the manuscript.

Assumptions & free parameters 8 free parameters · 5 assumptions · 2 invented entities

The method depends on several manually chosen thresholds and loss weights that are not derived from data or theory, plus assumptions about sensor accuracy, robot calibration, and the availability of object meshes. The virtual springs are optimization constructs rather than claimed physical entities. The ground-truth annotation for the self-collected dataset is the most consequential unstated element.

free parameters (8)
  • contact threshold theta_c = 0.05
    Manual threshold for the contact constraint (Eq. 1). Not derived from data or theory.
  • penetration threshold theta_p = 0.008
    Manual allowable overlap threshold for the penetration constraint (Eq. 2).
  • kinematic threshold theta_d = 0.03
    Manual distance threshold for the kinematic feasibility constraint (Eq. 3).
  • attractive spring weight k_a = 1
    Weight on the attractive potential in the test-time optimization (Eq. 5).
  • repulsive spring weight k_r = 1000
    Weight on the repulsive potential in the test-time optimization (Eq. 5).
  • L2 regularization weight lambda = 1000
    Weight on the L2 regularization term in the optimization objective (Eq. 5).
  • optimizer learning rate = 1e-3
    Adam learning rate for the test-time optimization, stated in Section V.
  • number of optimization iterations = 10
    Empirically tuned iteration count for the Adam optimizer, stated in Section V.
assumptions (5)
  • domain assumption Tactile sensors only make contact with the object and not with other parts of the robot.
    Stated in Section III (Sensor): 'we assume that the sensors will only make contact with the object O and no self-contact.' This is needed for the contact constraint and attractive spring to be meaningful.
  • domain assumption The object is rigid and its 3D mesh is available.
    Stated in Section III: 'we assume the 3D mesh M_O of the object O is either given or reconstructed.' The method samples points on the mesh for contact checking and optimization.
  • domain assumption Robot kinematic model and hand-eye calibration are accurate.
    Stated in Section III: the kinematic model and part meshes are given, and the camera-to-end-effector transform is obtained via hand-eye calibration. Tactile point cloud positions are computed from forward kinematics, so calibration errors propagate into the refinement.
  • domain assumption The visual backbone provides a sufficiently accurate initial pose, and the object is visible in the initial frame.
    Stated in Section III: 'The object O is observed using an RGB-D sensor and visible for at least the initial frame.' The local spring-mass optimization can only refine near the visual estimate; a grossly wrong initial pose could prevent convergence.
  • domain assumption Tactile point cloud positions computed from forward kinematics accurately represent the true contact points.
    In Section IV-B, tactile signals are represented as a point cloud 'whose positions could be obtained through forward kinematics.' For FSR sensors, binary thresholding discretizes contact; the paper does not analyze the resulting positional error, which directly affects the attractive spring.
invented entities (2)
  • Virtual attractive spring from tactile contacts
    purpose: Pulls the object pose toward the tactile point cloud in test-time optimization (E_a term in Eq. 5).
    A modeling construct, not a physical force. Its strength is set by the free parameter k_a and it has no external falsifiable handle outside the optimization objective.
  • Virtual repulsive spring from robot model
    purpose: Pushes the object pose out of penetration with the robot mesh (E_r term in Eq. 5).
    A modeling construct, not a physical force. Its strength is set by the free parameter k_r and it has no external falsifiable handle outside the optimization objective.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation." pith.science (2026). https://pith.science/paper/VIFM455P

@misc{pith2026250413179,
  author       = {Pith},
  title        = {Pith review of: ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VIFM455P}},
  note         = {Machine review of arXiv:2504.13179}
}
read the original abstract

Object 6D pose estimation is a critical challenge in robotics, particularly for manipulation tasks. While prior research combining visual and tactile (visuotactile) information has shown promise, these approaches often struggle with generalization due to the limited availability of visuotactile data. In this paper, we introduce ViTa-Zero, a zero-shot visuotactile pose estimation framework. Our key innovation lies in leveraging a visual model as its backbone and performing feasibility checking and test-time optimization based on physical constraints derived from tactile and proprioceptive observations. Specifically, we model the gripper-object interaction as a spring-mass system, where tactile sensors induce attractive forces, and proprioception generates repulsive forces. We validate our framework through experiments on a real-world robot setup, demonstrating its effectiveness across representative visual backbones and manipulation scenarios, including grasping, object picking, and bimanual handover. Compared to the visual models, our approach overcomes some drastic failure modes while tracking the in-hand object pose. In our experiments, our approach shows an average increase of 55% in AUC of ADD-S and 60% in ADD, along with an 80% lower position error compared to FoundationPose.

Figures

Figures reproduced from arXiv: 2504.13179 by the authors.

Figure 1
Figure 1. FoundationPose [16] (left) fails due to errors that are [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of ViTa-Zero. Red fingers in the robot hand model represent activated fingertip tactile sensors. the sensors will only make contact with the object O and no self-contact. Although we focus on taxel-based sensors in this paper, our framework can extend to other types of sensors, which we will discuss in the next section. IV. METHODOLOGY Our framework consists of three modules: visual es￾timation, feasibility… view at source ↗
Figure 3
Figure 3. Qualitative results. We demonstrate the performance during the object picking and bimanual handover tasks with “camera” and “eyedrop” objects. During these manipulation tasks, it is common to encounter scenarios where the object is highly occluded while moving, as illustrated in the figure. Visual approaches, like FoundationPose, can lose tracking and fail due to the absence of visual information. In contrast, our m… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Our robot platform. Our setup consists of two Universal Robots UR 5e arms and PSYONIC Ability hands. hand has five fingers, each equipped with six FSR sensors on the fingertip to provide tactile sensing. More details regarding this platform are available in [46]. Visua…
Figure 5
Figure 5. Figure 5: Comparison with FoundationPose (FP) and Mega￾Pose (MP). The performance is measured by the AUC of ADD-S and ADD (the higher, the better) and position error (PE) (the lower, the better). We evaluate the performance using position error (PE) and the area under the curve …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

72 extracted references · 60 canonical work pages

  1. [1]

    A Review of Robot Learn- ing for Manipulation: Challenges, Representations, and Algorithms,

    O. Kroemer, S. Niekum, and G. Konidaris, “A Review of Robot Learn- ing for Manipulation: Challenges, Representations, and Algorithms,” Journal of Machine Learning Research , vol. 22, no. 30, pp. 1–82, 2021

  2. [2]

    Visual dexterity: In-hand reorientation of novel and complex object shapes,

    T. Chen, M. Tippur, S. Wu, V . Kumar, E. Adelson, and P. Agrawal, “Visual dexterity: In-hand reorientation of novel and complex object shapes,” Science Robotics , vol. 8, no. 84, p. eadc9244, Nov. 2023, publisher: American Association for the Advancement of Science

  3. [3]

    A System for General In-Hand Object Re-Orientation,

    T. Chen, J. Xu, and P. Agrawal, “A System for General In-Hand Object Re-Orientation,” in Proceedings of the 5th Conference on Robot Learning. PMLR, Jan. 2022, pp. 297–307, iSSN: 2640-3498

  4. [4]

    General In-hand Object Rotation with Vision and Touch,

    H. Qi, B. Yi, S. Suresh, M. Lambeta, Y . Ma, R. Calandra, and J. Malik, “General In-hand Object Rotation with Vision and Touch,” Aug. 2023

  5. [5]

    Challenges for Monocular 6D Object Pose Estima- tion in Robotics,

    t. Thalhammer, D. Bauer, P. H ¨onig, J.-B. Weibel, J. Garc´ıa-Rodr´ıguez, and M. Vincze, “Challenges for Monocular 6D Object Pose Estima- tion in Robotics,” IEEE Transactions on Robotics , pp. 1–20, 2024, conference Name: IEEE Transactions on Robotics

  6. [6]

    Solving Rubik’s Cube with a Robot Hand,

    OpenAI, I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, J. Schneider, N. Tezak, J. Tworek, P. Welinder, L. Weng, Q. Yuan, W. Zaremba, and L. Zhang, “Solving Rubik’s Cube with a Robot Hand,” Oct. 2019, arXiv:1910.07113 [cs, stat]

  7. [7]

    Learning dexterous in-hand manipulation,

    O. M. Andrychowicz, B. Baker, M. Chociej, R. J ´ozefowicz, B. Mc- Grew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba, “Learning dexterous in-hand manipulation,” The International Journal of Robotics Research , vol. 39, no. 1, pp. 3–20, Jan. 2020, publisher: SAGE Publications Ltd STM

  8. [8]

    DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality,

    A. Handa, A. Allshire, V . Makoviychuk, A. Petrenko, R. Singh, J. Liu, D. Makoviichuk, K. Van Wyk, A. Zhurkevich, B. Sundaralingam, and Y . Narang, “DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) , May 2023, pp. 5977–5984

Show all 72 references
  1. [9]

    PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes,

    Y . Xiang, T. Schmidt, V . Narayanan, and D. Fox, “PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes,” vol. 14, Jun. 2018

  2. [10]

    DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion,

    C. Wang, D. Xu, Y . Zhu, R. Martin-Martin, C. Lu, L. Fei-Fei, and S. Savarese, “DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion,” 2019, pp. 3343–3352

  3. [11]

    Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose Estimation,

    K. Park, T. Patten, and M. Vincze, “Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose Estimation,” 2019, pp. 7668–7677

  4. [12]

    MRC-Net: 6-DoF Pose Estimation with MultiScale Residual Correlation,

    Y . Li, Y . Mao, R. Bala, and S. Hadap, “MRC-Net: 6-DoF Pose Estimation with MultiScale Residual Correlation,” 2024, pp. 10 476– 10 486

  5. [13]

    Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation,

    H. Wang, S. Sridhar, J. Huang, J. Valentin, S. Song, and L. J. Guibas, “Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation,” 2019, pp. 2642–2651

  6. [14]

    Learning Canonical Shape Space for Category-Level 6D Object Pose and Size Estimation,

    D. Chen, J. Li, Z. Wang, and K. Xu, “Learning Canonical Shape Space for Category-Level 6D Object Pose and Size Estimation,” 2020, pp. 11 973–11 982

  7. [15]

    TTA-COPE: Test-Time Adaptation for Category-Level Object Pose Estimation,

    T. Lee, J. Tremblay, V . Blukis, B. Wen, B.-U. Lee, I. Shin, S. Birch- field, I. S. Kweon, and K.-J. Yoon, “TTA-COPE: Test-Time Adaptation for Category-Level Object Pose Estimation,” 2023, pp. 21 285–21 295

  8. [16]

    FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects,

    B. Wen, W. Yang, J. Kautz, and S. Birchfield, “FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects,” 2024, pp. 17 868–17 879

  9. [17]

    MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare,

    Y . Labb´e, L. Manuelli, A. Mousavian, S. Tyree, S. Birchfield, J. Trem- blay, J. Carpentier, M. Aubry, D. Fox, and J. Sivic, “MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare,” Aug. 2022

  10. [18]

    OnePose++: Keypoint-Free One-Shot Object Pose Estimation without CAD Mod- els,

    X. He, J. Sun, Y . Wang, D. Huang, H. Bao, and X. Zhou, “OnePose++: Keypoint-Free One-Shot Object Pose Estimation without CAD Mod- els,” Advances in Neural Information Processing Systems , vol. 35, pp. 35 103–35 115, Dec. 2022

  11. [19]

    Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGB Images,

    Y . Liu, Y . Wen, S. Peng, C. Lin, X. Long, T. Komura, and W. Wang, “Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGB Images,” in Computer Vision – ECCV 2022 , S. Avidan, G. Brostow, M. Ciss ´e, G. M. Farinella, and T. Hassner, Eds. Cham: Springer Nature S...

  12. [20]

    SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation,

    J. Lin, L. Liu, D. Lu, and K. Jia, “SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation,” 2024, pp. 27 906– 27 916

  13. [21]

    Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,

    C. Chi, S. Feng, Y . Du, Z. Xu, E. Cousineau, B. C. Burchfiel, and S. Song, “Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,” vol. 19, Jul. 2023

  14. [22]

    3D Diffusion Policy,

    Y . Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu, “3D Diffusion Policy,” Mar. 2024, arXiv:2403.03954 [cs]

  15. [23]

    NeuralFeels with neural fields: Visuotactile perception for in-hand manipulation,

    S. Suresh, H. Qi, T. Wu, T. Fan, L. Pineda, M. Lambeta, J. Malik, M. Kalakrishnan, R. Calandra, M. Kaess, J. Ortiz, and M. Mukadam, “NeuralFeels with neural fields: Visuotactile perception for in-hand manipulation,” Science Robotics , Nov. 2024, publisher: American Association...

  16. [24]

    ViHOPE: Visuotactile In- Hand Object 6D Pose Estimation With Shape Completion,

    H. Li, S. Dikhale, S. Iba, and N. Jamali, “ViHOPE: Visuotactile In- Hand Object 6D Pose Estimation With Shape Completion,” IEEE Robotics and Automation Letters , vol. 8, no. 11, pp. 6963–6970, Nov. 2023

  17. [25]

    VisuoTactile 6D Pose Estimation of an In-Hand Object Using Vision and Tactile Sensor Data,

    S. Dikhale, K. Patel, D. Dhingra, I. Naramura, A. Hayashi, S. Iba, and N. Jamali, “VisuoTactile 6D Pose Estimation of an In-Hand Object Using Vision and Tactile Sensor Data,”IEEE Robotics and Automation Letters, vol. 7, no. 2, pp. 2148–2155, Apr. 2022

  18. [26]

    VinT-6D: A Large- Scale Object-in-hand Dataset from Vision, Touch and Proprioception,

    Z. Wan, Y . Ling, S. Yi, L. Qi, W. W. Lee, M. Lu, S. Yang, X. Teng, P. Lu, X. Yang, M.-H. Yang, and H. Cheng, “VinT-6D: A Large- Scale Object-in-hand Dataset from Vision, Touch and Proprioception,” in Proceedings of the 41st International Conference on Machine Learning. PMLR, ...

  19. [27]

    Hierarchical Graph Neural Networks for Proprioceptive 6D Pose Estimation of In-hand Objects,

    A. Rezazadeh, S. Dikhale, S. Iba, and N. Jamali, “Hierarchical Graph Neural Networks for Proprioceptive 6D Pose Estimation of In-hand Objects,” in 2023 IEEE International Conference on Robotics and Automation (ICRA), May 2023, pp. 2884–2890

  20. [28]

    PoseFusion: Robust Object-in-Hand Pose Estimation with SelectLSTM,

    Y . Tu, J. Jiang, S. Li, N. Hendrich, M. Li, and J. Zhang, “PoseFusion: Robust Object-in-Hand Pose Estimation with SelectLSTM,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Oct. 2023, pp. 6839–6846, iSSN: 2153-0866

  21. [29]

    Enhancing Generalizable 6D Pose Tracking of an In-Hand Object with Tactile Sensing,

    Y . Liu, X. Xu, W. Chen, H. Yuan, H. Wang, J. Xu, R. Chen, and L. Yi, “Enhancing Generalizable 6D Pose Tracking of an In-Hand Object with Tactile Sensing,” IEEE Robotics and Automation Letters , vol. 9, no. 2, pp. 1106–1113, Feb. 2024, arXiv:2210.04026 [cs]

  22. [30]

    In-Hand Pose Estimation Using Hand-Mounted RGB Cameras and Visuotactile Sensors,

    Y . Gao, S. Matsuoka, W. Wan, T. Kiyokawa, K. Koyama, and K. Harada, “In-Hand Pose Estimation Using Hand-Mounted RGB Cameras and Visuotactile Sensors,” IEEE Access, vol. 11, pp. 17 218– 17 232, 2023, conference Name: IEEE Access

  23. [31]

    HyperTaxel: Hyper-Resolution for Taxel-Based Tactile Signals Through Contrastive Learning,

    H. Li, S. Dikhale, J. Cui, S. Iba, and N. Jamali, “HyperTaxel: Hyper-Resolution for Taxel-Based Tactile Signals Through Contrastive Learning,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , Oct. 2024, pp. 7499–7506, iSSN: 2153- 0866

  24. [32]

    V-HOP: Visuo-Haptic 6D Object Pose Tracking,

    H. Li, M. Jia, T. Akbulut, Y . Xiang, G. Konidaris, and S. Srid- har, “V-HOP: Visuo-Haptic 6D Object Pose Tracking,” Feb. 2025, arXiv:2502.17434 [cs]

  25. [33]

    Touch and Go: Learning from Human-Collected Vision and Touch,

    F. Yang, C. Ma, J. Zhang, J. Zhu, W. Yuan, and A. Owens, “Touch and Go: Learning from Human-Collected Vision and Touch,” Jun. 2022

  26. [34]

    TACTO: A Fast, Flexible, and Open-Source Simulator for High-Resolution Vision- Based Tactile Sensors,

    S. Wang, M. Lambeta, P.-W. Chou, and R. Calandra, “TACTO: A Fast, Flexible, and Open-Source Simulator for High-Resolution Vision- Based Tactile Sensors,” IEEE Robotics and Automation Letters, vol. 7, no. 2, pp. 3930–3937, Apr. 2022, conference Name: IEEE Robotics and Automatio...

  27. [35]

    Fast Model-Based Contact Patch and Pose Estimation for Highly Deformable Dense-Geometry Tactile Sensors,

    N. Kuppuswamy, A. Castro, C. Phillips-Grafflin, A. Alspach, and R. Tedrake, “Fast Model-Based Contact Patch and Pose Estimation for Highly Deformable Dense-Geometry Tactile Sensors,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 1811–1818, Apr. 2020, conference Nam...

  28. [36]

    GelSlim 3.0: High- Resolution Measurement of Shape, Force and Slip in a Compact Tactile-Sensing Finger,

    I. H. Taylor, S. Dong, and A. Rodriguez, “GelSlim 3.0: High- Resolution Measurement of Shape, Force and Slip in a Compact Tactile-Sensing Finger,” in 2022 International Conference on Robotics and Automation (ICRA) , May 2022, pp. 10 781–10 787

  29. [37]

    BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects,

    B. Wen, J. Tremblay, V . Blukis, S. Tyree, T. Muller, A. Evans, D. Fox, J. Kautz, and S. Birchfield, “BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects,” Mar. 2023, arXiv:2303.14158 [cs]

  30. [38]

    Instant Neural Graphics Primitives with a Multiresolution Hash Encoding,

    T. M ¨uller, A. Evans, C. Schied, and A. Keller, “Instant Neural Graphics Primitives with a Multiresolution Hash Encoding,” ACM Trans. Graph. , vol. 41, no. 4, pp. 102:1–102:15, Jul. 2022, place: New York, NY , USA Publisher: ACM

  31. [39]

    ViSP for visual servoing: a generic software platform with a wide class of robot control skills,

    E. Marchand, F. Spindler, and F. Chaumette, “ViSP for visual servoing: a generic software platform with a wide class of robot control skills,” IEEE Robotics & Automation Magazine, vol. 12, no. 4, pp. 40–52, Dec. 2005, conference Name: IEEE Robotics & Automation Magazine

  32. [40]

    Tactile Object Pose Estimation from the First Touch with Geometric Contact Rendering,

    M. B. Villalonga, A. Rodriguez, B. Lim, E. Valls, and T. Sechopoulos, “Tactile Object Pose Estimation from the First Touch with Geometric Contact Rendering,” in Proceedings of the 2020 Conference on Robot Learning. PMLR, Oct. 2021, pp. 1015–1029, iSSN: 2640-3498

  33. [41]

    Collision- aware In-hand 6D Object Pose Estimation using Multiple Vision-based Tactile Sensors,

    G. M. Caddeo, N. A. Piga, F. Bottarel, and L. Natale, “Collision- aware In-hand 6D Object Pose Estimation using Multiple Vision-based Tactile Sensors,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) , May 2023, pp. 719–725

  34. [42]

    Dexterity from Touch: Self-Supervised Pre-Training of Tactile Representations with Robotic Play,

    I. Guzey, B. Evans, S. Chintala, and L. Pinto, “Dexterity from Touch: Self-Supervised Pre-Training of Tactile Representations with Robotic Play,” in Proceedings of The 7th Conference on Robot Learning . PMLR, Dec. 2023, pp. 3142–3166, iSSN: 2640-3498

  35. [43]

    Binding Touch to Everything: Learning Unified Multimodal Tactile Representations,

    F. Yang, C. Feng, Z. Chen, H. Park, D. Wang, Y . Dou, Z. Zeng, X. Chen, R. Gangopadhyay, A. Owens, and A. Wong, “Binding Touch to Everything: Learning Unified Multimodal Tactile Representations,” Jan. 2024, arXiv:2401.18084 null

  36. [44]

    Tactile- Augmented Radiance Fields,

    Y . Dou, F. Yang, Y . Liu, A. Loquercio, and A. Owens, “Tactile- Augmented Radiance Fields,” May 2024, arXiv:2405.04534 [cs]

  37. [45]

    Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks,

    J. Zhao, Y . Ma, L. Wang, and E. H. Adelson, “Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks,” Jun. 2024, arXiv:2406.13640 [cs]

  38. [46]

    Learning Visuotactile Skills with Two Multifingered Hands,

    T. Lin, Y . Zhang, Q. Li, H. Qi, B. Yi, S. Levine, and J. Malik, “Learning Visuotactile Skills with Two Multifingered Hands,” Apr. 2024, arXiv:2404.16823 [cs]

  39. [47]

    Rotating without Seeing: Towards In-hand Dexterity through Touch,

    Z.-H. Yin, B. Huang, Y . Qin, Q. Chen, and X. Wang, “Rotating without Seeing: Towards In-hand Dexterity through Touch,” vol. 19, Jul. 2023

  40. [48]

    ArrayBot: Reinforcement Learning for Generalizable Distributed Ma- nipulation through Touch,

    Z. Xue, H. Zhang, J. Cheng, Z. He, Y . Ju, C. Lin, G. Zhang, and H. Xu, “ArrayBot: Reinforcement Learning for Generalizable Distributed Ma- nipulation through Touch,” Jun. 2023, arXiv:2306.16857 [cs]

  41. [49]

    DIGIT: A Novel Design for a Low-Cost Compact High- Resolution Tactile Sensor With Application to In-Hand Manipulation,

    M. Lambeta, P.-W. Chou, S. Tian, B. Yang, B. Maloon, V . R. Most, D. Stroud, R. Santos, A. Byagowi, G. Kammerer, D. Jayaraman, and R. Calandra, “DIGIT: A Novel Design for a Low-Cost Compact High- Resolution Tactile Sensor With Application to In-Hand Manipulation,” IEEE Robotic...

  42. [50]

    GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and Force,

    W. Yuan, S. Dong, and E. H. Adelson, “GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and Force,” Sensors, vol. 17, no. 12, p. 2762, Dec. 2017

  43. [51]

    GelSlim: A High-Resolution, Compact, Robust, and Calibrated Tactile-sensing Finger,

    E. Donlon, S. Dong, M. Liu, J. Li, E. Adelson, and A. Rodriguez, “GelSlim: A High-Resolution, Compact, Robust, and Calibrated Tactile-sensing Finger,” May 2018, arXiv:1803.00628 [cs]

  44. [52]

    Tactile Mapping and Local- ization from High-Resolution Tactile Imprints,

    M. Bauza, O. Canal, and A. Rodriguez, “Tactile Mapping and Local- ization from High-Resolution Tactile Imprints,” in 2019 International Conference on Robotics and Automation (ICRA), May 2019, pp. 3811– 3817, iSSN: 2577-087X

  45. [53]

    Midas- Touch: Monte-Carlo inference over distributions across sliding touch,

    S. Suresh, Z. Si, S. Anderson, M. Kaess, and M. Mukadam, “Midas- Touch: Monte-Carlo inference over distributions across sliding touch,” in Proceedings of The 6th Conference on Robot Learning . PMLR, Mar. 2023, pp. 319–331, iSSN: 2640-3498

  46. [54]

    ShapeMap 3-D: Efficient shape mapping through dense touch and vision,

    S. Suresh, Z. Si, J. G. Mangelson, W. Yuan, and M. Kaess, “ShapeMap 3-D: Efficient shape mapping through dense touch and vision,” in 2022 International Conference on Robotics and Automation (ICRA) , May 2022, pp. 7073–7080

  47. [55]

    Monocular Depth Estimation for Soft Visuotactile Sensors,

    R. Ambrus, V . Guizilini, N. Kuppuswamy, A. Beaulieu, A. Gaidon, and A. Alspach, “Monocular Depth Estimation for Soft Visuotactile Sensors,” in 2021 IEEE 4th International Conference on Soft Robotics (RoboSoft), Apr. 2021, pp. 643–649

  48. [56]

    Making Sense of Vision and Touch: Learning Multimodal Representations for Contact-Rich Tasks,

    M. A. Lee, Y . Zhu, P. Zachares, M. Tan, K. Srinivasan, S. Savarese, L. Fei-Fei, A. Garg, and J. Bohg, “Making Sense of Vision and Touch: Learning Multimodal Representations for Contact-Rich Tasks,” IEEE Transactions on Robotics , vol. 36, no. 3, pp. 582–596, Jun. 2020, confer...

  49. [57]

    CPF: Learning a Contact Potential Field to Model the Hand-Object Interaction,

    L. Yang, X. Zhan, K. Li, W. Xu, J. Li, and C. Lu, “CPF: Learning a Contact Potential Field to Model the Hand-Object Interaction,” in2021 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, pp. 11 077–11 086, iSSN: 2380-7504

  50. [58]

    Learning a Contact Potential Field for Modeling the Hand-Object Interaction,

    L. Yang, X. Zhan, K. Li, W. Xu, J. Zhang, J. Li, and C. Lu, “Learning a Contact Potential Field for Modeling the Hand-Object Interaction,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 8, pp. 5645–5662, Aug. 2024, conference Name: IEEE Transacti...

  51. [59]

    ContactGrasp: Func- tional Multi-finger Grasp Synthesis from Contact,

    S. Brahmbhatt, A. Handa, J. Hays, and D. Fox, “ContactGrasp: Func- tional Multi-finger Grasp Synthesis from Contact,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , Nov. 2019, pp. 2386–2393, iSSN: 2153-0866

  52. [60]

    Deep Regression on Manifolds: A 3D Rotation Case Study,

    R. Br ´egier, “Deep Regression on Manifolds: A 3D Rotation Case Study,” in 2021 International Conference on 3D Vision (3DV) , Dec. 2021, pp. 166–174, iSSN: 2475-7888

  53. [61]

    PyTorch: An Imperative Style, High-Performance Deep Learning Library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “PyTorch: An Imperative Style, High-Pe...

  54. [62]

    Adam: A Method for Stochastic Optimiza- tion,

    D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimiza- tion,” Jan. 2017, arXiv:1412.6980 [cs]

  55. [63]

    Objaverse: A Universe of Annotated 3D Objects,

    M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. Vander- Bilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi, “Objaverse: A Universe of Annotated 3D Objects,” 2023, pp. 13 142–13 153

  56. [64]

    Google Scanned Objects: A High-Quality Dataset of 3D Scanned Household Items,

    L. Downs, A. Francis, N. Koenig, B. Kinman, R. Hickman, K. Rey- mann, T. B. McHugh, and V . Vanhoucke, “Google Scanned Objects: A High-Quality Dataset of 3D Scanned Household Items,” in 2022 International Conference on Robotics and Automation (ICRA) , May 2022, pp. 2553–2560

  57. [65]

    Segment Anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollar, and R. Gir- shick, “Segment Anything,” 2023, pp. 4015–4026

  58. [66]

    Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection,

    S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, J. Zhu, and L. Zhang, “Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection,” Jul. 2024, arXiv:2303.05499 [cs]

  59. [67]

    BOP Challenge 2023 on Detection Segmentation and Pose Estimation of Seen and Unseen Rigid Objects,

    T. Hodan, M. Sundermeyer, Y . Labbe, V . N. Nguyen, G. Wang, E. Brachmann, B. Drost, V . Lepetit, C. Rother, and J. Matas, “BOP Challenge 2023 on Detection Segmentation and Pose Estimation of Seen and Unseen Rigid Objects,” 2024, pp. 5610–5619

  60. [68]

    Method for registration of 3-D shapes,

    P. J. Besl and N. D. McKay, “Method for registration of 3-D shapes,” in Sensor Fusion IV: Control Paradigms and Data Structures , vol

  61. [69]

    Stable, open-loop precision manipula- tion with underactuated hands,

    L. U. Odhner and A. M. Dollar, “Stable, open-loop precision manipula- tion with underactuated hands,” The International Journal of Robotics Research, vol. 34, no. 11, pp. 1347–1360, Sep. 2015, publisher: SAGE Publications Ltd STM

  62. [70]

    Fully body visual self-modeling of robot morphologies,

    B. Chen, R. Kwiatkowski, C. V ondrick, and H. Lipson, “Fully body visual self-modeling of robot morphologies,” Science Robotics, vol. 7, no. 68, p. eabn1944, Jul. 2022, publisher: American Association for the Advancement of Science

  63. [71]

    Unifying 3D Representation and Control of Diverse Robots with a Single Camera,

    S. L. Li, A. Zhang, B. Chen, H. Matusik, C. Liu, D. Rus, and V . Sitzmann, “Unifying 3D Representation and Control of Diverse Robots with a Single Camera,” Jul. 2024, arXiv:2407.08722 [cs]

  64. [1611]

    1992, pp

    SPIE, Apr. 1992, pp. 586–606

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.