Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

ARMADA: Augmented Reality for Robot Manipulation and Robot-Free Data Acquisition

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read ARMADA claims that real-time AR feedback of a simulated robot is critical for collecting robot-free demonstrations that replay directly on physical hardware, lifting average replay success from 1.3% to 71.1%.

desk verdict Solid systems paper with a real empirical result, but the headline causal claim is partly confounded by a fixed condition order and an instruction change; worth serious review. read the letter →

arxiv 2412.10631 v1 pith:5EF4HZ52 submitted 2024-12-14 cs.RO

classification cs.RO
keywords augmentedrealityrobot-freedatacollectionimitationlearningdigitaltwinteleoperationAppleVisionProreplaysuccesshumandemonstrations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ARMADA is a system for collecting robot-manipulation demonstrations without a physical robot: a demonstrator wears an Apple Vision Pro, performs tasks barehanded, and watches a simulated robot arm mirror their hand in real time in augmented reality. The paper's central claim is that this live feedback is what makes the collected trajectories usable, since direct replay on physical hardware succeeds 71.1% of the time on average with feedback versus 1.3% without it. It also reports that demonstrations performed after a feedback session, with the overlay hidden, still replay 33.3% of the time, suggesting some learned transfer. If true, this removes the teleoperation-hardware bottleneck and opens a route to large-scale human data collection for imitation learning.

What carries the argument

The load-bearing mechanism is the closed 30 Hz feedback loop: the headset's camera-based hand tracking estimates wrist and knuckle poses, an inverse-kinematics solver maps that pose to joint commands for a simulated six-degree-of-freedom arm, and the resulting robot state is rendered as an AR overlay in the headset. Gripper open/close is mapped from thumb-index distance, so the human hand becomes a two-finger pincher in simulation. The system also renders constraint cues — a yellow color shift near singularities, a red wall at workspace limits, and a 'Slow Down' alert — that train the demonstrator to stay inside robot-compatible motion.

What would settle it

Counterbalance the condition order so half the participants receive Feedback first, then No Feedback; if the feedback effect is causal, the No Feedback condition should still yield near-zero replay success regardless of prior practice, and the Feedback condition should still exceed 70%.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that a real-time augmented-reality digital twin of a robot, overlaid onto a human's bare hands, turns otherwise unusable human motion into hardware-compatible robot trajectories. With ARMADA, 15 participants gave 675 demonstrations across three tasks; average replay success rose from 1.3% without feedback to 71.1% with feedback, with per-task gains of 76, 48, and 85 percentage points. After feedback was removed, success remained at 33.3%, which the authors read as evidence that people internalize some of the robot's constraints from the feedback session. The failures that remain with feedback are mostly small position and orientation errors from AR depth perception.

Load-bearing premise

The paper assumes the jump from 1.3% to 71.1% replay success is caused by the AR feedback itself, but every participant experienced No Feedback first, then Feedback, then Post Feedback, so practice and growing familiarity with the tasks could also explain part of the gain.

Editorial extensions

If this is right

  • Anyone with a Vision Pro can contribute demonstrations without owning a robot, removing the hardware bottleneck that limits imitation-learning data collection.
  • Because trajectories replay on real arms directly, ARMADA sidesteps the embodiment-gap problem faced by human-video approaches.
  • The Post Feedback result implies a short AR training session can improve barehanded data collection even when the overlay is later switched off, potentially lowering the cost of large-scale collection.
  • The plug-and-play interface means the same system can serve different robot arms by swapping 3D models and reconfiguring the communication layer, not just the specific ViperX arms tested.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the causal role of live feedback holds up under counterbalancing, the feedback itself could be made adaptive — for instance, showing only corrective cues — to push replay success beyond the 71% reported here.
  • A natural next experiment is to train a policy on ARMADA-collected trajectories and compare it against a policy trained on teleoperation data of the same tasks, which would test whether robot-free data closes the quality gap in downstream imitation learning.
  • The Post Feedback retention suggests a two-stage pipeline: teach demonstrators with AR once, then collect barehanded without headsets, which could scale collection beyond the number of available Vision Pro headsets.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents ARMADA, an Apple Vision Pro application that overlays a real-time simulated robot (a digital twin) on the user's view of their own hands, so that a human can provide barehanded manipulation demonstrations that are meant to be compatible with a physical robot. The authors report a user study with 15 participants, 3 tasks (Pick Tissue, Declutter, Bimanual Wipe), 5 starting states per task, and 3 feedback conditions (No Feedback, Feedback, Post Feedback), yielding 675 demonstrations that are directly replayed on physical Interbotix ViperX 300 arms. The central claim is that live AR feedback of the virtual robot is critical for collecting robot-free demonstrations that are directly replayable on real hardware, supported by Table I, which shows average replay success of 1.3% for No Feedback, 71.1% for Feedback, and 33.3% for Post Feedback. The paper also reports survey data suggesting the visualization is intuitive and useful, and describes the system architecture, IK-based control from hand tracking, and AR constraint visualizations (singularity, workspace, speed).

Significance. If the central claim is correct, ARMADA would be a notable step toward scalable imitation-learning data collection without physical robot access: it uses a consumer headset, requires no robot hardware during data collection, and the reported 1.3% to 71.1% improvement in direct replay success is a striking and practically meaningful effect. The evaluation has real strengths: an external, physical-hardware benchmark with success criteria defined independently of the system, a within-subject design with 15 participants, and 675 total demonstrations. The inclusion of a Post Feedback condition and qualitative reports is also valuable. However, the causal claim that feedback itself is critical is currently undermined by two design confounds: the No Feedback condition changes both feedback and task instruction, and the fixed condition order without counterbalancing confounds feedback with practice and fatigue. The absence of statistical inference further weakens the quantitative claims. The potential significance is high, but the current evidence does not yet cleanly separate the effect of feedback from these confounds.

major comments (3)
  1. [IV-A] The comparison that carries the paper's central claim contrasts No Feedback with Feedback while simultaneously changing the task instruction: participants in No Feedback are told 'to demonstrate the task with natural human motion,' whereas participants in Feedback are told 'to demonstrate the task such that the virtual robot is controlled to execute the task.' The near-zero No Feedback replay success (1.3% average in Table I) could therefore reflect the instruction to behave naturally rather than the absence of visual feedback. Please add a control condition that holds the instruction constant, for example by instructing participants to move as if controlling a virtual robot while showing them no visualization, or by counterbalancing the instruction wording across feedback conditions.
  2. [IV-C] All participants experience the three conditions in the same fixed order (No Feedback, then Feedback, then Post Feedback), and the task order, although randomized per participant, is held fixed across conditions. As a result, Feedback is always performed after 15 prior demonstrations and Post Feedback after 30, so practice and fatigue effects are perfectly confounded with feedback type. The interpretation that Post Feedback improves over No Feedback due to learning from feedback is confounded by the additional demonstration experience, and the Feedback vs Post Feedback gap may be inflated by fatigue. A counterbalanced design, or a control group that performs repeated No Feedback blocks with the same number of demonstrations, is needed to support the causal role of feedback.
  3. [V-A, Table I] The text repeatedly states that Feedback 'significantly' outperforms No Feedback and Post Feedback, but no statistical tests are reported. With 15 participants and only 5 trials per condition-task cell, the standard deviations are large (e.g., 21.7% for Declutter under Feedback), so the observed differences may not be statistically reliable without a paired analysis. Please report within-participant paired tests (e.g., Wilcoxon signed-rank) or a mixed-effects model with participant random effects, and provide effect sizes. This is needed to support the quantitative 'significantly higher' claims in Section V-A.
minor comments (4)
  1. [II-A] The sentence 'their utility is bottlenecked the ability to translate' is missing the word 'by'; it should read 'bottlenecked by the ability to translate.'
  2. [III-A, III-C] Figure references are inconsistently capitalized: the text uses 'fig. 2' and 'fig. 3' in some places and 'Figure 4' and 'Figure 5' in others; please standardize.
  3. [I] The paper states that code 'will be available on the project website when finalized' but provides no repository or release timeline in the current version. For a systems paper whose contribution includes the data collection pipeline, making code and the 675 collected trajectories available at the time of publication would substantially aid reproducibility.
  4. [V-A] The survey results report Likert means with standard deviations but do not include the full distribution or the exact number of responses per item; a supplementary table with the per-item response counts would be clearer.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: ARMADA's central claim rests on direct physical-robot replay success, an external benchmark independent of the system's definitions.

full rationale

The paper's load-bearing assertion that real-time AR feedback improves robot-free demonstration quality is evaluated by direct trajectory replay on physical Interbotix ViperX 300 hardware (Section IV-D), with task-specific success criteria (Section IV-B) defined independently of the ARMADA system. No fitted parameters are used to produce the reported 1.3% vs 71.1% success rates, and no prediction is derived from the system's own definitions; the result is an empirical comparison across feedback conditions. The fixed condition order and differing task instructions between No Feedback and Feedback are internal-validity concerns about causal attribution, not circularity: they do not make the measured replay outcomes equivalent to the input by construction. Self-citations are absent as load-bearing evidence; related-work comparisons (AR2-D2, DART, ARCap) are external context. The appropriate finding is no significant circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

No new physical entities are introduced; the digital twin is a visualization, not a new force or particle. The central claim is empirical and comparative. The listed free parameters are system implementation thresholds, not fitted to the evaluation data. The key unstated premise is that condition order does not confound the result.

free parameters (4)
  • Gripper close/open distance threshold
    Distance between thumb and index finger that triggers gripper closing or opening (Section III-B). Chosen by hand; affects gripper timing during collection but not the relative feedback comparison.
  • Singularity determinant threshold
    Threshold on det(J^T J) below which the virtual robot turns yellow (Section III-C). Feedback-only, not load-bearing for the central claim.
  • Velocity limit for 'Slow Down' warning
    Hand speed above which a warning is shown (Section III-C). Feedback-only.
  • Wrist-to-robot position offset and angle
    Hand-chosen shift and rotation that place the control point between thumb and index (Section III-B). Affects absolute success but not the comparative claim.
assumptions (4)
  • domain assumption ARKit hand tracking provides sufficiently accurate 3D wrist and finger pose estimates for real-time robot control.
    The entire system depends on this without validation against ground truth (Section III-B).
  • domain assumption The average of the four knuckle positions is a valid proxy for the intended end-effector pose.
    Used for position and orientation control; no justification beyond usability (Section III-B).
  • domain assumption Direct replay success on the physical robot is a valid measure of demonstration quality for imitation learning.
    The evaluation uses replay success as the proxy for data quality (Section IV-D), without validating that policies trained on this data perform well.
  • ad hoc to paper The fixed condition order does not introduce practice effects that confound the feedback comparison.
    The protocol always runs No Feedback before Feedback before Post Feedback (Section IV-C), implicitly assuming order effects are negligible.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ARMADA: Augmented Reality for Robot Manipulation and Robot-Free Data Acquisition." pith.science (2026). https://pith.science/paper/5EF4HZ52

@misc{pith2026241210631,
  author       = {Pith},
  title        = {Pith review of: ARMADA: Augmented Reality for Robot Manipulation and Robot-Free Data Acquisition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5EF4HZ52}},
  note         = {Machine review of arXiv:2412.10631}
}
read the original abstract

Teleoperation for robot imitation learning is bottlenecked by hardware availability. Can high-quality robot data be collected without a physical robot? We present a system for augmenting Apple Vision Pro with real-time virtual robot feedback. By providing users with an intuitive understanding of how their actions translate to robot motions, we enable the collection of natural barehanded human data that is compatible with the limitations of physical robot hardware. We conducted a user study with 15 participants demonstrating 3 different tasks each under 3 different feedback conditions and directly replayed the collected trajectories on physical robot hardware. Results suggest live robot feedback dramatically improves the quality of the collected data, suggesting a new avenue for scalable human data collection without access to robot hardware. Videos and more are available at https://nataliya.dev/armada.

Figures

Figures reproduced from arXiv: 2412.10631 by the authors.

Figure 1
Figure 1. Overview. (A) Human demonstrators wearing Apple Vision Pro can collect data directly with their hands. (B) Egocentric view within Vision Pro shows real-time robot execution overlaid on the user’s hands with augmented reality. (C) High-quality demonstrations collected with this system can be directly replayed on physical robot hardware. reality with Apple Vision Pro. By overlaying a simulated robot on the high-resolu… view at source ↗
Figure 2
Figure 2. Overview of the system architecture described in Section III-A. Human skeletal data is sent over websockets to an external compute device, which [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Constraints. (A) The virtual robot gradually turns yellow as it approaches a singular configuration. (B) The background turns red when the robot moves beyond the Cartesian space boundaries. (C) The user is alerted with a virtual text overlay when their hand motion exceeds the robot’s velocity limits [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Tasks. The three tasks from Section IV-B. Left: egocentric view of initial states during data collection with Feedback. Middle: egocentric view of final states during data collection with Feedback. Right: mid-task robot execution during trajectory replay. I V. U S E R …
Figure 5
Figure 5. Figure 5: User Interface. Left: main menu with toggle options. Right, Top: hand placement in spheres to initialize demonstration. Right, Bottom: hand placement to release spheres and end demonstration. D. Robot System The robot system consists of a pair of 6 DOF Interbotix Viper…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning

    cs.RO 2026-07 conditional novelty 5.0 of 10

    AgenticFocus converts ordinary human FPV videos into object-preserving mixed-reality humanoid demos with lower trajectory error and smoother wrist motion (SPARC −5.18) than Masquerade and Do as I Do.

Reference graph

Works this paper leans on

62 extracted references · 39 canonical work pages · cited by 1 Pith paper

  1. [1]

    The user is instructed to demonstrate the task with natural human motion

    No Feedback: The user does not see any AR feedback of the robot. The user is instructed to demonstrate the task with natural human motion

  2. [2]

    The user is instructed to demonstrate the task such that the virtual robot is controlled to execute the task

    Feedback: The user sees an AR digital twin of the robot updated in real time. The user is instructed to demonstrate the task such that the virtual robot is controlled to execute the task

  3. [3]

    The user is once again instructed to demonstrate the task such that the virtual robot is controlled to execute the task

    Post Feedback: The user once again cannot see any AR feedback of the robot. The user is once again instructed to demonstrate the task such that the virtual robot is controlled to execute the task. However, since the robot is no longer visible, the user must make an educated guess based on their experience with the Feedback condition as to how the robot is...

  4. [4]

    Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation,

    Z. Fu, T. Z. Zhao, and C. Finn, “Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation,” in Conference on Robot Learning (CoRL) , 2024

  5. [5]

    Robot execution is deemed successful if the tissue has been removed from the box

    Pick Tissue: With one hand, grasp a tissue and pull it out of a tissue box. Robot execution is deemed successful if the tissue has been removed from the box

  6. [6]

    Robot execution is deemed successful if the toy has been placed into the cardboard box

    Declutter: With one hand, pick a soft toy from the tabletop and place it into a cardboard box. Robot execution is deemed successful if the toy has been placed into the cardboard box

  7. [7]

    The robot visualization was intuitive and easy to understand

    Bimanual Wipe: With two hands, wipe a cloth across the table surface in the direction away from the demon- strator. Robot execution is deemed successful if both end effectors push the cloth away from the respective bases of the robot arms. Each task has 5 distinct starting states that consist of a unique tissue box position, toy position, and cloth positi...

  8. [8]

    Imitation learning: A survey of learning methods,

    A. Hussein, M. M. Gaber, E. Elyan, and C. Jayne, “Imitation learning: A survey of learning methods,” ACM Comput. Surv. , vol. 50, no. 2, Apr. 2017. [Online]. Available: https://doi.org/10.1145/3054912

Show all 62 references
  1. [9]

    Diffusion policy: Visuomotor policy learning via action diffusion,

    C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” 2024. [Online]. Available: https://arxiv.org/abs/2303.04137

  2. [10]

    Learning fine-grained bimanual manipulation with low-cost hardware,

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn, “Learning fine-grained bimanual manipulation with low-cost hardware,” in Robotics: Science and Systems (RSS) , 2023

  3. [11]

    R3m: A universal visual representation for robot manipulation,

    S. Nair, A. Rajeswaran, V . Kumar, C. Finn, and A. Gupta, “R3m: A universal visual representation for robot manipulation,” in Conference on Robot Learning (CoRL) , 2022

  4. [12]

    Aloha unleashed: A simple recipe for robot dexterity,

    T. Z. Zhao, J. Tompson, D. Driess, P. Florence, K. Ghasemipour, C. Finn, and A. Wahid, “Aloha unleashed: A simple recipe for robot dexterity,” 2024. [Online]. Available: https://arxiv.org/abs/2410.13126

  5. [13]

    Llama: Open and efficient foundation language models,

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, et al. , “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023

  6. [14]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics ,

  7. [15]

    Octo: An open-source generalist robot policy,

    Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y . Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine, “Octo: An open-source generalist robot policy,” in Proceedings of Robotic...

  8. [16]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo,et al., “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 4015–4026

  9. [17]

    Maxvit: Multi-axis vision transformer,

    Z. Tu, H. Talebi, H. Zhang, F. Yang, P. Milanfar, A. Bovik, and Y . Li, “Maxvit: Multi-axis vision transformer,” in European conference on computer vision . Springer, 2022, pp. 459–479

  10. [18]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763

  11. [19]

    Enhancing dexterity in robotic manipulation via hierarchical contact exploration,

    X. Cheng, S. Patil, Z. Temel, O. Kroemer, and M. T. Mason, “Enhancing dexterity in robotic manipulation via hierarchical contact exploration,” IEEE Robotics and Automation Letters , vol. 9, no. 1, pp. 390–397, 2024

  12. [20]

    Real-world robot learning with masked visual pre-training,

    I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Dar- rell, “Real-world robot learning with masked visual pre-training,” in Conference on Robot Learning . PMLR, 2023, pp. 416–426

  13. [21]

    Vip: Towards universal visual reward and representation via value-implicit pre-training,

    Y . J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V . Kumar, and A. Zhang, “Vip: Towards universal visual reward and representation via value-implicit pre-training,” in International Conference on Learning Representations (ICLR), 2023

  14. [22]

    Open x-embodiment: Robotic learning datasets and rt-x models,

    A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al., “Open x-embodiment: Robotic learning datasets and rt-x models,” in IEEE International Conference on Robotics and Automation (ICRA) , 2024

  15. [23]

    Dexpilot: Vision-based teleopera- tion of dexterous robotic hand-arm system,

    A. Handa, K. Van Wyk, W. Yang, J. Liang, Y .-W. Chao, Q. Wan, S. Birchfield, N. Ratliff, and D. Fox, “Dexpilot: Vision-based teleopera- tion of dexterous robotic hand-arm system,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 9164–9170

  16. [24]

    Open- vla: An open-source vision-language-action model,

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. , “Open- vla: An open-source vision-language-action model,” arXiv preprint arXiv:2406.09246, 2024

  17. [25]

    Toward next-generation learned robot manipu- lation,

    J. Cui and J. Trinkle, “Toward next-generation learned robot manipu- lation,” Science robotics, vol. 6, no. 54, p. eabd9461, 2021

  18. [26]

    Efficient contact mode enumer- ation in 3d,

    E. Huang, X. Cheng, and M. T. Mason, “Efficient contact mode enumer- ation in 3d,” in Algorithmic F oundations of Robotics XIV: Proceedings of the F ourteenth Workshop on the Algorithmic F oundations of Robotics

  19. [27]

    Springer, 2021, pp. 485–501

  20. [28]

    Combining marker-based mocap and rgb-d camera for acquiring high-fidelity hand motion data,

    W. Zhao, J. Chai, and Y .-Q. Xu, “Combining marker-based mocap and rgb-d camera for acquiring high-fidelity hand motion data,” in Proceedings of the ACM SIGGRAPH/eurographics symposium on computer animation , 2012, pp. 33–42

  21. [29]

    Dexcap: Scalable and portable mocap data collection system for dexterous manipulation,

    C. Wang, H. Shi, W. Wang, R. Zhang, L. Fei-Fei, and C. K. Liu, “Dexcap: Scalable and portable mocap data collection system for dexterous manipulation,” in Robotics: Science and Systems (RSS) , 2024

  22. [30]

    A glove-based system for studying hand-object manipulation via joint pose and force sensing,

    H. Liu, X. Xie, M. Millar, M. Edmonds, F. Gao, Y . Zhu, V . J. Santos, B. Rothrock, and S.-C. Zhu, “A glove-based system for studying hand-object manipulation via joint pose and force sensing,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)....

  23. [31]

    High-fidelity grasping in virtual reality using a glove-based system,

    H. Liu, Z. Zhang, X. Xie, Y . Zhu, Y . Liu, Y . Wang, and S.-C. Zhu, “High-fidelity grasping in virtual reality using a glove-based system,” in IEEE International Conference on Robotics and Automation (ICRA) , 2019, pp. 5180–5186

  24. [32]

    Bunny-visionpro: Real-time bimanual dexterous teleoperation for imitation learning,

    R. Ding, Y . Qin, J. Zhu, C. Jia, S. Yang, R. Yang, X. Qi, and X. Wang, “Bunny-visionpro: Real-time bimanual dexterous teleoperation for imitation learning,” arXiv preprint arXiv:2407.03162 , 2024

  25. [33]

    Anyteleop: A general vision-based dexterous robot arm-hand teleoperation system,

    Y . Qin, W. Yang, B. Huang, K. Van Wyk, H. Su, X. Wang, Y .-W. Chao, and D. Fox, “Anyteleop: A general vision-based dexterous robot arm-hand teleoperation system,” Robotics: Science and Systems (RSS) , 2023

  26. [34]

    Robotic telekinesis: Learning a robotic hand imitator by watching humans on youtube,

    A. Sivakumar, K. Shaw, and D. Pathak, “Robotic telekinesis: Learning a robotic hand imitator by watching humans on youtube,” Robotics: Science and Systems (RSS) , 2022

  27. [35]

    A mobile robot hand-arm teleoperation system by vision and imu,

    S. Li, J. Jiang, P. Ruppel, H. Liang, X. Ma, N. Hendrich, F. Sun, and J. Zhang, “A mobile robot hand-arm teleoperation system by vision and imu,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2020, pp. 10 900–10 906

  28. [36]

    Learning visuotactile skills with two multifingered hands,

    T. Lin, Y . Zhang, Q. Li, H. Qi, B. Yi, S. Levine, and J. Malik, “Learning visuotactile skills with two multifingered hands,” arXiv preprint arXiv:2404.16823, 2024

  29. [37]

    Learn- ing periodic tasks from human demonstrations,

    J. Yang, J. Zhang, C. Settle, A. Rai, R. Antonova, and J. Bohg, “Learn- ing periodic tasks from human demonstrations,” in 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2022, pp. 8658–8665

  30. [38]

    Semi-supervised 3d hand- object poses estimation with interactions in time,

    S. Liu, H. Jiang, J. Xu, S. Liu, and X. Wang, “Semi-supervised 3d hand- object poses estimation with interactions in time,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 687–14 697

  31. [39]

    Gello: A general, low-cost, and intuitive teleoperation framework for robot manipulators,

    P. Wu, Y . Shentu, Z. Yi, X. Lin, and P. Abbeel, “Gello: A general, low-cost, and intuitive teleoperation framework for robot manipulators,” IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024

  32. [40]

    Open-television: Teleoperation with immersive active visual feedback,

    X. Cheng, J. Li, S. Yang, G. Yang, and X. Wang, “Open-television: Teleoperation with immersive active visual feedback,” Conference on Robot Learning (CoRL) , 2024

  33. [41]

    Recent advancements in augmented reality for robotic applications: A survey,

    J. Fu, A. Rota, S. Li, J. Zhao, Q. Liu, E. Iovene, G. Ferrigno, and E. De Momi, “Recent advancements in augmented reality for robotic applications: A survey,” in Actuators, vol. 12, no. 8. MDPI, 2023, p. 323

  34. [42]

    Ar2-d2: Training a robot without a robot,

    J. Duan, Y . R. Wang, M. Shridhar, D. Fox, and R. Krishna, “Ar2-d2: Training a robot without a robot,” Conference on Robot Learning (CoRL), 2023

  35. [43]

    Towards generalizable zero-shot manipulation via translating human interaction plans,

    H. Bharadhwaj, A. Gupta, V . Kumar, and S. Tulsiani, “Towards generalizable zero-shot manipulation via translating human interaction plans,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 6904–6911

  36. [44]

    Xskill: Cross embodiment skill discovery,

    M. Xu, Z. Xu, C. Chi, M. Veloso, and S. Song, “Xskill: Cross embodiment skill discovery,” inConference on Robot Learning. PMLR, 2023, pp. 3536–3555

  37. [45]

    Mimicplay: Long-horizon imitation learning by watching human play,

    C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y . Zhu, and A. Anandkumar, “Mimicplay: Long-horizon imitation learning by watching human play,” Conference on Robot Learning (CoRL) , 2023

  38. [46]

    Immertwin: A mixed reality framework for enhanced robotic arm teleoperation,

    F. P. Audonnet, I. G. Ramirez-Alpizar, and G. Aragon-Camarasa, “Immertwin: A mixed reality framework for enhanced robotic arm teleoperation,” arXiv preprint arXiv:2409.08964 , 2024

  39. [47]

    Robotube: Learning household manipulation from human videos with simulated twin environments,

    H. Xiong, H. Fu, J. Zhang, C. Bao, Q. Zhang, Y . Huang, W. Xu, A. Garg, and C. Lu, “Robotube: Learning household manipulation from human videos with simulated twin environments,” in Conference on Robot Learning . PMLR, 2023, pp. 1–10

  40. [48]

    Ego4d: Around the World in 3,000 Hours of Egocentric Video,

    K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, M. Martin, T. Nagarajan, I. Radosavovic, S. K. Ramakrishnan, F. Ryan, J. Sharma, M. Wray, M. Xu, E. Z. Xu, C. Zhao, S. Bansal, D. Batra, V . Cartillier, S. Crane, T. D...

  41. [49]

    The epic-kitchens dataset: Collection, challenges and baselines,

    D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray, “The epic-kitchens dataset: Collection, challenges and baselines,” 2020. [Online]. Available: https://arxiv.org/abs/2005.00343

  42. [50]

    Arcade: Scalable demonstration collection and generation via augmented reality for imitation learning,

    Y . Yang, B. Ikeda, G. Bertasius, and D. Szafir, “Arcade: Scalable demonstration collection and generation via augmented reality for imitation learning,” IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2024

  43. [51]

    Su, X.-Q

    Y .-P. Su, X.-Q. Chen, C. Zhou, L. H. Pearson, C. G. Pretty, and J. G. Chase, “Integrating virtual, mixed, and augmented reality into remote robotic applications: A brief review of extended reality-enhanced robotic systems for intuitive telemanipulation and telemanufacturing t...

  44. [52]

    Slow Down

    to compute inverse kinematics (IK) to solve for the joint positions given the desired end-effector pose (i.e., the human hand pose), and send the joint position command to the robot arm. In the case of IK solving failure, we use the last commanded joint position. The distance ...

  45. [53]

    Virtual kinesthetic teaching for bimanual telemanipulation,

    I. Jang, H. Niu, E. C. Collins, A. Weightman, J. Carrasco, and B. Lennox, “Virtual kinesthetic teaching for bimanual telemanipulation,” in 2021 IEEE/SICE International Symposium on System Integration (SII). IEEE, 2021, pp. 120–125

  46. [54]

    Puppeteer your robot: Augmented reality leader-follower teleoperation,

    J. van Haastregt, M. C. Welle, Y . Zhang, and D. Kragic, “Puppeteer your robot: Augmented reality leader-follower teleoperation,” 2024 IEEE-RAS 23rd International Conference on Humanoid Robots (Humanoids) , pp. 1019–1026, 2024. [Online]. Available: https://api.semanticscholar....

  47. [55]

    Virtual reality based robot teleoperation via human-scene interaction,

    L. Meng, J. Liu, W. Chai, J. Wang, and M. Q.-H. Meng, “Virtual reality based robot teleoperation via human-scene interaction,” Procedia Computer Science , vol. 226, pp. 141–148, 2023

  48. [56]

    Telesim: A modular and plug-and-play framework for robotic arm teleoperation using a digital twin,

    F. P. Audonnet, J. Grizou, A. Hamilton, and G. Aragon-Camarasa, “Telesim: A modular and plug-and-play framework for robotic arm teleoperation using a digital twin,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 17 770–17 777

  49. [57]

    An augmented reality interface for teleoperating robot manipulators: Reducing demonstrator task load through digital twin control,

    A. Smith and M. Kennedy III, “An augmented reality interface for teleoperating robot manipulators: Reducing demonstrator task load through digital twin control,” arXiv preprint arXiv:2409.18394 , 2024

  50. [58]

    Dexhub and dart: Towards internet scale robot data collection,

    Y . Park, J. S. Bhatia, L. Ankile, and P. Agrawal, “Dexhub and dart: Towards internet scale robot data collection,” 2024. [Online]. Available: https://arxiv.org/abs/2411.02214

  51. [60]

    Arcap: Collecting high-quality human demonstrations for robot learning with augmented reality feedback,

    S. Chen, C. Wang, K. Nguyen, L. Fei-Fei, and C. K. Liu, “Arcap: Collecting high-quality human demonstrations for robot learning with augmented reality feedback,” arXiv preprint arXiv:2410.08464 , 2024

  52. [61]

    Drake: Model-based design and verification for robotics,

    R. Tedrake and the Drake Development Team, “Drake: Model-based design and verification for robotics,” 2019. [Online]. Available: https://drake.mit.edu

  53. [62]

    Egomimic: Scaling imitation learning via egocentric video,

    S. Kareer, D. Patel, R. Punamiya, P. Mathur, S. Cheng, C. Wang, J. Hoffman, and D. Xu, “Egomimic: Scaling imitation learning via egocentric video,” 2024. [Online]. Available: https: //arxiv.org/abs/2410.24221

  54. [2019]

    Available: https://api.semanticscholar.org/CorpusID: 52967399

    [Online]. Available: https://api.semanticscholar.org/CorpusID: 52967399

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.