Pith. sign in

REVIEW 4 major objections 6 minor 75 references

MIRAGE: Multimodal Intention Recognition and Admittance-Guided Enhancement in VR-based Multi-object Teleoperation

T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Two implicit assists - gaze-reading grasp planning and a path-bending virtual force field - fix separate failures in VR robot teleoperation.

desk verdict The VA movement-efficiency finding stands on its own, but the MMIPN success-rate effect is confounded by a simultaneous change in grasp planning policy. read the letter →

arxiv 2509.01996 v1 pith:XTSB72I3 submitted 2025-09-02 cs.RO cs.HC

classification cs.ROcs.HC
keywords human-robotinteractionsharedcontrolvirtualrealityteleoperationmultimodalCNNgaze-basedintentionrecognitionartificialpotentialfieldteleoperatedgrasping
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the two hardest problems in VR teleoperation of a robot arm - knowing which object the operator wants to grasp, and getting the arm there efficiently without force feedback - can each be handled by a different implicit assist inside one shared-control framework. A virtual admittance (VA) model wraps each object in an artificial potential field, so the operator's commanded trajectory is gently bent toward nearby targets, shortening the path the operator must trace. A multimodal CNN (MMIPN) fuses the operator's gaze, the robot's recent motion, the camera view, and object positions to estimate the intended grasp position, and the robot plans its grasp toward that estimate instead of descending vertically from wherever the gripper happens to hover. In a 16-participant study, MMIPN significantly raised grasp success counts and rates, VA significantly cut movement distance and improved movement efficiency, and gaze proved to be the dominant input modality. If the results hold, accurate and less effortful multi-object tele-grasping in VR needs no haptic hardware, only visual guidance plus gaze-driven intent disambiguation.

What carries the argument

Two mechanisms act in sequence. The Virtual Admittance (VA) model is a mass-damper-spring admittance equation whose driving force is the sum of artificial potential field (APF) forces of visible objects; the field bends the operator-commanded trajectory toward nearby objects, turning the trajectory into a visual cue that guides movement without force feedback. The Multimodal-CNN Human Intention Perception Network (MMIPN) fuses four channels - camera image, object positions, a T=3 robot-pose time series, binocular gaze - and regresses the intended grasp position through a fully connected layer under mean-absolute-error loss. A third element is the two-phase split: movement is manual with VA a

What would settle it

Run the study's grasping phase with MMIPN's estimate replaced, in separate conditions, by the true target position and by a deliberately wrong fixed offset from the gripper's vertical projection. If grasp success is statistically indistinguishable between the true and wrong estimates, the gain comes from the robot leaving the vertical-descent path, not from recognizing the operator's intention.

Watch

Extended reading notes

Core claim

Two implicit assistances, each active in a different phase, beat manual teleoperation in multi-object grasping. The virtual admittance model adds artificial potential-field forces from candidate objects to a mass-damper-spring equation, bending the trajectory toward targets as a visual cue; this cut movement distance and raised efficiency. MMIPN, a CNN fusing binocular gaze, a three-step robot-motion sequence, the camera image, and object positions, regresses the intended grasp position; trained on 375 samples, it reached 15.2 mm mean error, and removing gaze raises that to 442.8 mm. In the 2x2 study (N=16), MMIPN raised grasp success while VA shortened paths.

Load-bearing premise

The intention network is trained on 375 recordings from three people and then used unchanged on all sixteen study participants, so the reported grasp-success gain assumes gaze-to-target mapping transfers across users without per-person calibration.

Editorial extensions

If this is right

  • Grasp success improves without per-user calibration: MMIPN ran with fixed parameters across all 16 participants and still raised success counts and rates.
  • VA delivers the efficiency gains of haptic shared control without haptic hardware: path length dropped, movement efficiency rose, movement velocity increased, with no significant change in perceived workload.
  • Gaze is necessary but not sufficient: removing gaze collapses estimation accuracy (15.2 mm to 442.8 mm mean error), but removing any other modality also costs accuracy, so the multimodal fusion itself is load-bearing.
  • Guiding the robot away from a vertical-descent grasp is what rescues parallax-induced failures: in the single-view VR setup operators misjudge depth, and the MMIPN-planned path compensates for that misalignment.
  • Reported presence rose significantly with MMIPN, indicating that intention-aware assistance reduces the cognitive dissonance caused by alignment error, not just the physical task load.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct follow-up would record MMIPN's per-participant grasp-position estimation error during the study and check whether grasp-success gains track estimation accuracy; if they do not, the benefit may come from abandoning vertical descent rather than from intention recognition.
  • The gaze-dominance result suggests a minimal-input design rule - eye tracking plus the robot's own motion history may suffice for intent disambiguation in cluttered VR scenes - but the paper only demonstrates this for multi-cube grasping, so stacked or heterogeneous objects are the natural stress test.
  • Because a few participants felt a loss of control under VA, making the potential-field strength adaptive to the user's movement speed or stated preference could keep the path-shortening benefit without the agency cost; the paper does not explore this.
  • MMIPN reads intention at the moment the grasp command fires, so the same architecture could extend from 'which object' to 'what to do with it', regressing the intended destination or placement pose in pick-and-place tasks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents MIRAGE, a shared-control framework for VR-based multi-object teleoperation that combines a virtual admittance (VA) model with a multimodal CNN-based intention perception network (MMIPN). VA uses artificial potential fields to implicitly guide operator motion during the manual movement phase, while MMIPN estimates the intended grasp position from gaze, robot motion, environmental images, and object positions during a semi-automatic grasping phase. A within-subject user study with 16 participants across four conditions (baseline, VA-only, MMIPN-only, and combined) reports that MMIPN significantly increased grasp success count and success rate, while VA significantly reduced movement distance and improved movement efficiency. An offline ablation study on a dataset collected from three subjects indicates that gaze is the most important input modality. The paper frames the work as addressing depth-perception and target-disambiguation challenges in VR teleoperation.

Significance. If the central claims hold, the paper offers a practical demonstration that implicit visual guidance and multimodal intent estimation can improve different aspects of VR teleoperation: VA for movement efficiency and MMIPN for grasp success. The combination of two assistance methods in distinct task phases is a sensible design, and the objective performance metrics go beyond subjective questionnaires. The ablation study, despite its small scale, provides an explicit comparison of modality contributions and identifies gaze as dominant, consistent with prior gaze-based intent work. However, the current evidence does not yet isolate the mechanism behind the MMIPN success-rate improvement: the manipulation changes both the intention estimator and the grasp-motion policy simultaneously, so the reported main effect may reflect autonomous lateral correction rather than intention recognition. The small, same-user training set also leaves cross-user generalization unestablished. These are load-bearing gaps for the paper's main interpretation.

major comments (4)
  1. [§4.2, Table 1] The MMIPN manipulation is confounded with a change in grasp motion policy. In the non-MMIPN condition, the robot 'will plan its path to descend vertically' from the current end-effector pose; with MMIPN, it plans a path to the estimated grasp position, which may be laterally offset. Thus the significant effects on successful grasp count (F(1,15)=12.31) and success rate (F(1,15)=9.63) compare a vertical-descent policy against an estimated-position-guided policy. The improvement could arise simply from allowing lateral correction before descent, independent of whether MMIPN reads human intention. The paper's own motivation (§1, §7) is that single-view VR causes gripper-target misalignment, so enabling horizontal movement before grasping would improve success even with a non-intentional target selector. A control condition using the same autonomous path planning but with a non-intention bas
  2. [§5.2] The MMIPN model is trained on 375 samples from only three subjects, and the ablation reports point estimates without error bars or cross-subject validation. The user study then applies 'consistent parameters across different conditions and participants, without individual training' (§5.1). If gaze-to-target mapping does not transfer across users, the user-study success improvement could be driven by the autonomous path planning rather than by intention recognition. The paper should report per-subject or leave-one-subject-out accuracy for the trained model, or otherwise demonstrate that the estimator generalizes to the 16 study participants. Without this, the offline MAE of 15.2 mm does not establish that the user-study effect is attributable to MMIPN's intention estimates.
  3. [§6, SUS results] The trial-count description is ambiguous. The text says '10 blocks, with each block containing four trials, resulting in a total of 40 trials per participant.' Because the study has four within-subject conditions, it is unclear whether '40 trials' is per condition (160 total per participant) or per participant across all conditions (10 per condition). This matters for interpreting the reported degrees of freedom and statistical power. Please clarify the per-condition trial count and the total number of grasp attempts/outcomes used in the ANOVA.
  4. [§4.2] The reported effect size for MMIPN on presence is η²=0.86 with F(1,15)=5.94. This is implausibly large for a within-subject comparison with N=16; a partial eta-squared of 0.86 would correspond to an enormous F for this design. Please verify whether this is partial eta-squared, generalized eta-squared, or an error in reporting, and correct the value if needed. This does not affect the main claims but is important for reporting accuracy.
minor comments (6)
  1. [§4.2] Typo: 'teleportation' should be 'teleoperation' in the sentence about joint-angle commands.
  2. [§3.1] The notation 'the vector t represents the previous T steps of the temporal segment from t' is unclear; it should be a time index set or window notation, not a vector. Please rephrase.
  3. [§6] The coordinate transforms T_r^h and T_r^v are introduced but not fully defined in terms of axes; a brief explanation or figure reference would improve reproducibility.
  4. [§4.2] The sentence 'because both factor has only two levels' has a subject-verb agreement error; also, the qualitative results report 'N=9' for multiple groups, but it is unclear whether these overlap across conditions.
  5. [§5.1] MAPE is reported as 880.71% for MMIPN-1. This is numerically possible, but the metric is dominated by small ground-truth values; consider reporting a bounded error metric or median absolute error to avoid misleadingly large percentages.
  6. [§5.1] The parameters km=0.3 and ki=0.1 appear to be chosen a priori; a sentence explaining how these values were selected (or a sensitivity analysis) would strengthen the VA claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: central results are external user-study outcomes; VA parameters are a priori; sole self-citation is non-load-bearing.

full rationale

The paper's central claims are evaluated by a 2x2 within-subject user study with 16 participants who were not part of the MMIPN training set. The VA parameters (km=0.3, ki=0.1) and MMIPN parameters are fixed a priori and not fitted to the user-study outcomes, so the reported improvements in grasp success and movement efficiency are external empirical findings rather than quantities reconstructed from the model inputs. The gaze-importance conclusion comes from an ablation study on the training dataset; this is a model-diagnostic comparison, not a claim that the user-study effect is derived from that ablation, and it does not make the user-study result circular. The only self-citation (Sun et al. [59]) appears in the Discussion as supporting prior work on impedance-based guidance and is not load-bearing for any equation, parameter choice, or reported result. The Limitations section explicitly acknowledges the small dataset and preliminary nature of the evidence. The skeptic's concern that MMIPN also changes the grasp policy from vertical descent to estimated-position planning is a potential confound threatening causal interpretation, but it is not a circularity: no result is equal to its input by construction, and no fitted parameter is renamed as a prediction.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The framework relies on two manually set control gains (km, ki), unspecified impedance matrices, and architectural hyperparameters for the neural network. The MMIPN weights are trained on a small dataset, and the ablation evidence for modality importance lacks variance. No new physical entities are introduced.

free parameters (4)
  • km (motion mapping scale) = 0.3
    Controls the gain from hand motion to robot position in Eq. (1); set once without sensitivity analysis.
  • ki (APF strength) = 0.1
    Attractive force strength per object in Eq. (3); chosen a priori and not varied.
  • M_d, C_d, K_d (admittance impedance parameters) = not reported
    The mass/damping/stiffness matrices in Eq. (2) are not given; their values determine the VA dynamics and the observed path reshaping.
  • MMIPN time window T and image size = T=3, 160x160
    Hyperparameters chosen without ablating their effect; the time window is very short (3 frames).
assumptions (3)
  • domain assumption Operator gaze and robot motion during the grasp command are informative of the intended target
    This is the core premise of MMIPN; validated only on the small 3-subject ablation dataset, not through a dedicated transfer experiment.
  • domain assumption Sum of attractive APFs from all objects provides useful guidance toward the intended object
    The VA model sums forces from all objects without knowing the target, assuming the operator's movement direction will be reinforced; no analysis of cases where multiple objects pull in conflicting directions.
  • standard math The analytical inverse kinematics solution exists and the workspace is reachable
    The system relies on the analytical IK from Diankov [18] and shortest-distance selection; singularities are not discussed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MIRAGE: Multimodal Intention Recognition and Admittance-Guided Enhancement in VR-based Multi-object Teleoperation." pith.science (2026). https://pith.science/paper/XTSB72I3

@misc{pith2026250901996,
  author       = {Pith},
  title        = {Pith review of: MIRAGE: Multimodal Intention Recognition and Admittance-Guided Enhancement in VR-based Multi-object Teleoperation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XTSB72I3}},
  note         = {Machine review of arXiv:2509.01996}
}
read the original abstract

Effective human-robot interaction (HRI) in multi-object teleoperation tasks faces significant challenges due to perceptual ambiguities in virtual reality (VR) environments and the limitations of single-modality intention recognition. This paper proposes a shared control framework that combines a virtual admittance (VA) model with a Multimodal-CNN-based Human Intention Perception Network (MMIPN) to enhance teleoperation performance and user experience. The VA model employs artificial potential fields to guide operators toward target objects by adjusting admittance force and optimizing motion trajectories. MMIPN processes multimodal inputs, including gaze movement, robot motions, and environmental context, to estimate human grasping intentions, helping to overcome depth perception challenges in VR. Our user study evaluated four conditions across two factors, and the results showed that MMIPN significantly improved grasp success rates, while the VA model enhanced movement efficiency by reducing path lengths. Gaze data emerged as the most crucial input modality. These findings demonstrate the effectiveness of combining multimodal cues with implicit guidance in VR-based teleoperation, providing a robust solution for multi-object grasping tasks and enabling more natural interactions across various applications in the future.

Figures

Figures reproduced from arXiv: 2509.01996 by the authors.

Figure 1
Figure 1. A pictorial description of the MIRAGE framework that enhances HRI tele-grasping capability for multiple objects in VR. MIRAGE divides the multi-object grasping task into two phases: movement (manual) and grasping (semi-automatic). Each phase has a specific assistance method designed in MIRAGE: In the movement (manual) phase, Virtual Admittance (VA) modifies the robot trajectory (b), comparing to the non-VA condition… view at source ↗
Figure 3
Figure 3. The overview of the implementation of the [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 2
Figure 2. The brief MMIPN structure: The network received four [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: The virtual-physical configurations: the comparable set [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: The task-related information is given in the prompt bar on [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Box plots of objective metrics (A-H) and subjective questionnaires (I-L) across each condition. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

75 extracted references · 73 canonical work pages

  1. [1]

    Abi-Farraj, C

    F. Abi-Farraj, C. Pacchierotti, O. Arenz, G. Neumann, and P. R. Gior- dano. A haptic shared-control architecture for guided multi-target robotic grasping. IEEE Transactions on Haptics , 13(2):270–285,

  2. [2]

    Abuduweili, S

    A. Abuduweili, S. Li, and C. Liu. Adaptable human intention and trajectory prediction for human-robot collaboration, 2019. 3

  3. [3]

    Adjigble, C

    M. Adjigble, C. de Farias, R. Stolkin, and N. Marturi. Spectgrasp: Robotic grasping by spectral correlation. In 2021 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS), pages 3987–3994, 2021. 2

  4. [4]

    Adjigble, N

    M. Adjigble, N. Marturi, V . Ortenzi, and R. Stolkin. An assisted tele- manipulation approach: combining autonomous grasp planning with haptic cues. In 2019 IEEE/RSJ International Conference on Intelli- gent Robots and Systems (IROS), pages 3164–3171, 2019. 2, 8

  5. [5]

    Alonso and P

    V . Alonso and P. De La Puente. System transparency in shared auton- omy: A mini review. Frontiers in neurorobotics, 12:83, 2018. 9

  6. [6]

    Arevalo Arboleda, F

    S. Arevalo Arboleda, F. R ¨ucker, T. Dierks, and J. Gerken. Assisting manipulation and grasping in robot teleoperation with augmented real- ity visual cues. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, New York, NY , USA, 2021. Association for Computing Machinery. 2

  7. [7]

    R. M. Aronson and E. S. Short. Intentional user adaptation to shared control assistance. In Proceedings of the 2024 ACM/IEEE Interna- tional Conference on Human-Robot Interaction, HRI ’24, page 4–12, New York, NY , USA, 2024. Association for Computing Machinery. 2

  8. [8]

    S. T. Arzo, D. Sikeridis, M. Devetsikiotis, F. Granelli, R. Fierro, M. Esmaeili, and Z. Akhavan. Essential technologies and concepts for massive space exploration: Challenges and opportunities. IEEE Transactions on Aerospace and Electronic Systems, 59(1):3–29, 2023. 1

Show all 75 references
  1. [9]

    Barentine, A

    C. Barentine, A. McNay, R. Pfaffenbichler, A. Smith, E. Rosen, and E. Phillips. A vr teleoperation suite with manipulation assist. In Com- panion of the 2021 ACM/IEEE International Conference on Human- Robot Interaction, HRI ’21 Companion, page 442–446, New York, NY , USA, 202...

  2. [10]

    M. D. Barrera Machuca and W. Stuerzlinger. The effect of stereo dis- play deficiencies on virtual hand pointing. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , CHI ’19, page 1–14, New York, NY , USA, 2019. Association for Computing Machinery. 9

  3. [11]

    T. Brooks. Human/computer control of undersea teleoperators. In NASA. Ames Res. Center The 14th Ann. Conf. on Manual Control ,

  4. [12]

    Burgner, D

    J. Burgner, D. C. Rucker, H. B. Gilbert, P. J. Swaney, P. T. Russell, K. D. Weaver, and R. J. Webster. A telerobotic system for transnasal surgery. IEEE/ASME Transactions on Mechatronics, 19(3):996–1006,

  5. [13]

    C. T. Chang, E. Rosen, T. R. Groechel, M. Walker, and J. Z. Forde. Virtual, augmented, and mixed reality for hri (vam-hri). In 2022 17th ACM/IEEE International Conference on Human-Robot Interac- tion (HRI), pages 1237–1240. IEEE, 2022. 2

  6. [14]

    Charness, U

    G. Charness, U. Gneezy, and M. A. Kuhn. Experimental methods: Between-subject and within-subject design. Journal of economic be- havior & organization, 81(1):1–8, 2012. 6

  7. [15]

    Chiou, G.-T

    M. Chiou, G.-T. Epsimos, G. Nikolaou, P. Pappas, G. Petousakis, S. M ¨uhl, and R. Stolkin. Robot-assisted nuclear disaster response: Report and insights from a field exercise. In 2022 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS), pages 4545–4552...

  8. [16]

    Darvish, L

    K. Darvish, L. Penco, J. Ramos, R. Cisneros, J. Pratt, E. Yoshida, S. Ivaldi, and D. Pucci. Teleoperation of humanoid robots: A survey. IEEE Transactions on Robotics, 39(3):1706–1727, 2023. 2

  9. [17]

    De Souza, S

    R. De Souza, S. El-Khoury, J. Santos-Victor, and A. Billard. Recog- nizing the grasp intention from human demonstration. Robotics and Autonomous Systems, 74:108–121, 2015. 3

  10. [18]

    R. Diankov. Automated construction of robotic manipulation pro- grams. PhD thesis, USA, 2010. AAI3448143. 3

  11. [19]

    Dutta and T

    V . Dutta and T. Zielinska. Predicting the intention of human activities for real-time human-robot interaction (hri). InSocial Robotics: 8th In- ternational Conference, ICSR 2016, Kansas City, MO, USA, Novem- ber 1-3, 2016 Proceedings 8, pages 723–734. Springer, 2016. 3

  12. [20]

    A. L. Edwards. Experimental design in psychological research. Rine- hart, 1950. 7

  13. [21]

    Eltanbouly, M

    S. Eltanbouly, M. Abdel-Ghani, H. Helmy, A. Iskandar, M. Al-Sada, T. Nakajima, and O. Halabi. Assistive telexistence system using mo- tion blending. In 2022 International Conference on Cyberworlds (CW), pages 102–109, 2022. 2

  14. [22]

    Filthaut, D

    L. Filthaut, D. Murray-Rust, M. L. Lupetti, N. Y . Lii, P. Schmaus, and D. Leidner. Cosmic troubleshooting: Exploring third-person view for error handling in telerobotic planetary infrastructure maintenance. In 2024 International Conference on Space Robotics (iSpaRo), pages 33...

  15. [23]

    Fonooni and T

    B. Fonooni and T. Hellstr ¨om. Applying a priming mechanism for intention recognition in shared control. In 2015 IEEE International Multi-Disciplinary Conference on Cognitive Methods in Situation Awareness and Decision, pages 35–41, 2015. 2

  16. [24]

    Franzluebbers and K

    A. Franzluebbers and K. Johnson. Remote robotic arm teleoperation through virtual reality. InSymposium on Spatial User Interaction, SUI ’19. Association for Computing Machinery, 2019. 2

  17. [25]

    J. Fu, G. Maimone, E. Iovene, J. Zhao, A. Redaelli, G. Ferrigno, and E. De Momi. Human-inspired active compliant and passive shared control framework for robotic contact-rich tasks in medical applica- tions. IEEE Transactions on Robotics, 41:2549–2568, 2025. 3

  18. [26]

    Gallagher

    S. Gallagher. Multiple aspects in the sense of agency1. New Ideas in Psychology, 30(1):15–31, 2012. 9

  19. [27]

    Gonz ´alez-D´ıaz, M

    I. Gonz ´alez-D´ıaz, M. Molina-Moreno, J. Benois-Pineau, and A. de Rugy. Asymmetric multi-task learning for interpretable gaze- driven grasping action forecasting. IEEE Journal of Biomedical and Health Informatics, 28(12):7517–7530, 2024. 3

  20. [28]

    S. M. Goza, R. O. Ambrose, M. A. Diftler, and I. M. Spain. Telep- resence control of the nasa/darpa robonaut on a mobility platform. In Proceedings of the SIGCHI Conference on Human Factors in Com- puting Systems, CHI ’04, page 623–629, New York, NY , USA, 2004. Association fo...

  21. [29]

    T. R. Groechel, M. E. Walker, C. T. Chang, E. Rosen, and J. Z. Forde. A tool for organizing key characteristics of virtual, augmented, and mixed reality for human–robot interaction systems: Synthesizing vam-hri trends and takeaways. IEEE Robotics & Automation Maga- zine, 29(1)...

  22. [30]

    Haerling (Adamson) and S

    K. Haerling (Adamson) and S. Prion. Two-by-two factorial design. Clinical Simulation in Nursing, 49:90–91, 2020. 6

  23. [31]

    S. G. Hart. Nasa-task load index (nasa-tlx); 20 years later. Proceed- ings of the Human Factors and Ergonomics Society Annual Meeting , 50(9):904–908, 2006. 7

  24. [32]

    D. Jin, R. Zhang, Y . Li, Y . Ban, and S. Warisawa. Mitigating la- tency effects on subjective experience in robot teleoperation using a vr-enabled virtual spring. In 2024 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pages 1276–1282, 2024. 2

  25. [33]

    D. Kent, C. Saldanha, and S. Chernova. A comparison of remote robot teleoperation interfaces for general object manipulation. In Pro- ceedings of the 2017 ACM/IEEE International Conference on Human- Robot Interaction , HRI ’17, page 371–379, New York, NY , USA,

  26. [34]

    Krupke, L

    D. Krupke, L. Einig, E. Langbehn, J. Zhang, and F. Steinicke. Im- mersive remote grasping: realtime gripper control by a heterogenous robot control system. In Proceedings of the 22nd ACM Conference on Virtual Reality Software and Technology, VRST ’16, page 337–338, New York, N...

  27. [35]

    Laugwitz, T

    B. Laugwitz, T. Held, and M. Schrepp. Construction and evaluation of a user experience questionnaire. In Symposium of the Austrian HCI and usability engineering group, pages 63–76. Springer, 2008. 7

  28. [36]

    LeMasurier, J

    G. LeMasurier, J. Tukpah, M. Wonsick, J. Allspaw, B. Hertel, J. Ep- stein, R. Azadeh, T. Padir, H. A. Yanco, and E. Phillips. Comparing a 2d keyboard and mouse interface to virtual reality for human-in- the-loop robot planning for mobile manipulation. In 2024 33rd IEEE Interna...

  29. [37]

    G. Li, Q. Li, C. Yang, Y . Su, Z. Yuan, and X. Wu. The classification and new trends of shared control strategies in telerobotic systems: A 10 This is the author’s version of the article. To appear in an IEEE ISMAR conference. survey. IEEE Transactions on Haptics, 16(2):118–13...

  30. [38]

    S. Li, X. Zhang, and J. D. Webb. 3-d-gaze-based robotic grasping through mimicking human visuomotor function for people with mo- tion impairments. IEEE Transactions on Biomedical Engineering , 64(12):2824–2835, 2017. 3, 8

  31. [39]

    Y . Li, R. Cui, W. Yan, S. Zhang, and C. Yang. Reconciling con- flicting intents: Bidirectional trust-based variable autonomy for mo- bile robots. IEEE Robotics and Automation Letters, 9(6):5615–5622,

  32. [40]

    C. Lin, C. Zhang, J. Xu, R. Liu, Y . Leng, and C. Fu. Neural correla- tion of eeg and eye movement in natural grasping intention estimation. IEEE Transactions on Neural Systems and Rehabilitation Engineer- ing, 31:4329–4337, 2023. 3

  33. [41]

    Lin and W.-K

    H.-I. Lin and W.-K. Chen. Human intention recognition using markov decision processes. In 2014 CACS International Automatic Control Conference (CACS 2014), pages 340–343, 2014. 3

  34. [42]

    T.-C. Lin, A. Unni Krishnan, and Z. Li. Shared autonomous interface for reducing physical effort in robot teleoperation via human motion mapping. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 9157–9163, 2020. 2

  35. [43]

    R. Luo, M. Zolotas, D. Moore, and T. Padır. User-customizable shared control for robot teleoperation via virtual reality. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 12196–12203, 2024. 2

  36. [44]

    H. Lyu, G. Yang, H. Zhou, X. Huang, H. Yang, and Z. Pang. Teleop- eration of collaborative robot for remote dementia care in home envi- ronments. IEEE Journal of Translational Engineering in Health and Medicine, 8:1–10, 2020. 1

  37. [45]

    Z. Ma, X. Duan, X. Liu, Y . Zhang, and P. Huang. Human-assisted regulation of the deployment of a tethered space robot via feasibil- ity condition optimization and fast logarithmic sliding mode. IEEE Transactions on Aerospace and Electronic Systems, 60(2):2001–2015,

  38. [46]

    Marturi, A

    N. Marturi, A. Rastegarpanah, C. Takahashi, M. Adjigble, R. Stolkin, S. Zurek, M. Kopicki, M. Talha, J. A. Kuo, and Y . Bekiroglu. Towards advanced robotic manipulation for nuclear decommissioning: A pilot study on tele-operation and autonomy. In 2016 International Con- ferenc...

  39. [47]

    Narayanan, B

    V . Narayanan, B. M. Manoghar, R. P. RV , and A. Bera. Ewarenet: Emotion-aware pedestrian intent prediction and adaptive spatial pro- file fusion for social robot navigation. In 2023 IEEE International Conference on Robotics and Automation (ICRA) , pages 7569–7575,

  40. [48]

    Nenna, D

    F. Nenna, D. Zanardi, and L. Gamberini. Enhanced interactivity in vr-based telerobotics: An eye-tracking investigation of human per- formance and workload. International Journal of Human-Computer Studies, 177:103079, 2023. 1, 2, 9

  41. [49]

    Y . Oh, M. Toussaint, and J. Mainprice. Learning to arbitrate human and robot control using disagreement between sub-policies. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5305–5311, 2021. 2

  42. [50]

    Compute box description, 2018

    OnRobot A/S. Compute box description, 2018. 4

  43. [51]

    Pan and J

    Y . Pan and J. Xu. Gaze-based human intention prediction in the hybrid foraging search task. Neurocomputing, 587:127648, 2024. 3, 8

  44. [52]

    Plopski, T

    A. Plopski, T. Hirzle, N. Norouzi, L. Qian, G. Bruder, and T. Langlotz. The eye in extended reality: A survey on gaze interaction and eye tracking in head-worn extended reality. ACM Comput. Surv., 55(3), Mar. 2022. 3

  45. [53]

    Rodehutskors, M

    T. Rodehutskors, M. Schwarz, and S. Behnke. Intuitive bimanual tele- manipulation under communication restrictions by immersive 3d visu- alization and motion tracking. In 2015 IEEE-RAS 15th International Conference on Humanoid Robots (Humanoids), pages 276–283, 2015. 2

  46. [54]

    Schrepp, A

    M. Schrepp, A. Hinderks, and J. Thomaschewski. Design and evalu- ation of a short version of the user experience questionnaire (ueq-s). International Journal of Interactive Multimedia and Artificial Intelli- gence, 4:103, 01 2017. 7

  47. [55]

    Schydlo, M

    P. Schydlo, M. Rakovic, L. Jamone, and J. Santos-Victor. Anticipation in human-robot cooperation: A recurrent neural network approach for multiple action sequences prediction. In 2018 IEEE International Conference on Robotics and Automation (ICRA) , pages 5909–5914,

  48. [56]

    Slater, A

    M. Slater, A. Steed, J. McCarthy, and F. Maringelli. The influence of body movement on subjective presence in virtual environments. Hu- man Factors, 40(3):469–477, 1998. PMID: 9849105. 7

  49. [57]

    P. Song, P. Li, E. Aertbeli ¨en, and R. Detry. Robot trajectron: Tra- jectory prediction-based shared control for robot manipulation. arXiv preprint arXiv:2402.02499, 2024. 3

  50. [58]

    Stoyanov, R

    T. Stoyanov, R. Krug, A. Kiselev, D. Sun, and A. Loutfi. Assisted telemanipulation: A stack-of-tasks approach to remote manipulator control. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1–9, 2018. 2

  51. [59]

    C. Sun, Z. Ma, L. Teng, M. Zhang, C.-Y . Tang, H. Zhang, and Z. Sun. Development of an impedance-based human-robot hybrid interaction system with rnn force estimator on visual tasks. In 2024 IEEE 19th Conference on Industrial Electronics and Applications (ICIEA), pages 1–6, 2024. 3, 8

  52. [60]

    D. Sun, A. Kiselev, Q. Liao, T. Stoyanov, and A. Loutfi. A new mixed- reality-based teleoperation system for telepresence and maneuverabil- ity enhancement. IEEE Transactions on Human-Machine Systems , 50(1):55–67, 2020. 2

  53. [61]

    M. Usoh, E. Catena, S. Arman, and M. Slater. Using presence ques- tionnaires in reality. Presence, 9(5):497–503, 2000. 8

  54. [62]

    C. R. Wagner and R. D. Howe. Force feedback benefit depends on experience in multiple degree of freedom robotic surgery task. IEEE Transactions on Robotics, 23(6):1235–1240, 2007. 1

  55. [63]

    Wagner, A

    U. Wagner, A. Asferg Jacobsen, T. Feuchtner, H. Gellersen, and K. Pfeuffer. Eye-hand movement of objects in near space extended reality. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, UIST ’24, New York, NY , USA,

  56. [64]

    Walker, T

    M. Walker, T. Phung, T. Chakraborti, T. Williams, and D. Szafir. Vir- tual, augmented, and mixed reality for human-robot interaction: A survey and virtual design element taxonomy. J. Hum.-Robot Interact., 12(4), July 2023. 2, 9

  57. [65]

    K. Wan, C. Li, F.-S. Lo, and P. Zheng. A virtual reality-based immer- sive teleoperation system for remote human-robot collaborative man- ufacturing. Manufacturing Letters, 41:43–50, 2024. 2

  58. [66]

    X. Wang, S. Guo, Z. Xu, Z. Zhang, Z. Sun, and Y . Xu. A robotic tele- operation system enhanced by augmented reality for natural human– robot interaction. Cyborg and Bionic Systems, 5:0098, 2024. 2

  59. [67]

    Whitney, E

    D. Whitney, E. Rosen, E. Phillips, G. Konidaris, and S. Tellex. Com- paring robot grasping teleoperation across desktop and virtual reality with ros reality. In Robotics Research: The 18th International Sympo- sium ISRR, pages 335–350. Springer, 2019. 1, 2

  60. [68]

    J. O. Wobbrock, L. Findlater, D. Gergle, and J. J. Higgins. The aligned rank transform for nonparametric factorial analyses using only anova procedures. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems , CHI ’11, page 143–146, New York, NY , USA, 2...

  61. [69]

    Wonsick and T

    M. Wonsick and T. Padir. A systematic review of virtual reality in- terfaces for controlling and interacting with robots. Applied Sciences, 10(24):9051, 2020. 2

  62. [70]

    S. Xu, S. Moore, and A. Cosgun. Shared-control robotic manipulation in virtual reality. In2022 International Congress on Human-Computer Interaction, Optimization and Robotic Applications (HORA), pages 1– 6, 2022. 2

  63. [71]

    B. Yang, X. Chen, X. Xiao, P. Yan, Y . Hasegawa, and J. Huang. Gaze and environmental context-guided deep neural network and sequential decision fusion for grasp intention recognition. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 31:3687–3698, 2023. 3

  64. [72]

    B. Yang, J. Huang, X. Chen, X. Li, and Y . Hasegawa. Natural grasp intention recognition based on gaze in human–robot interaction. IEEE Journal of Biomedical and Health Informatics , 27(4):2059– 2070, 2023. 3

  65. [73]

    Y . Zhu, B. Jiang, Q. Chen, T. Aoyama, and Y . Hasegawa. A shared control framework for enhanced grasping performance in teleopera- tion. IEEE Access, 11:69204–69215, 2023. 2 11

  66. [2017]

    Association for Computing Machinery. 2

  67. [2024]

    Association for Computing Machinery. 8

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.