REVIEW 2 major objections 5 minor 75 references
Towards Deeper Understanding of Natural User Interactions in Virtual Reality Based Assembly Tasks
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read When an assembly workspace offers no nameable landmarks, untrained users fall back on 'put that there' phrasing, longer pointing, and richer multimodal combinations, and the paper argues natural user interfaces should be built to expect…
desk verdict A useful new multimodal VR elicitation dataset with large behavioral differences between two tasks, but the central causal claim about spatial anchors is underdetermined by a confounded two-condition design. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing apparatus is a single-factor within-subjects elicitation design whose two conditions are meant to differ in one property: whether the workspace supplies concrete, nameable spatial anchors. The analytical core is an annotation scheme built on the put-that-there command paradigm, defined as a voice-plus-gesture pattern in which an utterance names an object and a gesture or gaze supplies the location. Each utterance is classified as explicit (descriptive) or implicit (spatially vague, relying on gestures or gaze), and a multimodal sequence miner assembles each voice command with any stationary gaze or pointing gesture occurring within five seconds. This machinery lets the authors compare not only utterance content but also the duration of points, the gaze direction, the number of modalities per instruction, and how these change over task completion phases.
What would settle it
Run the same circuit-board assembly with visually anchored locations, such as outlined slots or labelled positions, while keeping the board layout and part count identical; if users still produce mostly implicit utterances and long pointing times, the claim that spatial ambiguity drives the behavior is wrong. Alternatively, within the released dataset, compare average pointing duration on implicit versus explicit utterances: if pointing is not longer for implicit commands, the implied linkage between vague language and prolonged deictic compensation fails.
Extended reading notes
Core claim
The paper's central discovery is that untrained users spontaneously follow a put-that-there instruction pattern, an utterance that names an object plus a deictic indication of where it goes, but the linguistic and gestural balance of that pattern shifts with the availability of spatial anchors. In the brick task, where pieces provide nameable reference points, participants used significantly more explicit descriptive commands, pointed for shorter periods, looked more toward the reference model, and leaned on unimodal speech. In the circuit-board task, where the blank board offers no obvious landmarks, participants used significantly more implicit commands such as 'put a resistor here' and compensated with significantly longer pointing, more gaze use, and more trimodal combinations; instruction frequency also predicted completion time far more strongly in this task. These differences were not static: command style evolved across task phases, with brick-task users gradually adding implicit commands while circuit-board users stayed implicit throughout.
Load-bearing premise
The study assumes the two assembly tasks differ only in the factor being tested, whether the scene provides concrete spatial anchors, so that all observed differences in user behavior can be attributed to that factor, even though the tasks also differ in piece count, part types, board geometry, and interaction demands.
Editorial extensions
If this is right
- Natural multimodal interfaces for assembly should accept put-that-there style instructions that pair vague voice with gesture and gaze instead of requiring users to follow preset command grammars.
- In workspaces with concrete anchors, designers can expect explicit descriptive speech to suffice and can keep gesture interpretation lightweight, while in anchor-free workspaces the interface must tolerate long pointing episodes and combine eye and head gaze with speech.
- Task phase matters: instruction style changes as an anchored assembly progresses, so session-level assumptions about modality use should give way to phase-adaptive interpretation.
- The released annotated dataset, containing voice transcripts, hand tracking, eye gaze, head pose, and command labels, gives other researchers a basis for training models of natural multimodal assembly instruction.
Reading between the lines
- A direct test that varies only anchor availability, keeping the same board layout and part count while adding or removing visual placeholders, would separate the spatial-anchor explanation from other task differences such as piece count and board geometry; the current study conflates those factors.
- Because instruction frequency predicted completion time sharply in the ambiguous task, an interface that auto-completes repeated component types could cut instruction overhead more in layout-heavy tasks than in anchored ones.
- Modality mix could serve as a real-time uncertainty signal: rising implicit utterances combined with longer pointing would flag regions the user finds hard to describe, letting an adaptive system highlight candidate placements.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a Wizard-of-Oz elicitation study (N=34) in which participants instruct a virtual robot arm through a collaborative LEGO assembly task and an instructive PCB assembly task in VR. Voice, hand tracking, eye gaze, and head pose are captured and analyzed. The main empirical claims are that explicit/descriptive utterances dominate the LEGO task while implicit "put-that-there" language dominates the PCB task, that pointing durations are longer in the PCB task, and that users produce more trimodal commands in the PCB task and more unimodal commands in the LEGO task. The paper also reports correlations between utterance frequency and task completion time and contributes a publicly available annotated dataset.
Significance. If the descriptive findings are reliable, the dataset is a useful resource for future natural multimodal interface design, and the large effects on utterance ratio (t=8.14), pointing time, and trimodal usage are noteworthy. The study is unusual in capturing untrained users' raw multimodal input with a Wizard-of-Oz robot, and the public dataset is a concrete contribution. However, the paper's headline causal interpretation—that spatial-anchor availability causes the observed differences—is not supported by the operationalization, and the central utterance-ratio measure lacks demonstrated coding reliability. The work would be publishable as a descriptive, dataset-focused study once these issues are addressed.
major comments (2)
- [§3.1, §6.2, abstract] The study is framed as a single-factor comparison of "task nature," but the two conditions differ on many dimensions at once: the participant manually manipulates objects only in Task 1; the object sets, geometries, and board layouts differ; the piece counts are 25 versus 20 (§3.6); and the observed completion times differ as well (M=382.84 s versus 451.40 s, Table 3). The abstract and §6.2/§6.3 nevertheless attribute the differences to the availability of concrete spatial anchors ("due to concrete spatial anchors... due to a lack of spatial anchors"). Because no condition manipulates anchor availability while holding the other task properties fixed, the causal attribution in the abstract is underdetermined. The authors should either soften the claims to descriptive differences between two assembly scenarios or add a follow-up condition that isolates anchor availability; the current Limitations section does not acknowledge this confound.
- [§4.1, §5.1.1] The headline effect—explicit versus implicit utterance ratios differing across tasks with t33 = 8.14—rests entirely on manual annotation of transcripts into nine labels, but the paper provides no information about the number of annotators, the annotation protocol, or inter-rater agreement. Without coding reliability evidence, the large ratio difference could reflect a single annotator's interpretation of the Bolt-based scheme rather than a stable behavioral difference. The authors should report inter-rater reliability (e.g., Cohen's kappa or equivalent) or provide a robustness check such as re-annotation of a subset.
minor comments (5)
- [Tables 1–4 and §5.4.3] Statistical reporting needs alignment: §5.4.3 reports F3,30=5.296, R=0.530, R²=0.281 for the Task 1 multimodal correlation, whereas Table 2 reports F3,30=5.438, R=0.535, R²=0.287; please use consistent values and correct notation, since Wilcoxon results are reported with t statistics in places (e.g., §5.2.1 reports "t33=37.0").
- [§3.3] The paper states that participants "underwent training" for task objectives and manipulation mechanics, which appears to conflict with the repeated claim that interactions come from untrained users; please clarify what was trained and why this does not undermine the "untrained" framing.
- [§4.2] The pointing-detection thresholds (z-rotation differences of 10° and 30°, minimum duration 0.5 s) are stated without justification; please cite prior work or report a sensitivity analysis showing the findings are robust to these choices.
- [Figure 10] The task-sequence visualization is dense and the legend ("G – Gaze, P – Pointing Gesture, U – Utterance") does not explain how sequences are encoded; the figure would benefit from a concrete example sequence annotation.
- [§6.4] Minor typographical issues (e.g., "V oice ismost effective" in §6.4.1 and "cuessuch" in §6.4.2) should be corrected.
Circularity Check
No significant circularity: the paper is an observational elicitation study whose central results are descriptive statistical comparisons grounded in collected data, not derived from fitted parameters or self-citation chains.
full rationale
This paper reports a Wizard-of-Oz elicitation study in which participants completed two VR assembly tasks while voice, hand tracking, and gaze were recorded. The main findings are empirical: explicit utterances were more frequent in the LEGO task, implicit utterances and trimodal commands were more frequent in the PCB task, and average pointing time was longer in the PCB task. These results come from direct statistical comparisons (paired t-tests, Wilcoxon signed-rank tests, correlations) of the collected data, not from a derivation that assumes its own conclusion. The explicit/implicit annotation scheme follows Bolt's external 'put-that-there' work and labels utterances by their linguistic content, not by the task condition, so the measured difference is an empirical observation rather than a definitional artifact. The authors cite their own prior work (e.g., [24], [25]) only for background in related work; those citations are not load-bearing for the study's conclusions. The paper does not fit parameters to a subset and then 'predict' a related quantity, nor does it invoke a self-citation chain to force its interpretation. The confound between task nature and spatial anchor availability is a validity and generalizability concern, appropriately categorized as correctness risk rather than circularity, because the paper's causal language is an interpretive claim about a two-condition design, not a mathematical reduction. Under the hard rules requiring a quotable reduction of Eq. X to Eq. Y or a fitted parameter renamed as prediction, no circular step can be exhibited. The appropriate finding is therefore no significant circularity.
Assumptions & free parameters
free parameters (3)
- Pointing pose detection thresholds
- Sequence context window =
5 seconds
- Task phase boundaries =
0-20%, 20-80%, 80-100%
assumptions (4)
- domain assumption Whisper-Large V2 transcription is accurate enough for command annotation
- domain assumption Manual annotation labels are reliable despite a single annotator and no inter-rater agreement
- domain assumption Wizard-of-Oz behavior approximates an autonomous natural interface
- domain assumption Higher-level interaction behaviors transfer from VR to real-world assembly
Cite this review
Pith. "Pith review of Towards Deeper Understanding of Natural User Interactions in Virtual Reality Based Assembly Tasks." pith.science (2026). https://pith.science/paper/HL4CFZSL
@misc{pith2026250817124,
author = {Pith},
title = {Pith review of: Towards Deeper Understanding of Natural User Interactions in Virtual Reality Based Assembly Tasks},
year = {2026},
howpublished = {\url{https://pith.science/paper/HL4CFZSL}},
note = {Machine review of arXiv:2508.17124}
}
read the original abstract
We explore natural user interactions using a virtual reality simulation of a robot arm for assembly tasks. Using a Wizard-of-Oz study, participants completed collaborative LEGO and instructive PCB assembly tasks, with the robot responding under experimenter control. We collected voice, hand tracking, and gaze data from users. Statistical analyses revealed that instructive and collaborative scenarios elicit distinct behaviors and adopted strategies, particularly as tasks progress. Users tended to use put-that-there language in spatially ambiguous contexts and more descriptive instructions in spatially clear ones. Our contributions include the identification of natural interaction strategies through analyses of collected data, as well as the supporting dataset, to guide the understanding and design of natural multimodal user interfaces for instructive interaction with systems in virtual reality.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
A. J. Aubrey, D. Marshall, P. L. Rosin, J. Vendeventer, D. W. Cun- ningham, and C. Wallraven. Cardiff conversation database (ccdb): A database of natural dyadic conversations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 277–282, 2013. 2
work page 2013
- [3]
-
[4]
Multimodal Dataset of Human-Robot Hugging Interaction
K. Bagewadi, J. Campbell, and H. B. Amor. Multimodal dataset of human-robot hugging interaction. arXiv preprint arXiv:1909.07471,
work page Pith review arXiv 1909
-
[5]
A. Ben-Youssef, C. Clavel, S. Essid, M. Bilac, M. Chamoux, and A. Lim. Ue-hri: a new dataset for the study of user engagement in spontaneous human-robot interactions. In Proceedings of the 19th ACM international conference on multimodal interaction , pp. 464– 472, 2017. 2
work page 2017
-
[6]
S. Bilakhia, S. Petridis, A. Nijholt, and M. Pantic. The mahnob mimicry database: A database of naturalistic human interactions. Pat- tern recognition letters, 66:52–61, 2015. 2
work page 2015
-
[7]
R. A. Bolt. “put-that-there” voice and gesture at the graphics interface. In Proceedings of the 7th annual conference on Computer graphics and interactive techniques, pp. 262–270, 1980. 1, 3, 4, 8
work page 1980
-
[8]
S. Borghi, F. Zucchi, E. Prati, A. Ruo, V . Villani, L. Sabattini, and M. Peruzzini. Unlocking human-robot dynamics: Introducing sensec- obot, a novel multimodal dataset on industry 4.0. In Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot In- teraction, HRI ’24, p. 880–884. Association for Computing Machin- ery, New York, NY , USA,...
Show all 75 references
-
[9]
Camburn, V
B. Camburn, V . Viswanathan, J. Linsey, D. Anderson, D. Jensen, R. Crawford, K. Otto, and K. Wood. Design prototyping methods: state of the art in strategies, techniques, and guidelines. Design Sci- ence, 3:e13, 2017. 9
2017
-
[10]
Can ´evet, W
O. Can ´evet, W. He, P. Motlicek, and J.-M. Odobez. The mummer data set for robot perception in multi-party hri scenarios. In 2020 29th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), pp. 1294–1300, 2020. doi: 10.1109/RO -MAN47096.2020.9223340 2
2020
-
[11]
Carlson, A
P. Carlson, A. Peters, S. B. Gilbert, J. M. Vance, and A. Luse. Virtual training: Learning transfer of assembly tasks. IEEE transactions on visualization and computer graphics, 21(6):770–782, 2015. 3
2015
-
[12]
Celiktutan, E
O. Celiktutan, E. Skordos, and H. Gunes. Multimodal human-human- robot interactions (mhhri) dataset for studying personality and engage- ment. IEEE Transactions on Affective Computing , 10(4):484–497,
-
[13]
Chang, X
Y . Chang, X. Wang, J. Wang, Y . Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wang, et al. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology , 15(3):1–45, 2024. 10
2024
-
[14]
Chu and Y .-L
C.-H. Chu and Y .-L. Liu. Augmented reality user interface design and experimental evaluation for human-robot collaborative assembly. Journal of Manufacturing Systems, 68:313–324, 2023. 1
2023
-
[15]
Cohen, C
P. Cohen, C. Swindells, S. Oviatt, and A. Arthur. A high-performance dual-wizard infrastructure for designing speech, pen, and multimodal interfaces. In Proceedings of the 10th international conference on Multimodal interfaces, pp. 137–140, 2008. 2
2008
-
[16]
P. R. Cohen, M. Johnston, D. McGee, S. Oviatt, J. Pittman, I. Smith, L. Chen, and J. Clow. Quickset: Multimodal interaction for distributed applications. In Proceedings of the fifth ACM international conference on Multimedia, pp. 31–40, 1997. 2
1997
-
[17]
L. M. Daling and S. J. Schlittmeier. Effects of augmented reality-, virtual reality-, and mixed reality–based training on objective perfor- mance measures and subjective evaluations in manual assembly tasks: a scoping review. Human factors, 66(2):589–626, 2024. 3
2024
-
[18]
Devillers, S
L. Devillers, S. Rosset, G. D. Duplessis, M. A. Sehili, L. B ´echade, A. Delaborde, C. Gossart, V . Letard, F. Yang, Y . Yemez, et al. Mul- timodal data collection of human-robot humorous interactions in the joker project. In 2015 international conference on affective computin...
2015
-
[19]
Dong and H
G. Dong and H. Liu. Feature engineering for machine learning and data analytics. CRC press, 2018. 3
2018
-
[20]
J. Epps, S. Oviatt, and F. Chen. Integration of speech and gesture inputs during multimodal interaction. In Proc Aust. Int. Conf. on CHI,
-
[21]
Fechter, B
M. Fechter, B. Schleich, and S. Wartzack. Comparative evaluation of wimp and immersive natural finger interaction: A user study on cad assembly modeling. Virtual Reality, 26(1):143–158, 2022. 1
2022
-
[22]
Fiorentino, R
M. Fiorentino, R. Radkowski, C. Stritzke, A. E. Uva, and G. Monno. Design review of cad assemblies using bimanual natural interface. International Journal on Interactive Design and Manufacturing (IJI- DeM), 7:249–260, 2013. 1
2013
-
[23]
Geiger, E
A. Geiger, E. Brandenburg, and R. Stark. Natural virtual reality user interface to define assembly sequences for digital human models. Ap- plied System Innovation, 3(1):15, 2020. 1
2020
-
[24]
R. K. Ghamandi, Y . Hmaiti, T. T. Nguyen, A. Ghasemaghaei, R. K. Kattoju, E. M. Taranta, and J. J. LaViola. What and how together: a taxonomy on 30 years of collaborative human-centered xr tasks. In 2023 IEEE International Symposium on Mixed and Augmented Real- ity (ISMAR), pp...
2023
-
[25]
R. K. Ghamandi, R. K. Kattoju, Y . Hmaiti, M. Maslych, E. M. Taranta, R. P. McMahan, and J. LaViola. Unlocking understanding: An inves- tigation of multimodal communication in virtual reality collaboration. In Proceedings of the 2024 CHI Conference on Human Factors in Computin...
2024
-
[26]
Gottsacker, Y
M. Gottsacker, Y . Hmaiti, M. Maslych, G. Bruder, J. J. LaViola Jr, and G. F. Welch. Xr-first design for productivity: A conceptual 10 © 2025 IEEE. This is the author’s version of the article that has been published in the proceedings of IEEE Visualization conference. The fina...
2025 arXiv
-
[27]
Hmaiti, M
Y . Hmaiti, M. Maslych, A. Ghasemaghaei, R. K. Ghamandi, and J. J. LaViola Jr. Visual perceptual confidence: Exploring discrepancies between self-reported and actual distance perception in virtual reality. IEEE Transactions on Visualization and Computer Graphics, 2024. 2
2024
-
[28]
Hmaiti, M
Y . Hmaiti, M. Maslych, E. M. Taranta, and J. J. LaViola. An explo- ration of the effects of head-centric rest frames on egocentric distance judgments in vr. In 2023 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 263–272. IEEE, 2023. 2
2023
-
[29]
S. Howard. User interface design and hci: identifying the training needs of practitioners. SIGCHI Bull., 27(3):17–22, jul 1995. doi: 10. 1145/221296.221302 1
1995
-
[30]
Husainy, A
A. Husainy, A. Joshi, V . Chougule, R. Thomake, H. Kamat, and H. Jadhav. The impact of virtual reality integration in robotics: En- hancing efficiency, safety, and human-robot interaction.Asian Review of Mechanical Engineering, 12:28–34, 12 2023. doi: 10.70112/arme -2023.12.2.4228 3
2023 doi
-
[31]
Iftikhar, M
M. Iftikhar, M. Saqib, M. Zareen, and H. Mumtaz. Artificial in- telligence: revolutionizing robotic surgery. Annals of Medicine and Surgery, 86(9):5401–5409, 2024. 1, 9
2024
-
[32]
P. G. Ikonomov and E. D. Milkova. Virtual assembly/disassembly system using natural human interaction and control. Virtual and aug- mented reality applications in manufacturing, pp. 111–125, 2004. 1
2004
-
[33]
Inamura and Y
T. Inamura and Y . Mizuchi. Sigverse: A cloud-based vr platform for research on multimodal human-robot interaction. Frontiers in Robotics and AI, 8:549360, 2021. 2
2021
-
[34]
D. B. Jayagopi, S. Sheiki, D. Klotz, J. Wienke, J.-M. Odobez, S. Wrede, V . Khalidov, L. Nyugen, B. Wrede, and D. Gatica-Perez. The vernissage corpus: A conversational human-robot-interaction dataset. In 2013 8th ACM/IEEE International Conference on Human- Robot Interaction (H...
2013
-
[35]
Jiang, A
Y . Jiang, A. Gupta, Z. Zhang, G. Wang, Y . Dou, Y . Chen, L. Fei-Fei, A. Anandkumar, Y . Zhu, and L. Fan. Vima: General robot manip- ulation with multimodal prompts. arXiv preprint arXiv:2210.03094, 2(3):6, 2022. 2
2022 arXiv
-
[36]
Johnston, P
M. Johnston, P. R. Cohen, D. McGee, S. Oviatt, J. A. Pittman, and I. Smith. Unification-based multimodal integration. In 35th Annual Meeting of the Association for Computational Linguistics and 8th Conference of the European Chapter of the Association for Compu- tational Lingu...
1997
-
[37]
J. S. Joyner, M. Vaughn-Cooke, and H. L. Benz. Comparison of dex- terous task performance in virtual reality and real-world environments. Frontiers in Virtual Reality, 2:599274, 2021. 1
2021
-
[38]
Kesim, T
E. Kesim, T. Numanoglu, O. Bayramoglu, B. B. Turker, N. Hussain, M. Sezgin, Y . Yemez, and E. Erzin. The ehri database: a multimodal database of engagement in human–robot interactions. Language Re- sources and Evaluation, 57(3):985–1009, 2023. 2
2023
-
[39]
J. J. LaViola Jr, S. Buchanan, and C. Pittman. Multimodal input for perceptual user interfaces. Interactive Displays: Natural Human- Interface Technologies, pp. 285–312, 2014. 1
2014
-
[40]
J. J. LaViola Jr, E. Kruijff, R. P. McMahan, D. Bowman, and I. P. Poupyrev. 3D user interfaces: theory and practice . Addison-Wesley Professional, 2017. 3
2017
-
[41]
grip-that- there
K. Mahadevan, M. Sousa, A. Tang, and T. Grossman. “grip-that- there”: An investigation of explicit and implicit task allocation tech- niques for human-robot collaboration. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pp. 1–14, 2021. 2
2021
-
[42]
A. D. Marshall, P. L. Rosin, J. Vandeventer, and A. Aubrey. 4d cardiff conversation database (4d ccdb): A 4d database of natural, dyadic conversations. Auditory-Visual Speech Processing,{AVSP} 2015, pp. 157–162, 2015. 2
2015
-
[43]
Martin, S
D. Martin, S. Malpica, D. Gutierrez, B. Masia, and A. Serrano. Multimodality in vr: A survey. ACM Computing Surveys (CSUR) , 54(10s):1–36, 2022. 2
2022
-
[44]
Medell ´ın-Castillo, J
H. Medell ´ın-Castillo, J. Corney, J. Ritchie, R. Sung, and T. Lim. Virtual assembly rapid prototyping of near net shapes. Ingenier´ıa Mec´anica. Tecnolog´ıa y Desarrollo, 3:66–76, 03 2009. doi: 10.1115/ WINVR2009-723 1
2009
-
[45]
B. A. Newman, R. M. Aronson, S. S. Srinivasa, K. Kitani, and H. Ad- moni. Harmonic: A multimodal dataset of assistive human–robot col- laboration. The International Journal of Robotics Research, 41(1):3– 11, 2022. 2
2022
-
[46]
S. Oviatt. Mulitmodal interactive maps: Designing for human perfor- mance. Human–Computer Interaction, 12(1-2):93–129, 1997. 3
1997
-
[47]
S. Oviatt. Ten myths of multimodal interaction. Commun. ACM , 42(11):74–81, nov 1999. doi: 10.1145/319382.319398 2, 7, 8
1999
-
[48]
S. Oviatt. Advances in robust multimodal interface design. IEEE computer graphics and applications, 23(05):62–68, 2003. 2
2003
-
[49]
Oviatt and P
S. Oviatt and P. Cohen. Perceptual user interfaces: multimodal inter- faces that process what comes naturally.Communications of the ACM, 43(3):45–53, 2000. 2
2000
-
[50]
Oviatt, P
S. Oviatt, P. Cohen, L. Wu, L. Duncan, B. Suhm, J. Bers, T. Holzman, T. Winograd, J. Landay, J. Larson, et al. Designing the user inter- face for multimodal speech and pen-based gesture applications: State- of-the-art systems and future research directions. Human-computer inte...
2000
-
[51]
Oviatt, R
S. Oviatt, R. Coulston, and R. Lunsford. When do we interact multi- modally? cognitive load and multimodal communication patterns. In Proceedings of the 6th International Conference on Multimodal Inter- faces, ICMI ’04, p. 129–136. Association for Computing Machinery, New York...
2004
-
[52]
Oviatt, A
S. Oviatt, A. DeAngeli, and K. Kuhn. Integration and synchronization of input modes during multimodal human-computer interaction. In Proceedings of the ACM SIGCHI Conference on Human factors in computing systems, pp. 415–422, 1997. 7, 8
1997
-
[53]
Petersen and D
N. Petersen and D. Stricker. Continuous natural user interface: Re- ducing the gap between real and digital world. In ISMAR, vol. 1, pp. 23–26, 2009. 1
2009
-
[54]
Radford, J
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever. Robust speech recognition via large-scale weak supervi- sion. In International Conference on Machine Learning , pp. 28492– 28518. PMLR, 2023. 4
2023
-
[55]
Reddy, P
K. Reddy, P. Gharde, H. Tayade, M. Patil, L. S. Reddy, D. Surya, and L. srivani Reddy. Advancements in robotic surgery: a comprehen- sive overview of current utilizations and upcoming frontiers. Cureus, 15(12), 2023. 1, 9
2023
-
[56]
Ringeval, A
F. Ringeval, A. Sonderegger, J. Sauer, and D. Lalanne. Introducing the recola multimodal corpus of remote collaborative and affective inter- actions. In 2013 10th IEEE international conference and workshops on automatic face and gesture recognition (FG), pp. 1–8. IEEE, 2013. 2
2013
-
[57]
F. D. Rose, E. A. Attree, B. M. Brooks, D. M. Parslow, and P. R. Penn. Training in virtual environments: transfer to real world tasks and equivalence to real task training. Ergonomics, 43(4):494–511,
-
[58]
Rozgic, B
V . Rozgic, B. Xiao, A. Katsamanis, B. R. Baucom, P. G. Georgiou, and S. S. Narayanan. A new multichannel multi modal dyadic interaction database. In INTERSPEECH, pp. 1982–1985, 2010. 2
1982
-
[59]
Sahoo and C.-Y
S. Sahoo and C.-Y . Lo. Smart manufacturing powered by recent tech- nological advancements: A review. Journal of Manufacturing Sys- tems, 64:236–250, 2022. 9
2022
-
[60]
Sanchez-Cortes, O
D. Sanchez-Cortes, O. Aran, and D. Gatica-Perez. An audio visual corpus for emergent leader analysis. In Workshop on multimodal cor- pora for machine learning: taking stock and road mapping the future, ICMI-MLMI. Citeseer, 2011. 2
2011
-
[61]
Selfridge and W
M. Selfridge and W. Vannoy. A natural language interface to a robot assembly system. IEEE Journal on Robotics and Automation , 2(3):167–171, 1986. 1
1986
-
[62]
B. S. Shafique, A. Vayani, M. Maaz, H. A. Rasheed, D. Dissanayake, M. I. Kurpath, Y . Hmaiti, G. Inoue, J. Lahoud, M. S. Rashid, et al. A culturally-diverse multilingual multimodal video benchmark & model. arXiv preprint arXiv:2506.07032, 2025. 10
2025
-
[63]
Shrestha, Y
S. Shrestha, Y . Zha, S. Banagiri, G. Gao, Y . Aloimonos, and C. Fer- muller. Natsgd: A dataset with speech, gestures, and demonstrations for robot learning in natural human-robot interaction. arXiv preprint arXiv:2403.02274, 2024. 2, 3
2024 arXiv
-
[64]
Siltanen, M
S. Siltanen, M. Hakkarainen, O. Korkalo, T. Salonen, J. Saaski, C. Woodward, T. Kannetis, M. Perakakis, and A. Potamianos. Mul- 11 © 2025 IEEE. This is the author’s version of the article that has been published in the proceedings of IEEE Visualization conference. The final ve...
2025
-
[65]
H. Su, W. Qi, J. Chen, C. Yang, J. Sandoval, and M. A. Laribi. Recent advancements in multimodal human–robot interaction. Frontiers in Neurorobotics, 17:1084000, 2023. 2
2023
-
[66]
N. T. V . Tuyen, A. L. Georgescu, I. Di Giulio, and O. Celiktutan. A multimodal dataset for robot learning to imitate social human-human interaction. In Companion of the 2023 ACM/IEEE International Con- ference on Human-Robot Interaction, HRI ’23, p. 238–242. Associa- tion for...
2023
-
[67]
G. C. van der Veer, T. , Michael J., W. , Yvonne, , and B. v. Muylwijk. On the interaction between system and user characteristics.Behaviour & Information Technology, 4(4):289–308, Oct. 1985. Publisher: Tay- lor & Francis eprint: https://doi.org/10.1080/01449298508901809. doi:...
1985 doi
-
[68]
Vayani, D
A. Vayani, D. Dissanayake, H. Watawana, N. Ahsan, N. Sasikumar, O. Thawakar, H. B. Ademtew, Y . Hmaiti, A. Kumar, K. Kukreja, et al. All languages matter: Evaluating lmms on culturally diverse 100 lan- guages. In Proceedings of the Computer Vision and Pattern Recogni- tion Con...
2025
-
[69]
X. Wang, S. Ong, and A. Y .-C. Nee. Multi-modal augmented-reality assembly guidance based on bare-hand interface.Advanced Engineer- ing Informatics, 30(3):406–421, 2016. 1
2016
-
[70]
J. O. Wobbrock, M. R. Morris, and A. D. Wilson. User-defined gestures for surface computing. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems , CHI ’09, p. 1083–1092. Association for Computing Machinery, New York, NY , USA, 2009. doi: 10.1145/15187...
2009
-
[71]
C. Yu. Robot Behavior Generation and Human Behavior Understand- ing in Natural Human-Robot Interaction . Theses, Institut Polytech- nique de Paris, June 2021. 8
2021
-
[72]
Zachmann and A
G. Zachmann and A. Rettig. Natural and robust interaction in virtual assembly simulation. In Eighth ISPE International Conference on Concurrent Engineering: Research and Applications (ISPE/CE2001), vol. 1, pp. 425–434. Citeseer, 2001. 1
2001
-
[73]
Zhang, Q
Q. Zhang, Q. Liu, J. Duan, and J. Qin. Research on teleoperated virtual reality human–robot five-dimensional collaboration system. Biomimetics, 8(8):605, 2023. 3 12
2023
-
[2015]
doi: 10.1016/j.ifacol.2015.06.116 1
2015 doi
-
[2019]
doi: 10.1109/TAFFC.2017.2737019 2
2017
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.