REVIEW 3 major objections 6 minor 36 references
Practice Makes Perfect: A Study of Digital Twin Technology for Assembly and Problem-solving using Lunar Surface Telerobotics
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Practicing on a virtual-reality digital twin cut completion time on a physical lunar rover task by 28% and unrecoverable antenna flips by 85%.
desk verdict A transparent feasibility study with a plausible but confounded result: digital twin practice helped, but so might any extra practice, as the authors themselves concede. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the calibrated digital twin, a virtual replica of the 'Armstrong' rover whose arm-joint speeds, drive speeds, and rotation times were measured on the physical rover and matched in simulation using timed benchmarks, with pilot testing used to tune remaining differences. The experimental contrast is a two-group between-subjects design: Group A performs the antenna-alignment task once on the physical rover after a 10-minute familiarization, while Group B performs the identical task in the digital twin first and then on the physical rover. The VR headset and game-controller interface are the same in both conditions, and the task is identical, so any difference in outcome is attributed to the intervening twin rehearsal.
What would settle it
Run a third group that gets one untimed practice run on the physical rover before the timed run; if their times and flip counts match the digital-twin group, the benefit is practice, not the twin. A second check is to deliberately mis-calibrate the twin's joint speeds and see whether the transfer advantage shrinks or vanishes.
Extended reading notes
Core claim
The paper's central claim is a measured transfer effect: operating a digital twin before operating the physical rover makes operators faster and less error-prone on the physical system. Group differences favor the twin-trained group, with $t(22)=3.05$, $p=0.006$, $d=1.25$ for completion time (28% faster) and $t(22)=2.68$, $p=0.014$, $d=1.09$ for unrecoverable antenna flips (85% fewer). The twin-trained group also showed a 36% drop in self-reported time pressure and a 17% drop in frustration, while overall mental demand was essentially unchanged. The authors use the comparable System Usability Scale scores of the two platforms (71.45 vs 70) to argue that the objective gains come from the training content rather than from a usability gap between virtual and physical interfaces.
Load-bearing premise
The claim that the digital twin itself causes the improvement assumes that the extra practice run is not the real cause, since the trained group practiced once more than the control group, which had no practice run at all.
Editorial extensions
If this is right
- Future lunar operators could rehearse assembly and real-time troubleshooting in VR before touching flight hardware, reducing wear and risk to expensive rovers.
- The largest measured gains appear in the most difficult manipulation step, gripping the antenna, so twin training may be most valuable for error-prone manual subtasks.
- For missions that deploy hundreds of identical antennas, a twin practice run can standardize operator skill ahead of the mission instead of learning on the job.
- The same digital twin approach can be carried from the simple testbed to flight-ready rovers by tuning lunar environment characteristics, which the paper states as its next step.
Reading between the lines
- Editorial inference: because the twin-trained group received an extra practice trial, a three-arm study with a physical-practice control is needed to separate rehearsal effects from the digital twin medium; the paper itself identifies this as its largest limitation.
- Editorial inference: the 28% and 85% figures should be read as an upper bound for real lunar operations, where communication latency, unknown terrain, and mission stakes were not simulated.
- Editorial inference: replacing the VR headset with a conventional screen in the practice condition would test whether the immersive display does the work or whether any simulation of the task would transfer.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes the design of a physical rover ('Armstrong') and a matching Unity-based digital twin, and reports a between-subjects experiment (n=24, 12 per group) comparing a physical-only condition (Group A) with a condition in which participants first practiced in the digital twin and then performed the same physical antenna-alignment task (Group B). The authors report a 28% reduction in task completion time (t(22)=3.05, p=0.006, d=1.25) and an 85% reduction in unrecoverable antenna flips (t(22)=2.68, p=0.014, d=1.09) for Group B, alongside survey results on TLX, SUS, and Affect Grid. The paper argues that digital twin training improves operator performance and subjective experience for lunar telerobotic assembly, and discusses implications for missions such as FARSIDE and FarView. The authors explicitly acknowledge in Section 7 that the design compares a virtual practice run against no practice run at all, and that further work is needed to compare digital twin training to physical training.
Significance. If the central causal claim were supported, the paper would provide a useful feasibility result for VR digital twin training in planetary telerobotics, with potential practical value for future lunar array deployments and as a lower-cost complement to physical analog facilities. The authors deserve credit for building and calibrating a physical/digital twin pair with benchmark data (Tables 1-3), for reporting effect sizes alongside p-values, and for including standard subjective measures (TLX, SUS, Affect Grid). The paper is transparent about several limitations, including sample size, task simplicity, and mechanical differences between the twins. However, the experiment as designed cannot distinguish the effect of the digital twin medium from the effect of an additional practice run, and the manuscript's stated hypothesis and abstract claim more than the data can support. The significance of the paper therefore depends on whether the claims can be reframed or the design supplemented with an appropriate control condition.
major comments (3)
- [Section 4.2 and Section 7] The central comparison confounds the digital twin medium with an additional practice run. Group A performed the physical task once after a 10-minute familiarization; Group B performed a virtual practice run and then the physical task. The observed 28% reduction in completion time and 85% reduction in flips are therefore compatible with the explanation that any extra rehearsal of the task improves performance. The paper acknowledges this in Section 7: 'Our experiment shows that participants were able to improve performance in our task after going through a virtual practice run, when compared to participants that did not get a practice run at all.' As written, the abstract and Section 5.1 claim that digital twin training improves performance, which the design cannot identify. The manuscript needs either an additional physical-practice control group or a thorough reframing of the claims to something like 'practice in the digital twin improves performance over no practice,' with all causal language about the digital twin medium removed or explicitly bounded.
- [Section 5.2, paragraphs 2-3] The SUS and Affect Grid results are used to conclude that performance differences are 'in fact due to training with the digital twin, and not some difference in platform or usability.' Equal SUS scores only show comparable perceived usability between the two systems; they do not rule out the alternative explanation that the benefit came from the extra task rehearsal. For the same reason, the Affect Grid results showing similar mood across groups do not establish that the digital twin, rather than practice, caused the performance difference. These inferential statements should be removed or explicitly bounded to what the measures actually support.
- [Section 5.1] The paper reports multiple t-tests (completion time, flips, SUS, and several TLX subscales) without correction for multiple comparisons. With six or more tests, the flips result (p = 0.014) would not survive a Bonferroni correction at alpha = 0.05, while the completion-time result (p = 0.006) would. In addition, timing was recorded manually and the sample is small (n = 12 per group). The manuscript should report confidence intervals for the effect sizes and explicitly discuss the sensitivity of the conclusions to these statistical choices.
minor comments (6)
- [Section 4.2] The sentence 'This study is a within-participants design' is incorrect; participants were assigned to one of two groups and compared between groups, which is a between-subjects design. The error is worth correcting because the t-tests in Section 5.1 treat the groups as independent.
- [Section 4] The description of the design as a '2x1 study' is ambiguous; a two-condition between-subjects design would be clearer.
- [Section 1, third paragraph] The phrase 'allowing for and more repeatable simulation' is missing a word; it should read 'allowing for more repeatable simulation.'
- [Section 3.1] The phrase 'first generation Oculus Quest' should be hyphenated as 'first-generation Oculus Quest.'
- [Section 5.2 and Figure 8] The claims about time pressure and frustration report percentage changes without test statistics or confidence intervals. Either add inferential statistics or clearly present these as descriptive observations.
- [Figure 5] The statement that prior VR or video-game experience 'is not a serious bias' appears to be based on visual inspection of scatterplots; consider reporting a correlation or regression coefficient to support this claim.
Circularity Check
No circular derivation: the training-benefit claim is an empirical between-group comparison with independent outcome measures.
full rationale
The paper's central claim is an empirical result: participants who practiced in the digital twin before operating the physical Armstrong rover completed the antenna-alignment task faster and with fewer unrecoverable flips than participants who only operated the physical rover. There is no fitted model, no parameter estimated from the outcome data and then presented as a prediction, and no definitional identity between the digital-twin training condition and the physical-task outcome. The virtual rover's movement characteristics were calibrated to the physical rover via independent benchmarks (drive and rotation times, Section 3.2), and the comparison then measures transfer to the physical task; even if that calibration were imperfect, it does not make the observed 28% completion-time improvement equivalent to an input. The self-citations (Burns et al. 2019/2021, Polidan et al. 2024, Conlon et al. 2022/2024, Walker et al. 2023/2024) provide mission context and related work, not the evidence for the training effect. Section 7 explicitly concedes the design confound: 'Our experiment shows that participants were able to improve performance in our task after going through a virtual practice run, when compared to participants that did not get a practice run at all.' That is an internal-validity threat about practice versus medium, not a circularity: the conclusion does not reduce by construction to its inputs. The claim that SUS and Affect Grid results support medium-specific training is an inference about confounds, not a circular derivation. No circular step meets the evidentiary bar.
Assumptions & free parameters
free parameters (1)
- Unity environment tuning parameters (friction and related physics) =
not reported
assumptions (3)
- domain assumption The control group's 10-minute familiarization is assumed to provide negligible task-specific training compared with a full virtual practice run.
- domain assumption The digital twin is sufficiently faithful to the physical rover for skill transfer.
- domain assumption Experimenters' manual timing of completion is unbiased.
Cite this review
Pith. "Pith review of Practice Makes Perfect: A Study of Digital Twin Technology for Assembly and Problem-solving using Lunar Surface Telerobotics." pith.science (2026). https://pith.science/paper/55JK5GWH
@misc{pith2026250513722,
author = {Pith},
title = {Pith review of: Practice Makes Perfect: A Study of Digital Twin Technology for Assembly and Problem-solving using Lunar Surface Telerobotics},
year = {2026},
howpublished = {\url{https://pith.science/paper/55JK5GWH}},
note = {Machine review of arXiv:2505.13722}
}
read the original abstract
Robotic systems that can traverse planetary or lunar surfaces to collect environmental data and perform physical manipulation tasks, such as assembling equipment or conducting mining operations, are envisioned to form the backbone of future human activities in space. However, the environmental conditions in which these robots, or "rovers," operate present challenges toward achieving fully autonomous solutions, meaning that rover missions will require some degree of human teleoperation or supervision for the foreseeable future. As a result, human operators require training to successfully direct rovers and avoid costly errors or mission failures, as well as the ability to recover from any issues that arise on the fly during mission activities. While analog environments, such as JPL's Mars Yard, can help with such training by simulating surface environments in the real world, access to such resources may be rare and expensive. As an alternative or supplement to such physical analogs, we explore the design and evaluation of a virtual reality digital twin system to train human teleoperation of robotic rovers with mechanical arms for space mission activities. We conducted an experiment with 24 human operators to investigate how our digital twin system can support human teleoperation of rovers in both pre-mission training and in real-time problem solving in a mock lunar mission in which users directed a physical rover in the context of deploying dipole radio antennas. We found that operators who first trained with the digital twin showed a 28% decrease in mission completion time, an 85% decrease in unrecoverable errors, as well as improved mental markers, including decreased cognitive load and increased situation awareness.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize "" * " " * ...
-
[2]
author Bale, S. D. , author Bassett, N. , author Burns, J. O. et al. ( year 2023 ). title Lusee'night': The lunar surface electromagnetics experiment . journal arXiv preprint arXiv:2301.10345 \/ ,
arXiv 2023
-
[3]
author Boslaugh, S. ( year 2012 ). title Statistics in a Nutshell, 2nd Edition \/ . publisher O'Reilly Media, Incorporated . https://books.google.com/books?id=s1llAQAACAAJ
work page 2012
-
[4]
author Brooke, J. ( year 1996 ). title Sus: A “quick and dirty” usability scale . journal Usability Evaluation in INdustry/Taylor and Francis \/ ,
work page 1996
-
[5]
author Burns, J. O. , author Hallinan, G. , author Lux, J. et al. ( year 2019 ). title Nasa probe study report: Farside array for radio science investigations of the dark ages and exoplanets (farside) . journal arXiv preprint arXiv:1911.08649 \/ ,
arXiv 2019
-
[6]
author Burns, J. O. , author MacDowall, R. , author Bale, S. et al. ( year 2021 ). title Low radio frequency observations from the moon enabled by nasa landed payload missions . journal The Planetary Science Journal \/ , volume 2 \/ issue (2) , pages 44
work page 2021
-
[7]
author Community, B. O. ( year 2018 ). title Blender - a 3D modelling and rendering package \/ . organization Blender Foundation address Stichting Blender Foundation, Amsterdam . http://www.blender.org
work page 2018
-
[8]
author Conlon, N. , author Ahmed, N. R. , & author Szafir, D. ( year 2024 ). title A survey of algorithmic methods for competency self-assessments in human-autonomy teaming . journal ACM Computing Surveys \/ , volume 56 \/ issue (7) , pages 1--31
work page 2024
Show all 36 references
-
[9]
i'm confident this will end poorly
author Conlon, N. , author Szafir, D. , & author Ahmed, N. ( year 2022 ). title “i'm confident this will end poorly”: robot proficiency self-assessment in human-robot teaming . In booktitle IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) \/ (pp. page...
2022
-
[10]
, author Zhang, Y
author Ding, Y. , author Zhang, Y. , & author Huang, X. ( year 2023 ). title Intelligent emergency digital twin system for monitoring building fire evacuation . journal Journal of Building Engineering \/ , volume 77 \/ , pages 107416
2023
-
[11]
, & author Chien, S
author Gao, Y. , & author Chien, S. ( year 2017 ). title Review on space robotics: Toward top-level science through space exploration . journal Science Robotics \/ , volume 2 \/ issue (7)
2017
-
[12]
author Haas, J. K. ( year 2014 ). title A history of the unity game engine ,
2014
-
[13]
author Hart, S. G. , & author Staveland, L. E. ( year 1988 ). title Development of nasa-tlx (task load index): Results of empirical and theoretical research . In booktitle Advances in psychology \/ (pp. pages 139--183 ). publisher Elsevier volume volume 52
1988
-
[14]
author Hibbard, J. J. , author Burns, J. O. , author MacDowall, R. et al. ( year 2025 ). title Results from nasa's first radio telescope on the moon: Terrestrial technosignatures and the low-frequency galactic background observed by rolses-1 onboard the odysseus lander . journ...
2025 arXiv
-
[15]
, & author Oron-Gilad, T
author Honig, S. , & author Oron-Gilad, T. ( year 2018 ). title Understanding and resolving failures in human-robot interaction: Literature review and model development . journal Frontiers in Psychology \/ , volume 9 \/ , pages 861
2018
-
[16]
title What is cognitive load? journal Interaction Design Foundation - IxDF \/ ,
author Interaction Design Foundation - IxDF ( year 2016 ). title What is cognitive load? journal Interaction Design Foundation - IxDF \/ , . https://www.interaction-design.org/literature/topics/cognitive-load
2016
-
[17]
, & author Scheutz, M
author Law, T. , & author Scheutz, M. ( year 2021 ). title Trust: Recent concepts and evaluations in human-robot interaction . journal Trust in human-robot interaction \/ , (pp. pages 27--57 )
2021
-
[18]
, author Wang, D
author Leng, J. , author Wang, D. , author Shen, W. et al. ( year 2021 ). title Digital twins-based smart manufacturing system design in industry 4.0: A review . journal Journal of manufacturing systems \/ , volume 60 \/ , pages 119--137
2021
-
[19]
, & author Challinor, A
author Lewis, A. , & author Challinor, A. ( year 2007 ). title 21 cm angular-power spectrum from the dark ages . journal Physical Review D—Particles, Fields, Gravitation, and Cosmology \/ , volume 76 \/ issue (8) , pages 083005
2007
-
[20]
, & author Furlanetto, S
author Loeb, A. , & author Furlanetto, S. R. ( year 2012 ). title The first galaxies
2012
-
[21]
, author Liu, C
author Lu, Y. , author Liu, C. , author Kevin, I. et al. ( year 2020 ). title Digital twin-driven smart manufacturing: Connotation, reference model, applications and research issues . journal Robotics and computer-integrated manufacturing \/ , volume 61 \/ , pages 101837
2020
-
[22]
, & author Harvey, C
author Matulis, M. , & author Harvey, C. ( year 2021 ). title A robot arm digital twin utilising reinforcement learning . journal Computers & Graphics \/ , volume 95 \/ , pages 106--114
2021
-
[23]
, author Yaqoob, M
author Mihai, S. , author Yaqoob, M. , author Hung, D. V. et al. ( year 2022 ). title Digital twins: A survey on enabling technologies, challenges, trends and future prospects . journal IEEE Communications Surveys & Tutorials \/ , volume 24 \/ issue (4) , pages 2255--2291
2022
-
[24]
author Nesnas, I. A. , author Fesq, L. M. , & author Volpe, R. A. ( year 2021 ). title Autonomy for space robots: Past, present, and future . journal Current Robotics Reports \/ , volume 2 \/ issue (3) , pages 251--263
2021
-
[25]
Salmon, G
author Paul M. Salmon, G. H. W. C. B. D. P. J. R. M., Neville A. Stanton , & author Young, M. S. ( year 2008 ). title What really is going on? review of situation awareness models for individuals and teams . journal Theoretical Issues in Ergonomics Science \/ , volume 9 \/ iss...
2008 doi
-
[26]
author Polidan, R. S. , author Burns, J. O. , author Ignatiev, A. et al. ( year 2024 ). title Farview: An in-situ manufactured lunar far side radio array concept for 21-cm dark ages cosmology . journal Advances in Space Research \/ , volume 74 \/ issue (1) , pages 528--546
2024
-
[27]
, author Dagar, A
author Rajasekhar, R. , author Dagar, A. K. , author Nagori, R. et al. ( year 2024 ). title Comprehensive analysis of chandrayaan-3 landing site region focussing on morphology, hydration and gravity anomalies . journal Icarus \/ , volume 415 \/ , pages 116074
2024
-
[28]
( year n.d
author Robotics, J. ( year n.d. ). title The mars yard iii . https://www-robotics.jpl.nasa.gov/how-we-do-it/facilities/marsyard-iii/
-
[29]
, author Weiss, A
author Russell, J. , author Weiss, A. , & author Mendelsohn, G. ( year 1989 ). title Affect grid: A single-item scale of pleasure and arousal . journal Journal of Personality and Social Psychology \/ , volume 57 \/ , pages 493--502 . :10.1037/0022-3514.57.3.493
1989 doi
-
[30]
, author Medicine et al
author National Academies of Sciences, E. , author Medicine et al. ( year 2021 ). title Decadal survey on astronomy and astrophysics 2020
2021
-
[31]
, author Nakagawa, K
author Shiomi, M. , author Nakagawa, K. , & author Hagita, N. ( year 2013 ). title Design of a gaze behavior at a small mistake moment for a robot . journal Interaction Studies \/ , volume 14 \/ issue (3) , pages 317--328
2013
-
[32]
author Stanford Artificial Intelligence Laboratory et al. (). title Robotic operating system . https://www.ros.org
-
[33]
, author Phung, T
author Walker, M. , author Phung, T. , author Chakraborti, T. et al. ( year 2023 ). title Virtual, augmented, and mixed reality for human-robot interaction: A survey and virtual design element taxonomy . journal ACM Transactions on Human-Robot Interaction \/ , volume 12 \/ iss...
2023
-
[34]
author Walker, M. E. , author Gramopadhye, M. , author Ikeda, B. et al. ( year 2024 ). title The cyber-physical control room: A mixed reality interface for mobile robot teleoperation and human-robot teaming . In booktitle Proceedings of the ACM/IEEE International Conference on...
2024
-
[35]
, & author Ru winkel, N
author Weidemann, A. , & author Ru winkel, N. ( year 2021 ). title The role of frustration in human--robot interaction--what is needed for a successful collaboration? journal Frontiers in psychology \/ , volume 12 \/ , pages 640186
2021
-
[36]
, author Liu, D
author Zeng, X. , author Liu, D. , author Chen, Y. et al. ( year 2023 ). title Landing site of the chang’e-6 lunar farside sample return mission from the apollo basin . journal Nature Astronomy \/ , volume 7 \/ issue (10) , pages 1188--1197
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.