{"id":"d1a299fd-6b11-4eef-8e5d-898c22e7cae2","arxiv_id":"2506.02622","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"HORUS, a Mixed Reality interface for managing teams of mobile robots, was tested with 20 users and beat standard single-robot teleoperation on speed, usability, and frustration.","lead":"A new mixed reality interface, HORUS, lets one operator monitor and control a team of mobile robots through a shared mini-map, task panels, and two teleoperation modes. In a 20-person user study, operators using HORUS found five hidden markers faster, with higher usability scores and less frustration, than operators driving the robots with a plain teleoperation screen.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"HORUS's reported speed advantage is confounded by capability: the HORUS condition had autonomous goal-setting and a shared map, while Teleop-Only had neither, so the interface-specific benefit is not established.","rationale":"The reader's weakest assumption targets exactly the same construct-validity issue: the HORUS condition bundles autonomy and shared mapping, while Teleop-Only does not. I treat this as the single most load-bearing threat because the paper's headline claim is about interface effectiveness, and the experimental contrast cannot isolate the interface. Other concerns—small sample size, unreported raw data, mislabeled statistical test, and training differences—are real but secondary; they affect reliability and reporting, not construct validity. If a 2D baseline with identical autonomy shows no HORUS advantage, the central claim fails; if HORUS still wins, the core conclusion survives. The reader's CONDITIONAL verdict already captures this requirement, so no change to the verdict is warranted. I agree with the reader's identification of the weakest assumption.","tokens_in":8877,"tokens_out":2815,"duration_ms":29326,"concrete_test":"Run a follow-up user study with a third condition: a desktop 2D interface that provides the same autonomous goal-setting, waypoints, shared map, and ArUco counter as HORUS, with matched training time. If HORUS does not significantly beat this non-MR baseline on task completion time and SUS, the reported advantage is attributable to autonomy rather than to the MR interface.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that HORUS 'validated the versatility and effectiveness' of the MR interface in multi-robot coordination rests on the task-time comparison in Table I (8:42 vs 11:17, t(18)=4.32, p<0.001). However, the two conditions in §IV.A do not differ only in interface modality. HORUS participants could assign autonomous navigation goals, set waypoints, and build a shared SLAM map, enabling parallel search patterns; Teleop-Only participants could only switch between two robots' video streams and drive them manually, one at a time. Thus the independent variable is a bundled treatment: MR visualization plus autonomy plus shared mapping versus plain teleoperation. The observed speedup could be entirely due to the autonomy and map, not to the MR presentation or interaction design. The training-time imbalance (18 vs 7 minutes, §IV.C.1) and the absence of raw data or a third condition prevent disentangling these factors. This is a construct-validity problem, not a mathematical error: the p-value may be correct for the bundled comparison, but it does not support the paper's interface-specific conclusion. The statistical reporting oddity (calling Mann-Whitney U 'parametric' while reporting t-tests, §IV.C) is secondary but reinforces the need for transparent data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents HORUS, a Unity-based Mixed Reality interface for the Meta Quest 3 that combines a mini-map ground station, per-robot status and sensor panels, task assignment (goal poses, waypoints, labeling, drawn paths), and two teleoperation modes for managing a team of ROSbot 2.0 robots. The authors describe the system architecture (multi-master ROS, map merging, TEB navigation) and report a between-subjects user study (n=10 per group) comparing a HORUS condition with a Teleop-Only condition on an ArUco-tag search-and-rescue inspired task. They report faster task completion (8:42 vs 11:17), higher SUS (82.3 vs 68.5), lower frustration, and no significant SSQ differences, concluding that HORUS is validated as an effective multi-robot coordination interface.","tokens_in":9074,"tokens_out":4669,"duration_ms":45453,"significance":"If the reported results were attributable to the MR interface itself, the study would provide useful evidence for MR-based multi-robot team management on real robots. The system contribution is substantial: it integrates a shared mini-map, multi-robot SLAM, task assignment, and teleoperation in one deployable MR headset, and it is evaluated on physical robots rather than in simulation. The reported effect sizes are large, and the task-time and SUS outcomes are directionally consistent. However, the empirical design as reported does not isolate the MR interface from the autonomy and mapping capabilities included in the HORUS condition, so the paper's central claim needs to be scaled back or supplemented. The absence of raw data and the inconsistent statistical labeling further limit the strength of the quantitative conclusions.","major_comments":[{"comment":"The HORUS condition bundles the MR interface with autonomous goal-setting, shared SLAM map building, and a mini-map, whereas the Teleop-Only condition provides only manual velocity control with camera switching. Consequently, the significant task-time difference (8:42 vs 11:17, d=1.93) cannot be attributed to the MR visualization or interaction design; it could be produced entirely by the autonomous navigation and shared map. A third condition (e.g., autonomous goal-setting with a conventional 2D interface) or an autonomy-only baseline is needed to support the conclusion that HORUS's MR features, rather than the added capabilities, drive the improvement.","section":"IV.A, Table I"},{"comment":"The training-time imbalance (18 minutes for HORUS vs 7 minutes for Teleop-Only) is a confound: the HORUS group received more than twice as much hands-on practice, which could inflate its performance independent of interface quality. The manuscript mentions this difference but does not analyze or control for it. At minimum, the authors should report whether task time correlates with training time and discuss the direction of the potential bias.","section":"IV.C.1"},{"comment":"The statistical reporting is internally inconsistent: the text states that a 'parametric test (i.e., Mann-Whitney U)' was used, but Mann-Whitney U is nonparametric, and the reported statistics are t-values with df=18, which correspond to an independent-samples t-test. The authors should state exactly which test was used for each outcome, report the corresponding test statistic (e.g., U or t with exact p), and avoid the mislabeling. This is necessary for the quantitative claims to be verifiable.","section":"IV.C"},{"comment":"The frustration difference (p=0.04) is reported without correction for the fact that six TLX dimensions were tested. Under a Bonferroni correction for six comparisons, p=0.04 would not reach significance. The paper should either apply a multiplicity correction, or explicitly identify frustration as a targeted hypothesis with justification, and adjust the language accordingly.","section":"IV.C.4, Table III"}],"minor_comments":[{"comment":"References [7] and [13] are the same work (Chen et al., 'A 3D mixed reality interface for human-robot teaming') cited twice with different venues, and references [8] and [12] duplicate Kennel-Maushart et al.; consolidate these citations.","section":"References"},{"comment":"The sentence introducing the statistical analysis says 'parametric test (i.e., Mann-Whitney U)'; this is a factual mischaracterization, as Mann-Whitney U is a nonparametric test.","section":"IV.C"},{"comment":"There is a stray period before 'As qualitative data' at the beginning of the qualitative paragraph.","section":"IV.C.5"},{"comment":"The HORUS condition is described as 'full HORUS application, excluding the semi-immersive teleoperation feature'; the abstract and conclusions should be precise about which teleoperation modes were evaluated.","section":"IV.A.1"},{"comment":"In the sentence about egocentric command inputs, 'fostersegocentric' is missing a space.","section":"II"}],"recommendation":"major_revision","confidential_remarks":"The paper is a system-plus-user-study contribution. The bundled-treatment issue is the main obstacle: the current design cannot support the interface-specific claims in the abstract and conclusion. If the authors can reframe the contribution as an evaluation of the complete HORUS system (including autonomy and mapping) and clearly state this limitation, a revised version could be acceptable; a new baseline would be stronger but may be beyond the manuscript's current scope. The statistical mislabeling should be corrected before any further review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a read if you work on mixed reality for multi-robot teams. It is a system paper with a user study, not a theoretical result. What is actually new: a real-robot comparison of a full MR team-management interface against plain teleoperation, run on two ROSbots with a Quest 3, 20 participants, and five ArUco-clue search task. The interface itself is a reasonable consolidation of prior MR work (Chen et al., Kennel-Maushart et al.) into a single system: mini-map, robot manager panel, goal setting, waypoints, label poses, and two teleop modes. That is a solid engineering contribution, and the study reports large, directionally consistent effects: faster task time (8:42 vs 11:17, d=1.93), higher SUS, lower frustration. I believe those results as reported for the bundled conditions.\n\nThe soft spots are real and mostly match the stress-test note. The two conditions differ in capability, not just interface: HORUS participants could assign autonomous goals, set waypoints, and build a shared map; Teleop-Only participants could only drive one robot at a time. The speed advantage could come entirely from the autonomy and shared map, not from the MR presentation. The authors themselves acknowledge this in the discussion, but the abstract overstates it by claiming the experiments validated HORUS's effectiveness. The statistical reporting is sloppy: they call Mann-Whitney U a parametric test and then report t-statistics, which is contradictory. No raw data or code are shared. The sample is small (n=10 per group), and the training-time imbalance (18 vs 7 minutes) is another confound. None of these are fatal, but they mean the paper should be framed as comparing a full mixed-reality management system to a bare teleop baseline, not as isolating the contribution of the MR interface.\n\nWho gets value: researchers working on HRI, mixed-reality teleoperation, or multi-robot interfaces. The paper is a decent systems contribution, not a paradigm shift. It deserves a serious referee who can push for corrected statistics, clearer framing, and ideally a third condition or raw data. I recommend sending it to peer review rather than desk-rejecting it.","headline":"A working MR multi-robot system with a real user study, but the headline speed advantage is bundled with autonomous navigation and shared mapping, so the interface-specific claim is not yet established.","tokens_in":9650,"tokens_out":2153,"would_cite":false,"duration_ms":22764,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a mixed-reality mini-map interface lets a single operator manage a team of mobile robots faster and with less frustration than teleoperating them one at a time, and reports a user study supporting this.","keywords":["mixed reality","multi-robot systems","human-robot interaction","teleoperation","task allocation","search and rescue","user study","mobile robots"],"falsifier":"Repeat the five-marker search with a control group that also has autonomous goal-setting and a shared map, with matched training time, and check whether HORUS still finishes faster and scores higher on usability; if the gap disappears, the claimed benefit comes from the added capabilities rather than the mixed-reality presentation.","tokens_in":8632,"feed_emoji":"🤖","tokens_out":7030,"duration_ms":66197,"temperature":0.7,"pith_summary":"The paper tries to establish that a single operator can manage a team of mobile robots more effectively with HORUS, a mixed-reality interface centered on a shared mini-map, than by teleoperating robots one at a time. In a search-and-rescue-style task with two robots and five hidden markers, HORUS users finished in 8:42 on average versus 11:17 for teleop-only users, a difference the paper reports as statistically significant with a large effect size. HORUS users also rated the system more usable and reported less frustration, even though they trained longer. This matters because multi-robot search and rescue is bottlenecked by the operator's cognitive load, and the paper argues that spatial, headset-anchored task assignment relieves that bottleneck.","feed_headline":"HORUS: one mixed-reality map speeds robot-team search by 23%","feed_subtitle":"Operators using HORUS found five hidden markers in 8:42 on average versus 11:17 with pure teleoperation.","key_machinery":"The load-bearing mechanism is the Mini-Map Ground Station, a spatially registered 3D map in the headset where every robot appears as a holographic model with a Robot Manager panel containing Status, Data Viz, Tasks, and Teleoperation tabs. Each robot builds a local occupancy grid, and a custom map-merging step (coarse TF alignment refined by phase correlation) produces one shared map that the operator uses to assign goal poses, waypoints, labels, and drawn navigation plans and to switch to direct teleoperation. The shared map is what lets one operator act on the whole team at once rather than fusing separate camera and sensor feeds mentally.","core_discovery":"The central claim is that combining goal-based task assignment, a live merged map, and per-robot teleoperation in one mixed-reality view lets a novice operator coordinate a small robot team better than pure first-person teleoperation. On the study's measures, HORUS users were faster (mean 8:42 vs 11:17, $t(18)=4.32$, $p<0.001$, Cohen's $d=1.93$), rated usability higher (SUS 82.3 vs 68.5, $p=0.006$), and reported lower frustration ($p=0.04$), with no significant difference in overall workload or simulator sickness. The paper concludes that HORUS validates mixed-reality interfaces as a practical tool for multi-robot coordination on real hardware.","pith_inferences":["A fairer test of the mixed-reality contribution would give the control condition the same autonomous goal-setting and shared map through a conventional 2D screen; until then, part of the 23% advantage may come from the added capabilities rather than from the headset presentation.","The study's strategy shift suggests a testable extension: log each robot's path and room coverage to quantify how much of the speedup comes from parallel search assignment rather than from faster control of any single robot.","If the mini-map merges maps from more than two robots, the same interface could be extended to heterogeneous ground-and-aerial teams, with the operator assigning each platform by its role rather than by its stream."],"forward_implications":["Operators can search an environment in parallel by assigning different rooms to different robots, which is what made the HORUS group faster in the study.","A mixed-reality team interface can score as 'excellent' on usability even when its training session is longer, because the interaction model corresponds to how operators think about the mission.","Keeping a mini-map visible while teleoperating one robot lets the operator preserve awareness of the rest of the team, reducing the need to switch contexts.","If the per-robot overhead stays flat as robots are added, the same Ground Station pattern could support larger teams and remote operation with minimal extra operator training."],"supporting_citations":[{"why":"Supplies the ArUco fiducial markers that define the five-target search task used to compare HORUS against teleoperation.","marker":"[10]"},{"why":"Supplies the Timed Elastic Band local planner that lets robots follow goal poses, waypoints, and drawn paths assigned through the mini-map.","marker":"[9]"},{"why":"Establishes the motivation that multi-robot interfaces and operator situational awareness matter, the gap HORUS addresses.","marker":"[2]"},{"why":"A prior 3D mixed-reality human-robot teaming interface with avatar co-localization, which HORUS extends by adding per-robot management and teleoperation.","marker":"[7]"},{"why":"A prior mixed-reality multi-robot task-allocation system whose board-game metaphor HORUS builds on with a unified mini-map Ground Station.","marker":"[8]"},{"why":"Cited as the statistical test used for the study's comparisons between the HORUS and teleoperation-only groups.","marker":"[11]"}],"fun_headline_variants":["HORUS MR interface speeds robot-team search by 23%","Mixed-reality HORUS cuts multi-robot search time by 23%","HORUS MR view lets one operator coordinate faster robot search","HORUS: one MR map makes robot-team search 23% faster","HORUS MR tool improves multi-robot coordination speed by 23%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the teleoperation-only condition is a fair baseline for individual robot teleoperation, since the two conditions differ in available capabilities, training time, and group assignment, not just in the interface.","fun_headline_variants_meta":{"raw":{"variants":["HORUS MR interface speeds robot-team search by 23%","Mixed-reality HORUS cuts multi-robot search time by 23%","HORUS MR view lets one operator coordinate faster robot search","HORUS: one MR map makes robot-team search 23% faster","HORUS MR tool improves multi-robot coordination speed by 23%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000616,"raw_usage":{"total_tokens":2854,"prompt_tokens":934,"completion_tokens":1920,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":1823}},"tokens_in":550,"tokens_out":1920,"duration_ms":13732,"temperature":1.0,"reasoning_tokens":1823,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:19:01.399503+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the five-marker search with a control group that also has autonomous goal-setting and a shared map, with matched training time, and check whether HORUS still finishes faster and scores higher on usability; if the gap disappears, the claimed benefit comes from the added capabilities rather than the mixed-reality presentation.","supporting_citations":[{"cited_title":"Multi-robot interfaces and operator situational awareness: Study of the impact of immersion and prediction,","cited_arxiv_id":null,"evidence_quote":"Establishes the motivation that multi-robot interfaces and operator situational awareness matter, the gap HORUS addresses."},{"cited_title":"A 3D mixed reality interface for human-robot teaming,","cited_arxiv_id":null,"evidence_quote":"A prior 3D mixed-reality human-robot teaming interface with avatar co-localization, which HORUS extends by adding per-robot management and teleoperation."},{"cited_title":"Interacting with multi- robot systems via mixed reality,","cited_arxiv_id":null,"evidence_quote":"A prior mixed-reality multi-robot task-allocation system whose board-game metaphor HORUS builds on with a unified mini-map Ground Station."},{"cited_title":"Mann-whitney u test,","cited_arxiv_id":null,"evidence_quote":"Cited as the statistical test used for the study's comparisons between the HORUS and teleoperation-only groups."}],"review_version":1}