{"id":"e18290c6-6b83-4a4a-9ab0-ea01b1f9ae6f","arxiv_id":"2501.02304","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"ARTHUR is an authoring system that combines 20 feedback types, 10 actions, and 18 conditions in a hybrid desktop/tablet/AR workflow, demonstrated by replicating prior HRC systems and a five-participant expert evaluation.","lead":"The paper presents ARTHUR, an open-source tool for authoring augmented reality feedback and controls for human-robot collaboration using a hybrid combination of desktop, tablet, and head-mounted displays. It is notable because it offers an integrated, testable approach to a long-standing practical problem: creating and refining AR guidance for collaborative robots.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'general applicability' claim rests on an unvalidated component inventory: the 20/10/18 items are asserted from literature but never mapped against the cited taxonomies, and three demonstrations cannot establish coverage.","rationale":"The reader's weakest assumption matches my main concern; I agree. The component inventory is the 'language' in which all authored HRC interfaces are expressed. The paper's evidence for coverage is (i) an assertion that review papers informed the set (Section 2.2), (ii) three demonstration scenarios (Section 4.1), and (iii) an expert evaluation of one replicated scenario (Section 4.2). None of these constitutes a systematic coverage analysis. Because extensibility requires programming (Section 3.3), the tool's end-user authoring promise is bounded by the 48 components. I also considered an internal inconsistency: Section 1 claims in-situ authoring of actions and conditions, while Section 3.2.3 says conditions are generally created in the web interface. That is a real overclaim and should be fixed, but it does not undermine the central hybrid-UI feasibility claim as much as inventory incompleteness does. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":19206,"tokens_out":6539,"duration_ms":65952,"concrete_test":"Construct a coverage matrix that maps every design component in the cited taxonomies (e.g., Suzuki et al. CHI'22 groups UIs/widgets, spatial references, embedded effects; Walker et al. virtual design elements) to ARTHUR's 20 feedback, 10 action, and 18 condition types, marking each as covered, partially covered, or missing. Then attempt to replicate 10 additional scenarios from those taxonomies using only ARTHUR's current component set. If the matrix reveals a category with no implemented component and that category is needed by a scenario, the general-applicability claim should be explicitly narrowed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of general applicability (Sections 1, 4.1, 6) depends on the expressive completeness of the component inventory: 20 feedback types, 10 actions, and 18 conditions (Section 3.2). The paper states these were informed by survey taxonomies (Section 2.2, citing Suzuki et al. and Walker et al.), but it never provides a mapping from those taxonomies to the implemented set, nor does it analyze what was omitted. The three demonstration scenarios (Section 4.1) are implemented by the authors on a single robot arm (UR5e via one adapter, Section 3.3), so they cannot falsify inventory incompleteness. The paper's own Sections 5.1 and 5.2 acknowledge limitations and extensions, and Section 3.2.2 admits some actions 'might not be applicable or possible on all types of robots.' Since adding a missing component requires code changes in two places (Section 3.3), an end-user cannot repair a coverage gap. If a substantial class of HRC scenarios requires feedback, action, or condition types absent from the 48-item set, the 'complete authoring' and 'general applicability' claims fail even though the hybrid UI itself may be sound.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"ARTHUR is an open-source authoring tool for augmented reality-supported human-robot collaboration. The paper describes a system that provides 20 feedback types, 10 actions, and 18 conditions as design components, and a hybrid user interface combining a web front end on PC/tablet with an AR head-mounted display for in-situ testing and refinement. The authoring workflow is organized into configuration, refinement, and operation phases. The authors demonstrate the system by reconstructing three scenarios from prior work (a replicated HoloLens interaction setup, a motion-intent visualization comparison, and a sensor-value visualization) and report a qualitative usage evaluation in which five AR/robotics experts replicated the first scenario and gave favorable subjective feedback. The paper claims that ARTHUR enables complete AR-based HRC interfaces and has general applicability, while acknowledging limitations and the need for further studies in Sections 5 and 5.1.","tokens_in":19440,"tokens_out":6959,"duration_ms":64730,"significance":"If the claims are supported, ARTHUR would be a valuable systems contribution: it is one of few open-source authoring tools that span feedback, actions, and conditions for AR-based HRC, and its hybrid-interface design directly addresses known usability problems of in-situ AR authoring. The paper's strengths include a public repository, supplementary videos, demonstrations anchored to externally published scenarios, and transparent acknowledgment of its own limitations in Section 5.1. The main weakness is that the strongest claims—general applicability and the effectiveness of the hybrid approach—rest on an unvalidated component inventory and a small, self-selected, single-condition qualitative study. The contribution is a promising feasibility demonstration rather than a fully validated general-purpose authoring solution.","major_comments":[{"comment":"The general-applicability claim is load-bearing and not adequately supported. The abstract and Section 6 describe ARTHUR as enabling 'complete AR-based HRC interfaces' with 'general applicability,' and Section 3.2 presents the 20/10/18 inventory as the mechanism for that coverage. However, the paper never maps the cited taxonomies (Suzuki et al. and Walker et al., Section 2.2) onto the implemented inventory, nor does it state which taxonomy items were intentionally omitted. The three demonstration scenarios in Section 4.1 were authored by the system's developers using only the UR5e adapter, so they cannot falsify inventory incompleteness. Section 3.3 makes the coverage cost concrete: adding a component requires code changes in both the authoring service and the AR interface, so an end-user cannot repair a coverage gap. I recommend either adding a systematic coverage analysis (e.g., a table mapping taxonomy categories to the 48 components, with explicit exclusions) or replacing 'complete' and 'general applicability' with scoped claims about the demonstrated scenario classes.","section":"Sections 3.2, 4.1, and 6"},{"comment":"The hybrid-UI effectiveness claim rests on a small, self-selected, single-condition qualitative study. Five participants (all male, with AR/robotics expertise) completed the authoring task, but there is no baseline or comparison condition (e.g., desktop-only or HMD-only authoring), and the outcome measures are self-reported. The paper itself in Section 5 notes that 'further studies are obviously necessary to confirm this claim' and in Section 5.1 calls for evaluations with robotics engineers, UX designers, and workers. Given these acknowledgments, the evidence supports a formative feasibility statement rather than the abstract's characterization of the hybrid approach as 'effective.' I ask the authors to either clearly frame the evaluation as a feasibility/reflection study and soften the effectiveness wording, or add a comparative condition and a more diverse participant sample.","section":"Sections 4.2, 5, and 5.1"},{"comment":"The replication demonstrations are not verified to the fidelity claimed. Scenario 1 is described as replicating Hietanen et al., yet it uses seven visual elements instead of six 'to achieve the same result,' and the safety zone is not natively supported; it requires an external service publishing points over MQTT. The paper does not report whether the resulting behavior was compared with the original system or whether the fidelity was assessed in any structured way. Since these demonstrations are the main evidence for the capability to 'replicate representative examples from prior work' and for general applicability, the paper should specify the intended fidelity level and explicitly list deviations from each original system.","section":"Section 4.1.1"}],"minor_comments":[{"comment":"In Section 3.1, the sentence 'ARTHUR can be extended with dynamic localization of tools and parts ... though this is not currently not included in the system' contains a double negative; it should read 'not currently included.'","section":"Section 3.1"},{"comment":"In Section 3.2.3, '18conditions' is missing a space, and the sentence says '5 different categories' but then lists six categories (spatial, operator, robot, environment, task, and logic); the count should be corrected.","section":"Section 3.2.3"},{"comment":"In Section 4.2.2, 'all participants where able to complete' should be 'all participants were able to complete.'","section":"Section 4.2.2"},{"comment":"In Section 5, 'simmultaneous use' should be 'simultaneous use.'","section":"Section 5"},{"comment":"In Section 5.1, 'whom is expected to be responsible' should be 'who is expected to be responsible.'","section":"Section 5.1"},{"comment":"In Section 3.2.1, 'allows to customization' should be 'allows customization.'","section":"Section 3.2.1"},{"comment":"The qualitative analysis description in Section 4.2 is underspecified; it mentions thematic analysis of the experimenter's written protocol refined by video and audio inspection, but it does not describe the coding procedure, intercoder agreement, or how themes were derived. Adding this detail would make the evaluation easier to interpret.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"This is a promising systems paper with a transparent limitations section and a genuinely useful open-source artifact. The main risk is overclaiming general applicability and hybrid-interface effectiveness from an unvalidated component inventory and a five-participant qualitative study. I recommend major revision focused on either scoping the claims or adding the missing coverage analysis and a stronger evaluation. I saw no integrity or novelty disclosure concerns; the authors clearly identify their prior work (RoboVisAR) and provide access to code and videos."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"ARTHUR is a genuine systems contribution: an open-source authoring tool that combines feedback, actions, and conditions in a desktop/tablet/HMD workflow. The paper demonstrates the system by replicating three published scenarios and includes a small expert evaluation. The integration of bidirectional interaction, in-situ refinement, and a hybrid interface in one tool is new, relative to prior work such as the authors' own RoboVisAR.\n\nCredit where earned: the system is implemented and open-source; the demonstrations reproduce setups from Hietanen et al. and De Franco et al., which gives external grounding; the writing is clear; and the limitations are stated candidly, including the call for further studies in Section 5.1. The five-participant qualitative evaluation is appropriate for an HCI systems paper at this stage and supports the feasibility claim, not more.\n\nThe soft spots are real but fixable. The claim of 'general applicability' is not backed up. The 20/10/18 inventory is asserted as informed by survey taxonomies, but the paper never maps those taxonomies to the implemented set, so readers cannot see what was omitted. Three scenarios, all on a single UR5e with one adapter, cannot establish coverage. The stress-test note is right: adding a missing component requires code changes in two places, so an end-user can't repair a gap. The abstract's 'general applicability' overstates what a replication of three examples shows. The evaluation is self-selected experts, no baseline, and self-reported, which the paper already concedes.\n\nNone of this sinks the paper. The tool is real, the component set is a reasonable first cut, and the authors are honest about the evidence. They should temper 'general applicability' to something like 'flexibility across representative scenarios' and ideally add a mapping from the cited taxonomies to the inventory. That would turn a good systems paper into a more defensible one.\n\nFor whom: researchers in AR-supported HRC and HRI who want a reusable platform or a concrete starting point for authoring tools. This deserves a serious referee; it should go to review rather than desk rejection. I would accept with revisions focusing on claim calibration and, optionally, a coverage analysis, rather than on the artifact itself.","headline":"A real, open-source AR-HRC authoring system whose 'general applicability' claim outruns the evidence, but the artifact and honest limitations make it worth a serious look.","tokens_in":19969,"tokens_out":4293,"would_cite":true,"duration_ms":37794,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ARTHUR is an open-source authoring tool that lets users create complete augmented-reality interfaces for human-robot collaboration by combining feedback, actions, and conditions across a desktop/tablet and head-mounted display.","keywords":["augmented reality","human-robot collaboration","authoring tool","hybrid user interface","in-situ refinement","head-mounted display","assembly guidance","industrial manufacturing"],"falsifier":"Author a realistic HRC scenario outside the three demonstrated (for example, a mobile robot navigating a shared warehouse) using only ARTHUR's shipped components; if the scenario requires a feedback, action, or condition that is not in the inventory and can only be added by writing a new component in code, the general-applicability claim fails. A lighter check is to map the component sets in published AR-HRI taxonomies onto ARTHUR's 20/10/18 inventory and count missing entries that correspond to common scenarios.","tokens_in":19029,"feed_emoji":"🤖","tokens_out":6592,"duration_ms":60675,"temperature":0.7,"pith_summary":"The paper presents ARTHUR, an open-source authoring tool that lets people build complete augmented-reality interfaces for human-robot collaboration without writing code. The central idea is that any such interface can be assembled from three kinds of building blocks: feedback that tells the worker what the robot and system are doing, actions that let the worker command the robot, and conditions that decide when feedback appears or actions fire. To make authoring practical, ARTHUR splits the workflow into configuration on a PC, refinement with a tablet and AR headset in the workspace, and operation on the headset, with changes syncing instantly across devices. The authors replicate three representative systems from prior work and report that all five expert participants completed the replication task and found the hybrid interface easy and beneficial. If this holds, a single tool could cover the design, tuning, and daily use of AR guidance for multiple robot tasks.","feed_headline":"One tool authors complete AR robot-collaboration interfaces","feed_subtitle":"ARTHUR combines 20 feedback types, 10 actions, 18 conditions across desktop, tablet, and headset for in-situ refinement.","key_machinery":"The central mechanism is the component model of feedback, actions, and conditions plus a hybrid interface spanning web clients and the headset. Feedback items such as robot paths, task instructions, zones, icons, and audio convey state; actions such as play/pause robot, confirm task, and sending messages let the operator command the system; conditions in spatial, operator, robot, environment, task, and logic categories gate when feedback is shown or actions are triggered. Trackers and anchors place virtual content in the real world, and the authoring service stores configuration so edits made in the web interface are instantly reflected in AR. The division into configuration, refinement, and operation phases lets the same authored setup be reused and modified in place.","core_discovery":"The paper claims that authoring AR support for human-robot collaboration reduces to composing three kinds of design components: 20 feedback types for robot, task, and system state; 10 actions for user control; and 18 conditions that customize feedback and trigger actions. Around this inventory, ARTHUR provides a hybrid user interface: a web interface on PC or tablet for configuration and precise edits, and an AR head-mounted display for in-situ placement, testing, and refinement, all synchronized through a publish-subscribe service. The authors argue this breaks the sequential design-deploy-refine loop, because changes made on any device appear immediately in AR and the same setup can transition fluidly into live operation. As evidence, they replicate prior systems including a button-and-safety-zone setup and a sensor-value display, and in an evaluation with five participants every participant reproduced the replicated scenario and reported that distributing authoring across devices was beneficial.","pith_inferences":["If the 20/10/18 inventory is as representative as claimed, ARTHUR could become a common testbed that turns scattered single-visualization studies into comparable head-to-head evaluations; the paper itself does not run such a comparison.","The component abstraction looks portable beyond robot arms: adding a robot adapter service could reuse most feedback, actions, and conditions for mobile robots, though the paper notes that tracking on moving platforms remains a hardware limitation.","The hand-selected inventory could be stress-tested by systematically mapping published AR-HRI taxonomies onto ARTHUR's components and counting gaps; that coverage analysis is not in the paper.","A recommendation engine that suggests compatible groups of feedback and actions, or auto-detection of connected sensors and buttons, could reduce authoring errors; the authors list these as future work, and the architecture supports them via publish-subscribe topics."],"forward_implications":["An author can build a working AR guidance interface for a robot cell in one sitting, then switch to live operation without recompiling or redeploying the application.","Because changes sync immediately across devices, a worker can fine-tune content position, color, and thresholds in the workspace rather than alternating between a simulated environment and the real one.","The action components give the operator a way to command the robot from the authored interface, making information flow bidirectional rather than one-way robot-to-user.","The same generalized visualizations can serve many similar tasks loaded from a bill of process, avoiding per-instruction authoring for step-by-step guidance.","Other researchers can use ARTHUR to recreate and directly compare different motion-intent and safety visualizations from prior work in one physical setup."],"supporting_citations":[{"why":"Supplies the taxonomy of AR-HRI design components and interaction levels that informed the choice of ARTHUR's feedback, actions, and conditions.","marker":"[5]"},{"why":"Prior in-situ authoring tool for robot visualizations; ARTHUR extends it with assembly instructions, user input, and a hybrid interface.","marker":"[15]"},{"why":"The AR button, safety-zone, and status-panel setup that ARTHUR replicates as Scenario 1.","marker":"[7]"},{"why":"Introduces hybrid user interfaces, the concept ARTHUR applies by combining desktop/tablet and head-mounted displays.","marker":"[43]"},{"why":"Shows asynchronous switching between a 2D desktop interface and an immersive view, the pattern ARTHUR adapts for HRC authoring.","marker":"[52]"},{"why":"Documents the cognitive overhead of switching displays, the risk ARTHUR's design tries to mitigate by keeping switching fluid.","marker":"[55]"},{"why":"Provides the virtual design element taxonomy that informed which visual feedback components ARTHUR implements.","marker":"[38]"},{"why":"Supplies the usage-evaluation strategy for toolkit research that the five-participant expert study follows.","marker":"[63]"}],"fun_headline_variants":["ARTHUR: one tool for AR robot collab with hybrid UI","Author AR robot collab on PC, tablet, then tweak in AR","20 feedback, 10 actions, 18 conditions: ARTHUR for AR collab authoring","Hybrid authoring breaks AR robot collab design loop"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the hand-selected set of 20 feedback types, 10 actions, and 18 conditions is rich enough to author a broad range of real human-robot collaboration scenarios, even though coverage was only demonstrated on three prior setups and not measured systematically.","fun_headline_variants_meta":{"raw":{"variants":["ARTHUR: one tool for AR robot collab with hybrid UI","Author AR robot collab on PC, tablet, then tweak in AR","20 feedback, 10 actions, 18 conditions: ARTHUR for AR collab authoring","Hybrid authoring breaks AR robot collab design loop"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001119,"raw_usage":{"total_tokens":4642,"prompt_tokens":913,"completion_tokens":3729,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":3646}},"tokens_in":529,"tokens_out":3729,"duration_ms":25100,"temperature":1.0,"reasoning_tokens":3646,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:13:38.016857+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Author a realistic HRC scenario outside the three demonstrated (for example, a mobile robot navigating a shared warehouse) using only ARTHUR's shipped components; if the scenario requires a feedback, action, or condition that is not in the inventory and can only be added by writing a new component in code, the general-applicability claim fails. A lighter check is to map the component sets in published AR-HRI taxonomies onto ARTHUR's 20/10/18 inventory and count missing entries that correspond to common scenarios.","supporting_citations":[{"cited_title":"In: HRI’24 (2024)","cited_arxiv_id":null,"evidence_quote":"Prior in-situ authoring tool for robot visualizations; ARTHUR extends it with assembly instructions, user input, and a hybrid interface."},{"cited_title":"Robotics and Computer-Integrated Manufacturing 63, 101891 (2020) https://doi.org/10.1016/ j.rcim.2019.101891","cited_arxiv_id":null,"evidence_quote":"The AR button, safety-zone, and status-panel setup that ARTHUR replicates as Scenario 1."},{"cited_title":"In: Proceedings of the 4th Annual ACM Symposium on User Interface Software and Technology","cited_arxiv_id":null,"evidence_quote":"Introduces hybrid user interfaces, the concept ARTHUR applies by combining desktop/tablet and head-mounted displays."},{"cited_title":"In: ISS’21 Workshop Proceedings: ”Transitional Interfaces in Mixed and Cross-Reality: A new frontier?” (2021)","cited_arxiv_id":null,"evidence_quote":"Shows asynchronous switching between a 2D desktop interface and an immersive view, the pattern ARTHUR adapts for HRC authoring."},{"cited_title":"https:// mosquitto.org","cited_arxiv_id":null,"evidence_quote":"Supplies the usage-evaluation strategy for toolkit research that the five-participant expert study follows."}],"review_version":1}