{"id":"c6f121d1-04ef-4c35-84fd-35508e98d0db","arxiv_id":"2504.21563","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A phase-adaptive teleoperation GUI, showing fewer elements during autonomous driving, scored higher on usability and task time than a static GUI in a click-dummy study.","lead":"This paper designs and tests a screen interface for people who remotely operate or assist driverless cars. A version that shows fewer details while the car is driving itself performed better in a 36-person click-through test, though the test order may explain part of the difference.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fixed presentation order confounds the static-vs-dynamic comparison: the dynamic GUI was always evaluated second, so the reported SUS and task-time advantages may reflect practice rather than the interface itself.","rationale":"The reader correctly identifies the click-dummy evaluation as a key limitation and notes the fixed-order design in the rationale. I agree that the empirical claims are not fully established, but I see the fixed presentation order as the single most load-bearing concern because it confounds the central static-vs-dynamic comparison even before considering whether the click-dummy generalizes to real remote assistance. The paper itself acknowledges this in Section 4.3, yet the abstract retains a causal formulation. The expert-interview process and element catalogue are genuinely useful contributions, and the card-sorting results are informative descriptively, but they do not rescue the headline comparative claim. A counterbalanced or between-subjects replication is a feasible and decisive check. Since the reader already assigned CONDITIONAL, my analysis does not require a different verdict; it sharpens the reason for that condition. I therefore recommend keeping the conditional verdict rather than accepting the abstract's causal claim as supported by the current data.","tokens_in":18461,"tokens_out":3063,"duration_ms":36724,"concrete_test":"Run a preregistered replication with N = 36 where half the participants receive the dynamic GUI first and half receive the static GUI first, or use a between-subjects design. The critical analysis is to compare only first-trial data: SUS and task completion time for participants who first saw the static GUI versus those who first saw the dynamic GUI. If the dynamic advantage disappears or reverses in this first-exposure comparison, the original effects are attributable to learning or familiarity. If the dynamic advantage survives the order-balanced design, the fixed-order limitation is resolved and the headline claim becomes credible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim, stated in the abstract as \"the dynamic GUI significantly outperforming the static version in both usability and task completion time,\" is threatened by a within-subject order confound. In Section 2.2, Task I and Task II, every participant first interacted with the static GUI and then with the dynamic GUI, using the same scenario. Section 4.3 explicitly acknowledges this: \"This result may be influenced by the presentation order, as participants were always shown the static GUI first, followed by the dynamic version\" and \"the faster task completion times observed with the dynamic GUI may be attributed to the fixed order of presentation and potential learning effects.\" The abstract nevertheless presents the comparison as a causal advantage of the dynamic GUI. Because GUI version and temporal position are perfectly confounded, the significant SUS and task-completion differences cannot be attributed to the information-density manipulation: any second-trial benefit from task familiarization, UI practice, or reduced novelty would produce the same pattern. The observation that one-third of participants reported no difference between the variants is consistent with practice effects rather than a perceived interface improvement. This concern is more load-bearing than the click-dummy proxy issue because it threatens internal validity even under the study's own assumptions; the click-dummy limitation affects generalization, whereas the order confound affects whether the headline comparison is interpretable at all.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper follows a user-centered design process to develop a Graphical User Interface (GUI) for teleoperation of automated vehicles, covering both Remote Driving and Remote Assistance. Nine teleoperation experts were interviewed to define teleoperation process steps and to rate the relevance of 57 informational elements derived from a literature review. The resulting GUI was implemented as a click-dummy in a static variant and a dynamic variant that adapts the displayed elements according to the teleoperation phase, and was evaluated in an online study with N=36 participants. The authors report that the dynamic GUI achieved significantly higher System Usability Scale (SUS) ratings and faster task completion times than the static GUI, with no significant overall UEQ difference. A card-sorting task compared participants' ratings of informational elements with the experts' ratings. The paper concludes that an adaptive, phase-dependent display is beneficial and recommends follow-up evaluation with a real vehicle.","tokens_in":18707,"tokens_out":3228,"duration_ms":35976,"significance":"The work offers a structured method for deriving teleoperation GUI content from expert-defined process steps and provides empirical heat maps of informational-element relevance that can inform future interface designs. The explicit comparison of static versus dynamic information presentation directly addresses the research question of when supplemental information should be displayed. If the reported dynamic-vs-static advantage were reliable, it would support phase-adaptive GUI design. However, the central comparison is undermined by a fixed presentation order that the authors themselves acknowledge in Section 4.3, so the main empirical claim is currently not established. The paper is honest about its limitations, which is a strength, but the abstract and conclusions still present the confounded results as definitive.","major_comments":[{"comment":"The within-subject design always presents the static GUI first and the dynamic GUI second, perfectly confounding GUI variant with trial order. As Section 4.3 notes, the significantly higher SUS and faster task completion for the dynamic GUI could be entirely due to practice or familiarization effects. The abstract's claim that the dynamic GUI 'significantly outperforms' the static version in usability and task completion time is therefore not supported by the data as analyzed. Please reframe the abstract, results, and conclusions to present these outcomes as fixed-order pilot results, or alternatively report an analysis that accounts for learning effects, and temper the causal language.","section":"Section 2.2, Table 5, Abstract"},{"comment":"The evaluation uses a click-dummy with still photographs, non-professional participants (N=36), and a simplified waypoint-clicking task that omits real vehicle interaction, communication channels, and dynamic traffic. The authors themselves note in Section 4.3 that many informational elements were not needed in the click task and that usability scores are influenced by the prototype interaction. This does not invalidate the study as an early-stage interface screening, but it does mean that the informational-element ratings and the perceived GUI differences cannot be generalized to real remote assistance workstations. The paper should explicitly restrict its claims to click-dummy evaluation and avoid implying that the evaluated GUI is ready for operational use.","section":"Section 2.2 Online Study, Section 4.3"},{"comment":"The expert sample of N=9 was randomly split into Remote Driving (n=5) and Remote Assistance (n=4) groups. As acknowledged in Section 4.2, this random assignment may have produced sector imbalances that confound the comparison of the two teleoperation concepts. Because the heat-map differences between Remote Driving and Remote Assistance are said to be small (0.22 on average), the reader should be able to assess this risk. Please report the experts' professional domains per group and discuss whether any observed differences are plausibly due to domain expertise rather than the teleoperation concept itself.","section":"Section 2.1, Section 4.2, Table 4"}],"minor_comments":[{"comment":"The test name is misspelled as 'Saphiro-Wilk'; it should be 'Shapiro-Wilk'.","section":"Section 3.3"},{"comment":"The text 'Vise versa' should be 'Vice versa'.","section":"Section 3.2"},{"comment":"The row label 'Fastended Seat Belts' should be 'Fastened Seat Belts'.","section":"Table 4"},{"comment":"The text 'Intervene transistion if something goes wrong' contains a typo: 'transistion' should be 'transition'.","section":"Table 3"},{"comment":"The study reports multiple significance tests (SUS, task time, UEQ overall, pragmatic, hedonic) without correction for multiple comparisons; given the small sample and the exploratory nature, at least a footnote acknowledging this would be appropriate.","section":"Section 3.3, Table 5"},{"comment":"The manuscript references the two GUI variants but does not include the actual images in the provided text; please ensure the final version includes clear, high-resolution figures, and consider annotating which elements are shown or hidden in the dynamic variant.","section":"Figures 2 and 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is squarely within the scope of a human-factors/automotive UI venue and presents a systematic, reproducible design process. The main issue is the order confound in the core comparison; because the authors already acknowledge it, the necessary revision is primarily about reframing the claims and possibly adding a counterbalanced follow-up or re-analysis. The expert-sample imbalance is a secondary concern that can be addressed in the text. I believe the paper can become acceptable after these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, honestly-written HCI design study that gives you a useful catalogue of 57 informational elements for AV teleoperation, a process model for Remote Driving vs Remote Assistance, and a first pass at a phase-adaptive GUI. The empirical headline—dynamic GUI beats static on SUS and task time—is real in the data but not interpretable as a causal advantage because every participant saw static first, dynamic second. The authors know this; they say so plainly in Section 4.3. The abstract still states the comparison as if it were a win, which is the paper's biggest overreach.\n\nWhat's genuinely new: no one has systematically worked out which informational elements are needed at which phase of Remote Assistance. The expert interview heat maps (Table 4) and the expert-vs-participant card sorting (Table 6) are valuable. The finding that maps, routes, and trip info matter more for assistance while traffic-sign and object highlighting matter more for driving is plausible and useful. The GUI itself, with the hood-dashboard idea and phase-dependent element reduction, is a concrete design pattern others can test.\n\nSoft spots, in order of severity:\n\n1. The order confound. Static always first, dynamic always second, same scenario. SUS and time improvements could just be practice. The effect sizes (r≈.66-.69) are large, but so is the learning effect from repeating the same task. One-third of participants saw no difference between the versions, which undercuts the idea that the manipulation was obviously meaningful. This is not a minor point: it undermines the internal validity of the headline claim.\n\n2. The click-dummy proxy. 36 non-professionals clicking waypoints on photos is a long way from handling a real AV request. The authors acknowledge this. It limits generalization, but it's a lesser sin because the study is explicitly a first evaluation.\n\n3. Small expert sample (N=9) randomly split into two groups. The authors note the sector imbalance risk. That's fair, but it's another reason to treat the interview results as suggestive, not definitive.\n\nThe paper is not deceptive: the limitations are right there. But the abstract should not have led with the causal comparison. As a reviewer, I'd ask for a counterbalanced or between-subjects replication before accepting the claims; the process model and element set can stand.\n\nWho this is for: anyone working on teleoperation interfaces, especially for remote assistance. The tables alone are worth the download. I'd bring it to reading group and I'd probably cite the element catalogue.\n\nRecommendation: send to peer review. The work is important to the subfield and the flaws are fixable; desk rejection would be a waste.","headline":"Useful element catalogue and process model for teleoperation GUIs, but the static-vs-dynamic comparison is confounded by fixed presentation order and should be read as a pilot, not proof.","tokens_in":19204,"tokens_out":2385,"would_cite":true,"duration_ms":23357,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A teleoperation GUI that adapts to the task beats a static full-information interface, with significantly higher usability scores and faster task completion in a click-dummy study of remote assistance.","keywords":["teleoperation","automated vehicles","graphical user interface","situational awareness","remote assistance","remote driving","usability evaluation","phase-adaptive interface"],"falsifier":"Run the static and dynamic GUIs with trained remote operators controlling a real vehicle (or a high-fidelity simulator with live video and realistic latency); if the dynamic GUI no longer yields higher SUS scores and faster task completion, the central advantage claim is refuted.","tokens_in":18306,"feed_emoji":"🚗","tokens_out":5301,"duration_ms":52861,"temperature":0.7,"pith_summary":"This paper develops a graphical user interface for remotely operating and assisting automated vehicles, and tests whether showing fewer information elements during autonomous driving improves the interface. The authors first map the teleoperation process through interviews with nine experts, who rated the relevance of 57 informational elements across five process phases. From those ratings they built two click-dummy GUIs: a static version that always shows the full information set, and a dynamic version that hides low-relevance elements while the vehicle drives itself. In an online study with 36 participants, the dynamic GUI scored significantly higher on the System Usability Scale and produced significantly faster waypoint-based task completion than the static GUI. The result suggests that phase-adaptive information display is a workable design principle for teleoperation interfaces, with follow-up testing in a real vehicle still needed.","feed_headline":"Adaptive teleoperation GUI outperforms static interface","feed_subtitle":"In a click-dummy study, dynamic display scored 76.6 vs 68.5 on usability and cut task time by 16 percent.","key_machinery":"The load-bearing object is the 'dynamic GUI': a phase-adaptive interface that changes which supplemental informational elements are displayed as the process moves through five teleoperation phases (autonomous driving, transition to teleoperation, teleoperation, transition back, autonomous driving). During autonomous monitoring it shows a reduced set of elements rated as merely nice-to-have; during active teleoperation it shows the full set rated as necessary. The element set itself comes from a morphological box of 57 informational categories, filtered by expert ratings on a 0-2 necessity scale, and implemented as an interactive click-dummy in which participants advance a simulated vehicle by clicking waypoints on still photos.","core_discovery":"The central claim is that a teleoperation GUI which adapts its displayed information to the current phase of the human-vehicle interaction—showing a rich set of elements only during active teleoperation and a reduced set while the vehicle drives autonomously—is more usable and more efficient than a static GUI that keeps the full set visible at all times. Evidence comes from a click-dummy study of Remote Assistance: the dynamic GUI received a mean System Usability Scale (SUS) score of 76.6 versus 68.5 for the static GUI (Wilcoxon z = -4.11, p = .00004, r = .69), and median task completion time dropped from 61.73 s to 41.25 s (Wilcoxon z = -3.97, p = .00007, r = .66), a 16% reduction in mean time. The paper also reports that expert-identified information requirements differ between Remote Driving and Remote Assistance, with map, route, and trip information valued more for assistance and traffic-sign and object highlighting valued more for driving.","pith_inferences":["A counterbalanced or randomized presentation would separate the genuine effect of phase-adaptive display from learning effects, because the static GUI was always shown first.","The same phase-adaptation principle could extend to other human-supervision contexts, such as fleet monitoring, telemedicine, or remote operation of construction and mining equipment, where operators switch between passive monitoring and active control.","The disagreement between experts and lay participants (participants wanted more vehicle parameters; experts valued latency, network quality, map, and control mode) suggests that personalized or role-adaptive element sets may be worth testing.","Testing under real network latency and with live video would reveal whether the reduced GUI still alerts operators to state changes quickly enough; missing such changes could offset the measured speed and usability gains."],"forward_implications":["Teleoperation interfaces for automated vehicles can be designed around a phase model rather than a single always-on display.","Reducing information during autonomous monitoring phases can improve both perceived usability and objective task speed.","Information requirements for Remote Assistance and Remote Driving overlap substantially, but their relative priorities differ (map/route/trip vs traffic-sign/object highlighting).","The click-dummy results support further investment in real-vehicle studies of dynamic GUIs, since hedonic quality remained neutral."],"supporting_citations":[{"why":"Supplies the 80 general information requirements (reduced to 20 essential elements) that the morphological box of 57 elements builds on.","marker":"[20]"},{"why":"Provides a prior comparison of full versus reduced display configurations during orientation and navigation phases, motivating the static versus dynamic comparison here.","marker":"[21]"},{"why":"Reports a click-dummy evaluation of a remote-assistance control-center GUI, providing the baseline methodology for click-dummy testing.","marker":"[24]"},{"why":"Offers a recent command-based teleoperation GUI online study used as a comparison point for click-dummy evaluation of Remote Assistance.","marker":"[26]"},{"why":"Defines the SUS adjective scale ('ok', 'good', 'acceptable') used to interpret the usability ratings.","marker":"[37]"},{"why":"Provides the User Experience Questionnaire used to measure pragmatic and hedonic quality.","marker":"[38]"},{"why":"Defines the Waypoint Guidance interaction concept that the click-dummy scenario implements.","marker":"[39]"},{"why":"Provides the effect-size conventions used to interpret the reported r values.","marker":"[40]"}],"fun_headline_variants":["Dynamic GUI outperforms static for remote driving","Phase-aware GUI improves teleoperation efficiency","Adaptive info display: 16% faster teleop tasks","Contextual GUI boosts usability in remote assistance","Teleop GUI that adapts beats one-size-fits-all"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The click-dummy, in which 36 non-professional participants advance through still photos by clicking waypoints, is a valid enough proxy for real remote assistance that usability and timing measurements transfer to actual operators.","fun_headline_variants_meta":{"raw":{"variants":["Dynamic GUI outperforms static for remote driving","Phase-aware GUI improves teleoperation efficiency","Adaptive info display: 16% faster teleop tasks","Contextual GUI boosts usability in remote assistance","Teleop GUI that adapts beats one-size-fits-all"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000282,"raw_usage":{"total_tokens":1687,"prompt_tokens":981,"completion_tokens":706,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":633}},"tokens_in":597,"tokens_out":706,"duration_ms":7721,"temperature":1.0,"reasoning_tokens":633,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:59:12.781039+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the static and dynamic GUIs with trained remote operators controlling a real vehicle (or a high-fidelity simulator with live video and realistic latency); if the dynamic GUI no longer yields higher SUS scores and faster task completion, the central advantage claim is refuted.","supporting_citations":[{"cited_title":"Effective remote automated vehicle operation: a mixed reality contextual comparison study","cited_arxiv_id":null,"evidence_quote":"Provides a prior comparison of full versus reduced display configurations during orientation and navigation phases, motivating the static versus dynamic comparison here."},{"cited_title":"Guiding, not Driving: Design and Evaluation of a Command-Based User Interface for Teleoperation of Autonomous Vehicles","cited_arxiv_id":"2502.00750","evidence_quote":"Offers a recent command-based teleoperation GUI online study used as a comparison point for click-dummy evaluation of Remote Assistance."},{"cited_title":"Determining what individual SUS scores mean: adding an adjective rating scale","cited_arxiv_id":null,"evidence_quote":"Defines the SUS adjective scale ('ok', 'good', 'acceptable') used to interpret the usability ratings."},{"cited_title":"User experience questionnaire","cited_arxiv_id":null,"evidence_quote":"Provides the User Experience Questionnaire used to measure pragmatic and hedonic quality."},{"cited_title":"A power primer","cited_arxiv_id":null,"evidence_quote":"Provides the effect-size conventions used to interpret the reported r values."}],"review_version":1}