{"id":"758d1ad6-cec6-4b27-b9f3-26e25392d05f","arxiv_id":"2608.11401","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In a haptic point-to-point study, a shared controller that preserves task-irrelevant movement variability produced comparable task performance but significantly higher perceived usability than a variability-constraining controller.","lead":"This paper tested whether a robot assistant that preserves a person's natural movement variability feels better to use than one that smooths movements out, without hurting task accuracy. In a 41-person haptic pointing study, participants rated the variability-preserving mode higher on usability, especially ease of use, while task performance matched the conventional assistant.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The usability benefit attributed to preserved variability is confounded with a lower feedback-gain profile in highVar; an assistance-matched control is needed to sustain the causal claim.","rationale":"The reader's weakest_assumption correctly identifies the confound between variability preservation and assistance strength. This is the single most load-bearing issue because the paper's central contribution is a causal interpretation: preserving task-irrelevant variability enhances perceived usability. If the independent variable is actually variability plus lighter assistance, the experiment does not separate the proposed mechanism from a mundane preference for weaker force. The design is otherwise reasonable: within-subject, neutral labels, 41 participants, a significant manipulation check for H1, and multiple experience measures. The selective effect on usability subscales (ease of use, not perceived usefulness) is interesting but itself consistent with the force-profile explanation, since lower forces can feel easier. The non-significant performance comparisons would need formal equivalence bounds before 'maintaining task performance' is accepted, but the gain confound is the more direct threat to the causal claim. I therefore concur with the reader's CONDITIONAL verdict; the additional control condition or mediation analysis is necessary, not optional, to support the stated causal claim.","tokens_in":19261,"tokens_out":3354,"duration_ms":31138,"concrete_test":"Add a fourth within-subject condition (highVar-matched) that uses the lowVar gain matrix Lp=Ld=150 throughout but injects zero-mean lateral force noise calibrated to reproduce the highVar maximum variance (Vmax around 0.65e-4 m^2) during the mid-movement phase. If usability and ease-of-use ratings in this condition remain significantly higher than in lowVar and comparable to highVar, the variability attribution survives; if they match lowVar, the reported effect is driven by feedback-gain magnitude and scheduling. As a cheaper first pass, re-analyze the existing data with a within-subject mediation model predicting perceived ease of use from time-averaged automation force magnitude and Vmax jointly; if the force term absorbs the highVar advantage, the confound is present.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim—that preserving task-irrelevant variability improves perceived usability—requires that highVar and lowVar differ only in variability treatment while providing matched task-relevant assistance. Section 3.2.3 (Eqs. 5–6, Fig. 2) shows they do not: lowVar uses constant gains Lp=Ld=150, whereas highVar uses Lp=75 and Ld=20 over the first two-thirds of the movement, ramping only to Lp=100 and Ld=150 within the final third. Thus highVar applies substantially lower assistive and damping force for most of each trial. Section 6.1 further concedes that the highVar parametrization deliberately included a damping component to ensure comparable performance, so the manipulation is two-dimensional: variability structure and feedback-gain magnitude and scheduling are covaried. The observed usability advantage (p=.004, d_z=.76), driven mainly by perceived ease of use (p<.001, d_z=1.00), could plausibly reflect the lighter, less forceful assistance rather than preserved variability per se. The authors' reliance on non-significant performance differences (Section 5.1) to infer equivalence also needs an equivalence margin, but that is secondary; even if performance is matched, the force-profile confound remains. Without an assistance-matched or gain-matched control condition, the causal attribution to variability structure is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a within-subject haptic interaction experiment comparing three conditions: unsupported movement (noSup), a variability-constraining shared controller (lowVar), and a variability-respecting shared controller (highVar). The authors report that highVar preserves significantly more task-irrelevant positional variance than lowVar, that the two assisted modes do not differ significantly on settling time, endpoint error, or endpoint variance, and that highVar yields significantly higher usability ratings, driven mainly by perceived ease of use. The conclusions assert a causal link between variability structure and interaction experience in physical human-machine interaction.","tokens_in":19535,"tokens_out":3679,"duration_ms":33648,"significance":"If the central causal claim were established, the work would be a meaningful contribution to shared-control and haptic-assistance design, since it would show that preserving task-irrelevant movement variability can improve perceived usability without a measurable performance cost. The study is carefully reported in many respects: the within-subject design is appropriate, the variance measures are operationally defined, several validated questionnaires are used, and effect sizes are reported alongside p-values. The main difficulty is that the highVar and lowVar manipulations differ along two dimensions at once, so the paper's headline attribution of the usability benefit to variability structure is not currently identified.","major_comments":[{"comment":"The highVar and lowVar modes do not differ only in how they treat task-irrelevant variability; they also differ substantially in feedback-gain magnitude and scheduling. lowVar uses constant gains Lp=Ld=150, whereas highVar begins at Lp=75 and Ld=20 and increases to Lp=100 and Ld=150 only once the position error falls below one third of the start distance. Over most of each movement, highVar therefore supplies considerably less assistive and damping force. Section 6.1 additionally concedes that the highVar parametrization deliberately included a damping component to ensure comparable performance. The observed usability advantage (p=.004, dz=.76; perceived ease of use p<.001, dz=1.00) could plausibly reflect the lighter, less forceful assistance profile rather than preserved variability per se. An assistance-matched or gain-matched control condition, or an analysis that equates the applied actuator force across modes, is required to support the causal attribution made in the title and conclusions.","section":"Section 3.2.3, Fig. 2"},{"comment":"The paper states that Hypothesis H2b is a non-inferiority hypothesis, but no non-inferiority analysis is performed. The conclusion that task-relevant performance 'does not differ' between highVar and lowVar is based only on non-significant pairwise tests (settling time p=.409, endpoint error p=.949, endpoint variance p=.866). There is no prespecified equivalence margin, no confidence interval for the differences, and no TOST or similar equivalence test. Absence of a significant difference is not evidence of equivalence, so later statements that performance was 'matched' or 'the same level' (Sections 6.1 and 6.2) overstate what the data establish.","section":"Section 5.1, Hypothesis H2b"},{"comment":"The sentence 'Critically, the same level of task performance between highVar and lowVar allows us to attribute differences in user experience specifically to the variability manipulation' conflates two issues. First, as noted in the previous comment, performance equivalence is not established. Second, even if performance were equivalent, the highVar and lowVar modes still differ in feedback-gain magnitude and scheduling, so the variability manipulation is not isolated. The causal attribution to variability structure alone is therefore not identified by the current experimental design.","section":"Section 6.1"}],"minor_comments":[{"comment":"The maximum-variance effect size for the lowVar-highVar comparison is reported as dz=-0.90 in Table 1 but as dz=0.98 in the text of Section 5.1; these values should be reconciled and the sign convention clarified.","section":"Table 1 vs Section 5.1"},{"comment":"The manuscript says each participant completed 82 movements per mode while also stating there were 12 repetitions of each of four key movement directions, which would total 48 movements. The relationship between these numbers should be clarified.","section":"Section 3.3 vs Section 4.2"},{"comment":"The reference list includes citations that appear unrelated to the physical-HMI context, such as [12] (GPTs are GPTs) and [30] (U.S. workers' AI exposure), where they are cited in support of claims about healthcare, rehabilitation, and skilled manual work. These citations should be checked and replaced or repositioned.","section":"References"},{"comment":"The limitation discussion would benefit from explicitly acknowledging the gain-scheduling confound described in Section 3.2.3, since the current text frames the highVar design choice only as a deliberate trade-off for performance matching.","section":"Section 6.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is well structured and experimentally careful, but the central causal claim is not identifiable from the current design because the variability manipulation is confounded with assistance strength and scheduling. I do not think rejection is warranted; the required control condition is a feasible additional experimental arm or re-analysis that equates assistance force across modes. Before acceptance, the authors should either add such a condition or substantially soften the causal language in the title, abstract, and conclusions. Also, the citation mismatches ([12], [30]) should be corrected before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Kai—\n\nQuick take: this is the first experimental run of the HVROC controller and the first real attempt to tie preserved task-irrelevant variability to user experience in a tightly coupled haptic task. The study is well put together: 41 participants, within-subject, three conditions, validated questionnaires, and the statistics are reported clearly enough to follow exactly what was done. The manipulation check works—highVar produced significantly more lateral variance than lowVar (p=.004). The pattern in the experience measures is also interesting: usability and perceived ease of use favored highVar, while agency, flow, and self-efficacy tracked performance rather than variability. That dissociation is a genuinely new observation.\n\nThe soft spot is the one you'd expect from the design. The two assisted modes don't differ only in variability treatment; they differ in feedback-gain magnitude and scheduling. lowVar holds Lp=Ld=150 throughout; highVar starts at Lp=75, Ld=20 and only ramps up in the last third (Eqs. 5–6, Fig. 2). So highVar applies considerably less assistive force for most of the movement. The paper's own Section 6.1 admits the highVar parametrization deliberately included a damping component to match performance. That means the usability benefit—especially the ease-of-use jump (p<.001)—could just as plausibly be an effect of a lighter, less forceful assistance profile as of preserved variability per se. The causal claim in the abstract and conclusions is therefore a step ahead of what the experiment isolates.\n\nTwo secondary issues: the performance-equivalence argument rests on non-significance alone, so an equivalence margin (or a Bayesian test) would make it solid. And the data are only available on request, which limits reproducibility, though the questionnaire items and controller parameters are described in enough detail that the setup could be replicated.\n\nHonestly, these are fixable. An assistance-matched control (e.g., lowVar with the same gain schedule but variance-suppressing objective) would settle the attribution, and the authors themselves acknowledge the damping choice in the limitations. The study deserves peer review; I'd send it out with a request to address the confound or to soften the causal wording. If you work on pHMI or shared control, this is worth having on your radar.\n\nBest","headline":"First empirical test of variability-preserving haptic control—worth reading for its clean within-subject design and large effects, but the headline causal claim is confounded by unevenly matched assistance strength.","tokens_in":20023,"tokens_out":3898,"would_cite":true,"duration_ms":34772,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Preserving variability in shared control raises perceived usability without hurting task performance.","keywords":["shared control","human-machine interaction","movement variability","user experience","sense of agency","haptic assistance","usability","optimal control"],"falsifier":"Run a follow-up with two controllers that have identical gain schedules and force magnitudes, with artificial lateral variability injected in only one; the central claim predicts the variable condition is rated easier to use, whereas a force-magnitude account predicts no difference.","tokens_in":19099,"feed_emoji":"🤖","tokens_out":5345,"duration_ms":44490,"temperature":0.7,"pith_summary":"This paper asks whether a shared controller that deliberately keeps the natural variability of human movement can feel better to use while still performing as well as a conventional one. In a haptic point-to-point task with 41 participants, the authors compared an unsupported condition, a variability-constraining assist mode, and a variability-preserving assist mode. They report that the variability-preserving mode produced significantly higher perceived usability than the constraining mode, with the difference driven mainly by perceived ease of use, while measures of settling time, endpoint error, and endpoint variance did not differ between the two assisted modes. The authors take this as evidence that variability structure itself shapes interaction experience independently of objective performance.","feed_headline":"Preserving variability makes shared control feel easier to use","feed_subtitle":"Users rate a variability-preserving assistant easier to use than a constraining one, with identical task performance.","key_machinery":"The machinery is Human-Variability-Respecting Optimal Control (HVROC), a controller that builds a stochastic model of human movement variability into the optimal feedback law. Here it is implemented with an error-dependent gain schedule: the variability-preserving mode keeps gains low over most of the movement and increases damping only near the target, while the variability-constraining mode applies constant high gains throughout. The operational measure of task-irrelevant variability is the maximum across-repetition positional variance perpendicular to the main movement direction. This design lets the authors manipulate variability structure while holding task-relevant performance approximately constant, which is what allows the usability difference to be attributed to variability rather than to performance.","core_discovery":"The central claim is that preserving task-irrelevant movement variability in a shared controller improves perceived usability without sacrificing objective task performance. In the reported experiment, usability ratings were significantly higher in the highVar condition than in the lowVar condition (p = .004, d_z = 0.76), and the effect appeared primarily in perceived ease of use, with emotional response showing a trend in the same direction. Meanwhile, the two assisted modes did not differ significantly in settling time, endpoint error, or endpoint variance, and the highVar mode did produce significantly more task-irrelevant variability, measured as maximum lateral positional variance. The authors interpret these results as establishing a causal link between variability structure and interaction experience in tightly coupled physical human-machine interaction.","pith_inferences":["The highVar mode applies considerably less assistive force over most of the movement (gains of 75/20 versus 150/150 until the final third), so the usability gain may owe to lighter, less intrusive assistance rather than to preserved variability itself; a clean test would match force magnitude while varying only lateral variability.","The same design logic could extend beyond lateral deviation to other redundant dimensions, such as grip-force variation or timing jitter, where users also form sensorimotor predictions about how the interaction should feel.","Personalizing the variability target to each user's own variability profile might strengthen the effect further, since the study used identical gain schedules for all participants.","If the effect truly stems from variability structure, comparable usability gains should appear in tasks that allow task-irrelevant variability without any assistive force, such as unassisted movement with visual or auditory feedback that reshapes variability; an absence of such gains would point to force-magnitude as the active ingredient."],"forward_implications":["Control policies that are equivalent in objective task performance can feel measurably different to users, so interaction experience should be treated as an explicit design objective, not a byproduct.","Variability-preserving assistance may support long-term engagement and acceptance in rehabilitation and assistive robotics, where sustained use matters as much as task success.","Preserving task-irrelevant variability does not inherently cost efficiency or accuracy, opening a design space where multiple controllers are functionally equivalent but experientially distinct.","Because the usability gain was driven by perceived ease of use, future controller tuning can target the felt naturalness of assistance rather than only error-based metrics.","The dissociation between perceived usefulness and ease of use suggests users judge usefulness by task success but judge ease of use by how assistance feels, so both dimensions should be measured separately in HMI evaluation."],"supporting_citations":[{"why":"Supplies the Human-Variability-Respecting Optimal Control method that the study applies and evaluates.","marker":"[29]"},{"why":"Provides the optimal feedback control theory distinguishing task-relevant from task-irrelevant variability, which motivates the manipulation.","marker":"[47]"},{"why":"Documents trajectory variability in point-to-point movements, grounding the construct of natural movement variability.","marker":"[1]"},{"why":"Establishes signal-dependent noise as a source of stochastic motor behavior, underpinning the random component the controller is designed to respect.","marker":"[21]"},{"why":"Provides the QUEAD usability questionnaire used to measure the central experience outcome.","marker":"[42]"},{"why":"Supplies the sense of agency scale used to assess perceived authorship under the different control modes.","marker":"[45]"},{"why":"Gives the German validation of the sense of agency scale administered to participants.","marker":"[3]"},{"why":"Supplies the self-efficacy questionnaire used to measure perceived competence under assistance.","marker":"[23]"},{"why":"Provides the Flow Short Scale used to measure engagement and challenge-skill balance across modes.","marker":"[39]"}],"fun_headline_variants":["Variability-aware shared control improves perceived usability","Preserving natural movement variability improves usability in HMI","Shared control that keeps human variability feels more usable","Human-variability-aware control enhances interaction experience","Variability in shared control: better usability, same performance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study treats the two assisted modes as differing only in how they handle task-irrelevant variability, but the variability-preserving mode also uses much lower feedback gains over most of the movement, so the usability gain could stem from a lighter assistive force rather than from preserved variability itself.","fun_headline_variants_meta":{"raw":{"variants":["Variability-aware shared control improves perceived usability","Preserving natural movement variability improves usability in HMI","Shared control that keeps human variability feels more usable","Human-variability-aware control enhances interaction experience","Variability in shared control: better usability, same performance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000548,"raw_usage":{"total_tokens":2564,"prompt_tokens":839,"completion_tokens":1725,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":455,"completion_tokens_details":{"reasoning_tokens":1652}},"tokens_in":455,"tokens_out":1725,"duration_ms":21653,"temperature":1.0,"reasoning_tokens":1652,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:11:48.063390+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a follow-up with two controllers that have identical gain schedules and force magnitudes, with artificial lateral variability injected in only one; the central claim predicts the variable condition is rated easier to use, whereas a force-magnitude account predicts no difference.","supporting_citations":[{"cited_title":"Human-Variability-Respecting Optimal Control for Physical Human- MachineInteraction,in:202433rdIEEEInternationalConferenceonRobotandHumanInteractiveCommunication(ROMAN),LosAngeles","cited_arxiv_id":null,"evidence_quote":"Supplies the Human-Variability-Respecting Optimal Control method that the study applies and evaluates."},{"cited_title":"Human arm trajectory formation","cited_arxiv_id":null,"evidence_quote":"Documents trajectory variability in point-to-point movements, grounding the construct of natural movement variability."},{"cited_title":"2017IEEEInternationalConferenceonSystems,Man,andCybernetics (SMC) doi:10.1109/SMC.2017.8122720","cited_arxiv_id":null,"evidence_quote":"Provides the QUEAD usability questionnaire used to measure the central experience outcome."},{"cited_title":"The sense of agency scale: A measure of consciously perceived control over one’s mind, body, and the immediate environment","cited_arxiv_id":null,"evidence_quote":"Supplies the sense of agency scale used to assess perceived authorship under the different control modes."},{"cited_title":"A German translation and validation of the sense of agency scale","cited_arxiv_id":null,"evidence_quote":"Gives the German validation of the sense of agency scale administered to participants."},{"cited_title":"User control and its many facets: A study of perceived control in human-computer interaction","cited_arxiv_id":null,"evidence_quote":"Supplies the self-efficacy questionnaire used to measure perceived competence under assistance."},{"cited_title":"Die Erfassung des Flow-Erlebens","cited_arxiv_id":null,"evidence_quote":"Provides the Flow Short Scale used to measure engagement and challenge-skill balance across modes."}],"review_version":1}