{"id":"6974d7aa-24c4-42b0-b612-79722ef6ae0f","arxiv_id":"2608.05723","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An upper-limb exoskeleton framework generates muscle-based torque references with reinforcement learning, refines them online, and delivers them via a passivity-preserving torque controller, with a pilot EMG study showing reduced muscle activity.","lead":"This paper builds an upper-limb exoskeleton controller that learns assistance torques from a simulated arm with muscles, then refines and delivers them while preserving interaction safety. It matters because it targets varied, non-repeating arm movements, and a small EMG pilot reports up to 48% lower target-muscle activity than moving without the exoskeleton.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's tracking guarantee depends on unmeasured environment passivity (Eq. 53), which the tank-replenishment experiments appear to violate; 'resumes tracking' is thus empirical, not proven.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing issue: Theorem 1's asymptotic stability and torque-tracking conclusion rely on Eq. (53), an unmeasured and unenforced passivity assumption about the human/environment. I agree with that reading. The concern is genuinely load-bearing because the paper's contribution includes a 'rigorous theoretical analysis' that guarantees torque tracking under passivity, and the proof does not hold if the wearer actively injects energy. The passivity inequality itself is established from the tank construction and is not the problem; the safety claim is therefore largely intact. The empirical torque-tracking results (RMSE 0.14-0.27 Nm on real joints) and the energy-tank behavior are valuable independent support, and the paper's Limitations section acknowledges that assistance restoration depends on wearer interaction. These are addressable with additional modeling or explicit relabeling of the guarantee, so the appropriate verdict remains CONDITIONAL rather than REJECT or ACCEPT. I also noted a possible gap in Proposition 1's proof regarding the boundedness of M(q) qdd_tilde, but the environment-passivity assumption is the more direct and more consequential threat to the central claim, so I did not elevate the Proposition issue to headline status.","tokens_in":22494,"tokens_out":14011,"duration_ms":122037,"concrete_test":"","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central theoretical claim is that the interaction-torque controller guarantees accurate torque tracking while preserving passivity (Theorem 1). The passivity inequality Hdot <= qdot^T tau_e is derived from Table III and appears internally consistent; the load-bearing problem is the stability/tracking part of Theorem 1. Its proof invokes Eq. (53), Hdot_env(x_env) <= -qdot^T tau_e, justified only as 'dissipation can always be presumed in the environment.' This is a behavioral assumption about the human, not a property of the robot, and it is never measured or enforced. The experiments appear to violate it: in the tank-replenishment phases (green regions in Figs. 14-15), the wearer intentionally moves the joint so that the interaction torque is aligned with joint rotation, i.e., qdot^T tau_e > 0, doing positive work on the exoskeleton. Under (53), that would require the human arm to be dissipating mechanical energy at that same moment. If the wearer is instead an active energy source, Vdot in (52) is not negative, and Theorem 1 does not establish asymptotic stability or torque-tracking convergence. The observed 'resumes tracking after tank replenishment' is therefore an empirical result, not a theoretical guarantee. This does not invalidate the passivity-based safety claim, but it does undercut the 'rigorous theoretical analysis guarantees closed-loop torque tracking' contribution as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ATP, a three-stage framework for upper-limb exoskeleton assistance: a reinforcement-learned musculoskeletal muscle controller that generates anatomical reference torques from a MuJoCo/MyoArm model, an online torque-refinement module that smooths the reference and incorporates a learned anomaly score, and an interaction-torque controller with energy-tank-based passivity preservation for a cable-driven SEA exoskeleton. The authors report simulation results, real-hardware torque-tracking experiments on two direct-drive shoulder joints and one cable-driven elbow joint, and an EMG study with five participants showing reduced target-muscle activity during static and dynamic tasks, with reductions up to 48% relative to unassisted movement. The central claims are that ATP generalizes to complex nonperiodic upper-limb movements without predefined trajectories and that the interaction-torque controller provides rigorous closed-loop torque-tracking guarantees subject to passivity.","tokens_in":22799,"tokens_out":7003,"duration_ms":65295,"significance":"If the results hold, the paper would make a useful contribution by combining learned biomechanical torque references with passivity-based interaction control for upper-limb exoskeletons. The open-sourced GPU-accelerated musculoskeletal training environment, the real-world torque-tracking RMSEs of 0.14-0.27 Nm, the explicit passivity demonstrations during tank switching, and the ablation of anomaly-guided torque refinement are concrete strengths that go beyond a purely simulation-based study. However, the paper's theoretical torque-tracking guarantee currently depends on an unvalidated environment-passivity assumption, and the EMG evaluation lacks inferential statistics. These issues do not invalidate the empirical results, but they require the theoretical claims to be restated or supported with additional evidence.","major_comments":[{"comment":"The proof of Theorem 1 relies on Eq. (53), which asserts that the human/environment is dissipative: Hdot_env(x_env) <= -qdot^T tau_e, justified only by the statement 'Dissipation can always be presumed in the environment.' This is a behavioral assumption about the wearer, not a property of the robot controller, and it is not measured or enforced in the experiments. Indeed, the tank-replenishment phases in Figs. 14 and 15 (green shaded regions) show the wearer moving the joint so that the interaction torque is aligned with the joint rotation, i.e., qdot^T tau_e > 0, which means the wearer is injecting energy into the exoskeleton. Under Eq. (53), such an interaction would require the human arm to be dissipating mechanical energy at the same moment. If the wearer is instead an active energy source, Vdot in Eq. (52) is not necessarily negative, so Theorem 1 does not establish asymptotic stability or torque-tracking convergence. The observed 'resumes tracking after tank replenishment' is therefore an empirical result, not a theoretical guarantee. The authors should either measure or enforce environment passivity, or explicitly restrict the claim to passive environments and soften the abstract/introduction statement that 'rigorous theoretical analysis guarantees closed-loop torque tracking subject to passivity.'","section":"Section V, Theorem 1 and Eq. (53)"},{"comment":"The proof of Proposition 1 asserts, immediately before Eq. (23), that 'By choosing a sufficiently large beta_q and a correspondingly larger K_q, there exists a finite mu such that lim sup_{t->inf} ||M(q) qddot~|| <= mu whenever ||tau_tilde_f|| <= zeta.' This bound on ||M(q) qddot~|| is load-bearing for the subsequent uniform ultimate boundedness argument in Eqs. (24)-(27), but no derivation is provided. The bound should follow from the observer error dynamics (22) with explicit conditions on beta_q, K_q, and the boundedness properties of M(q), C(qdot,q), and g(q); as written, the proof leaves a gap between the stability of qtilde and the required bound on the second derivative. Please supply this derivation or state the additional assumptions needed.","section":"Section V, Proposition 1"},{"comment":"The EMG study is presented as a central validation of the assistance benefit, but no inferential statistics are reported. The text says the data were 'statistically analyzed across all trials,' yet the results in Fig. 19 are summarized only as mean reductions (e.g., 64%, 45%, 59%, 33%, 48%, 47%) for a cohort of five participants. Because the static-task results are very similar between ATP and gravity compensation, and the dynamic-task effects vary across muscles and tasks, the claim that ATP reduces target-muscle activity compared with the baselines needs repeated-measures statistical analysis (or at least confidence intervals and effect sizes) to be supported. Without such analysis, the reported percentages may not be robust to individual-subject variability.","section":"Section VII-C, Fig. 19"}],"minor_comments":[{"comment":"The anomaly-score dynamics s(t+1) = s(t) + (partial f_a / partial tau_e)^T u_d Delta t assume that the diffusion-model-based anomaly score is differentiable with respect to tau_e; please clarify whether this gradient is computed analytically or numerically, and whether s is treated as a scalar throughout.","section":"Section IV, Eq. (16)"},{"comment":"The tank dynamics contain divisions by t_f and t_o. If the tanks can reach exactly zero energy, the right-hand sides are singular. The definitions of L_{f,o} and delta_{f,o} suggest the tanks are prevented from reaching zero, but this should be stated explicitly, together with the initialization of t_f and t_o.","section":"Section V, Eqs. (41) and (44)"},{"comment":"The entries for Lafan1 and Ours have no standard deviations, while the other datasets do; the meaning of the 'Ratio' column (e.g., 1.55 (64.5%)) should be explained more clearly, since the percentage appears to be a performance-retention measure rather than a ratio.","section":"Table V"},{"comment":"The ethics approval text contains the placeholder 'XXX' ('approved by the ethics committee of XXX'); the institution should be named.","section":"Section VI"},{"comment":"The source code is mentioned as 'available at ATP_muscle_controller,' but no URL or repository identifier is given; please provide a complete reference.","section":"Abstract and Section VIII"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid integration of learned biomechanical references with passivity-based control, and the empirical results are credible. The main reason for major revision is that Theorem 1's tracking guarantee depends on an unvalidated environment-passivity assumption that the paper's own tank-replenishment experiments appear to violate. This is fixable by restating the theorem with explicit 'if the environment is passive' conditions and by being more careful in the abstract/introduction about what is proven versus empirically demonstrated. The Proposition 1 bound on ||M(q) qddot~|| also needs a proper derivation. The novelty claim about being 'the first' framework should be double-checked against the cited torque-profile and passivity-based SEA literature, though no obvious misconduct is apparent."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely integrative exoskeleton paper with real hardware results, but the theoretical tracking guarantee leans on an unmeasured passivity assumption about the human, and the supporting evidence (Proposition 1, EMG pilot, code link) is weaker than the main text implies. I'd send it to review, but would ask for the theory claims to be aligned with what is proven.\n\nWhat's new: they combine a GPU-trained RL muscle controller on MyoArm with an anomaly-score-constrained torque refiner and an energy-tank passivity-based backstepping controller for a cable-driven SEA upper-limb exoskeleton. I haven't seen that particular pipeline before. The hardware results are credible: torque tracking RMSE 0.14–0.27 Nm on shoulder and elbow joints, tank behavior reproduced in simulation and on the real arm, and the muscle controller generalizes to free movements. The passivity argument for the robot side, via two virtual tanks and Table III, is internally consistent as far as I checked. The open-sourcing of the MuJoCo Playground environment is a plus if the repository is actually reachable.\n\nSoft spots, in order of seriousness. First, Theorem 1's tracking guarantee depends on Eq. (53), the assumption that the environment (human) is dissipative. That's asserted, not measured, and the tank-replenishment experiments look like they violate it: the wearer intentionally moves with the interaction torque, injecting energy. So 'resumes tracking after replenishment' is an empirical observation, not a consequence of the theorem. The paper should either weaken the claim to 'tracking resumes in our tests' or add a secondary guarantee that does not rely on human passivity. Second, Proposition 1's proof asserts the bound mu on ||M(q)qddot~|| with no derivation; it's probably true under the observer assumptions, but it's a handwave in a proof of a claimed rigorous guarantee. Third, the EMG study is five participants, no statistics, and the 'without exoskeleton' comparison is confounded by the exoskeleton's 7.78 kg mass. The claimed up-to-48% reduction is suggestive, not conclusive. Fourth, the source code link is given as 'ATP_muscle_controller' with no URL; if the authors want the open-source claim to count, it needs to be accessible.\n\nNone of this kills the paper. The engineering contribution is real and the passivity safety claim on the robot side appears sound. It just needs an honest recalibration of what is proven versus what is demonstrated empirically.\n\nRecommendation: send to peer review, conditionally. The unified pipeline and hardware validation deserve referee time, but the theoretical section needs revision and the EMG claims need tempering.","headline":"Solid integration of learned muscle-derived torque references with passivity-based delivery; the tracking guarantee is shakier than the abstract suggests.","tokens_in":23325,"tokens_out":2667,"would_cite":true,"duration_ms":22156,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A unified framework turns anatomical muscle torques into safe, trajectory-free assistance for upper-limb exoskeletons.","keywords":["upper-limb exoskeleton","anatomical assistance","passivity-based control","energy tank","musculoskeletal model","reinforcement learning","torque refinement","electromyography evaluation"],"falsifier":"Have a participant deliberately push against the assisted joints so that interaction power $\\dot{q}^T\\tau_e$ is positive over a sustained interval, and record whether the closed loop remains passive and torque tracking is preserved. If the system still tracks torque while net energy flows from the human into the robot, the dissipative-environment inequality (53) used in Theorem 1 is violated; if tracking degrades exactly when tank energy is exhausted, the guarantee holds only under the paper's stated assumption.","tokens_in":22275,"feed_emoji":"🦾","tokens_out":8032,"duration_ms":68491,"temperature":0.7,"pith_summary":"The paper proposes ATP, a framework that aims to let an upper-limb exoskeleton assist natural, unscripted arm movements by converting anatomical, muscle-derived torques into physical support. It claims to be the first exoskeleton assistance framework to generate anatomical assistance for multi-joint, complex, nonperiodic upper-limb movements without relying on predefined motion trajectories. The pipeline trains a unified muscle controller in a scalable musculoskeletal simulation, refines its torque output online using safety and comfort signals, and delivers the result through a backstepping controller with energy tanks that preserve passivity. If the central claim holds, exoskeletons could provide movement-dependent, task-general support in daily life rather than only for scripted tasks. Simulation, real-hardware, and EMG experiments with five participants report accurate torque tracking and up to 48% reduction in target-muscle activity.","feed_headline":"Anatomical torque plus passivity control assists any upper-limb motion","feed_subtitle":"A learned muscle model and energy-tank safety let an exoskeleton support natural arm movements without preset paths.","key_machinery":"The load-bearing mechanism is the three-stage ATP chain. First, a unified muscle controller, trained by reinforcement learning in a GPU-accelerated musculoskeletal simulation, maps wearer kinematics to anatomical torque references and generalizes across diverse movements. Second, an online torque-refinement optimization repeatedly minimizes the distance from the reference while penalizing torque-rate and a learned anomaly score, which suppresses tendon-induced spikes and reshapes assistance during collisions or near kinematic singularities. Third, an interaction-torque controller built with backstepping for the cable-driven series-elastic exoskeleton uses two virtual energy tanks—reservoirs that store or release interaction energy—so that when a tank is depleted the controller relaxes the tracking objective to keep the human-robot port passive, and resumes tracking once replenished. Theorem 1 gives asymptotic stability and accurate torque tracking under the condition that the environment dissipates energy (inequality (53)) and the gain and parameter conditions (37)–(39) hold.","core_discovery":"The central discovery, stated on the paper's own terms, is that a unified reinforcement-learning muscle controller trained in a scalable musculoskeletal simulation can produce real-time anatomical joint-torque references that generalize across many upper-limb motions, and that these references can be made safe and deliverable by online refinement plus a passivity-preserving interaction-torque controller. The controller tracks the refined torque on a cable-driven series-elastic exoskeleton without constraining the wearer to a preset trajectory, and the energy-tank corrections guarantee passivity and resume torque tracking after the tank is replenished. The pilot EMG evidence supports the claim that assistance reduces the activity of primary target muscles by up to 48% relative to moving without the exoskeleton, and reduces it relative to gravity compensation and open-loop assistance in the tested tasks.","pith_inferences":["Because the passivity proof relies on a dissipative human/environment, field deployments would need to detect active user effort or otherwise guarantee safe behavior when the wearer pushes energy into the system.","The same learned-muscle-reference plus energy-tank delivery pattern could transfer to other compliant wearable robots; a direct next test is whether metabolic cost, not just EMG, decreases as assistance is scaled.","Because the anomaly score guides refinement in only two tested scenarios, the framework invites extension to other anomalies, such as payload handling or involuntary spasms, which the paper explicitly leaves for future work."],"forward_implications":["Upper-limb exoskeletons can deliver assist-as-needed torque support during unscripted, multi-joint movements without restricting the wearer's motion freedom.","A muscle controller trained jointly on diverse motion datasets retains most of its per-task tracking performance and generalizes to real-time, unseen movements measured by IMUs.","The online refinement module can detect and respond to anomalous interaction conditions, reducing assistance during simulated collisions and generating corrective torque near kinematic singularities.","The energy-tank controller preserves passivity even when tracking is temporarily suspended, and automatically returns to accurate torque tracking after the tank is replenished.","In the pilot EMG study, ATP reduced primary target-muscle activity by up to 48% in a dynamic multi-joint task compared with moving without the exoskeleton, and matched or improved on gravity-compensation and open-loop assistance."],"supporting_citations":[{"why":"Supplies the GPU-accelerated simulation platform used for large-scale muscle-controller training.","marker":"[9]"},{"why":"Provides the physiological upper-limb musculoskeletal model used to generate anatomical torque references.","marker":"[10]"},{"why":"Justifies the tank-based passivity approach and the dissipative-environment inequality used in the stability proof.","marker":"[42]"},{"why":"Provides the anomaly-score detector and intention predictor used in online refinement and look-ahead torque selection.","marker":"[46]"},{"why":"Provides the backstepping adaptive controller for series-elastic actuators that the interaction controller extends.","marker":"[48]"},{"why":"Introduces the task-energy tank concept used to passify the torque controller.","marker":"[50]"},{"why":"Extends energy-tank passification with valve-based tanks, which inspires the two-tank design.","marker":"[51]"},{"why":"Supports the assumption that the environment always dissipates energy in the interaction port.","marker":"[52]"}],"fun_headline_variants":["Unified muscle model enables exoskeleton aid for any arm motion","Passivity-guaranteed control assists diverse upper-limb movements","RL-learned anatomical torque adapts exoskeleton to real-time motions","Exoskeleton with anatomical torque reduces muscle activity by 48%","Cable-driven exoskeleton with passivity control assists without preset paths"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The safety and stability guarantee assumes the wearer's arm and surroundings only absorb energy and never actively push energy back into the exoskeleton, so a person who deliberately resists or drives the motion lies outside the guarantee.","fun_headline_variants_meta":{"raw":{"variants":["Unified muscle model enables exoskeleton aid for any arm motion","Passivity-guaranteed control assists diverse upper-limb movements","RL-learned anatomical torque adapts exoskeleton to real-time motions","Exoskeleton with anatomical torque reduces muscle activity by 48%","Cable-driven exoskeleton with passivity control assists without preset paths"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000328,"raw_usage":{"total_tokens":1856,"prompt_tokens":991,"completion_tokens":865,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":773}},"tokens_in":607,"tokens_out":865,"duration_ms":6920,"temperature":1.0,"reasoning_tokens":773,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:36:27.449175+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a participant deliberately push against the assisted joints so that interaction power $\\dot{q}^T\\tau_e$ is positive over a sustained interval, and record whether the closed loop remains passive and torque tracking is preserved. If the system still tracks torque while net energy flows from the human into the robot, the dissipative-environment inequality (53) used in Theorem 1 is violated; if tracking degrades exactly when tank energy is exhausted, the guarantee holds only under the paper's stated assumption.","supporting_citations":[{"cited_title":"Myosim: Fast and physiologically realistic mujoco models for mus- culoskeletal and exoskeletal studies,","cited_arxiv_id":null,"evidence_quote":"Provides the physiological upper-limb musculoskeletal model used to generate anatomical torque references."},{"cited_title":"Unified force-impedance control,","cited_arxiv_id":null,"evidence_quote":"Justifies the tank-based passivity approach and the dissipative-environment inequality used in the stability proof."},{"cited_title":"Upper- limb rehabilitation with a dual-mode individualized exoskeleton robot: A generative-model-based solution,","cited_arxiv_id":null,"evidence_quote":"Provides the anomaly-score detector and intention predictor used in online refinement and look-ahead torque selection."},{"cited_title":"Adaptive human–robot interaction control for robots driven by series elastic actuators,","cited_arxiv_id":null,"evidence_quote":"Provides the backstepping adaptive controller for series-elastic actuators that the interaction controller extends."},{"cited_title":"Unified passivity-based cartesian force/impedance control for rigid and flexible joint robots via task- energy tanks,","cited_arxiv_id":null,"evidence_quote":"Introduces the task-energy tank concept used to passify the torque controller."},{"cited_title":"Valve-based virtual energy tanks: A framework to simultaneously passify controls and embed control objectives,","cited_arxiv_id":null,"evidence_quote":"Extends energy-tank passification with valve-based tanks, which inspires the two-tank design."},{"cited_title":"Ott,Cartesian impedance control of redundant and flexible-joint robots","cited_arxiv_id":null,"evidence_quote":"Supports the assumption that the environment always dissipates energy in the interaction port."}],"review_version":2}