Real-Time Reinforcement Learning for Dynamic Tasks with a Parallel Soft Robot

Allison Pinosky; Helena Young; Jake Ketchum; James Avtges; Millicent Schlafly; Ryan L. Truby; Taekyoung Kim; Todd D. Murphey

arxiv: 2509.19525 · v1 · pith:WLGFEMFOnew · submitted 2025-09-23 · 💻 cs.RO

Real-Time Reinforcement Learning for Dynamic Tasks with a Parallel Soft Robot

James Avtges , Jake Ketchum , Millicent Schlafly , Helena Young , Taekyoung Kim , Allison Pinosky , Ryan L. Truby , Todd D. Murphey This is my paper

classification 💻 cs.RO

keywords softlearningactuatorscontroldynamicbalancingcapabledemonstrate

0 comments

read the original abstract

Closed-loop control remains an open challenge in soft robotics. The nonlinear responses of soft actuators under dynamic loading conditions limit the use of analytic models for soft robot control. Traditional methods of controlling soft robots underutilize their configuration spaces to avoid nonlinearity, hysteresis, large deformations, and the risk of actuator damage. Furthermore, episodic data-driven control approaches such as reinforcement learning (RL) are traditionally limited by sample efficiency and inconsistency across initializations. In this work, we demonstrate RL for reliably learning control policies for dynamic balancing tasks in real-time single-shot hardware deployments. We use a deformable Stewart platform constructed using parallel, 3D-printed soft actuators based on motorized handed shearing auxetic (HSA) structures. By introducing a curriculum learning approach based on expanding neighborhoods of a known equilibrium, we achieve reliable single-deployment balancing at arbitrary coordinates. In addition to benchmarking the performance of model-based and model-free methods, we demonstrate that in a single deployment, Maximum Diffusion RL is capable of learning dynamic balancing after half of the actuators are effectively disabled, by inducing buckling and by breaking actuators with bolt cutters. Training occurs with no prior data, in as fast as 15 minutes, with performance nearly identical to the fully-intact platform. Single-shot learning on hardware facilitates soft robotic systems reliably learning in the real world and will enable more diverse and capable soft robots.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Damage Adaptation in Seconds for Architected Materials
cs.RO 2026-06 unverdicted novelty 5.0

LEAP enables real-time proprioceptive adaptation to unseen damage in a 6DoF soft wrist using HSA actuators by combining latent damage representations with a robust ensemble method, with conditions identified for linea...