In a simulated low-Reynolds-number swimmer with 2 to 4 rigid paddle pairs, reinforcement learning recovers the biologically common back-to-front metachronal wave as the most efficient stroke, while front-to-back or paired strokes can be faster at wide spacings.
Deep reinforcement learning for tracking a moving target in jellyfish-like swimming
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We develop a deep reinforcement learning method for training a jellyfish-like swimmer to effectively track a moving target in a two-dimensional flow. This swimmer is a flexible object equipped with a muscle model based on torsional springs. We employ a deep Q-network (DQN) that takes the swimmer's geometry and dynamic parameters as inputs, and outputs actions which are the forces applied to the swimmer. In particular, we introduce an action regulation to mitigate the interference from complex fluid-structure interactions. The goal of these actions is to navigate the swimmer to a target point in the shortest possible time. In the DQN training, the data on the swimmer's motions are obtained from simulations conducted using the immersed boundary method. During tracking a moving target, there is an inherent delay between the application of forces and the corresponding response of the swimmer's body due to hydrodynamic interactions between the shedding vortices and the swimmer's own locomotion. Our tests demonstrate that the swimmer, with the DQN agent and action regulation, is able to dynamically adjust its course based on its instantaneous state. This work extends the application scope of machine learning in controlling flexible objects within fluid environments.
citation-role summary
citation-polarity summary
fields
physics.flu-dyn 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Optimizing Metachronal Paddling with Reinforcement Learning at Low Reynolds Number
In a simulated low-Reynolds-number swimmer with 2 to 4 rigid paddle pairs, reinforcement learning recovers the biologically common back-to-front metachronal wave as the most efficient stroke, while front-to-back or paired strokes can be faster at wide spacings.