Pith. sign in

Model-Based Reinforcement Learning with SINDy

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We draw on the latest advancements in the physics community to propose a novel method for discovering the governing non-linear dynamics of physical systems in reinforcement learning (RL). We establish that this method is capable of discovering the underlying dynamics using significantly fewer trajectories (as little as one rollout with $\leq 30$ time steps) than state of the art model learning algorithms. Further, the technique learns a model that is accurate enough to induce near-optimal policies given significantly fewer trajectories than those required by model-free algorithms. It brings the benefits of model-based RL without requiring a model to be developed in advance, for systems that have physics-based dynamics. To establish the validity and applicability of this algorithm, we conduct experiments on four classic control tasks. We found that an optimal policy trained on the discovered dynamics of the underlying system can generalize well. Further, the learned policy performs well when deployed on the actual physical system, thus bridging the model to real system gap. We further compare our method to state-of-the-art model-based and model-free approaches, and show that our method requires fewer trajectories sampled on the true physical system compared other methods. Additionally, we explored approximate dynamics models and found that they also can perform well.

fields

cs.LG 1

years

2025 1

verdicts

REJECT 1

representative citing papers

Learning from Less: SINDy Surrogates in RL

cs.LG · 2025-04-25 · reject · novelty 3.0

SINDy-fitted surrogate environments for Mountain Car and Lunar Lander reproduce state dynamics from 75 to 1,000 samples and train RL agents with fewer steps, though policy transfer to the real environments is unquantified.

citing papers explorer

Showing 1 of 1 citing paper.

  • Learning from Less: SINDy Surrogates in RL cs.LG · 2025-04-25 · reject · none · ref 2 · internal anchor

    SINDy-fitted surrogate environments for Mountain Car and Lunar Lander reproduce state dynamics from 75 to 1,000 samples and train RL agents with fewer steps, though policy transfer to the real environments is unquantified.