SINDy-fitted surrogate environments for Mountain Car and Lunar Lander reproduce state dynamics from 75 to 1,000 samples and train RL agents with fewer steps, though policy transfer to the real environments is unquantified.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Learning from Less: SINDy Surrogates in RL
SINDy-fitted surrogate environments for Mountain Car and Lunar Lander reproduce state dynamics from 75 to 1,000 samples and train RL agents with fewer steps, though policy transfer to the real environments is unquantified.