Pith. sign in

Model-based Policy Optimization using Symbolic World Model

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The application of learning-based control methods in robotics presents significant challenges. One is that model-free reinforcement learning algorithms use observation data with low sample efficiency. To address this challenge, a prevalent approach is model-based reinforcement learning, which involves employing an environment dynamics model. We suggest approximating transition dynamics with symbolic expressions, which are generated via symbolic regression. Approximation of a mechanical system with a symbolic model has fewer parameters than approximation with neural networks, which can potentially lead to higher accuracy and quality of extrapolation. We use a symbolic dynamics model to generate trajectories in model-based policy optimization to improve the sample efficiency of the learning algorithm. We evaluate our approach across various tasks within simulated environments. Our method demonstrates superior sample efficiency in these tasks compared to model-free and model-based baseline methods.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

background 1

representative citing papers

M3PO: Massively Multi-Task Model-Based Policy Optimization

cs.LG · 2025-06-26 · reject · novelty 4.0

A hybrid model-based RL algorithm combining TDMPC2-style world models, PPO clipped updates, and POME-style exploration bonuses reports strong benchmark results but omits reproducible evidence.

citing papers explorer

Showing 1 of 1 citing paper.

  • M3PO: Massively Multi-Task Model-Based Policy Optimization cs.LG · 2025-06-26 · reject · none · ref 15 · internal anchor

    A hybrid model-based RL algorithm combining TDMPC2-style world models, PPO clipped updates, and POME-style exploration bonuses reports strong benchmark results but omits reproducible evidence.