A framework that recovers an agent's assumed dynamics, rewards, and observation noise in continuous partially observable tasks by training policies over a model space and maximizing the likelihood of observed actions.
Maximum likelihood from incomplete data via the em algorithm
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Inverse Rational Control with Partially Observable Continuous Nonlinear Dynamics
A framework that recovers an agent's assumed dynamics, rewards, and observation noise in continuous partially observable tasks by training policies over a model space and maximizing the likelihood of observed actions.