The paper gives the first provable polynomial-sample and quasi-polynomial-time guarantees for expert distillation and belief-weighted asymmetric actor-critic in POMDPs with privileged state information.
Learning in observable POMDPs, without computationally intractable oracles
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Provable Partially Observable Reinforcement Learning with Privileged Information
The paper gives the first provable polynomial-sample and quasi-polynomial-time guarantees for expert distillation and belief-weighted asymmetric actor-critic in POMDPs with privileged state information.