EXPLAIN trains an interpretable linear policy from an expert's offline trajectories by combining advantage-weighted policy gradients with a behavioral cloning regularizer.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
"So, Tell Me About Your Policy...": Distillation of interpretable policies from Deep Reinforcement Learning agents
EXPLAIN trains an interpretable linear policy from an expert's offline trajectories by combining advantage-weighted policy gradients with a behavioral cloning regularizer.