REVIEW 2 cited by
Optimization Methods for Interpretable Differentiable Decision Trees in Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Optimization Methods for Interpretable Differentiable Decision Trees in Reinforcement Learning
abstract
Decision trees are ubiquitous in machine learning for their ease of use and interpretability. Yet, these models are not typically employed in reinforcement learning as they cannot be updated online via stochastic gradient descent. We overcome this limitation by allowing for a gradient update over the entire tree that improves sample complexity affords interpretable policy extraction. First, we include theoretical motivation on the need for policy-gradient learning by examining the properties of gradient descent over differentiable decision trees. Second, we demonstrate that our approach equals or outperforms a neural network on all domains and can learn discrete decision trees online with average rewards up to 7x higher than a batch-trained decision tree. Third, we conduct a user study to quantify the interpretability of a decision tree, rule list, and a neural network with statistically significant results ($p < 0.001$).
Forward citations
Cited by 2 Pith papers
-
Explainable AI for Next-Generation Wireless Physical Layer: Basics, State-of-the-Art, and Open Challenges
A survey formalizing responsibility-oriented goals for wireless XAI, developing a taxonomy of explainability approaches, reviewing PHY layer applications, and discussing open challenges including performance tradeoffs...
-
A Geometric Theory of Cognition for Machine Intelligence
Cognition is re-described as Riemannian gradient flow on a manifold; an anisotropic metric makes fast intuitive and slow deliberative timescales emerge from a single equation.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.