REVIEW 3 major objections 4 minor 48 references
From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read A frozen neural policy can be rewritten as a Prolog program that plays the task as well as or better than the network.
desk verdict Finite-MDP part is solid and genuinely new; continuous-case guarantees rest on an unverified Assumption 4, plus two small reporting fixes (Acrobot 'matches' label, stale 'ten of ten' in abstract). read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing objects are the ordered decision list over a first-order predicate vocabulary (with a default clause, making the student a total deterministic policy); the performance-difference identity, which expresses the return gap between teacher and student as the teacher's discounted visitation integral of the advantage of the teacher's action under the student's value function; and the advantage-gap certificate, which bounds that gap by the advantage-weighted disagreement instead of the worst-case advantage. For continuous observations, the threshold decision-list class at resolution B — clauses built from quantile threshold tests on each coordinate — is the object that yields the O
What would settle it
Run the threshold conversion on a continuous task whose teacher has a fuzzy or stochastic decision rule (so the boundary has positive measure) or whose visited states lie on a lower-dimensional manifold; if the disagreement rate does not fall toward zero as resolution grows, or falls far slower than 1/B, Assumption 4 is violated and Theorem 2's guarantee fails. Alternatively, in the finite setting, construct an MDP where the student does not weakly dominate the teacher on disagreeing states and check whether the advantage-gap certificate remains non-vacuous; if it ever understates the exact ga
Extended reading notes
Core claim
The central claim is that a frozen PPO policy can be distilled into an ordered first-order decision list emitted as a Prolog program, then expanded under an exact-return oracle until it matches or surpasses the network. The authors establish four formal results: a return-loss bound making the distilled program a machine-checkable certificate; an advantage-gap refinement that removes a horizon factor and is exact when the student weakly dominates the teacher; monotone termination of the expansion loop; and a resolution theory showing the propositional threshold conversion has disagreement O(1/B) and a matching Ω(B^{d-1}) lower bound on rules for an oblique boundary. Empirically the expanded p
Load-bearing premise
The proof that continuous conversion is arbitrarily faithful requires the teacher's decision regions to be piecewise constant with a boundary of measure zero, and the teacher's discounted state visitation to be absolutely continuous with bounded density and bi-Lipschitz marginals; if the visits concentrate on lower-dimensional sets, the guarantees lapse.
Editorial extensions
If this is right
- Distilled policies can be certified: the exact return gap is bounded by an exactly computed quantity in finite MDPs, checked rather than estimated.
- The expansion loop guarantees monotone improvement, so the symbolic student can surpass its neural teacher wherever the teacher is imperfect.
- Relational clause structure gives size-generalization: a program induced on one MiniGrid layout transfers to unseen sizes, while coordinate-based trees collapse.
- The resolution theory sets a limit: continuous control in high dimension cannot be faithfully converted to axis-aligned threshold rules without exponential cost, explaining the LunarLander ceiling.
Reading between the lines
- The expansion oracle makes the rule proposer interchangeable; any source of candidate edits (including a generative model) is audited by exact policy evaluation, so the certificate is the enforcement mechanism.
- The advantage-gap certificate could be applied outside this pipeline, e.g., to measure where any imitator's errors actually matter under the teacher's visitation, possibly guiding data collection in imitation learning.
- The exponential lower bound suggests that for high-dimensional control, the pipeline needs relational or object-centric features to break the curse; this is testable by combining the first-order form with automatically extracted object predicates.
- The O(1/B) rate should degrade visibly when the teacher's decision boundary is not piecewise smooth or when state visitation concentrates on lower-dimensional sets; a targeted experiment on such a task would directly probe Assumption 4.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a three-stage post-hoc pipeline that turns a frozen PPO policy into an executable first-order Prolog decision list: extraction labels a reachable-state census (or DAgger samples), FOIL-style induction builds an ordered rule list, and an 'expand' stage hill-climbs on exact return, accepting an edit only when an exact Bellman-solve oracle certifies a return increase. The theory section proves a performance-difference return-loss bound (Theorem 1), an advantage-localized certificate (Proposition 3), monotone termination of the expansion loop (Proposition 1), exact finite-MDP computation (Proposition 2), and, under Assumption 4, an O(1/B) disagreement/return-gap conversion rate (Theorem 2, Corollary 3) with an Ω(B^{d-1}) lower bound (Proposition 4). Experiments cover an exact 16,944-state KeyDoor MDP (expanded Prolog attains exact optimal return and beats the capped teacher), MiniGrid DoorKey cross-size transfer, and CartPole/Acrobot/LunarLander continuous control with an explicit baseline comparison.
Significance. If the finite-MDP and relational-transfer results stand, the paper makes a meaningful contribution: it produces readable, executable artifacts with exact certificates, demonstrates monotone improvement via an exact-return oracle, and it transparently reports where the relational representation wins (DoorKey) and loses (propositional control). The exact KeyDoor certificates, the pinned reproducible pipeline, and the honest negative results are clear strengths. However, the continuous-conversion claims are conditional on a regularity assumption that is not checked in the experiments, and the lower-bound proof has a gap; the continuous part therefore needs revision before the full claim is established.
major comments (3)
- [Assumption 4, Theorem 2/Corollary 3, Section 6] The continuous headline guarantees, Eqs. (14) and (18), are conditional on Assumption 4, but the Section 6 experiments never verify it. The discounted occupancy μ of a deterministic greedy policy under deterministic Gymnasium dynamics must be absolutely continuous with bounded density and bi-Lipschitz marginal CDFs, and the decision boundary must have finite (d-1)-Hausdorff content with a linear Minkowski-content tube bound. None of these is estimated: support dimension, marginal density lower bounds, and boundary measure P are not reported. The parenthetical in Assumption 4 asserting that 'the visitation measures of Section 6 satisfy this' is not demonstrated, and CartPole/Acrobot observation spaces are not compact boxes because velocity coordinates are unbounded. Figure 9's monotone fidelity does not establish the O(1/B) rate. Please either verify a sufficient condition on the actual t
- [Proposition 4] The proof lower-bounds the disagreement rate of any cell-majority policy by c/B and shows that Ω(B^{d-1}) grid cells are crossed by the oblique hyperplane. But the theorem's conclusion — that a decision list must 'distinguish' Ω((1/ε)^{d-1}) cells and hence the cost is exponential in d — does not follow from counting cells. A single axis-aligned clause can cover many cells, and no argument is given that a short list can approximate the halfspace only with error Ω(1/B). The sentence 'each of which the optimal list must devote a distinct cell-label to' is asserted. This step is load-bearing for the abstract's 'matching lower bound' and for the interpretation of the LunarLander ceiling. A rigorous lower bound on the number of clauses in L_B, or a corrected statement (e.g., a lower bound on the number of grid cells rather than rules), is needed.
- [Table 2 and Section 7] The Acrobot rows for CART and VIPER report our-list-minus-baseline = -15.0 and -14.0 with P(improve)=0.00 and Holm-corrected p = 3.7×10^{-4}, yet the verdict column reads 'matches'. A significant negative delta in all fifteen seeds is a loss, not a match. The same inconsistency appears in the Section 7 sentence 'matches rather than beats them on Acrobot.' This does not change the overall negative-result message, but the statistical conclusion should be corrected.
minor comments (4)
- [Listing 1] Clauses 6 and 7 are both labeled default with empty bodies; under Definition 4 the first empty-body clause fires at every state and the second is dead. Please clarify whether the listing is abridged or the program has a redundant clause.
- [Table 1] The R2 expanded row reports 14/15 non-vacuous seeds, while the text says 'below unity in all fifteen seeds'. Reconcile the table with the prose.
- [Assumption 4] The claim that the box-shaped observation spaces of Section 6 satisfy the compact-support and marginal-CDF conditions is inaccurate: CartPole and Acrobot velocity coordinates are unbounded. If the empirical support is compact in practice, say so with evidence.
- [Figure 9] The text says held-out fidelity 'rises monotonically' with B; the curves increase overall but appear to have small non-monotonicities at low B. Please state the precise monotonicity claim or plot individual points.
Circularity Check
No significant circularity: the finite-MDP bounds follow from standard performance-difference identities, the continuous guarantees are grid-approximation theorems with unverified assumptions (a correctness risk, not circularity), and the sole self-citation is motivational rather than load-bearing.
full rationale
Walking the derivation chain: Theorem 1 is proved from Lemma 1 (the Kakade-Langford performance-difference lemma) plus the bounded-advantage range; the disagreement rate epsilon-dagger is the same measure defined in Definition 6 and computed exactly by Proposition 2, but the inequality is not fitted to the data it bounds. Proposition 3 is an exact identity (Lemma 2) followed by triangle-inequality and maximum bounds; its non-vacuity condition (all disagreement advantages non-positive) is checked per seed, and in that regime the certificate coincides with the exact gap only because the triangle inequality becomes equality. That is a mathematical consequence, not a disguised prediction. Proposition 1's monotone improvement and finite termination follow immediately from Assumption 2's acceptance rule (edits are accepted only when exact return increases by at least tau, with finite candidate sets); this is definitional rather than an empirical discovery, but the paper is transparent about it, and the KeyDoor exact-optimal-return result is an empirical hill-climb outcome, not a consequence of that tautology. Theorem 2 and Proposition 4 are grid-approximation arguments with no fitted constants; they rely on Assumption 4 (absolute continuity, bounded density, bi-Lipschitz marginals, finite boundary content), which Section 6 does not verify on CartPole, Acrobot, or LunarLander. That is an unverified-assumption/validity risk, not circularity: the theorems do not assume their conclusions. The only self-citation (Garrido-Merchan and Puente, 2025) appears in the introduction and discussion as motivation and as a proposed generative continuation; it is not used in any proof, bound, or benchmark. No central claim reduces, by the paper's own equations or by self-citation, to its inputs: the bounds are derived, the expansion guarantee is an explicit definitional consequence, and the continuous results are conditional on stated regularity assumptions.
Assumptions & free parameters
free parameters (4)
- resolution B =
swept {2,3,4,6,8,12}
- acceptance margin tau in EXPAND =
unspecified positive
- threshold grid t_{j,i} =
empirical quantiles of teacher-visited values per feature
- maximum clause literals in FOIL specialization =
3
assumptions (8)
- domain assumption Assumption 1: Finite discounted MDP with bounded reward.
- domain assumption Assumption 2: Exact evaluation oracle and finite candidate proposals for EXPAND.
- domain assumption Assumption 3: Exact model access with enumerable reachable set.
- domain assumption Assumption 4: Continuous-observation regularity (piecewise-constant teacher, zero-measure boundary, finite Hausdorff boundary measure, absolutely continuous occupancy with bounded density, bi-Lipschitz marginals).
- standard math Performance-difference lemma (Kakade-Langford).
- standard math Neumann series and spectral-radius argument for I - gamma P^pi invertibility.
- standard math Existence of an optimal deterministic stationary policy in a finite discounted MDP.
- standard math Markov property and tower rule in trajectory expectations.
Cite this review
Pith. "Pith review of From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems." pith.science (2026). https://pith.science/paper/GE33UQC3
@misc{pith2026260715459,
author = {Pith},
title = {Pith review of: From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/GE33UQC3}},
note = {Machine review of arXiv:2607.15459}
}
read the original abstract
A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit. We present a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in the manner of classical relational learning, and emits the result as a Prolog program whose every decision is executed by an off-the-shelf logic engine; a subsequent expansion stage edits the rule base and accepts an edit only when policy evaluation certifies a return increase. We prove four guarantees. A return-loss bound makes the distilled program a machine-checkable certificate in a finite Markov decision process, and the expansion loop improves monotonically and terminates. For the continuous-observation setting we answer whether the conversion is possible at all: the propositional threshold instantiation converts the network to arbitrary fidelity as the resolution B grows, with disagreement O(1/B) and a return gap that closes at the same rate, and a matching lower bound shows the cost is exponential in the observation dimension for an oblique decision boundary. Empirically, on a two-room key-and-door task with 16,944 reachable states the expanded Prolog program attains exact optimal return in every seed and, in a budget-capped regime, exceeds the stochastic teacher on exact return in ten of ten seeds. On three continuous-control tasks the emitted program substitutes the network, matching the neural teacher within noise on Acrobot with eleven clauses and recovering about 97% of its return on CartPole, while on the finer-control LunarLander it recovers only partially, exactly the ceiling the exponential lower bound predicts.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Courville, and Marc G
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville, and Marc G. Bellemare. Deep reinforcement learning at the edge of the statistical precipice. In Marc'Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, editors, Advances in Neural Information Processing Systems 34: Annual Conference on Neu...
2021
-
[2]
Verifiable reinforcement learning via policy extraction
Osbert Bastani, Yewen Pu, and Armando Solar - Lezama. Verifiable reinforcement learning via policy extraction. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol \` o Cesa - Bianchi, and Roman Garnett, editors, Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, Neur...
2018
-
[3]
Leo Breiman, J. H. Friedman, Richard A. Olshen, and C. J. Stone. Classification and Regression Trees. Wadsworth, 1984. ISBN 0-534-98053-8
1984
-
[4]
Maxime Chevalier - Boisvert, Bolun Dai, Mark Towers, Rodrigo Perez - Vicente, Lucas Willems, Salem Lahlou, Suman Pal, Pablo Samuel Castro, and Jordan K. Terry. M inigrid & M iniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levin...
2023
-
[5]
Youri Coppens, Denis Steckelmacher, Catholijn M. Jonker, and Ann Nowé. Synthesising reinforcement learning policies through set-valued inductive rule learning. arXiv preprint arXiv:2106.06009, 2021
arXiv 2021
-
[6]
Interpretable and explainable logical policies via neurally guided symbolic abstraction
Quentin Delfosse, Hikaru Shindo, Devendra Singh Dhami, and Kristian Kersting. Interpretable and explainable logical policies via neurally guided symbolic abstraction. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information Processing Systems 36: Annual Conference on Neural Informa...
2023
-
[7]
Interpretable concept bottlenecks to align reinforcement learning agents
Quentin Delfosse, Sebastian Sztwiertnia, Mark Rothermel, Wolfgang Stammer, and Kristian Kersting. Interpretable concept bottlenecks to align reinforcement learning agents. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 37: Annual...
2024
-
[8]
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608, 2017
arXiv 2017
Show all 48 references
-
[10]
Garrido-Merchán and Cristina Puente
Eduardo C. Garrido-Merchán and Cristina Puente. GOFAI meets G enerative AI : Development of expert systems by means of L arge L anguage M odels. arXiv preprint arXiv:2507.13550, 2025
2025
-
[11]
Neural logic reinforcement learning
Zhengyao Jiang and Shan Luo. Neural logic reinforcement learning. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA , volume 97 of Proceedings of...
2019
-
[12]
Kakade and John Langford
Sham M. Kakade and John Langford. Approximately optimal approximate reinforcement learning. In Claude Sammut and Achim G. Hoffmann, editors, Machine Learning, Proceedings of the Nineteenth International Conference (ICML 2002), University of New South Wales, Sydney, Australia, ...
2002
-
[13]
Interpretable and editable programmatic tree policies for reinforcement learning
Hector Kohler, Quentin Delfosse, Riad Akrour, Kristian Kersting, and Philippe Preux. Interpretable and editable programmatic tree policies for reinforcement learning. arXiv preprint arXiv:2405.14956, 2024
2024 arXiv
-
[14]
Learning finite state representations of recurrent policy networks
Anurag Koul, Alan Fern, and Sam Greydanus. Learning finite state representations of recurrent policy networks. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. URL https://openreview.net/forum?i...
2019
-
[15]
Petersen, Sookyung Kim, Cl \' a udio P
Mikel Landajuela, Brenden K. Petersen, Sookyung Kim, Cl \' a udio P. Santiago, Ruben Glatt, T. Nathan Mundhenk, Jacob F. Pettit, and Daniel M. Faissol. Discovering symbolic policies with deep reinforcement learning. In Marina Meila and Tong Zhang, editors, Proceedings of the 3...
2021
-
[16]
Toward interpretable deep reinforcement learning with linear model U - T rees
Guiliang Liu, Oliver Schulte, Wang Zhu, and Qingcan Li. Toward interpretable deep reinforcement learning with linear model U - T rees. In Michele Berlingerio, Francesco Bonchi, Thomas G \" a rtner, Neil Hurley, and Georgiana Ifrim, editors, Machine Learning and Knowledge Disco...
2018 doi
-
[19]
Gordon, and Drew Bagnell
St \' e phane Ross, Geoffrey J. Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Geoffrey J. Gordon, David B. Dunson, and Miroslav Dud \' k, editors, Proceedings of the Fourteenth International Conference on...
2011
-
[20]
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[21]
EXPIL : Explanatory predicate invention for learning in games
Jingyuan Sha, Hikaru Shindo, Quentin Delfosse, Kristian Kersting, and Devendra Singh Dhami. EXPIL : Explanatory predicate invention for learning in games. arXiv preprint arXiv:2406.06107, 2024
2024 arXiv
-
[22]
B lend RL : A framework for merging symbolic and neural policy learning
Hikaru Shindo, Quentin Delfosse, Devendra Singh Dhami, and Kristian Kersting. B lend RL : A framework for merging symbolic and neural policy learning. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025 . OpenReview.n...
2025
-
[23]
Mark Towers, Ariel Kwiatkowski, Jordan Terry, John U. Balis, Gianluca De Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Markus Krimmel, Arjun KG, Rodrigo Perez-Vicente, Andrea Pierré, Sander Schulhoff, Jun Jet Tai, Hannah Tan, and Omar G. Younis. Gymnasium: A standar...
2024 arXiv
-
[24]
Programmatically interpretable reinforcement learning
Abhinav Verma, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, and Swarat Chaudhuri. Programmatically interpretable reinforcement learning. In Jennifer G. Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Sto...
2018
-
[25]
Imitation-projected programmatic reinforcement learning
Abhinav Verma, Hoang Minh Le, Yisong Yue, and Swarat Chaudhuri. Imitation-projected programmatic reinforcement learning. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d'Alch \' e - Buc, Emily B. Fox, and Roman Garnett, editors, Advances in Neural Informatio...
2019
-
[27]
Kakade and John Langford , editor =
Sham M. Kakade and John Langford , editor =. Approximately Optimal Approximate Reinforcement Learning , booktitle =. 2002 , timestamp =
2002
-
[28]
Verifiable Reinforcement Learning via Policy Extraction , booktitle =
Osbert Bastani and Yewen Pu and Armando Solar. Verifiable Reinforcement Learning via Policy Extraction , booktitle =. 2018 , url =
2018
-
[29]
Programmatically Interpretable Reinforcement Learning , booktitle =
Abhinav Verma and Vijayaraghavan Murali and Rishabh Singh and Pushmeet Kohli and Swarat Chaudhuri , editor =. Programmatically Interpretable Reinforcement Learning , booktitle =. 2018 , url =
2018
-
[30]
Imitation-Projected Programmatic Reinforcement Learning , booktitle =
Abhinav Verma and Hoang Minh Le and Yisong Yue and Swarat Chaudhuri , editor =. Imitation-Projected Programmatic Reinforcement Learning , booktitle =. 2019 , url =
2019
-
[31]
Neural Logic Reinforcement Learning , booktitle =
Zhengyao Jiang and Shan Luo , editor =. Neural Logic Reinforcement Learning , booktitle =. 2019 , url =
2019
-
[32]
Interpretable and Explainable Logical Policies via Neurally Guided Symbolic Abstraction , booktitle =
Quentin Delfosse and Hikaru Shindo and Devendra Singh Dhami and Kristian Kersting , editor =. Interpretable and Explainable Logical Policies via Neurally Guided Symbolic Abstraction , booktitle =. 2023 , url =
2023
-
[33]
The Thirteenth International Conference on Learning Representations,
Hikaru Shindo and Quentin Delfosse and Devendra Singh Dhami and Kristian Kersting , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =
2025
-
[34]
Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents , booktitle =
Quentin Delfosse and Sebastian Sztwiertnia and Mark Rothermel and Wolfgang Stammer and Kristian Kersting , editor =. Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents , booktitle =. 2024 , url =
2024
-
[35]
Petersen and Sookyung Kim and Cl
Mikel Landajuela and Brenden K. Petersen and Sookyung Kim and Cl. Discovering symbolic policies with deep reinforcement learning , booktitle =. 2021 , url =
2021
-
[36]
2024 , url =
Stephanie Milani and Nicholay Topin and Manuela Veloso and Fei Fang , title =. 2024 , url =. doi:10.1145/3616864 , timestamp =
2024 doi
-
[37]
Courville and Marc G
Rishabh Agarwal and Max Schwarzer and Pablo Samuel Castro and Aaron C. Courville and Marc G. Bellemare , editor =. Deep Reinforcement Learning at the Edge of the Statistical Precipice , booktitle =. 2021 , url =
2021
-
[38]
Ross Quinlan , title =
J. Ross Quinlan , title =. Mach. Learn. , volume =. 1990 , url =. doi:10.1007/BF00117105 , timestamp =
1990 doi
-
[39]
Saso Dzeroski and Luc De Raedt and Kurt Driessens , title =. Mach. Learn. , volume =. 2001 , url =. doi:10.1023/A:1007694015589 , timestamp =
2001 doi
-
[40]
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning , booktitle =
St. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning , booktitle =. 2011 , url =
2011
-
[41]
7th International Conference on Learning Representations,
Anurag Koul and Alan Fern and Sam Greydanus , title =. 7th International Conference on Learning Representations,. 2019 , url =
2019
-
[42]
Theory Pract
Jan Wielemaker and Tom Schrijvers and Markus Triska and Torbj. Theory Pract. Log. Program. , volume =. 2012 , url =. doi:10.1017/S1471068411000494 , timestamp =
2012 doi
-
[43]
Toward Interpretable Deep Reinforcement Learning with Linear Model
Guiliang Liu and Oliver Schulte and Wang Zhu and Qingcan Li , editor =. Toward Interpretable Deep Reinforcement Learning with Linear Model. Machine Learning and Knowledge Discovery in Databases - European Conference,. 2018 , url =. doi:10.1007/978-3-030-10928-8\_25 , timestamp =
2018 doi
-
[44]
Leo Breiman and J. H. Friedman and Richard A. Olshen and C. J. Stone , title =. 1984 , isbn =
1984
-
[45]
Maxime Chevalier. Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , year =
2023
-
[46]
2017 , eprint =
John Schulman and Filip Wolski and Prafulla Dhariwal and Alec Radford and Oleg Klimov , title =. 2017 , eprint =
2017
-
[47]
Jonker and Ann Nowé , title =
Youri Coppens and Denis Steckelmacher and Catholijn M. Jonker and Ann Nowé , title =. 2021 , eprint =
2021
-
[48]
2024 , eprint =
Jingyuan Sha and Hikaru Shindo and Quentin Delfosse and Kristian Kersting and Devendra Singh Dhami , title =. 2024 , eprint =
2024
-
[49]
2024 , eprint =
Hector Kohler and Quentin Delfosse and Riad Akrour and Kristian Kersting and Philippe Preux , title =. 2024 , eprint =
2024
-
[50]
Garrido-Merchán and Cristina Puente , title =
Eduardo C. Garrido-Merchán and Cristina Puente , title =. 2025 , eprint =
2025
-
[51]
Mark Towers and Ariel Kwiatkowski and Jordan Terry and John U. Balis and Gianluca De Cola and Tristan Deleu and Manuel Goulão and Andreas Kallinteris and Markus Krimmel and Arjun KG and Rodrigo Perez-Vicente and Andrea Pierré and Sander Schulhoff and Jun Jet Tai and Hannah Tan...
2024
-
[52]
2017 , eprint =
Finale Doshi-Velez and Been Kim , title =. 2017 , eprint =
2017
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.