REVIEW 2 major objections 5 minor 58 references
This paper treats the explainability of a reinforcement-learning agent as a measurable property of the logical rules learned from its behavior, not as a matter of user preference, and proposes four objective metrics to quantify that propert
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 04:31 UTC pith:BEBDPBDO
load-bearing objection A useful packaged set of metrics for ILP-based policy explanations, but the headline objectivity claim rests on an in-sample circularity that needs a held-out fix. the 2 major comments →
Explaining Reinforcement Learning Agents via Inductive Logic Programming
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Central claim: a trained agent's decisions can be summarized as a finite rule set, and four numbers computed on that rule set surface what return curves hide. Activation rate = how often a rule fires when the agent actually performed that action, read as per-action reliability. Feature coverage = the share of rule bodies mentioning a given feature, an action-level analogue of feature importance. Syntactic distance = one minus the intersection-over-union overlap between rule bodies for the same action across two policies. Semantic distance = the same overlap computed on the conclusions the two rule sets draw from identical observations, capturing behavioral rather than structural difference.
What carries the argument
The load-bearing object is the logical policy approximation: an Answer Set Programming rule set (a rule-based logic formalism in which rules derive actions from observed features), learned from state-action traces, turning a black-box policy into an inspectable artifact. The four metrics are defined on it: activation rate counts how often rules reproduce observed actions; feature coverage counts feature mentions normalized by total body size; syntactic distance is an intersection-over-union distance between rule-body sets; semantic distance is the same overlap on the conclusions produced for identical contexts. This is what lets a qualitative judgment (is this explanation any good?) be repla
Load-bearing premise
The framework assumes the shortest rule set induced from observed decisions is a faithful, stable surrogate for the agent's policy — the assumption enters when traces become learning examples and when the rule set is read as the explanation — and in the warehouse experiments the rules reproduce only about 30% of decisions, so low activation there may reflect approximation failure rather than agent uncertainty.
What would settle it
Compute the activation rate of a converged rule set on both the traces it was learned from and on a fresh set of traces collected from the same agent in the same environment; if the rate drops sharply out of the training traces while staying high on them, the metric is measuring trace memorization, not policy alignment, and the framework's core diagnostic would be false.
If this is right
- Activation rate works as a diagnostic that separates actions the agent has learned confidently from actions whose rules barely fire, catching instability or over-specialization that a flat return curve misses.
- Feature coverage gives action-level feature relevance and lets a practitioner spot features that are effectively unused, supporting principled pruning of the feature set.
- Syntactic and semantic distances track policy evolution during training and between agents, revealing coordination, specialization, and continued policy change even after mean return has plateaued.
- Activation rate computed in altered environments measures how far learned rules generalize and points to which action-level skills are worth transferring.
- Together the metrics position the logical rule set as a debugging artifact for the whole training loop, not just a final explanation.
Where Pith is reading between the lines
- Editorial inference: the same numbers could drive data collection — traces where activation is low are precisely the states where the current rule set is wrong, so they are candidates for more exploration or for curriculum design.
- Editorial inference: because all four metrics are computed on a user-chosen feature map, they double as a quantitative audit of that map: features that never appear in rules or whose permutation barely changes activation are candidates for removal, suggesting a principled way to co-design explanation and representation.
- Editorial inference: if semantic distance keeps moving after return plateaus, it offers a testable stopping criterion or a stability monitor for deployment, since a policy can look converged in reward while still reshaping its behavior.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a framework for evaluating symbolic, logic-based explanations of RL and MARL policies. It uses ILASP to induce ASP rule sets from execution traces and then defines four metrics: activation rate, feature coverage, syntactic distance, and semantic distance. These are intended to quantify action-specific policy confidence, feature relevance, policy evolution during training, inter-agent agreement, and transferability. The experiments cover the Intersection single-agent domain, the RWARE cooperative MARL domain, and the Simple Adversary contrastive MARL domain, and each is organized around four research questions. The central claim is that the metrics provide objective, user-independent measurements of policy properties that go beyond return curves.
Significance. If validated, this would be a useful contribution to explainable RL: the metrics are formally defined, the code is released, and the multi-domain study gives concrete evidence that symbolic approximations can reveal action-level and agent-level structure not visible in global returns. The out-of-distribution activation-rate experiments in §5.6 are a particularly good step toward testing generalization of the learned rules. However, the paper's stronger interpretative claims—that activation rate measures agent confidence, or that semantic distance measures behavioral convergence—depend on the ILASP hypothesis being a faithful surrogate of the neural policy. The current evidence is largely in-sample, and one metric definition is ambiguous. These issues are fixable but need to be addressed before the central claims are fully supported.
major comments (2)
- [§4.2 / §4.3.1, Eqs. (3), (7)–(8)] Activation rate is computed on the same execution traces used to construct the ILP examples. Since ILASP is asked to find H that covers exactly those state–action pairs, α(a) is a train-set coverage score of the fitted symbolic model, not an independent measurement of how faithfully H represents the neural policy. The RWARE result in Fig. 6a (α(H) ≈ 0.30 at convergence) is interpreted in §5.3 as evidence about action-specific confidence, but an equally plausible reading is that ILASP's search space and feature map cannot fit the policy on 70% of the training traces. RQ1 and RQ4 therefore need a validation step: held-out traces, a random-policy or random-rule baseline, or a precision/recall decomposition. The OOD experiments in §5.6 partly address transfer but do not validate the training-phase interpretation.
- [§4.3.4, Eq. (11)] AS_H(c) is not defined precisely. If, as in ASP, the context C = F_F(s) is part of the program, the answer sets contain all context atoms, so the Jaccard index in Eq. (11) is inflated by the shared context and may be largely insensitive to behavioral differences. If only action atoms are intended, that restriction must be stated and used consistently in the experiments (Fig. 7b–d). Without this, the semantic-distance conclusions about inter-agent convergence and specialization rest on an ambiguous quantity.
minor comments (5)
- [§4.3.1, Eq. (7)] The notation AS_{r_k} is used but never formally defined; it should be clarified whether this is the set of answer sets of H in which rule r_k is applicable/fires. The definition should be explicit to make Eq. (7) reproducible.
- [§4.3.1, Eq. (8) and text after it] The 'activation rate of the whole theory' is described as a sum of per-action rates. This is not the overall fraction of matched actions unless actions occur equally often; the sum weights actions equally. The paper should define α(H) explicitly (e.g., as a weighted average) and state which quantity is plotted in Fig. 6a.
- [§4.1–§4.3] The 'objective, user-independent' claim is qualified by the fact that the whole pipeline depends on the user-defined feature map F_F and on the chosen value discretizations (Dist sets, Dir sets). This should be acknowledged explicitly in the limitations section, because different feature maps can lead to different metric values.
- [§5.3, Fig. 6c] The interpretation that activation for move_up/move_down goes to zero because the adversary 'rarely moves up or down' conflates low action frequency with low rule fidelity. The denominator in Eq. (8) conditions on the action being performed, so a zero value means that whenever the action was performed the rule did not fire, which is a different statement. The text should be reworded to avoid this confusion.
- [General] Several claims, especially in the abstract and Section 6, state that the metrics are 'planning-oriented' without demonstrating a connection to planning or sequential decision-making. Either define this term or soften the claim.
Circularity Check
Activation rate is computed on the same traces used to induce the ILASP theory, so it measures in-sample fit rather than independent alignment with the neural policy; RQ1/RQ4 conclusions conflate approximation fidelity with agent policy properties.
specific steps
-
fitted input called prediction
[§4.2 (Eq. 3) and §4.3.1 (Def. 1, Eqs. 7–8); interpreted in §5.3 (RQ1) and §5.6 (RQ4)]
"Given a set of RL execution traces 𝑇={𝑡1,...,𝑡𝑁}, where each trace 𝑡𝑖 is a sequence of state-action pairs ⟨𝑠,𝑎⟩ ... To construct the ILASP task, we encode each state-action pair ⟨𝑠,𝑎⟩ into a CDPI ... with 𝐶=𝐹F(𝑠), 𝑒𝑖𝑛𝑐=𝐹A(𝑎) and 𝑒𝑒𝑥𝑐={𝐹A(𝑎′)|𝑎′∈𝐴\𝑎} ... we invoke ILASP to infer a hypothesis 𝐻⊆𝑆𝑀 that satisfies Equation 3. ... Let 𝐻 be a logical policy approximation, and 𝑇={𝑡1,𝑡2,...,𝑡𝑁} a set of 𝑁 execution traces collected from an RL agent. The activation rate ... is computed as [Eq. 7]."
The T in Def. 1 is the same T used to build the ILP examples E: every observed state-action pair is turned into a CDPI whose included atom is the executed action and whose excluded atoms are the other actions, and Eq. 3 asks ILASP to find H that has an answer set containing that atom and excluding the others for every such example. The numerator of Eqs. 7–8 checks exactly that containment, so α(T,H) counts how often the induced theory covers its own training examples. It is therefore an in-sample goodness-of-fit statistic. Interpreting it as an 'objective measure' of how the symbolic rules 'align with the agent's behavior', and using it to support RQ1/RQ4 conclusions about action-specific confidence, learning dynamics, and transferability, treats a training-set fit measure as an independen
full rationale
The central circularity concerns the activation rate. The ILASP hypothesis H is induced from execution traces T (§4.2, Eq. 3), and then the activation rate α is defined over the same T and interpreted as alignment with the agent's behavior (§4.3.1, Eqs. 7–8). Because the numerator of α is precisely the condition ILASP was asked to satisfy on those very examples, α is an in-sample fidelity score, not an independent measurement of the neural policy. This infects the RQ1 training-dynamics claims and the RQ4 transferability readings that rely on α. The other three metrics—feature coverage, syntactic distance, and semantic distance—are straightforward functions of H and are not circular in the same way. The RQ4 transfer experiments evaluate theories on held-out environment variants, which is genuine out-of-distribution evidence, and the constant-permutation analysis is a sensitivity check rather than a circular prediction. Thus the circularity is limited to the activation-rate-as-policy-diagnostic claim, giving a partial-circularity score of 6 rather than a fully forced derivation.
Axiom & Free-Parameter Ledger
free parameters (3)
- User-defined feature vocabulary and discretization (F, Dir, Dist sets) =
Intersection: Dist∈[0,100], Id∈[0,14]; RWARE: Dir∈{N,S,E,W}, Dist∈{0,...,3}; Simple Adversary: Dir∈{N,S,E,W}, Dist∈{0,5,
- ILASP hypothesis H, including learned numeric thresholds in rule bodies =
e.g., 'obs_is_close(V1,35)', 'V1 <= 1', 'V2 <= 2'
- Training convergence point (gray vertical line) =
hand-picked per domain
axioms (4)
- standard math Stable model semantics of ASP as implemented in ILASP; each learned program yields a single relevant answer set per context.
- domain assumption The feature map F_F captures the decision-relevant state information (Assumptions 1–2).
- domain assumption Execution traces are representative of the agent's policy and contain no contradictory examples; ILASP can find a hypothesis satisfying Eq. 3.
- ad hoc to paper Answer sets used in semantic distance contain only derived action atoms, or context atoms can be safely ignored.
read the original abstract
Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user studies, thus targeting the needs of a specific audience and lacking shared evaluation metrics. On the other hand, logic-based approaches within eXplainable Artificial Intelligence (XAI) provide compact, human-readable abstractions of decision-making. However, the systematic quantification of the explainability degree of logical representations remains an open problem. This work aims to advance the state of the art in XRL by introducing objective and planning-oriented metrics for policy explainability in RL settings. At the same time, it contributes to the field of logic for XAI by providing a principled way to quantify the explainability of logical rules, moving beyond common-sense assessments and simple propositional fragments. We employ Inductive Logic Programming (ILP) to extract symbolic representations of RL policies and define a novel set of explainability metrics, including activation rate, feature coverage, syntactic distance and semantic distance. These metrics quantify alignment between symbolic rules and agent behavior, the role of features in decision-making, and the evolution of policies during training and across agents in single and multi-agent RL. Experiments across different RL domains show that the proposed metrics highlight action-specific learning dynamics beyond global return, provide fine-grained insights into domain features beyond classical approaches for global feature importance estimation, and uncover coordination, specialization, and adaptation patterns in MARL. Moreover, they provide crucial insights for the transfer and generalization of action-specific policies.
Figures
Reference graph
Works this paper leans on
-
[1]
Ofra Amir, Finale Doshi-Velez, and David Sarne. 2019. Summarizing agent strategies.Autonomous Agents and Multi-Agent Systems33, 5 (2019), 628–644
2019
-
[2]
Elia Amparore, Giuseppe Contissa, Francesco De Benedetti, Valerio Lagnese, Giovanni Livraga, Daniele Malerba, Monica Palmirani, Giovanni Sartor, Franco Turini, Carlo Zavattari, et al. 2021. To trust or not to trust an explanation using LEAF. InProceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society. ACM, 336–346. https://doi.org/10.1145/346...
arXiv 2021
-
[3]
Ziyan An, Hendrik Baier, Abhishek Dubey, Ayan Mukhopadhyay, and Meiyi Ma. 2024. Enabling MCTS Explainability for Sequential Planning Through Computation Tree Logic. InECAI 2024 (Frontiers in Artificial Intelligence and Applications), Ulle Endriss, Francisco S. Melo, Kerstin Bach, Alberto Bugarín-Diz, José M. Alonso-Moral, Senén Barro, and Fredrik Heintz (...
-
[4]
Andrew Anderson, Jonathan Dodge, Amrita Sadarangani, Zoe Juozapaitis, Evan Newman, Jed Irvine, Souti Chattopadhyay, Alan Fern, and Margaret Burnett. 2019. Explaining Reinforcement Learning to Mere Mortals: An Empirical Study. InProceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI). 1328–1334. https://doi.org/10.24963/ij...
-
[5]
Osbert Bastani, Yewen Pu, and Armando Solar-Lezama. 2018. Verifiable Reinforcement Learning via Policy Extraction. InAdvances in Neural Information Processing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2018/file/e6d8545...
2018
-
[6]
Kayla Boggess. 2025. Explanations for Multi-Agent Reinforcement Learning. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 29245–29246
2025
-
[7]
Nadia Burkart and Marco F Huber. 2021. A survey on the explainability of supervised machine learning.Journal of Artificial Intelligence Research70 (2021), 245–317
2021
-
[8]
Luciano Caroprese, Ester Zumpano, and Domenico Ursino. 2026. Reinforcement Learning Meets Logic Programming: Towards Explainable AI. InLogics in Artificial Intelligence, Giovanni Casini, Besik Dundua, and Temur Kutsia (Eds.). Springer Nature Switzerland, Cham, 13–27
2026
-
[9]
Andrew Cropper and Sebastijan Dumančić. 2022. Inductive logic programming at 30: a new introduction.Journal of Artificial Intelligence Research74 (2022), 765–850
2022
-
[10]
Adnan Darwiche and Pierre Marquis. 2022. On Quantifying Literals in Boolean Logic and its Applications to Explainable AI.Journal of Artificial Intelligence Research72 (Jan. 2022), 285–328. https://doi.org/10.1613/jair.1.12756
-
[11]
Artur d’Avila Garcez and Luís C. Lamb. 2023. Neurosymbolic AI: the 3rd wave.Artif. Intell. Rev.56, 11 (2023), 12387–12406. https: //doi.org/10.1007/S10462-023-10448-W
-
[12]
Patil, Matthias Dorfer, Patrick M
Marius-Constantin Dinu, Markus Hofmarcher, Vihang P. Patil, Matthias Dorfer, Patrick M. Blies, Johannes Brandstetter, Jose A. Arjona- Medina, and Sepp Hochreiter. 2022. XAI and Strategy Extraction via Reward Redistribution. InProceedings of the International Workshop on Extending Explainable AI Beyond Deep Models and Classifiers. 177–205
2022
-
[13]
Rudresh Dwivedi, Devam Dave, Het Naik, Smiti Singhal, Rana Omer, Pankesh Patel, Bin Qian, Zhenyu Wen, Tejal Shah, Graham Morgan, et al. 2023. Explainable AI (XAI): Core ideas, techniques, and solutions.Comput. Surveys55, 9 (2023), 1–33
2023
-
[14]
Richard Evans and Edward Grefenstette. 2018. Learning explanatory rules from noisy data.J. Artif. Int. Res.61, 1 (Jan. 2018), 1–64
2018
-
[15]
Jakob Foerster, Ioannis Alexandros Assael, Nando De Freitas, and Shimon Whiteson. 2016. Learning to communicate with deep multi-agent reinforcement learning.Advances in neural information processing systems29 (2016)
2016
-
[16]
Daniel Furelos-Blanco, Mark Law, Anders Jonsson, Krysia Broda, and Alessandra Russo. 2021. Induction and exploitation of subgoal automata for reinforcement learning.Journal of Artificial Intelligence Research70 (2021), 1031–1116. JAIR, Vol. 1, Article . Publication date: July 2026. ILP for XRL•27
2021
-
[17]
Michael Gelfond and Vladimir Lifschitz. 1988. The Stable Model Semantics for Logic Programming. InICLP/SLP. https://api. semanticscholar.org/CorpusID:261517573
1988
-
[18]
Samuel Greydanus, Anurag Koul, Jonathan Dodge, and Alan Fern. 2018. Visualizing and Understanding Atari Agents. InProceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 1792–1801
2018
-
[19]
Mohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham, and Daniel Kroening. 2021. DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement Learning. InProceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI ’21), Vol. 2. 36
2021
-
[20]
Hado van Hasselt, Arthur Guez, and David Silver. 2016. Deep reinforcement learning with double Q-Learning. InProceedings of the Thirtieth AAAI Conference on Artificial Intelligence(Phoenix, Arizona)(AAAI’16). AAAI Press, 2094–2100
2016
-
[21]
Huang, Kush Bhatia, Pieter Abbeel, and Anca D
Sandy H. Huang, Kush Bhatia, Pieter Abbeel, and Anca D. Dragan. 2018. Establishing Appropriate Trust via Critical States. In2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE Press, 3929–3936. https://doi.org/10.1109/IROS.2018.8593649
arXiv 2018
-
[22]
Alexis David Jacq, Johan Ferret, Matthieu Geist, and Olivier Pietquin. 2022. Lazy-MDPs: Towards Interpretable RL by Learning When to Act. InAutonomous Agents and Multi-Agent Systems
2022
-
[23]
Zoe Juozapaitis, Anurag Koul, Alan Fern, Martin Erwig, and Finale Doshi-Velez. 2019. Explainable Reinforcement Learning via Reward Decomposition. InProceedings of the 28th International Joint Conference on Artificial Intelligence Workshop on Explainable Artificial Intelligence
2019
-
[24]
Alex Kendall, Jeffrey Hawke, David Janz, Przemyslaw Mazur, Daniele Reda, John-Mark Allen, Vinh-Dieu Lam, Alex Bewley, and Amar Shah. 2019. Learning to Drive in a Day. In2019 International Conference on Robotics and Automation (ICRA). 8248–8254. https://doi.org/10.1109/ICRA.2019.8793742
arXiv 2019
-
[25]
Anurag Koul, Sam Greydanus, and Alan Fern. 2018. Learning Finite State Representations of Recurrent Policy Networks.ArXiv abs/1811.12530 (2018). https://api.semanticscholar.org/CorpusID:54434799
Pith/arXiv arXiv 2018
-
[26]
Mark Law, Alessandra Russo, and Krysia Broda. 2015. Learning weak constraints in answer set programming.Theory and Practice of Logic Programming15, 4-5 (2015), 511–525
2015
-
[27]
Edouard Leurent. 2018. An Environment for Autonomous Driving Decision-Making. https://github.com/eleurent/highway-env
2018
-
[28]
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. 2016. End-to-end training of deep visuomotor policies.J. Mach. Learn. Res.17, 1 (Jan. 2016), 1334–1373
2016
-
[29]
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017. Multi-Agent Actor-Critic for Mixed Cooperative- Competitive Environments.Neural Information Processing Systems (NIPS)(2017)
2017
-
[30]
Prashan Madumal, Tim Miller, Liz Sonenberg, and Frank Vetere. 2020. Explainable Reinforcement Learning through a Causal Lens. Proceedings of the AAAI Conference on Artificial Intelligence34, 03 (Apr. 2020), 2493–2500. https://doi.org/10.1609/aaai.v34i03.5631
-
[31]
Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B. Tenenbaum, and Jiajun Wu. 2019. The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences from Natural Supervision.arXiv preprint arXiv:1904.12584(2019)
Pith/arXiv arXiv 2019
-
[32]
Joao Marques-Silva. 2025. Logic-Based Explainability: Past, Present and Future. InLeveraging Applications of Formal Methods, Verification and Validation. Software Engineering Methodologies, Tiziana Margaria and Bernhard Steffen (Eds.). Springer Nature Switzerland, Cham, 181–204
2025
-
[33]
Joe McCalmon, Thai Le, Sarra Alqahtani, and Dongwon Lee. 2022. CAPS: Comprehensible Abstract Policy Summaries for Explaining Reinforcement Learning Agents. InProceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems. 889–897
2022
-
[34]
Daniele Meli, Alberto Castellini, and Alessandro Farinelli. 2024. Learning logic specifications for policy guidance in pomdps: an inductive logic programming approach.Journal of Artificial Intelligence Research79 (2024), 725–776
2024
-
[35]
Daniele Meli and Paolo Fiorini. 2025. Inductive learning of robot task knowledge from raw data and online expert feedback.Machine Learning114, 4 (2025), 91
2025
-
[36]
Daniele Meli, Hirenkumar Nakawala, and Paolo Fiorini. 2023. Logic programming for deliberative robotic task planning.Artificial Intelligence Review56, 9 (2023), 9011–9049
2023
-
[37]
Stephanie Milani, Nicholay Topin, Manuela Veloso, and Fei Fang. 2024. Explainable Reinforcement Learning: A Survey and Comparative Review.ACM Comput. Surv.56, 7 (2024). https://doi.org/10.1145/3616864
doi:10.1145/3616864 2024
-
[38]
Igor Mordatch and Pieter Abbeel. 2017. Emergence of Grounded Compositional Language in Multi-Agent Populations.arXiv preprint arXiv:1703.04908(2017)
Pith/arXiv arXiv 2017
-
[39]
Muggleton
S. Muggleton. 1991. Inductive logic programming.New Generation Computing8, 4 (1991), 295–318
1991
-
[40]
Stephen Muggleton. 1995. Inverse entailment and Progol.New generation computing13 (1995), 245–286
1995
-
[41]
Meike Nauta, Jasper van der Waa, Anouk Leufkens, Guszti Eiben, Christin Seifert, Mireia Ribera, and Virginia Dignum. 2023. From anecdotal evidence to quantitative evaluation methods: A systematic review on evaluating explainable AI.Artificial Intelligence313 (2023), 103824. https://doi.org/10.1016/j.artint.2022.103824 JAIR, Vol. 1, Article . Publication d...
arXiv 2023
-
[42]
Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht. 2021. Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks. InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS). http://arxiv.org/abs/2006.07869
Pith/arXiv arXiv 2021
-
[43]
Alan Perotti, Claudio Borile, Arianna Miola, Francesco Paolo Nerini, Paolo Baracco, and André Panisson. 2024. Explainability, Quantified: Benchmarking XAI Techniques. InExplainable Artificial Intelligence. Springer Nature Switzerland, 421–444
2024
-
[44]
Gabrielle Ras, Ning Xie, Marcel van Gerven, and Derek Doran. 2022. Explainable Deep Learning: A Field Guide for the Uninitiated. Journal of Artificial Intelligence Research73 (May 2022), 68 pages. https://doi.org/10.1613/jair.1.13200
-
[45]
Carl Orge Retzlaff, Srijita Das, Christabel Wayllace, Payam Mousavi, Mohammad Afshari, Tianpei Yang, Anna Saranti, Alessa Angerschmid, Matthew E. Taylor, and Andreas Holzinger. 2024. Human-in-the-Loop Reinforcement Learning: A Survey and Position on Requirements, Challenges, and Opportunities.Journal of Artificial Intelligence Research79 (April 2024), 57 ...
-
[46]
Finn Rietz, Sven Magg, Fredrik Heintz, Todor Stoyanov, Stefan Wermter, and Johannes A. Stork. 2022. Hierarchical Goals Contextualize Local Reward Decomposition Explanations.Neural Computing and Applications(2022). https://doi.org/10.1007/s00521-022-07171-x Published online, May 12, 2022
-
[47]
Tim Rocktäschel and Sebastian Riedel. 2017. End-to-end Differentiable Proving. InAdvances in Neural Information Processing Systems (NeurIPS). 3788–3800
2017
-
[48]
Pedro Sequeira and Melinda Gervasio. 2020. Interestingness elements for explainable reinforcement learning: Understanding agents’ capabilities and limitations.Artificial Intelligence288 (2020), 103367. https://doi.org/10.1016/j.artint.2020.103367
arXiv 2020
-
[49]
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al. 2016. Mastering the game of Go with deep neural networks and tree search. Nature529, 7587 (2016), 484–489. https://doi.org/10.1038/nature16961
-
[50]
Sarath Sreedharan, Siddharth Srivastava, and Subbarao Kambhampati. 2020. TLdR: Policy Summarization for Factored SSP Problems Using Temporal Abstractions. InProceedings of the 30th International Conference on Automated Planning and Scheduling
2020
-
[51]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto. 2018.Reinforcement Learning: An Introduction(2nd ed.). MIT Press
2018
-
[52]
Nicholay Topin, Stephanie Milani, Fei Fang, and Manuela Veloso. 2021. Iterative bounding mdps: Learning interpretable policies via non-interpretable methods. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 9923–9931
2021
-
[53]
Abhinav Verma, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, and Swarat Chaudhuri. 2018. Programmatically Interpretable Reinforcement Learning. InProceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 5045–5054
2018
-
[54]
Rodríguez-Aguilar
Celeste Veronese, Daniele Meli, Filippo Bistaffa, Manel Rodríguez-Sot, Alessandro Farinelli, and Juan A. Rodríguez-Aguilar. 2023. Inductive Logic Programming For Transparent Alignment With Multiple Moral Values.CEUR Workshop Proceedings3615 (2023), 84 – 88
2023
-
[55]
Celeste Veronese, Daniele Meli, and Alessandro Farinelli. 2025. Online Inductive Learning from Answer Sets for Efficient Reinforcement Learning Exploration. InHybrid Models for Coupling Deductive and Inductive Reasoning. Springer Nature Switzerland, 93–106
2025
-
[56]
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. 2022. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. InProceedings of the 36th International Conference on Neural Information Processing Systems(New Orleans, LA, USA)(NeurIPS ’22). Curran Associates Inc., Red Hook, NY, USA, Article 1787, 14 pages
2022
-
[57]
Tom Zahavy, Nir Ben-Zrihem, and Shie Mannor. 2016. Graying the black box: Understanding DQNs. InProceedings of The 33rd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 48), Maria Florina Balcan and Kilian Q. Weinberger (Eds.). PMLR, New York, New York, USA, 1899–1908
2016
-
[58]
Zheng Zhang, Levent Yilmaz, and Bo Liu. 2024. A Critical Review of Inductive Logic Programming Techniques for Explainable AI.IEEE Transactions on Neural Networks and Learning Systems35, 8 (2024), 10220–10236. https://doi.org/10.1109/TNNLS.2023.3246980 JAIR, Vol. 1, Article . Publication date: July 2026. ILP for XRL•29 A Complete results for the experiment...
arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.