Pith. sign in

REVIEW 1 cited by

EXPIL: Explanatory Predicate Invention for Learning in Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.06107 v1 pith:PAZQZL3V submitted 2024-06-10 cs.AI

classification cs.AI
keywords agentsgamesbackgroundexpilknowledgelearningneuralagent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning (RL) has proven to be a powerful tool for training agents that excel in various games. However, the black-box nature of neural network models often hinders our ability to understand the reasoning behind the agent's actions. Recent research has attempted to address this issue by using the guidance of pretrained neural agents to encode logic-based policies, allowing for interpretable decisions. A drawback of such approaches is the requirement of large amounts of predefined background knowledge in the form of predicates, limiting its applicability and scalability. In this work, we propose a novel approach, Explanatory Predicate Invention for Learning in Games (EXPIL), that identifies and extracts predicates from a pretrained neural agent, later used in the logic-based agents, reducing the dependency on predefined background knowledge. Our experimental evaluation on various games demonstrate the effectiveness of EXPIL in achieving explainable behavior in logic agents while requiring less background knowledge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A trained PPO policy is distilled into an executable first-order Prolog decision list that, after exact-return expansion, can match or exceed the teacher, with certified return loss and an O(1/B) fidelity/resolution theory.

Pith tools