Pith. sign in

Incorporating Relational Background Knowledge into Reinforcement Learning via Differentiable Inductive Logic Programming

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Relational Reinforcement Learning (RRL) can offers various desirable features. Most importantly, it allows for incorporating expert knowledge into the learning, and hence leading to much faster learning and better generalization compared to the standard deep reinforcement learning. However, most of the existing RRL approaches are either incapable of incorporating expert background knowledge (e.g., in the form of explicit predicate language) or are not able to learn directly from non-relational data such as image. In this paper, we propose a novel deep RRL based on a differentiable Inductive Logic Programming (ILP) that can effectively learn relational information from image and present the state of the environment as first order logic predicates. Additionally, it can take the expert background knowledge and incorporate it into the learning problem using appropriate predicates. The differentiable ILP allows an end to end optimization of the entire framework for learning the policy in RRL. We show the efficacy of this novel RRL framework using environments such as BoxWorld, GridWorld as well as relational reasoning for the Sort-of-CLEVR dataset.

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Neural Logic Networks for Interpretable Classification

cs.LG · 2025-08-11 · conditional · novelty 7.0

A generalized neural logic network with negative weights, unobserved-data biases, and a factorized rule structure learns sparse IF-THEN rules and recovers Boolean network ground truth from partial data.

citing papers explorer

Showing 1 of 1 citing paper.

  • Neural Logic Networks for Interpretable Classification cs.LG · 2025-08-11 · conditional · none · ref 2020 · internal anchor

    A generalized neural logic network with negative weights, unobserved-data biases, and a factorized rule structure learns sparse IF-THEN rules and recovers Boolean network ground truth from partial data.