Pith. sign in

Graph neural induction of value iteration

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Many reinforcement learning tasks can benefit from explicit planning based on an internal model of the environment. Previously, such planning components have been incorporated through a neural network that partially aligns with the computational graph of value iteration. Such network have so far been focused on restrictive environments (e.g. grid-worlds), and modelled the planning procedure only indirectly. We relax these constraints, proposing a graph neural network (GNN) that executes the value iteration (VI) algorithm, across arbitrary environment models, with direct supervision on the intermediate steps of VI. The results indicate that GNNs are able to model value iteration accurately, recovering favourable metrics and policies across a variety of out-of-distribution tests. This suggests that GNN executors with strong supervision are a viable component within deep reinforcement learning systems.

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Unrolling Dynamic Programming via Graph Filters

cs.AI · 2025-07-29 · conditional · novelty 5.0

BellNet learns graph-filter coefficients for truncated policy iteration and approximates optimal policies in fewer steps than classical DP on grid-world tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Unrolling Dynamic Programming via Graph Filters cs.AI · 2025-07-29 · conditional · none · ref 13 · internal anchor

    BellNet learns graph-filter coefficients for truncated policy iteration and approximates optimal policies in fewer steps than classical DP on grid-world tasks.