Pith. sign in

REVIEW 3 cited by

Relational Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.01830 v2 pith:YF5BEVD4 submitted 2018-06-05 cs.LG stat.ML

Relational Deep Reinforcement Learning

classification cs.LG stat.ML
keywords deeplearningperformanceagentreinforcementrelationaluponability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce an approach for deep reinforcement learning (RL) that improves upon the efficiency, generalization capacity, and interpretability of conventional approaches through structured perception and relational reasoning. It uses self-attention to iteratively reason about the relations between entities in a scene and to guide a model-free policy. Our results show that in a novel navigation and planning task called Box-World, our agent finds interpretable solutions that improve upon baselines in terms of sample complexity, ability to generalize to more complex scenes than experienced during training, and overall performance. In the StarCraft II Learning Environment, our agent achieves state-of-the-art performance on six mini-games -- surpassing human grandmaster performance on four. By considering architectural inductive biases, our work opens new directions for overcoming important, but stubborn, challenges in deep RL.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States

    cs.LG 2026-07 conditional novelty 6.0

    Relational hidden states anchored to environment states are what let a model-free RL agent plan, and a free-slot control without that anchoring shows no planning signatures.

  2. Graph-Enhanced Policy Optimization in LLM Agent Training

    cs.AI 2025-10 conditional novelty 6.0

    GEPO adds graph-centrality-based intrinsic rewards, dynamic discounts, and two-level advantage shaping to group-based RL, improving LLM agent success on ALFWorld, WebShop, and a private Workbench benchmark.

  3. Graph Neural Based End-to-end Data Association Framework for Online Multiple-Object Tracking

    cs.CV 2019-07 unverdicted novelty 6.0

    A graph neural network framework learns affinities from appearance and motion then solves bipartite matching for online multiple-object tracking.