Pith. sign in

REVIEW 2 cited by

Generalized Planning With Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.02305 v1 pith:HPVQSPGH submitted 2020-05-05 cs.AI cs.LG

classification cs.AIcs.LG
keywords generalizedinstancesplanningprinciplesdeepdomainlargerlearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A hallmark of intelligence is the ability to deduce general principles from examples, which are correct beyond the range of those observed. Generalized Planning deals with finding such principles for a class of planning problems, so that principles discovered using small instances of a domain can be used to solve much larger instances of the same domain. In this work we study the use of Deep Reinforcement Learning and Graph Neural Networks to learn such generalized policies and demonstrate that they can generalize to instances that are orders of magnitude larger than those they were trained on.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Per-Domain Generalizing Policies: On Validation Instances and Scaling Behavior

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Dynamically generating larger validation instances during training selects planning policies that generalize to larger instances better than fixed validation sets, across all 9 domains tested.

  2. Relational GNNs Cannot Learn $C_2$ Features for Planning

    cs.AI 2025-06 conditional novelty 4.0 of 10

    Relational GNNs cannot represent C2 logic features for planning value functions, as shown by an indistinguishability counterexample and experiments, while PLOI-style graph encodings can.

Pith tools