Pith. sign in

REVIEW 1 cited by

SAGE: Generating Symbolic Goals for Myopic Models in Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.05079 v1 pith:PARTMM2P submitted 2022-03-09 cs.LG cs.AI

SAGE: Generating Symbolic Goals for Myopic Models in Deep Reinforcement Learning

classification cs.LG cs.AI
keywords learningmodelsdomainsplanningapproachesenvironmentincompletemany
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Model-based reinforcement learning algorithms are typically more sample efficient than their model-free counterparts, especially in sparse reward problems. Unfortunately, many interesting domains are too complex to specify the complete models required by traditional model-based approaches. Learning a model takes a large number of environment samples, and may not capture critical information if the environment is hard to explore. If we could specify an incomplete model and allow the agent to learn how best to use it, we could take advantage of our partial understanding of many domains. Existing hybrid planning and learning systems which address this problem often impose highly restrictive assumptions on the sorts of models which can be used, limiting their applicability to a wide range of domains. In this work we propose SAGE, an algorithm combining learning and planning to exploit a previously unusable class of incomplete models. This combines the strengths of symbolic planning and neural learning approaches in a novel way that outperforms competing methods on variations of taxi world and Minecraft.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Dependency-Guided Code Generation: Structured Matrix Decomposition and Consistency-Guided Refinement

    cs.SE 2026-07 conditional novelty 4.0

    A dependency-aware code generation method that decomposes code-dependency matrices into quantized and low-rank components and uses them in a consistency-guided retrieval-refinement loop.