Pith. sign in

REVIEW 3 cited by

Knowledge-Guided Exploration in Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.15670 v1 pith:DEBG4XDZ submitted 2022-10-26 cs.LG cs.AI

classification cs.LGcs.AI
keywords actiondeeppermissibleagentpermissibilitystatedecideexploration
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

This paper proposes a new method to drastically speed up deep reinforcement learning (deep RL) training for problems that have the property of state-action permissibility (SAP). Two types of permissibility are defined under SAP. The first type says that after an action $a_t$ is performed in a state $s_t$ and the agent has reached the new state $s_{t+1}$, the agent can decide whether $a_t$ is permissible or not permissible in $s_t$. The second type says that even without performing $a_t$ in $s_t$, the agent can already decide whether $a_t$ is permissible or not in $s_t$. An action is not permissible in a state if the action can never lead to an optimal solution and thus should not be tried (over and over again). We incorporate the proposed SAP property and encode action permissibility knowledge into two state-of-the-art deep RL algorithms to guide their state-action exploration together with a virtual stopping strategy. Results show that the SAP-based guidance can markedly speed up RL training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

    cs.AI 2026-07 unverdicted novelty 6.0 of 10

    KGRL prunes invalid parametrized actions via a Datalog knowledge base and refines continuous parameters by gradient descent, reporting better sample efficiency and return than PAMDP RL baselines.

  2. Evolution, Future of AI, and Singularity

    cs.OH 2025-06 conditional novelty 6.0 of 10

    The paper translates principles from evolutionary developmental biology into a new AI design paradigm that promises continual learning, structured representations, and a grounded path to technological singularity.

  3. On the Parallels Between Evolutionary Theory and the State of AI

    q-bio.NC 2025-05 conditional novelty 5.0 of 10

    The authors propose that evolutionary developmental biology's principles, such as encapsulation, regulatory control, and local variation-selection, can overcome continual learning and explainability limits of current ...

Pith tools