Pith. sign in

REVIEW 1 cited by

Intrinsic Motivation in Model-based Reinforcement Learning: A Brief Review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.10067 v1 pith:SFJBFLHS submitted 2023-01-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords intrinsicagentmotivationlearningmethodsmodelresearchworld
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The reinforcement learning research area contains a wide range of methods for solving the problems of intelligent agent control. Despite the progress that has been made, the task of creating a highly autonomous agent is still a significant challenge. One potential solution to this problem is intrinsic motivation, a concept derived from developmental psychology. This review considers the existing methods for determining intrinsic motivation based on the world model obtained by the agent. We propose a systematic approach to current research in this field, which consists of three categories of methods, distinguished by the way they utilize a world model in the agent's components: complementary intrinsic reward, exploration policy, and intrinsically motivated goals. The proposed unified framework describes the architecture of agents using a world model and intrinsic motivation to improve learning. The potential for developing new techniques in this area of research is also examined.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bounded Exploration with World Model Uncertainty in Soft Actor-Critic Reinforcement Learning Algorithm

    cs.LG 2024-12 reject novelty 4.0 of 10

    Proposes bounded exploration, selecting high world-model uncertainty actions from SAC's sampled candidates, and reports mixed, statistically weak results on MuJoCo benchmarks.

Pith tools