Pith. sign in

REVIEW 1 cited by

Hieros: Hierarchical Imagination on Structured State Space Sequence World Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.05167 v3 pith:5UM2DXQW submitted 2023-10-08 cs.AI

classification cs.AI
keywords worldimaginationduringhierosmodelmodelstrainingapproaches
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

One of the biggest challenges to modern deep reinforcement learning (DRL) algorithms is sample efficiency. Many approaches learn a world model in order to train an agent entirely in imagination, eliminating the need for direct environment interaction during training. However, these methods often suffer from either a lack of imagination accuracy, exploration capabilities, or runtime efficiency. We propose Hieros, a hierarchical policy that learns time abstracted world representations and imagines trajectories at multiple time scales in latent space. Hieros uses an S5 layer-based world model, which predicts next world states in parallel during training and iteratively during environment interaction. Due to the special properties of S5 layers, our method can train in parallel and predict next world states iteratively during imagination. This allows for more efficient training than RNN-based world models and more efficient imagination than Transformer-based world models. We show that our approach outperforms the state of the art in terms of mean and median normalized human score on the Atari 100k benchmark, and that our proposed world model is able to predict complex dynamics very accurately. We also show that Hieros displays superior exploration capabilities compared to existing approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GLAM: Global-Local Variation Awareness in Mamba-based World Model

    cs.LG 2025-01 conditional novelty 6.0 of 10

    GLAM improves world model prediction in model-based RL by feeding state differences into two parallel Mamba modules and training agents on imagined variation-aware trajectories.

Pith tools