Pith. sign in

REVIEW 1 cited by

Hierarchical Reinforcement Learning via Advantage-Weighted Information Maximization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.01365 v2 pith:7MR74TPL submitted 2019-01-05 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords policyhierarchicaloptionapproachlearnlearningmethodpolicies
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Real-world tasks are often highly structured. Hierarchical reinforcement learning (HRL) has attracted research interest as an approach for leveraging the hierarchical structure of a given task in reinforcement learning (RL). However, identifying the hierarchical policy structure that enhances the performance of RL is not a trivial task. In this paper, we propose an HRL method that learns a latent variable of a hierarchical policy using mutual information maximization. Our approach can be interpreted as a way to learn a discrete and latent representation of the state-action space. To learn option policies that correspond to modes of the advantage function, we introduce advantage-weighted importance sampling. In our HRL method, the gating policy learns to select option policies based on an option-value function, and these option policies are optimized based on the deterministic policy gradient method. This framework is derived by leveraging the analogy between a monolithic policy in standard RL and a hierarchical policy in HRL by using a deterministic option policy. Experimental results indicate that our HRL approach can learn a diversity of options and that it can enhance the performance of RL in continuous control tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Temporal Abstractions via Variational Homomorphisms in Option-Induced Abstract MDPs

    cs.AI 2025-07 reject novelty 6.0 of 10

    A variational option-critic algorithm with latent option embeddings and an implicit chain-of-thought cold-start is presented; the central optimality-preservation proof has a gap and some reported benchmark wins are in...

Pith tools