Pith. sign in

REVIEW 2 cited by

k-Means Maximum Entropy Exploration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.15623 v4 pith:4I7VUJEE submitted 2022-05-31 cs.LG

classification cs.LG
keywords explorationentropyrewardscontinuousdistributionhigh-dimensionallearningreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Exploration in high-dimensional, continuous spaces with sparse rewards is an open problem in reinforcement learning. Artificial curiosity algorithms address this by creating rewards that lead to exploration. Given a reinforcement learning algorithm capable of maximizing rewards, the problem reduces to finding an optimization objective consistent with exploration. Maximum entropy exploration uses the entropy of the state visitation distribution as such an objective. However, efficiently estimating the entropy of the state visitation distribution is challenging in high-dimensional, continuous spaces. We introduce an artificial curiosity algorithm based on lower bounding an approximation to the entropy of the state visitation distribution. The bound relies on a result we prove for non-parametric density estimation in arbitrary dimensions using k-means. We show that our approach is both computationally efficient and competitive on benchmarks for exploration in high-dimensional, continuous spaces, especially on tasks where reinforcement learning algorithms are unable to find rewards.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy

    cs.LG 2024-12 conditional novelty 6.0 of 10

    ELEMENT combines an average episodic state entropy reward with a kNN-graph lifelong entropy reward for reward-free RL exploration.

  2. Enhancing Diversity in Parallel Agents: A Maximum State Entropy Exploration Story

    cs.LG 2025-05 reject novelty 4.0 of 10

    A centralized policy gradient for parallel state entropy maximization improves state coverage on small gridworlds, but the paper's concentration-rate proof is invalid.

Pith tools