Pith. sign in

REVIEW 1 cited by

Unsupervised Skill-Discovery and Skill-Learning in Minecraft

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.08398 v1 pith:QEA7W4TF submitted 2021-07-18 cs.AI cs.LG

classification cs.AIcs.LG
keywords learningagentslearnmapsminecraftpixelsrepresentationsresults
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-training Reinforcement Learning agents in a task-agnostic manner has shown promising results. However, previous works still struggle in learning and discovering meaningful skills in high-dimensional state-spaces, such as pixel-spaces. We approach the problem by leveraging unsupervised skill discovery and self-supervised learning of state representations. In our work, we learn a compact latent representation by making use of variational and contrastive techniques. We demonstrate that both enable RL agents to learn a set of basic navigation skills by maximizing an information theoretic objective. We assess our method in Minecraft 3D pixel maps with different complexities. Our results show that representations and conditioned policies learned from pixels are enough for toy examples, but do not scale to realistic and complex maps. To overcome these limitations, we explore alternative input observations such as the relative position of the agent along with the raw pixels.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unsupervised Skill Discovery through Skill Regions Differentiation

    cs.LG 2025-06 conditional novelty 5.0 of 10

    SD3 separates skills by maximizing each skill's state density deviation from other skills and adds a VAE-based latent-space exploration reward, achieving modest benchmark gains.

Pith tools