MASP meta-learns a similarity matrix over macro-actions and regularizes Q-values so that similar actions move together, improving exploration and performance in augmented-action-space RL.
The Atari Grand Challenge Dataset
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Recent progress in Reinforcement Learning (RL), fueled by its combination, with Deep Learning has enabled impressive results in learning to interact with complex virtual environments, yet real-world applications of RL are still scarce. A key limitation is data efficiency, with current state-of-the-art approaches requiring millions of training samples. A promising way to tackle this problem is to augment RL with learning from human demonstrations. However, human demonstration data is not yet readily available. This hinders progress in this direction. The present work addresses this problem as follows. We (i) collect and describe a large dataset of human Atari 2600 replays -- the largest and most diverse such data set publicly released to date, (ii) illustrate an example use of this dataset by analyzing the relation between demonstration quality and imitation learning performance, and (iii) outline possible research directions that are opened up by our work.
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Meta-learning how to Share Credit among Macro-Actions
MASP meta-learns a similarity matrix over macro-actions and regularizes Q-values so that similar actions move together, improving exploration and performance in augmented-action-space RL.