REVIEW 7 cited by
MineRL: A Large-Scale Dataset of Minecraft Demonstrations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The sample inefficiency of standard deep reinforcement learning methods precludes their application to many real-world problems. Methods which leverage human demonstrations require fewer samples but have been researched less. As demonstrated in the computer vision and natural language processing communities, large-scale datasets have the capacity to facilitate research by serving as an experimental and benchmarking platform for new methods. However, existing datasets compatible with reinforcement learning simulators do not have sufficient scale, structure, and quality to enable the further development and evaluation of methods focused on using human examples. Therefore, we introduce a comprehensive, large-scale, simulator-paired dataset of human demonstrations: MineRL. The dataset consists of over 60 million automatically annotated state-action pairs across a variety of related tasks in Minecraft, a dynamic, 3D, open-world environment. We present a novel data collection scheme which allows for the ongoing introduction of new tasks and the gathering of complete state information suitable for a variety of methods. We demonstrate the hierarchality, diversity, and scale of the MineRL dataset. Further, we show the difficulty of the Minecraft domain along with the potential of MineRL in developing techniques to solve key research challenges within it.
Forward citations
Cited by 7 Pith papers
-
Mitigating Compounding Error via Video Representation Regularization
Compounding error in autoregressive video diffusion tracks effective-rank collapse of DiT hidden states, and representation regularization (SigReg/Unif) stabilizes long rollouts where data scaling does not.
-
BuilderBench: The Building Blocks of Intelligent Agents
BuilderBench is a fast, open-source 3D block-building benchmark where current RL and LLM agents fail at all non-trivial construction tasks, exposing weak open-ended exploration.
-
PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
A new open Minecraft benchmark for 2v2 LLM-agent competition, and a system, TactiCrafter, that beats its baselines on points and win rate.
-
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
StarDojo is a 1,000-task benchmark in Stardew Valley combining production and social activities, and the best tested MLLM (GPT-4.1) achieves only 12.7% success on its 100-task subset.
-
Hierarchical Reinforcement Learning with Targeted Causal Interventions
HRC learns the causal structure among subgoals and prioritizes interventions on the subgoals that matter most for the final goal, lowering training cost in hierarchical RL.
-
Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
Augmenting a diffusion video transformer with an RNN memory block and frame-wise overlapping attention improves long-horizon consistency, with simple LSTM matching newer Mamba2 and TTT memory blocks.
-
A Comprehensive Review of Multi-Agent Reinforcement Learning in Video Games
A survey of multi-agent reinforcement learning in video games, plus a proposed five-dimension, MDP-based classification for comparing game complexity.
Discussion (0). Sign in to comment.