Pith. sign in

hub

Minigrid & mini- world: Modular & customizable reinforcement learning environments for goal-oriented tasks

19 Pith papers cite this work, alongside 3 external citations. Polarity classification is still indexing.

19 Pith papers citing it
3 external citations · Pith
abstract

We present the Minigrid and Miniworld libraries which provide a suite of goal-oriented 2D and 3D environments. The libraries were explicitly created with a minimalistic design paradigm to allow users to rapidly develop new environments for a wide range of research-specific needs. As a result, both have received widescale adoption by the RL community, facilitating research in a wide range of areas. In this paper, we outline the design philosophy, environment details, and their world generation API. We also showcase the additional capabilities brought by the unified API between Minigrid and Miniworld through case studies on transfer learning (for both RL agents and humans) between the different observation spaces. The source code of Minigrid and Miniworld can be found at https://github.com/Farama-Foundation/{Minigrid, Miniworld} along with their documentation at https://{minigrid, miniworld}.farama.org/.

hub tools

citation-role summary

dataset 2

citation-polarity summary

roles

dataset 2

representative citing papers

Delay-Empowered Causal Hierarchical Reinforcement Learning

cs.LG · 2026-05-12 · unverdicted · novelty 6.0

DECHRL models causal structures and stochastic delay distributions within hierarchical RL and incorporates them into a delay-aware empowerment objective to improve performance under temporal uncertainty.

Sample-efficient Neuro-symbolic Proximal Policy Optimization

cs.AI · 2026-04-28 · unverdicted · novelty 6.0

H-PPO-Product and H-PPO-SymLoss achieve faster learning and higher final returns than standard PPO and Reward Machine baselines on OfficeWorld, WaterWorld, and DoorKey by transferring imperfect logical policy specifications from easier to harder instances.

Sample-Efficient Neurosymbolic Deep Reinforcement Learning

cs.AI · 2026-01-06 · unverdicted · novelty 5.0

A neuro-symbolic DRL approach transfers partial policies as logical rules to bias exploration and rescale Q-values, showing improved performance over reward machine baselines in gridworld environments.

Polychromic Objectives for Reinforcement Learning

cs.LG · 2025-09-29 · unverdicted · novelty 5.0

Introduces polychromic objectives adapted into PPO via vine sampling and modified advantages, showing higher success rates and better coverage under perturbations on BabyAI, Minigrid, and algorithmic tasks.

citing papers explorer

Showing 19 of 19 citing papers.