UED is reframed as an entropy-regularized minimax problem with two-timescale gradient convergence guarantees for zero-sum scores, and a generalized learnability score improves robustness on three benchmarks, with the best empirical variants lying outside the guaranteed regime.
Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
An Optimisation Framework for Unsupervised Environment Design
UED is reframed as an entropy-regularized minimax problem with two-timescale gradient convergence guarantees for zero-sum scores, and a generalized learnability score improves robustness on three benchmarks, with the best empirical variants lying outside the guaranteed regime.