Pith. sign in

REVIEW 1 cited by

Refining Minimax Regret for Unsupervised Environment Design

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12284 v2 pith:SUVGANQJ submitted 2024-02-19 cs.LG cs.AI

classification cs.LGcs.AI
keywords regretlevelslearningminimaxobjectiveadversaryenvironmentpolicy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In unsupervised environment design, reinforcement learning agents are trained on environment configurations (levels) generated by an adversary that maximises some objective. Regret is a commonly used objective that theoretically results in a minimax regret (MMR) policy with desirable robustness guarantees; in particular, the agent's maximum regret is bounded. However, once the agent reaches this regret bound on all levels, the adversary will only sample levels where regret cannot be further reduced. Although there are possible performance improvements to be made outside of these regret-maximising levels, learning stagnates. In this work, we introduce Bayesian level-perfect MMR (BLP), a refinement of the minimax regret objective that overcomes this limitation. We formally show that solving for this objective results in a subset of MMR policies, and that BLP policies act consistently with a Perfect Bayesian policy over all levels. We further introduce an algorithm, ReMiDi, that results in a BLP policy at convergence. We empirically demonstrate that training on levels from a minimax regret adversary causes learning to prematurely stagnate, but that ReMiDi continues learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predicting Task Difficulty Without Rollouts

    cs.LG 2026-08 conditional novelty 6.0 of 10

    Pre-rollout task difficulty for agentic benchmarks is predictable from token-level entropy features, with Spearman rho=0.399 in-distribution and 0.225 out-of-distribution.

Pith tools