Pith. sign in

REVIEW 6 cited by

Leveraging Procedural Generation to Benchmark Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.01588 v2 pith:TS2PYK7T submitted 2019-12-03 cs.LG stat.ML

classification cs.LGstat.ML
keywords benchmarkefficiencyenvironmentsgeneralizationgenerationlearningproceduralreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Procgen Benchmark, a suite of 16 procedurally generated game-like environments designed to benchmark both sample efficiency and generalization in reinforcement learning. We believe that the community will benefit from increased access to high quality training environments, and we provide detailed experimental protocols for using this benchmark. We empirically demonstrate that diverse environment distributions are essential to adequately train and evaluate RL agents, thereby motivating the extensive use of procedural content generation. We then use this benchmark to investigate the effects of scaling model size, finding that larger models significantly improve both sample efficiency and generalization.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 171 citations worldwide. Full citation record

  1. Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics

    cs.AI 2025-07 conditional novelty 7.0 of 10

    Assistax provides fast JAX-based assistive robotics environments with trainable humanoid partners, and shows current RL baselines have a coordination gap when facing unseen human preferences.

  2. Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A new collection of 50 procedural generators designed for completion-supervised fine-tuning beats three existing procedural collections and a no-procedural baseline on reasoning benchmarks at 3B scale in mean scores.

  3. On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning

    cs.LG 2026-02 conditional novelty 6.0 of 10

    VM-IDM and IDM labeling coincide at infinite unlabeled data, and IDM learning is more label-efficient than BC because the ground-truth IDM is typically less complex and less stochastic than the expert policy.

  4. Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A JAX benchmark suite of nine memory-improvable partially observable RL environments, with evidence that recurrent and transformer agents beat memoryless baselines but fall short of state-augmented ceilings.

  5. Dyn-O: Building Structured World Models with Object-Centric Representations

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Dyn-O learns object-centric world models directly from pixels in complex Procgen games, using SAM2-guided slot attention and Mamba state-space dynamics, and reports better rollout prediction than DreamerV3.

  6. An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models

    cs.LG 2025-06 conditional novelty 4.0 of 10

    MultiNet provides an open-source benchmark, data SDK, evaluation harness, and adapted VLA models for assessing generalization across vision, language, and action tasks.

Pith tools