Pith. sign in

REVIEW 6 cited by

URLB: Unsupervised Reinforcement Learning Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.15191 v1 pith:AC6JJ5VU submitted 2021-10-28 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords unsupervisedurlbbenchmarkcontrollearningreinforcementtasksadaptation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep Reinforcement Learning (RL) has emerged as a powerful paradigm to solve a range of complex yet specific control tasks. Yet training generalist agents that can quickly adapt to new tasks remains an outstanding challenge. Recent advances in unsupervised RL have shown that pre-training RL agents with self-supervised intrinsic rewards can result in efficient adaptation. However, these algorithms have been hard to compare and develop due to the lack of a unified benchmark. To this end, we introduce the Unsupervised Reinforcement Learning Benchmark (URLB). URLB consists of two phases: reward-free pre-training and downstream task adaptation with extrinsic rewards. Building on the DeepMind Control Suite, we provide twelve continuous control tasks from three domains for evaluation and open-source code for eight leading unsupervised RL methods. We find that the implemented baselines make progress but are not able to solve URLB and propose directions for future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Ring Attention with Blockwise Transformers for Near-Infinite Context

    cs.CL 2023-10 unverdicted novelty 7.0 of 10

    Ring Attention uses blockwise computation and ring communication to let Transformers process sequences up to device-count times longer than prior memory-efficient methods.

  2. Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    ROVER pretrains transferable exploration policies by maximizing occupancy coverage with a learned resolvent world model and virtual sink state, outperforming baselines on sparse navigation tasks.

  3. NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control

    cs.GR 2026-04 unverdicted novelty 6.0 of 10

    NaP-Control uses RL to directly predict optimized diffusion noise from a task-agnostic prior, enabling fast inference and higher success rates for versatile whole-body character control while preserving motion quality.

  4. NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control

    cs.GR 2026-04 conditional novelty 5.0 of 10

    A latent-noise navigation policy trained with RL steers a frozen diffusion character-control prior, achieving fast and robust task execution without test-time guidance.

  5. Test-Time Alignment via Hypothesis Reweighting

    cs.LG 2024-12 unverdicted novelty 5.0 of 10

    HyRe personalizes reward models at test time by reweighting an ensemble of heads trained on aggregate preferences, using few target examples to outperform uniform averaging and prior methods on RewardBench and 32 tasks.

  6. Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A survey that categorizes deep reinforcement learning scaling strategies into data, network, and training budget dimensions and outlines challenges for scaling DRL systems.

Pith tools