Pith. sign in

REVIEW 4 cited by

Deep Reinforcement Learning at the Edge of the Statistical Precipice

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.13264 v4 pith:Y5POH3PW submitted 2021-08-30 cs.LG cs.AIstat.MEstat.ML

classification cs.LGcs.AIstat.MEstat.ML
keywords performanceresultsdeepstatisticalestimatespointuncertaintyaggregate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep reinforcement learning (RL) algorithms are predominantly evaluated by comparing their relative performance on a large suite of tasks. Most published results on deep RL benchmarks compare point estimates of aggregate performance such as mean and median scores across tasks, ignoring the statistical uncertainty implied by the use of a finite number of training runs. Beginning with the Arcade Learning Environment (ALE), the shift towards computationally-demanding benchmarks has led to the practice of evaluating only a small number of runs per task, exacerbating the statistical uncertainty in point estimates. In this paper, we argue that reliable evaluation in the few run deep RL regime cannot ignore the uncertainty in results without running the risk of slowing down progress in the field. We illustrate this point using a case study on the Atari 100k benchmark, where we find substantial discrepancies between conclusions drawn from point estimates alone versus a more thorough statistical analysis. With the aim of increasing the field's confidence in reported results with a handful of runs, we advocate for reporting interval estimates of aggregate performance and propose performance profiles to account for the variability in results, as well as present more robust and efficient aggregate metrics, such as interquartile mean scores, to achieve small uncertainty in results. Using such statistical tools, we scrutinize performance evaluations of existing algorithms on other widely used RL benchmarks including the ALE, Procgen, and the DeepMind Control Suite, again revealing discrepancies in prior comparisons. Our findings call for a change in how we evaluate performance in deep RL, for which we present a more rigorous evaluation methodology, accompanied with an open-source library rliable, to prevent unreliable results from stagnating the field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 41 citations worldwide. Full citation record

  1. THBKG: A Temporal Biomedical Knowledge Graph for Decision-Aligned Clinical Advancement Prediction

    cs.LG 2026-08 conditional novelty 7.0 of 10

    A timestamped biomedical knowledge graph lets clinical advancement be predicted from evidence available at decision time; graph propagation outperforms direct-evidence baselines at the top of the ranking.

  2. Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control

    cs.LG 2026-08 conditional novelty 6.0 of 10

    POGP trains a prefix value function over the diffusion denoising chain, giving a learned early-stopping rule that cuts average denoising steps about 2.7x with 98-99% of full-chain return and a small gain in final task...

  3. Social-spatial dependencies for learning visual navigation

    cs.NE 2026-07 conditional novelty 6.0 of 10

    Neural-network agents trained in social environments learn hybrid navigation strategies that combine individual landmark use with social following, with strategy shifts driven by the ratio of skilled to unskilled soci...

  4. Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

    cs.AI 2025-06 conditional novelty 5.0 of 10

    An offline RL policy can be improved at test time by inferring a latent belief over environment dynamics from past transitions and planning with model-based rollouts averaged over that belief.

Pith tools