Pith. sign in

REVIEW 1 cited by

Fast and Data-Efficient Training of Rainbow: an Experimental Study on Atari

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.10247 v1 pith:RWJSQJHR submitted 2021-11-19 cs.LG

classification cs.LG
keywords rainbowdataperformancetrainingarcadecompetitiveenvironmentimproved
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Across the Arcade Learning Environment, Rainbow achieves a level of performance competitive with humans and modern RL algorithms. However, attaining this level of performance requires large amounts of data and hardware resources, making research in this area computationally expensive and use in practical applications often infeasible. This paper's contribution is threefold: We (1) propose an improved version of Rainbow, seeking to drastically reduce Rainbow's data, training time, and compute requirements while maintaining its competitive performance; (2) we empirically demonstrate the effectiveness of our approach through experiments on the Arcade Learning Environment, and (3) we conduct a number of ablation studies to investigate the effect of the individual proposed modifications. Our improved version of Rainbow reaches a median human normalized score close to classic Rainbow's, while using 20 times less data and requiring only 7.5 hours of training time on a single GPU. We also provide our full implementation including pre-trained models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DeGuV: Depth-Guided Visual Reinforcement Learning for Generalization and Interpretability in Manipulation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    Depth-guided masking improves visual RL generalization, sample efficiency, and interpretability on manipulation tasks.

Pith tools