Pith. sign in

REVIEW 2 cited by

EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.00564 v2 pith:OJJKBIC4 submitted 2024-03-01 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords efficientzerocontroldiversetasksacrossalgorithmscontinuousdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sample efficiency remains a crucial challenge in applying Reinforcement Learning (RL) to real-world tasks. While recent algorithms have made significant strides in improving sample efficiency, none have achieved consistently superior performance across diverse domains. In this paper, we introduce EfficientZero V2, a general framework designed for sample-efficient RL algorithms. We have expanded the performance of EfficientZero to multiple domains, encompassing both continuous and discrete actions, as well as visual and low-dimensional inputs. With a series of improvements we propose, EfficientZero V2 outperforms the current state-of-the-art (SOTA) by a significant margin in diverse tasks under the limited data setting. EfficientZero V2 exhibits a notable advancement over the prevailing general algorithm, DreamerV3, achieving superior outcomes in 50 of 66 evaluated tasks across diverse benchmarks, such as Atari 100k, Proprio Control, and Vision Control.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach

    cs.AI 2025-10 conditional novelty 6.0 of 10

    A sample-efficient SAC-based method with replay-ratio resets, offline data bootstrapping, and expert-driven fine-tuning produces a goalkeeper that outperforms the built-in AI in EA SPORTS FC 25.

  2. Sample-Efficient Reinforcement Learning Controller for Deep Brain Stimulation in Parkinson's Disease

    cs.LG 2025-07 reject novelty 5.0 of 10

    A DDPG-based adaptive DBS controller with a predictive reward model and Gumbel-Softmax exploration suppresses beta power faster than standard DDPG in simulation and survives FP16 quantization.

Pith tools