Pith. sign in

REVIEW 4 cited by

Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.16158 v3 pith:7TH3VAOB submitted 2024-05-25 cs.LG

classification cs.LG
keywords scalingoptimisticalgorithmbiggercontrolenhancementsmodel-freeregularized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sample efficiency in Reinforcement Learning (RL) has traditionally been driven by algorithmic enhancements. In this work, we demonstrate that scaling can also lead to substantial improvements. We conduct a thorough investigation into the interplay of scaling model capacity and domain-specific RL enhancements. These empirical findings inform the design choices underlying our proposed BRO (Bigger, Regularized, Optimistic) algorithm. The key innovation behind BRO is that strong regularization allows for effective scaling of the critic networks, which, paired with optimistic exploration, leads to superior performance. BRO achieves state-of-the-art results, significantly outperforming the leading model-based and model-free algorithms across 40 complex tasks from the DeepMind Control, MetaWorld, and MyoSuite benchmarks. BRO is the first model-free algorithm to achieve near-optimal policies in the notoriously challenging Dog and Humanoid tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A constrained rational activation with denominator degree one larger than numerator and no constant term stabilizes high-UTD continuous control, while trading off long-term plasticity.

  2. The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Introduces LEAST, an adaptive early-episode-stopping rule for off-policy deep RL that improves learning efficiency on MuJoCo and DeepMind Control benchmarks.

  3. A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Forget and Grow (FoG) combines decaying replay weights for old experiences with progressive critic-network expansion to improve continuous-control reinforcement learning, beating BRO, SimBa, and TD-MPC2 on most of 41 ...

  4. Is Exploration or Optimization the Problem for Deep Reinforcement Learning?

    cs.LG 2025-08 reject novelty 4.0 of 10

    Deep RL agents' best experienced trajectories are 2-3 times better than their learned policy's average return, suggesting exploitation and optimization issues dominate exploration challenges.

Pith tools