Pith. sign in

REVIEW 2 cited by

Making Reinforcement Learning Work on Swimmer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.07587 v3 pith:GDDWJC2D submitted 2022-08-16 cs.LG

classification cs.LG
keywords methodsswimmerperformancealgorithmsdirecthyper-parameterlearningoften
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The SWIMMER environment is a standard benchmark in reinforcement learning (RL). In particular, it is often used in papers comparing or combining RL methods with direct policy search methods such as genetic algorithms or evolution strategies. A lot of these papers report poor performance on SWIMMER from RL methods and much better performance from direct policy search methods. In this technical report we show that the low performance of RL methods on SWIMMER simply comes from the inadequate tuning of an important hyper-parameter, the discount factor. Furthermore we show that, by setting this hyper-parameter to a correct value, the issue can be easily fixed. Finally, for a set of often used RL algorithms, we provide a set of successful hyper-parameters obtained with the Stable Baselines3 library and its RL Zoo.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

    cs.LG 2026-08 conditional novelty 6.0 of 10

    SP3O is a reward-model-free, critic-free, gradient-based RL algorithm that optimizes policies from segment-level preferences in stochastic MDPs via off-policy importance sampling and PPO-style clipping.

  2. Average-Reward Soft Actor-Critic

    cs.LG 2025-01 reject novelty 4.0 of 10

    ASAC extends soft actor-critic to the entropy-regularized average-reward setting with a policy improvement theorem, but its claimed novelty is undermined by the earlier RVI-SAC algorithm.

Pith tools