Pith. sign in

REVIEW 1 cited by

On the consistency of hyper-parameter selection in value-based deep reinforcement learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17523 v3 pith:K6W7OJYM submitted 2024-06-25 cs.LG cs.AI

classification cs.LGcs.AI
keywords deephyper-parameteralgorithmichyper-parameterslearningreinforcementselectionchoices
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep reinforcement learning (deep RL) has achieved tremendous success on various domains through a combination of algorithmic design and careful selection of hyper-parameters. Algorithmic improvements are often the result of iterative enhancements built upon prior approaches, while hyper-parameter choices are typically inherited from previous methods or fine-tuned specifically for the proposed technique. Despite their crucial impact on performance, hyper-parameter choices are frequently overshadowed by algorithmic advancements. This paper conducts an extensive empirical study focusing on the reliability of hyper-parameter selection for value-based deep reinforcement learning agents, including the introduction of a new score to quantify the consistency and reliability of various hyper-parameters. Our findings not only help establish which hyper-parameters are most critical to tune, but also help clarify which tunings remain consistent across different training regimes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts

    cs.CY 2024-12 conditional novelty 5.0 of 10

    A framework that systematizes, operationalizes, and applies amounts, concepts, instances, and populations for valid GenAI measurement, extending Adcock and Collier's measurement theory.

Pith tools