Pith. sign in

REVIEW 2 cited by

Accuracy-based Curriculum Learning in Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.09614 v2 pith:542UXZLW submitted 2018-06-25 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningcurriculumaccuracyrequirementsaccuracy-basedadaptiveagentdeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we investigate a new form of automated curriculum learning based on adaptive selection of accuracy requirements, called accuracy-based curriculum learning. Using a reinforcement learning agent based on the Deep Deterministic Policy Gradient algorithm and addressing the Reacher environment, we first show that an agent trained with various accuracy requirements sampled randomly learns more efficiently than when asked to be very accurate at all times. Then we show that adaptive selection of accuracy requirements, based on a local measure of competence progress, automatically generates a curriculum where difficulty progressively increases, resulting in a better learning efficiency than sampling randomly.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Skill Discovery via Regret-Aware Optimization

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A regret-aware skill discovery algorithm, RSD, improves sample efficiency and zero-shot goal-reaching in high-dimensional continuous control by focusing exploration on unmastered skills.

  2. Scenario Generation for Risk-Aware Reinforcement Learning with Probably Approximately Safe Guarantees

    cs.LG 2026-06 unverdicted novelty 4.0 of 10

    Approximates encountered state distribution via VAE and constructs dual bound barrier certificates to provide probably approximately safe guarantees in RL by optimizing the non-robust region.

Pith tools