Pith. sign in

REVIEW 11 cited by

Automatic Curriculum Learning For Deep RL: A Short Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.04664 v2 pith:DEZN4FAC submitted 2020-03-10 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningautomaticcurriculumdeepencouragerecentaccessibleadapted
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic Curriculum Learning (ACL) has become a cornerstone of recent successes in Deep Reinforcement Learning (DRL).These methods shape the learning trajectories of agents by challenging them with tasks adapted to their capacities. In recent years, they have been used to improve sample efficiency and asymptotic performance, to organize exploration, to encourage generalization or to solve sparse reward problems, among others. The ambition of this work is dual: 1) to present a compact and accessible introduction to the Automatic Curriculum Learning literature and 2) to draw a bigger picture of the current state of the art in ACL to encourage the cross-breeding of existing concepts and the emergence of new ideas.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Solver-Aware Decompositions for Programming-by-Example: When Dividing Requires Knowing how to Conquer

    cs.AI 2026-08 conditional novelty 7.0 of 10

    Solver-aware training of a PBE decomposer with a frozen synthesizer's loss outperforms supervised imitation of ground-truth subgoals, solving tasks that a ground-truth decomposition oracle fails.

  2. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

    cs.AI 2026-07 conditional novelty 6.5 of 10

    Deep, capability-targeted, co-evolving synthetic environments raise a 9B computer-use agent from 36.5% to 67.1% and enable RL gains the live web cannot supply.

  3. LEACL: LLM-Enhanced Automatic Curriculum Learning for Reinforcement Learning in Long-Horizon Manipulation Tasks

    cs.RO 2026-07 conditional novelty 6.0 of 10

    LLMs generate subtask decompositions and ACL task spaces so sparse-reward RL solves long-horizon manipulation better than dense human rewards on five LIBERO tasks.

  4. Self-Modifying Lean Proof Agents with Verifier-Grounded Benchmark Coevolution

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A verifier-grounded self-evolving Lean proof agent with a champion-driven, self-hardening benchmark reached 45.1% held-out miniF2F solve rate versus 32.0% for a fixed-benchmark baseline.

  5. Exploring Multimodal AI Reasoning for Meteorological Forecasting from Skew-T Diagrams

    physics.ao-ph 2025-08 conditional novelty 6.0 of 10

    A 250M-parameter vision-language model fine-tuned on Skew-T diagrams achieves CSI comparable to IFS-HRES for 3-hour precipitation probability in South Korean summer.

  6. Efficient Skill Discovery via Regret-Aware Optimization

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A regret-aware skill discovery algorithm, RSD, improves sample efficiency and zero-shot goal-reaching in high-dimensional continuous control by focusing exploration on unmastered skills.

  7. Energy-Based Transfer for Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    An energy-based out-of-distribution score gates a teacher's action advice in reinforcement learning, so guidance is given only in states the teacher has seen during training.

  8. EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making

    cs.AI 2025-08 reject novelty 5.0 of 10

    EvoCurr couples an LLM curriculum designer with an LLM code-generating solver, but its only reported success is 1 of 5 runs and no direct baseline is shown.

  9. GACL: Grounded Adaptive Curriculum Learning with Active Task and Performance Monitoring

    cs.RO 2025-08 conditional novelty 5.0 of 10

    GACL adds domain grounding via alternating reference and synthetic task sampling to a VAE-based regret-driven curriculum teacher, reporting higher success rates than CLUTR on BARN navigation and quadruped locomotion i...

  10. ADEPT: Adaptive Diffusion Environment for Policy Transfer Sim-to-Real

    cs.RO 2025-06 conditional novelty 5.0 of 10

    An adaptive curriculum that steers a pretrained diffusion terrain generator via policy success-weighted latent blending improves zero-shot sim-to-real off-road navigation performance.

  11. Reactive Aerobatic Flight via Reinforcement Learning

    cs.RO 2025-05 conditional novelty 5.0 of 10

    An RL policy trained with an automated curriculum and domain randomization lets a quadrotor perform continuous inverted flight while reactively navigating a moving gate.

Pith tools