Pith. sign in

REVIEW 10 cited by

A Survey of Exploration Methods in Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.00157 v2 pith:4VJDD6XE submitted 2021-09-01 cs.LG cs.AI

A Survey of Exploration Methods in Reinforcement Learning

classification cs.LG cs.AI
keywords learningexplorationreinforcementmethodsagentssurveyalgorithmsarticle
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Exploration is an essential component of reinforcement learning algorithms, where agents need to learn how to predict and control unknown and often stochastic environments. Reinforcement learning agents depend crucially on exploration to obtain informative data for the learning process as the lack of enough information could hinder effective learning. In this article, we provide a survey of modern exploration methods in (Sequential) reinforcement learning, as well as a taxonomy of exploration methods.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Optimal Semiparametric Dynamic Pricing with Feature Diversity

    stat.ME 2026-05 unverdicted novelty 7.0

    A stagewise greedy algorithm for semiparametric contextual dynamic pricing achieves regret T to the max of 1/2 and 3 over (2 beta plus 1) for linear m, with a matching lower bound proving optimality.

  2. SpecRL: Reinforcement Learning with Test-Based Completeness Rewards for Formal Specification Synthesis

    cs.SE 2026-04 unverdicted novelty 6.0

    Reinforcement learning with spectest completeness rewards lifts a 7B model’s Dafny specification verification success and completeness over supervised fine-tuning by about 50% and 26%.

  3. LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

    cs.AI 2026-07 conditional novelty 5.0

    LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.

  4. Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning

    cs.LG 2026-06 unverdicted novelty 5.0

    Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.

  5. SpecRL: Reinforcement Learning with Test-Based Completeness Rewards for Formal Specification Synthesis

    cs.SE 2026-04 unverdicted novelty 5.0

    SpecRL uses the fraction of negative tests rejected by candidate specifications as a reward signal in RL training to produce stronger and more verifiable formal specifications than prior methods.

  6. Smart Walkers in Discrete Space

    cond-mat.stat-mech 2026-01 unverdicted novelty 5.0

    Configuration entropy serves as a reliable proxy for the learned skills of reinforcement learning agents performing tasks in discrete space, validated through walker encounters and chess engine tests.

  7. CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models

    cs.CL 2025-09 conditional novelty 5.0

    Adding actor perplexity and multi-head critic variance as intrinsic exploration bonuses improves RLVR math reasoning accuracy by roughly +2 to +3 points on AIME benchmarks.

  8. Flexible Empowerment at Reasoning with Extended Best-of-N Sampling

    cs.LG 2026-04 unverdicted novelty 4.0

    Extended best-of-N sampling with Tsallis statistics allows flexible empowerment in RL reasoning, balancing exploration-exploitation and improving locomotion task performance.

  9. Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring

    cs.RO 2026-04 conditional novelty 4.0

    Contextual multi-task DDQN learns one AUV policy for multiple simulated reef-monitoring tasks that matches mixture-of-experts performance and generalizes better on a discrete toy domain.

  10. Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring

    cs.RO 2026-04 unverdicted novelty 4.0

    A context-dependent multi-task RL policy is trained and evaluated in HoloOcean simulation to solve multiple reef monitoring tasks with claimed improvements in sample efficiency, zero-shot generalization, and robustnes...