REVIEW 10 cited by
A Survey of Exploration Methods in Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
A Survey of Exploration Methods in Reinforcement Learning
read the original abstract
Exploration is an essential component of reinforcement learning algorithms, where agents need to learn how to predict and control unknown and often stochastic environments. Reinforcement learning agents depend crucially on exploration to obtain informative data for the learning process as the lack of enough information could hinder effective learning. In this article, we provide a survey of modern exploration methods in (Sequential) reinforcement learning, as well as a taxonomy of exploration methods.
Forward citations
Cited by 10 Pith papers
-
Optimal Semiparametric Dynamic Pricing with Feature Diversity
A stagewise greedy algorithm for semiparametric contextual dynamic pricing achieves regret T to the max of 1/2 and 3 over (2 beta plus 1) for linear m, with a matching lower bound proving optimality.
-
SpecRL: Reinforcement Learning with Test-Based Completeness Rewards for Formal Specification Synthesis
Reinforcement learning with spectest completeness rewards lifts a 7B model’s Dafny specification verification success and completeness over supervised fine-tuning by about 50% and 26%.
-
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.
-
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.
-
SpecRL: Reinforcement Learning with Test-Based Completeness Rewards for Formal Specification Synthesis
SpecRL uses the fraction of negative tests rejected by candidate specifications as a reward signal in RL training to produce stronger and more verifiable formal specifications than prior methods.
-
Smart Walkers in Discrete Space
Configuration entropy serves as a reliable proxy for the learned skills of reinforcement learning agents performing tasks in discrete space, validated through walker encounters and chess engine tests.
-
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Adding actor perplexity and multi-head critic variance as intrinsic exploration bonuses improves RLVR math reasoning accuracy by roughly +2 to +3 points on AIME benchmarks.
-
Flexible Empowerment at Reasoning with Extended Best-of-N Sampling
Extended best-of-N sampling with Tsallis statistics allows flexible empowerment in RL reasoning, balancing exploration-exploitation and improving locomotion task performance.
-
Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring
Contextual multi-task DDQN learns one AUV policy for multiple simulated reef-monitoring tasks that matches mixture-of-experts performance and generalizes better on a discrete toy domain.
-
Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring
A context-dependent multi-task RL policy is trained and evaluated in HoloOcean simulation to solve multiple reef monitoring tasks with claimed improvements in sample efficiency, zero-shot generalization, and robustnes...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.