Pith. sign in

REVIEW 3 cited by

Exploration versus exploitation in reinforcement learning: a stochastic control approach

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.01552 v3 pith:Y74VH3AY submitted 2018-12-04 math.OC cs.LG

classification math.OCcs.LG
keywords explorationproblemexploitationlearningcontrolgaussianclassicaldistribution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider reinforcement learning (RL) in continuous time and study the problem of achieving the best trade-off between exploration of a black box environment and exploitation of current knowledge. We propose an entropy-regularized reward function involving the differential entropy of the distributions of actions, and motivate and devise an exploratory formulation for the feature dynamics that captures repetitive learning under exploration. The resulting optimization problem is a revitalization of the classical relaxed stochastic control. We carry out a complete analysis of the problem in the linear--quadratic (LQ) setting and deduce that the optimal feedback control distribution for balancing exploitation and exploration is Gaussian. This in turn interprets and justifies the widely adopted Gaussian exploration in RL, beyond its simplicity for sampling. Moreover, the exploitation and exploration are captured, respectively and mutual-exclusively, by the mean and variance of the Gaussian distribution. We also find that a more random environment contains more learning opportunities in the sense that less exploration is needed. We characterize the cost of exploration, which, for the LQ case, is shown to be proportional to the entropy regularization weight and inversely proportional to the discount rate. Finally, as the weight of exploration decays to zero, we prove the convergence of the solution of the entropy-regularized LQ problem to the one of the classical LQ problem.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Relaxed Control with Entropy Regularization for It\^o Stochastic Systems with Input Delay

    math.OC 2026-05 unverdicted novelty 5.0 of 10

    Entropy-regularized relaxed control for infinite-horizon Itô stochastic systems with input delay yields Gaussian optimal distributions that converge to Dirac measures as the exploration weight vanishes.

  2. An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems

    cs.CL 2024-12 unverdicted novelty 3.0 of 10

    A survey and position paper that reviews LLM prompting, RAG, and RL techniques and argues they could support open-ended implementation generation, without presenting new results.

  3. A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

    cs.AI 2024-11 conditional novelty 2.0 of 10

    A comprehensive but flawed survey of RL algorithms that catalogs many methods and applications without rigorous comparative analysis.

Pith tools