The one- and two-timescale algorithms for discounted exponential-utility RL achieve O~(1/sqrt(n)) finite-time rates under Markovian sampling with parameter-free stepsizes.
Title resolution pending
1 Pith paper cite this work, alongside 18 external citations. Polarity classification is still indexing.
1
Pith paper citing it
18
external citations · OpenAlex
fields
cs.LG 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning
The one- and two-timescale algorithms for discounted exponential-utility RL achieve O~(1/sqrt(n)) finite-time rates under Markovian sampling with parameter-free stepsizes.