The AVE algorithm achieves O~(sqrt(M^2 A H^4 n log^3 |F|)) cumulative regret for episodic MDPs with low Bellman rank and realizable function approximation.
Approximate Dynamic Programming: Solving the curses of dim ensionality, volume
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2019 1verdicts
ACCEPT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank
The AVE algorithm achieves O~(sqrt(M^2 A H^4 n log^3 |F|)) cumulative regret for episodic MDPs with low Bellman rank and realizable function approximation.