A distributional framework for optimizing Lipschitz risk functionals in offline contextual bandits yields data-dependent suboptimality bounds of Õ(1/√n) that match risk-neutral rates and are minimax optimal.
Distributional reinforcement learning with quantile regression
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
C51 matches StreamQ in streaming RL on 55 Atari games while a new Adaptive Q(λ) algorithm based on bounded derivatives and variance-adjusted updates reaches nearly double the human baseline.
citing papers explorer
-
Pessimistic Risk-Aware Policy Learning in Contextual Bandits
A distributional framework for optimizing Lipschitz risk functionals in offline contextual bandits yields data-dependent suboptimality bounds of Õ(1/√n) that match risk-neutral rates and are minimax optimal.
-
Revisiting Adam for Streaming Reinforcement Learning
C51 matches StreamQ in streaming RL on 55 Atari games while a new Adaptive Q(λ) algorithm based on bounded derivatives and variance-adjusted updates reaches nearly double the human baseline.