Pith. sign in

Proceedings of the AAAI conference on artificial intelligence , volume=

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning

cs.LG · 2026-08-03 · conditional · novelty 6.0

Expectile n-step Q-learning (ENQ) applies an upper-expectile loss to the action-value TD error, reducing the pessimistic bias of multi-step returns in off-policy RL; the authors prove contraction and bias bounds and report competitive results across 27 tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning cs.LG · 2026-08-03 · conditional · none · ref 5

    Expectile n-step Q-learning (ENQ) applies an upper-expectile loss to the action-value TD error, reducing the pessimistic bias of multi-step returns in off-policy RL; the authors prove contraction and bias bounds and report competitive results across 27 tasks.