Proposes LDB-DF and NDB-DF algorithms for contextual dueling bandits with delayed feedback using an IPW estimator in the loss, with O(d sqrt(T)) regret for the linear case and sub-linear guarantees for the neural case.
Neural dueling bandits
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2representative citing papers
Neural variance-aware dueling bandit algorithms achieve sublinear regret with network width m = Omega~(T^6), an improvement over the previous Omega~(T^14), under both UCB and Thompson sampling.
citing papers explorer
-
Linear and Neural Dueling Bandits with Delayed Feedback
Proposes LDB-DF and NDB-DF algorithms for contextual dueling bandits with delayed feedback using an IPW estimator in the loss, with O(d sqrt(T)) regret for the linear case and sub-linear guarantees for the neural case.
-
Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration
Neural variance-aware dueling bandit algorithms achieve sublinear regret with network width m = Omega~(T^6), an improvement over the previous Omega~(T^14), under both UCB and Thompson sampling.