Neural-ADB claims an O~((d/T)^(1/2)) worst sub-optimality gap for active contextual dueling bandits with non-linear rewards, but the proof relies on a reversed matrix inequality.
Finite-time analysis of the multiarmed bandit problem
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Active Human Feedback Collection via Neural Contextual Dueling Bandits
Neural-ADB claims an O~((d/T)^(1/2)) worst sub-optimality gap for active contextual dueling bandits with non-linear rewards, but the proof relies on a reversed matrix inequality.