DBGD-NF and its blocking variant achieve regret bounds of O(n average-delay^{1/3} T^{2/3}) and O(n(T^{2/3} + sqrt(dT))) for online nonsubmodular optimization with delayed bandit feedback.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting
DBGD-NF and its blocking variant achieve regret bounds of O(n average-delay^{1/3} T^{2/3}) and O(n(T^{2/3} + sqrt(dT))) for online nonsubmodular optimization with delayed bandit feedback.