Proposes LDB-DF and NDB-DF algorithms for contextual dueling bandits with delayed feedback using an IPW estimator in the loss, with O(d sqrt(T)) regret for the linear case and sub-linear guarantees for the neural case.
Feel-good thompson sampling for contextual dueling bandits
4 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.LG 4representative citing papers
ActiveDPO is a theoretically grounded active data selection method for sample-efficient LLM alignment that parameterizes the reward model directly with the LLM being aligned.
Proposes e RCDP-UCB algorithm achieving regret bound ŵO(d(√T + C + D)) for linear dueling bandits under post-serving contexts, delays, and corruptions with additive corruption-delay cost.
Neural variance-aware dueling bandit algorithms achieve sublinear regret with network width m = Omega~(T^6), an improvement over the previous Omega~(T^14), under both UCB and Thompson sampling.
citing papers explorer
-
Linear and Neural Dueling Bandits with Delayed Feedback
Proposes LDB-DF and NDB-DF algorithms for contextual dueling bandits with delayed feedback using an IPW estimator in the loss, with O(d sqrt(T)) regret for the linear case and sub-linear guarantees for the neural case.
-
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
ActiveDPO is a theoretically grounded active data selection method for sample-efficient LLM alignment that parameterizes the reward model directly with the LLM being aligned.
-
Robust Linear Dueling Bandits with Post-serving Context under Unknown Delays and Adversarial Corruptions
Proposes e RCDP-UCB algorithm achieving regret bound ŵO(d(√T + C + D)) for linear dueling bandits under post-serving contexts, delays, and corruptions with additive corruption-delay cost.
-
Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration
Neural variance-aware dueling bandit algorithms achieve sublinear regret with network width m = Omega~(T^6), an improvement over the previous Omega~(T^14), under both UCB and Thompson sampling.