Best-action queries yield Õ(min{T/k, √(T-k)}) regret for i.i.d. stochastic rewards but only Ω(√(T-k)) regret for correlated stochastic or adversarial rewards in the bandit-feedback model.
ISBN 9781605585161
4 Pith papers cite this work, alongside 147 external citations. Polarity classification is still indexing.
years
2026 4representative citing papers
RankGuard is a decentralized OLTR system that filters model updates using local click data for poisoning resistance and supplies the first formal convergence guarantee for decentralized OLTR.
UKA is a gradient-free active dialogue learning framework using Theory-of-Mind uncertainty estimation to acquire user-aligned conversational knowledge, outperforming baselines in dialogue quality and user alignment across benchmarks.
A new TWCTV regularizer using weighted Schatten-p norms on gradients and adaptive sparse weighting in the M-product framework is proposed for robust tensor completion, with an ADMM solver and claimed superior performance on image tasks.
citing papers explorer
-
Multi-Armed Bandits With Best-Action Queries
Best-action queries yield Õ(min{T/k, √(T-k)}) regret for i.i.d. stochastic rewards but only Ω(√(T-k)) regret for correlated stochastic or adversarial rewards in the bandit-feedback model.
-
Efficient and Robust Online Learning to Rank in Decentralized Systems
RankGuard is a decentralized OLTR system that filters model updates using local click data for poisoning resistance and supplies the first formal convergence guarantee for decentralized OLTR.
-
User-Aware Active Knowledge Acquisition for Emotional Support Dialogue
UKA is a gradient-free active dialogue learning framework using Theory-of-Mind uncertainty estimation to acquire user-aligned conversational knowledge, outperforming baselines in dialogue quality and user alignment across benchmarks.
-
Robust Low-Rank Tensor Completion based on M-product with Weighted Correlated Total Variation and Sparse Regularization
A new TWCTV regularizer using weighted Schatten-p norms on gradients and adaptive sparse weighting in the M-product framework is proposed for robust tensor completion, with an ADMM solver and claimed superior performance on image tasks.