Pith. sign in

REVIEW 1 cited by

Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial Corruptions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.06811 v2 pith:ICT7BYMH submitted 2022-05-13 cs.LG stat.ML

classification cs.LGstat.ML
keywords algorithmcorruptionregretcasesnearlyachievesadversarialalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study the linear contextual bandit problem in the presence of adversarial corruption, where the reward at each round is corrupted by an adversary, and the corruption level (i.e., the sum of corruption magnitudes over the horizon) is $C\geq 0$. The best-known algorithms in this setting are limited in that they either are computationally inefficient or require a strong assumption on the corruption, or their regret is at least $C$ times worse than the regret without corruption. In this paper, to overcome these limitations, we propose a new algorithm based on the principle of optimism in the face of uncertainty. At the core of our algorithm is a weighted ridge regression where the weight of each chosen action depends on its confidence up to some threshold. We show that for both known $C$ and unknown $C$ cases, our algorithm with proper choice of hyperparameter achieves a regret that nearly matches the lower bounds. Thus, our algorithm is nearly optimal up to logarithmic factors for both cases. Notably, our algorithm achieves the near-optimal regret for both corrupted and uncorrupted cases ($C=0$) simultaneously.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cascading Bandits Robust to Adversarial Corruptions

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Cascading bandits can be made robust to adversarial click corruption using multi-instance position-based elimination, with regret logarithmic in time and linear in the corruption budget.

Pith tools