Pith. sign in

REVIEW 3 cited by

Robust Thompson Sampling Algorithms Against Reward Poisoning Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19705 v1 pith:X7MKUW2R submitted 2024-10-25 cs.LG

classification cs.LG
keywords algorithmssamplingthompsonrewardpoisoningrobustadversarialagent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Thompson sampling is one of the most popular learning algorithms for online sequential decision-making problems and has rich real-world applications. However, current Thompson sampling algorithms are limited by the assumption that the rewards received are uncorrupted, which may not be true in real-world applications where adversarial reward poisoning exists. To make Thompson sampling more reliable, we want to make it robust against adversarial reward poisoning. The main challenge is that one can no longer compute the actual posteriors for the true reward, as the agent can only observe the rewards after corruption. In this work, we solve this problem by computing pseudo-posteriors that are less likely to be manipulated by the attack. We propose robust algorithms based on Thompson sampling for the popular stochastic and contextual linear bandit settings in both cases where the agent is aware or unaware of the budget of the attacker. We theoretically show that our algorithms guarantee near-optimal regret under any attack strategy.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimism as a Vulnerability: Deceptive Stackelberg Control of UCB Bandit Followers

    cs.GT 2026-06 conditional novelty 6.5 of 10

    Under targetability and exploitability, a two-phase honeypot-then-trap leader strictly exceeds the classical SSE utility ceiling against a UCB follower at O(sqrt(T ln T)) signaling cost.

  2. Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Adversarially training a Decision-Pretrained Transformer against learned reward-poisoning attackers makes it robust to test-time reward corruption, outperforming robust bandit baselines in experiments.

  3. Periodic Bootstrap Thompson Sampling For Periodically Non-Stationary Bandit Problems

    cs.LG 2026-07 conditional novelty 3.0 of 10

    A Thompson Sampling variant that periodically resets its beliefs and runs forced exploration reduces cumulative regret versus plain Thompson Sampling in simulated periodic bandit problems.

Pith tools