Pith. sign in

REVIEW 1 cited by

Reward Poisoning in Reinforcement Learning: Attacks Against Unknown Learners in Unknown Environments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.08492 v1 pith:SVKZHN4J submitted 2021-02-16 cs.LG cs.AIcs.CR

classification cs.LGcs.AIcs.CR
keywords unknownadversaryattackblack-boxpoisoningrewardattacksenvironment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study black-box reward poisoning attacks against reinforcement learning (RL), in which an adversary aims to manipulate the rewards to mislead a sequence of RL agents with unknown algorithms to learn a nefarious policy in an environment unknown to the adversary a priori. That is, our attack makes minimum assumptions on the prior knowledge of the adversary: it has no initial knowledge of the environment or the learner, and neither does it observe the learner's internal mechanism except for its performed actions. We design a novel black-box attack, U2, that can provably achieve a near-matching performance to the state-of-the-art white-box attack, demonstrating the feasibility of reward poisoning even in the most challenging black-box setting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Poisoning Attack Against Reinforcement Learning under Black-box Environments

    cs.LG 2024-12 conditional novelty 6.0 of 10

    An online attacker knowing only which states are reachable can poison rewards and transitions to make a Q-learning agent follow a target policy in a maze.

Pith tools