REVIEW 2 cited by
A Definition of Non-Stationary Bandits
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Despite the subject of non-stationary bandit learning having attracted much recent attention, we have yet to identify a formal definition of non-stationarity that can consistently distinguish non-stationary bandits from stationary ones. Prior work has characterized non-stationary bandits as bandits for which the reward distribution changes over time. We demonstrate that this definition can ambiguously classify the same bandit as both stationary and non-stationary; this ambiguity arises in the existing definition's dependence on the latent sequence of reward distributions. Moreover, the definition has given rise to two widely used notions of regret: the dynamic regret and the weak regret. These notions are not indicative of qualitative agent performance in some bandits. Additionally, this definition of non-stationary bandits has led to the design of agents that explore excessively. We introduce a formal definition of non-stationary bandits that resolves these issues. Our new definition provides a unified approach, applicable seamlessly to both Bayesian and frequentist formulations of bandits. Furthermore, our definition ensures consistent classification of two bandits offering agents indistinguishable experiences, categorizing them as either both stationary or both non-stationary. This advancement provides a more robust framework for non-stationary bandit learning.
Forward citations
Cited by 2 Pith papers
-
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
A bandit-based surrogate selection method produces CLIP universal adversarial perturbations that transfer across datasets, models, and tasks, beating prior UAP baselines by large margins.
-
Natural Policy Gradient for Average Reward Non-Stationary RL
A natural actor-critic algorithm for non-stationary average-reward MDPs achieves dynamic regret O~(sqrt(|S||A|) Delta_T^{1/6} T^{5/6}).
Discussion (0). Continue with ORCID to comment.