REVIEW 3 cited by
Online and Bandit Algorithms for Nonstationary Stochastic Saddle-Point Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Saddle-point optimization problems are an important class of optimization problems with applications to game theory, multi-agent reinforcement learning and machine learning. A majority of the rich literature available for saddle-point optimization has focused on the offline setting. In this paper, we study nonstationary versions of stochastic, smooth, strongly-convex and strongly-concave saddle-point optimization problem, in both online (or first-order) and multi-point bandit (or zeroth-order) settings. We first propose natural notions of regret for such nonstationary saddle-point optimization problems. We then analyze extragradient and Frank-Wolfe algorithms, for the unconstrained and constrained settings respectively, for the above class of nonstationary saddle-point optimization problems. We establish sub-linear regret bounds on the proposed notions of regret in both the online and bandit setting.
Forward citations
Cited by 3 Pith papers
-
A Modular Algorithm for Non-Stationary Online Convex-Concave Optimization
A modular algorithm for online convex-concave optimization achieves near-optimal dynamic duality gap bounds by combining adaptive experts with a multi-predictor aggregator.
-
Forgetting-Factor Regret for Online Zero-Sum Games
A forgetting-factor regret metric with exponentially decaying weights is introduced for online zero-sum games, with tracking bounds proven for gradient, Frank-Wolfe, and gradient-free algorithms under time-varying payoffs.
-
Distributed Online Stochastic Convex-Concave Optimization: Dynamic Regret Analyses under Single and Multiple Consensus Steps
Two distributed online stochastic mirror descent algorithms achieve dynamic saddle point regret O(max{T^{θ1}, T^{θ2}(1+V_T)}) under Bregman divergence, generalizing earlier Euclidean results to stochastic, non-Euclide...
Discussion (0). Continue with ORCID to comment.