Pith. sign in

REVIEW

A Zeroth-Order Momentum Method for Risk-Averse Online Convex Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.02838 v1 pith:4CI7XD6K submitted 2022-09-06 cs.LG cs.GTstat.ML

classification cs.LGcs.GTstat.ML
keywords agentscostvaluesactionscvarriskrisk-aversealgorithm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider risk-averse learning in repeated unknown games where the goal of the agents is to minimize their individual risk of incurring significantly high cost. Specifically, the agents use the conditional value at risk (CVaR) as a risk measure and rely on bandit feedback in the form of the cost values of the selected actions at every episode to estimate their CVaR values and update their actions. A major challenge in using bandit feedback to estimate CVaR is that the agents can only access their own cost values, which, however, depend on the actions of all agents. To address this challenge, we propose a new risk-averse learning algorithm with momentum that utilizes the full historical information on the cost values. We show that this algorithm achieves sub-linear regret and matches the best known algorithms in the literature. We provide numerical experiments for a Cournot game that show that our method outperforms existing methods.

Discussion (0). Sign in to comment.

Pith tools