Pith. sign in

REVIEW

On the connection between Bregman divergence and value in regularized Markov decision processes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.12160 v4 pith:K6TVQ27D submitted 2022-10-21 cs.LG cs.AImath.OC

classification cs.LGcs.AImath.OC
keywords bregmancurrentdecisiondivergencefunctionlearningmarkovpolicy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this short note we derive a relationship between the Bregman divergence from the current policy to the optimal policy and the suboptimality of the current value function in a regularized Markov decision process. This result has implications for multi-task reinforcement learning, offline reinforcement learning, and regret analysis under function approximation, among others.

Discussion (0). Continue with ORCID to comment.

Pith tools