Pith. sign in

REVIEW

Multinomial Logit Bandit with Low Switching Cost

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.04876 v1 pith:GIXBFWBO submitted 2020-07-09 cs.LG stat.ML

classification cs.LGstat.ML
keywords costswitchingalgorithmassortmentadaptivityalmostbanditbound
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study multinomial logit bandit with limited adaptivity, where the algorithms change their exploration actions as infrequently as possible when achieving almost optimal minimax regret. We propose two measures of adaptivity: the assortment switching cost and the more fine-grained item switching cost. We present an anytime algorithm (AT-DUCB) with $O(N \log T)$ assortment switches, almost matching the lower bound $\Omega(\frac{N \log T}{ \log \log T})$. In the fixed-horizon setting, our algorithm FH-DUCB incurs $O(N \log \log T)$ assortment switches, matching the asymptotic lower bound. We also present the ESUCB algorithm with item switching cost $O(N \log^2 T)$.

Discussion (0). Continue with ORCID to comment.

Pith tools