Pith. sign in

REVIEW 1 cited by

Locally Differentially Private (Contextual) Bandits Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.00701 v4 pith:HMTFYZUB submitted 2020-06-01 cs.LG stat.ML

classification cs.LGstat.ML
keywords banditslearningprivatecontextualdifferentiallyframeworksblack-boxcontext-free
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study locally differentially private (LDP) bandits learning in this paper. First, we propose simple black-box reduction frameworks that can solve a large family of context-free bandits learning problems with LDP guarantee. Based on our frameworks, we can improve previous best results for private bandits learning with one-point feedback, such as private Bandits Convex Optimization, and obtain the first result for Bandits Convex Optimization (BCO) with multi-point feedback under LDP. LDP guarantee and black-box nature make our frameworks more attractive in real applications compared with previous specifically designed and relatively weaker differentially private (DP) context-free bandits algorithms. Further, we extend our $(\varepsilon, \delta)$-LDP algorithm to Generalized Linear Bandits, which enjoys a sub-linear regret $\tilde{O}(T^{3/4}/\varepsilon)$ and is conjectured to be nearly optimal. Note that given the existing $\Omega(T)$ lower bound for DP contextual linear bandits (Shariff & Sheffe, 2018), our result shows a fundamental difference between LDP and DP contextual bandits learning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Prophet Inequalities under Local Differential Privacy

    cs.GT 2026-06 unverdicted novelty 8.0 of 10

    Under LDP, optimal online stopping uses binary reports and achieves competitive ratio e^ε/(n-1+e^ε) vs the non-private online optimum and (1+e^{-ε})/2 vs the LDP prophet.

Pith tools