Pith. sign in

REVIEW 1 cited by

Distilled Thompson Sampling: Practical and Efficient Thompson Sampling via Imitation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.14266 v3 pith:XS353AA7 submitted 2020-11-29 cs.LG cs.AI

classification cs.LGcs.AI
keywords policyalgorithmimitationlatencysamplingthompsonallowingdeployment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Thompson sampling (TS) has emerged as a robust technique for contextual bandit problems. However, TS requires posterior inference and optimization for action generation, prohibiting its use in many online platforms where latency and ease of deployment are of concern. We operationalize TS by proposing a novel imitation-learning-based algorithm that distills a TS policy into an explicit policy representation, allowing fast decision-making and easy deployment in mobile and server-based environments. Using batched data collected under the imitation policy, our algorithm iteratively performs offline updates to the TS policy, and learns a new explicit policy representation to imitate it. Empirically, our imitation policy achieves performance comparable to batch TS while allowing more than an order of magnitude reduction in decision-time latency. Buoyed by low latency and simplicity of implementation, our algorithm has been successfully deployed in multiple video upload systems for Meta. Using a randomized controlled trial, we show our algorithm resulted in significant improvements in video quality and watch time.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Planning Framework for Adaptive Labeling

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A planning framework for adaptive labeling where a smoothed auto-differential policy gradient (Smoothed-Autodiff) selects batches to minimize final posterior uncertainty, outperforming active-learning heuristics and R...

Pith tools