Pith. sign in

REVIEW 3 cited by

Neural Contextual Bandits with Deep Representation and Shallow Exploration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.01780 v1 pith:MYYVSUR4 submitted 2020-12-03 cs.LG stat.ML

classification cs.LGstat.ML
keywords deepneuralcontextuallastlayerlearningalgorithmapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study a general class of contextual bandits, where each context-action pair is associated with a raw feature vector, but the reward generating function is unknown. We propose a novel learning algorithm that transforms the raw feature vector using the last hidden layer of a deep ReLU neural network (deep representation learning), and uses an upper confidence bound (UCB) approach to explore in the last linear layer (shallow exploration). We prove that under standard assumptions, our proposed algorithm achieves $\tilde{O}(\sqrt{T})$ finite-time regret, where $T$ is the learning time horizon. Compared with existing neural contextual bandit algorithms, our approach is computationally much more efficient since it only needs to explore in the last layer of the deep neural network.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PinEqualizer: Full Funnel Content Exploration and Debiasing System at Pinterest

    cs.IR 2026-07 conditional novelty 6.0 of 10

    A full-funnel exploration and debiasing system deployed at Pinterest is reported to increase fresh-content impressions by ~350% and lift user-engagement and content-provider metrics.

  2. Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration

    cs.LG 2025-06 unverdicted novelty 6.0 of 10

    Variance-aware neural dueling bandit algorithms achieve sublinear regret of order O(d sqrt(sum sigma_t^2) + sqrt(d T)) for wide networks on nonlinear utilities.

  3. In-Domain African Languages Translation Using LLMs and Multi-armed Bandits

    cs.CL 2025-05 reject novelty 4.0 of 10

    Bandit-based model selection matches or slightly improves on the best single NMT system for in-domain English-to-African translation, but the claimed high-confidence statistical support is absent.

Pith tools