Pith. sign in

REVIEW 2 cited by

Randomized Exploration in Generalized Linear Bandits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.08947 v3 pith:JLKJCECL submitted 2019-06-21 cs.LG stat.ML

Randomized Exploration in Generalized Linear Bandits

classification cs.LG stat.ML
keywords banditsgeneralizedglm-fpllinearalgorithmsexplorationfirstglm-tsl
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We study two randomized algorithms for generalized linear bandits. The first, GLM-TSL, samples a generalized linear model (GLM) from the Laplace approximation to the posterior distribution. The second, GLM-FPL, fits a GLM to a randomly perturbed history of past rewards. We analyze both algorithms and derive $\tilde{O}(d \sqrt{n \log K})$ upper bounds on their $n$-round regret, where $d$ is the number of features and $K$ is the number of arms. The former improves on prior work while the latter is the first for Gaussian noise perturbations in non-linear models. We empirically evaluate both GLM-TSL and GLM-FPL in logistic bandits, and apply GLM-FPL to neural network bandits. Our work showcases the role of randomization, beyond posterior sampling, in exploration.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness, and Safety

    stat.ML 2026-05 unverdicted novelty 6.0

    Introduces a regularized estimator achieving optimal MSE rates under a new relative balancedness condition while providing safety guarantees that match independent learning when tasks are unrelated.

  2. Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness, and Safety

    stat.ML 2026-05 unverdicted novelty 6.0

    Matrix-weighted regularization for robust multi-task regression achieves optimal MSE under weaker spectral assumptions and performs no worse than independent learning when balancedness is poor.