Pith. sign in

Best-Arm Identification in Linear Bandits

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it
abstract

We study the best-arm identification problem in linear bandit, where the rewards of the arms depend linearly on an unknown parameter $\theta^*$ and the objective is to return the arm with the largest reward. We characterize the complexity of the problem and introduce sample allocation strategies that pull arms to identify the best arm with a fixed confidence, while minimizing the sample budget. In particular, we show the importance of exploiting the global linear structure to improve the estimate of the reward of near-optimal arms. We analyze the proposed strategies and compare their empirical performance. Finally, as a by-product of our analysis, we point out the connection to the $G$-optimality criterion used in optimal experimental design.

years

2026 2

verdicts

UNVERDICTED 2

representative citing papers

Anytime-valid Optimal Policy Identification

stat.ME · 2026-06-16 · unverdicted · novelty 6.0

Constructs a time-indexed set S_t retaining the true optimal policy uniformly over time with high probability, enabling early stopping with sample complexity O((log |Π| + log log(1/Δ_min))/Δ_min²) when the optimum is unique.

Logging Policy Design for Off-Policy Evaluation

stat.ML · 2026-05-14 · unverdicted · novelty 5.0 · 2 refs

Derives optimal logging policies for minimizing off-policy evaluation error under known, unknown, and partially known target policies and reward distributions.

citing papers explorer

Showing 2 of 2 citing papers.

  • Anytime-valid Optimal Policy Identification stat.ME · 2026-06-16 · unverdicted · none · ref 26 · internal anchor

    Constructs a time-indexed set S_t retaining the true optimal policy uniformly over time with high probability, enabling early stopping with sample complexity O((log |Π| + log log(1/Δ_min))/Δ_min²) when the optimum is unique.

  • Logging Policy Design for Off-Policy Evaluation stat.ML · 2026-05-14 · unverdicted · none · ref 48 · 2 links · internal anchor

    Derives optimal logging policies for minimizing off-policy evaluation error under known, unknown, and partially known target policies and reward distributions.