Pith. sign in

REVIEW 2 cited by

Combinatorial Bandits for Maximum Value Reward Function under Max Value-Index Feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16074 v1 pith:6AQ4ZACJ submitted 2023-05-25 cs.LG math.STstat.TH

classification cs.LGmath.STstat.TH
keywords feedbackregretalgorithmboundmaximumrewardundervalue
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We consider a combinatorial multi-armed bandit problem for maximum value reward function under maximum value and index feedback. This is a new feedback structure that lies in between commonly studied semi-bandit and full-bandit feedback structures. We propose an algorithm and provide a regret bound for problem instances with stochastic arm outcomes according to arbitrary distributions with finite supports. The regret analysis rests on considering an extended set of arms, associated with values and probabilities of arm outcomes, and applying a smoothness condition. Our algorithm achieves a $O((k/\Delta)\log(T))$ distribution-dependent and a $\tilde{O}(\sqrt{T})$ distribution-independent regret where $k$ is the number of arms selected in each round, $\Delta$ is a distribution-dependent reward gap and $T$ is the horizon time. Perhaps surprisingly, the regret bound is comparable to previously-known bound under more informative semi-bandit feedback. We demonstrate the effectiveness of our algorithm through experimental results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Oracle-Efficient Combinatorial Semi-Bandits

    stat.ML 2025-10 conditional novelty 7.0 of 10

    New algorithms for combinatorial semi-bandits achieve near-optimal regret while reducing oracle calls from every round to doubly-logarithmically many.

  2. Offline Learning for Combinatorial Multi-armed Bandits

    cs.LG 2025-01 conditional novelty 7.0 of 10

    A pessimistic lower-confidence-bound algorithm achieves suboptimality bounds for offline combinatorial multi-armed bandits with probabilistically triggered arms, under coverage conditions requiring observation of each...

Pith tools