Pith. sign in

REVIEW 2 cited by

Offline Learning for Combinatorial Multi-armed Bandits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.19300 v2 pith:3NNUU2KE submitted 2025-01-31 cs.LG

classification cs.LG
keywords combinatorialofflineclcbdatasetsframeworklearningboundcmab
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The combinatorial multi-armed bandit (CMAB) is a fundamental sequential decision-making framework, extensively studied over the past decade. However, existing work primarily focuses on the online setting, overlooking the substantial costs of online interactions and the readily available offline datasets. To overcome these limitations, we introduce Off-CMAB, the first offline learning framework for CMAB. Central to our framework is the combinatorial lower confidence bound (CLCB) algorithm, which combines pessimistic reward estimations with combinatorial solvers. To characterize the quality of offline datasets, we propose two novel data coverage conditions and prove that, under these conditions, CLCB achieves a near-optimal suboptimality gap, matching the theoretical lower bound up to a logarithmic factor. We validate Off-CMAB through practical applications, including learning to rank, large language model (LLM) caching, and social influence maximization, showing its ability to handle nonlinear reward functions, general feedback models, and out-of-distribution action samples that excludes optimal or even feasible actions. Extensive experiments on synthetic and real-world datasets further highlight the superior performance of CLCB.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Best Arm Identification with Possibly Biased Offline Data

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LUCB-H adaptively combines offline and online data for best arm identification, matching or beating standard LUCB depending on whether the historical data is helpful or misleading.

  2. A Unified Online-Offline Framework for Co-Branding Campaign Recommendations

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A unified framework for co-branding learns partner success probabilities and market gains online, and allocates sub-brand budgets offline with a 1-1/e approximation guarantee.

Pith tools