REVIEW 2 cited by
Offline Learning for Combinatorial Multi-armed Bandits
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The combinatorial multi-armed bandit (CMAB) is a fundamental sequential decision-making framework, extensively studied over the past decade. However, existing work primarily focuses on the online setting, overlooking the substantial costs of online interactions and the readily available offline datasets. To overcome these limitations, we introduce Off-CMAB, the first offline learning framework for CMAB. Central to our framework is the combinatorial lower confidence bound (CLCB) algorithm, which combines pessimistic reward estimations with combinatorial solvers. To characterize the quality of offline datasets, we propose two novel data coverage conditions and prove that, under these conditions, CLCB achieves a near-optimal suboptimality gap, matching the theoretical lower bound up to a logarithmic factor. We validate Off-CMAB through practical applications, including learning to rank, large language model (LLM) caching, and social influence maximization, showing its ability to handle nonlinear reward functions, general feedback models, and out-of-distribution action samples that excludes optimal or even feasible actions. Extensive experiments on synthetic and real-world datasets further highlight the superior performance of CLCB.
Forward citations
Cited by 2 Pith papers
-
Best Arm Identification with Possibly Biased Offline Data
LUCB-H adaptively combines offline and online data for best arm identification, matching or beating standard LUCB depending on whether the historical data is helpful or misleading.
-
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
A unified framework for co-branding learns partner success probabilities and market gains online, and allocates sub-brand budgets offline with a 1-1/e approximation guarantee.
Discussion (0). Continue with ORCID to comment.