Pith. sign in

REVIEW 2 cited by

Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01386 v3 pith:5OTYYFYG submitted 2024-06-03 cs.LG

classification cs.LG
keywords multivariantcmabepisodiccmab-mtconditionframeworktriggeringapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We introduce a novel framework of combinatorial multi-armed bandits (CMAB) with multivariant and probabilistically triggering arms (CMAB-MT), where the outcome of each arm is a $d$-dimensional multivariant random variable and the feedback follows a general arm triggering process. Compared with existing CMAB works, CMAB-MT not only enhances the modeling power but also allows improved results by leveraging distinct statistical properties for multivariant random variables. For CMAB-MT, we propose a general 1-norm multivariant and triggering probability-modulated smoothness condition, and an optimistic CUCB-MT algorithm built upon this condition. Our framework can include many important problems as applications, such as episodic reinforcement learning (RL) and probabilistic maximum coverage for goods distribution, all of which meet the above smoothness condition and achieve matching or improved regret bounds compared to existing works. Through our new framework, we build the first connection between the episodic RL and CMAB literature, by offering a new angle to solve the episodic RL through the lens of CMAB, which may encourage more interactions between these two important directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Intelligent Channel Allocation for IEEE 802.11be Multi-Link Operation: When MAB Meets LLM

    cs.NI 2025-06 conditional novelty 6.0 of 10

    BAI-MCTS and an LLM-initialized variant solve the WiFi 7 channel allocation problem as a multi-armed bandit, converging faster than prior bandit-MCTS baselines.

  2. DOPPLER: Dual-Policy Learning for Device Assignment in Asynchronous Dataflow Graphs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    DOPPLER trains two cooperating neural policies, one that orders graph operations and one that maps them to GPUs, to reduce execution time in asynchronous work-conserving multi-GPU systems.

Pith tools