REVIEW 2 cited by
Linear Contextual Bandits with Interference
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Interference, a key concept in causal inference, extends the reward modeling process by accounting for the impact of one unit's actions on the rewards of others. In contextual bandit (CB) settings, where multiple units are present in the same round, potential interference can significantly affect the estimation of expected rewards for different arms, thereby influencing the decision-making process. Although some prior work has explored multi-agent and adversarial bandits in interference-aware settings, the effect of interference in CB, as well as the underlying theory, remains significantly underexplored. In this paper, we introduce a systematic framework to address interference in Linear CB (LinCB), bridging the gap between causal inference and online decision-making. We propose a series of algorithms that explicitly quantify the interference effect in the reward modeling process and provide comprehensive theoretical guarantees, including sublinear regret bounds, finite sample upper bounds, and asymptotic properties. The effectiveness of our approach is demonstrated through simulations and a synthetic data generated based on MovieLens data.
Forward citations
Cited by 2 Pith papers
-
Balancing Interference and Correlation in Spatial Experimental Designs: A Causal Graph Cut Approach
A surrogate for the ATE estimator's MSE is optimized with spectral graph cuts to produce cluster-randomized designs that adapt to the spatial covariance and accommodate moderate-to-large interference.
-
Learning Peer Influence Probabilities with Linear Contextual Bandits
In a linear contextual bandit setting with k network interventions per round, cumulative regret and influence-probability estimation error obey a rate trade-off, and a new algorithm, InfluenceCB, can attain any point ...
Discussion (0). Continue with ORCID to comment.