Introduces P2MLE-UCB and GP2-UCB round-based algorithms achieving optimal regret bounds for multiplicative and general position-aware MNL bandits, with efficient optimization subroutines.
Learning an optimal assortment policy under observational data
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
New algorithms for joint contextual MNL assortment and pricing deliver improved online regret bounds of order W sqrt(d T log N)/L0 and local suboptimality guarantees offline.
Pairing new products with top incumbents and exploring multiple new products simultaneously up to a potential-based threshold is optimal for minimizing regret in assortment-based social learning.
citing papers explorer
-
Learning in Position-Aware Multinomial Logit Bandits: From Multiplicative to General Position Effects
Introduces P2MLE-UCB and GP2-UCB round-based algorithms achieving optimal regret bounds for multiplicative and general position-aware MNL bandits, with efficient optimization subroutines.
-
Optimal Online and Offline Algorithms for Contextual MNL with Applications to Assortment and Pricing
New algorithms for joint contextual MNL assortment and pricing deliver improved online regret bounds of order W sqrt(d T log N)/L0 and local suboptimality guarantees offline.
-
Optimal Exploration of New Products under Assortment Decisions
Pairing new products with top incumbents and exploring multiple new products simultaneously up to a potential-based threshold is optimal for minimizing regret in assortment-based social learning.