REVIEW 2 cited by
Learning the Pareto Front Using Bootstrapped Observation Samples
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We consider Pareto front identification (PFI) for linear bandits (PFILin), i.e., the goal is to identify a set of arms with undominated mean reward vectors when the mean reward vector is a linear function of the context. PFILin includes the best arm identification problem and multi-objective active learning as special cases. The sample complexity of our proposed algorithm is optimal up to a logarithmic factor. In addition, the regret incurred by our algorithm during the estimation is within a logarithmic factor of the optimal regret among all algorithms that identify the Pareto front. Our key contribution is a new estimator that in every round updates the estimate for the unknown parameter along multiple context directions -- in contrast to the conventional estimator that only updates the parameter estimate along the chosen context. This allows us to use low-regret arms to collect information about Pareto optimal arms. Our key innovation is to reuse the exploration samples multiple times; in contrast to conventional estimators that use each sample only once. Numerical experiments demonstrate that the proposed algorithm successfully identifies the Pareto front while controlling the regret.
Forward citations
Cited by 2 Pith papers
-
Adaptive Data Augmentation for Thompson Sampling
A hypothetical-context estimator is claimed to make Thompson Sampling minimax optimal for linear contextual bandits under arbitrary contexts, though the proof has a critical gap.
-
Short-Term Pain for Long-Term Gain: Adaptive Experiment with Post-Commitment Reward Shift
RAEC’s predetermined reserved exploration achieves matching minimax regret for post-commitment reward-shift bandits across short-experiment, balanced, and short-commitment regimes.
Discussion (0). Continue with ORCID to comment.