REVIEW 4 major objections 4 minor 276 references
Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits
T0 review · 4 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Lexicographic low-rank matrix bandits are solved near-optimally with online updates.
desk verdict New combination of lexicographic bandits and low-rank structure with a real online speedup, but the main regret bound needs an oracle D_rr and an unstated scaling condition. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by three components. Objective-specific Stein-type subspace estimation solves a nuclear-norm regularized loss on score-transformed exploration samples to obtain estimates of each parameter matrix's row and column subspaces, with error controlled by $T_1$ and by $D_{rr}$, the smallest $r$-th singular value across all objective matrices. A transformed feature map $f_i(X)$ rotates the arm into the estimated subspaces and vectorizes four blocks, so the first $k=(d_1+d_2)r-r^2$ coordinates carry the low-rank signal and the remaining coordinates form the tail; anisotropic regularization $\Lambda=\mathrm{diag}(\lambda_0 I_k,\lambda_\perp I_{p-k})$ with $\lambda_\perp\gg\lambda_0$ pushes learning onto the signal coordinates. Finally, Lexi-LowGLM filters the candidate set objective by objective with tolerance $W_i\,c_{t,i_t}(X_t)$, keeping the lexicographically optimal arm inside while bounding each objective's gap, and updates estimators via an online Newton-type proximal step whose logarithmic potential sum yields the $O(T)$ update cost.
What would settle it
Run Lexi-LowGLM on a synthetic instance where the true minimum singular value $D_{rr}$ is ten times smaller than the value supplied to the algorithm, and track objective-wise regret and subspace estimation error over $T=10^5$; if the overestimate breaks the confidence intervals, the regret should stop following the claimed $\widetilde{O}(\sqrt{T})$ slope or the subspace error should exceed the theorem's bound.
Extended reading notes
Core claim
The central claim is that a lexicographic generalized low-rank matrix bandit admits a computationally light algorithm whose regret is uniformly sublinear on every objective. Concretely, under low-rank, bounded-link, bounded-noise, and lexicographic trade-off assumptions, Lexi-LowGLM first estimates each objective's row and column subspaces from $T_1$ exploration samples, then runs lexicographic candidate filtering on reduced features while maintaining each objective estimator by one proximal Newton step per round. Theorem 2 states that with probability at least $1-2\delta$, the cumulative regret on objective $i$ is $\widetilde{O}(W_i^{\mathrm{lex}}\sqrt{m}\,(d_1+d_2)r\sqrt{T})$, where $W_i^{\mathrm{lex}}=1+w+\cdots+w^{i-1}$; Theorem 1 gives the scalarized variant with factor $W^{\mathrm{sca}}$. The effective dimension replaces $d_1d_2$ with $(d_1+d_2)r$, and the online update reduces the cumulative estimator-update complexity over $T$ rounds from $O(T^2)$ to $O(T)$, matching the single-objective rate when $m=1$.
Load-bearing premise
The proof relies on knowing in advance the minimum strength, across objectives, of the $r$-th singular value of each unknown parameter matrix; an overestimate makes the exploration length and tail bound too small, and the confidence intervals that carry the regret bound need not hold.
Editorial extensions
If this is right
- Every objective's regret is sublinear in $T$ simultaneously, so the learner approaches the lexicographically optimal arm while keeping lower-priority losses controlled.
- The regret depends on $(d_1+d_2)r$ rather than $d_1d_2$, so matrix structure pays off even when outcomes are vector-valued and priority-ordered.
- Lexi-LowGLM's estimator-update cost is $O(T)$ over $T$ rounds, compared with $O(T^2)$ for batch refitting, making the method feasible for long horizons.
- When $w=0$, the lexicographic factor is $W_i^{\mathrm{lex}}=1$ for every objective, and the $\sqrt{m}$ factor improves on the scalarized baseline's factor $m$.
- With a single objective the bound reduces to $\widetilde{O}((d_1+d_2)r\sqrt{T})$, matching the rate of prior generalized low-rank matrix bandit methods.
- The confidence and potential arguments separate estimation from lexicographic decision making, so the online-update proof strategy can be reused in related bandit problems.
- When $w=0$, the lexicographic filtering cost disappears and the algorithm degrades gracefully to a multi-objective low-rank bandit with uniform guarantees.
Reading between the lines
- Editorial: A hidden practical cost is $D_{rr}$, the smallest $r$-th singular value across objectives; the theorem fixes this number when setting $T_1$ and $S_\perp$, so a fully parameter-free version of the algorithm would need to estimate $D_{rr}$ online and would likely pay an extra burn-in or multiplicative cost.
- Editorial: Because each estimator update consumes only the current observation rather than replaying all history, the online Newton step should extend naturally to non-stationary or delayed rewards without changing the $O(T)$ complexity.
- Editorial: The rank-two experiments show the third-objective regret still growing at $T=10{,}000$; the theory predicts eventual $\widetilde{O}(\sqrt{T})$ growth, so running the same instance well past $T=100{,}000$ would directly test whether the slope flattens as predicted.
- Editorial: The common confidence width used in the filtering tolerance may be conservative for arms with small uncertainty in all active objectives; exploiting each objective's estimated subspace geometry could sharpen the candidate elimination without changing the theoretical order.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces lexicographic generalized low-rank matrix bandits, a multi-objective bandit setting in which each objective has its own low-rank generalized linear model and the learner uses a lexicographic preference order. The authors propose two algorithms: Scalar-LowGLM, a scalarized batch-estimation baseline, and Lexi-LowGLM, which combines objective-specific subspace estimation, sequential lexicographic arm filtering, and online Newton-style estimator updates. For both algorithms they state objective-wise regret bounds of O~(W (d_1+d_2) r sqrt(T)) with W depending on the lexicographic trade-off parameter w, and they claim that Lexi-LowGLM reduces the cumulative estimator-update complexity from O(T^2) to O(T). The proofs follow a three-stage structure: subspace estimation via a nuclear-norm regularized Stein-type estimator, confidence sets for (batch or online) generalized linear estimators in a transformed feature space, and a lexicographic filtering argument. Numerical experiments on synthetic rank-one and rank-two instances compare the proposed methods with G-ESTT and MTLO.
Significance. If the stated guarantees were fully established, this would be a useful contribution: it extends low-rank matrix bandits to multiple lexicographically ordered objectives and obtains a dimension dependence on the intrinsic rank parameter (d_1+d_2)r instead of the ambient d_1 d_2, while also proposing an online estimator-update scheme that reduces the per-horizon computational cost relative to repeated batch refitting. The problem formulation and the algorithmic idea are clear and well motivated, and the paper explicitly provides regret theorems with a nontrivial filtering analysis for the lexicographic candidate set. However, the advertised low-rank rate is currently conditional on an unstated and practically unavailable oracle for D_rr (the minimum r-th singular value of the objective matrices) and on a 'standard scaling condition' that does not appear among Assumptions 1-6.
major comments (4)
- [Theorems 1 and 2 (Sections 4.2-4.3) and their proofs (Appendices B-C)] The parameter specification in Theorem 1 and the proof of Lemma 3 require the learner to know D_rr, the minimum r-th largest singular value of the objective parameter matrices, in order to set T1 and S_perp. D_rr is not an observable quantity, and no practical estimator or adaptive choice is provided. If D_rr is overestimated, S_perp is too small and the high-probability event in Lemma 3 fails, invalidating the confidence widths; if D_rr is underestimated, T1 may exceed T. Moreover, the final step of both proofs silently invokes the condition sqrt(M(d1+d2)r)/D_rr <= O((d1+d2)r) to drop the 1/D_rr term. This condition is not stated in Assumptions 1-6, and when it fails the derived bound is O~(W (sqrt(M(d1+d2)r)/D_rr + (d1+d2)r) sqrt(T)), which is not the advertised intrinsic-dimension rate. The theorems should either include an explicit lower-bound assumption on D_rr and the scaling condition, or state the regret bound with the full 1/D_rr dependence.
- [Lemma 4 (Appendix B, Eq. (7))] The confidence-bound proof for the batch estimator uses the equality gt,i(theta_hat) = sum_tau y_tau x_tau as the first-order optimality condition for the estimator defined in Eq. (7), which includes the constraint ||theta||_2 <= S. If the constrained minimizer lies on the boundary, the KKT condition contains a Lagrange multiplier term and the displayed equality is not valid. This affects the derivation of the confidence radius beta_t and consequently the regret bound. The authors should either argue that the unconstrained minimizer is always inside the feasible ball (e.g., by choosing S appropriately), or redo the argument with the proper KKT conditions.
- [Section 4.3, Eq. (10); Section 1 complexity claim] The claim that Lexi-LowGLM reduces the estimator-update complexity from O(T^2) to O(T) is not fully supported. Each round requires solving the constrained proximal problem in Eq. (10), a p-dimensional quadratic program with a ball constraint, yet the paper does not specify the solver, the number of iterations, or the per-round cost. If the proximal solve requires an iterative algorithm, the per-round cost may be polynomial in p and could dominate the T-round complexity. The O(T) claim is only justified if each update is shown to cost O(1) or O(p) with explicit constants; otherwise the computational contribution should be restated as a number-of-data-passes claim rather than a per-round complexity claim.
- [Assumptions 3-4] Assumptions 3 and 4 refer to constants c_mu, L_mu, U that are uniform 'over the relevant domain', but the relevant domain is never defined. Since the link functions are evaluated at fi(X)^T theta for estimated parameters and at <X, Theta*> for the true parameters, the domain should be specified (e.g., a bounded interval implied by ||X||_F <= 1 and ||Theta*_i||_F <= S, along with the estimator norm constraint). Without this, the existence of the constants is not well posed.
minor comments (4)
- [Theorems 1 and 2] The theorems state the regret bound with probability 1-2*delta, but the proofs use delta in the individual concentration steps and union bounds; it would be helpful to state explicitly how the two delta terms in the exploration phase and the online phase are combined.
- [Algorithm input and Theorem 1 parameter list] The algorithm inputs include S_perp, but the formula for S_perp appears only inside the proof of Theorem 1. For reproducibility, the theorem statement should list the full parameter assignments for T1, S_perp, lambda_perp, and lambda_0 explicitly.
- [Figures 1 and 2] The experiments are averaged over 10 trials but no error bars or standard deviations are shown, which makes it hard to judge the stability of the reported regret curves.
- [Various equations] There are several minor typos and inconsistencies in the text, e.g., the use of '~O' versus 'O~' for the big-O notation, and the definition of W_sca in Theorem 1 uses (1+w)^(i-1) while the proof uses (1+w)^(m-j); these should be harmonized.
Circularity Check
No significant circularity: the regret bounds are derived from concentration and elliptical-potential arguments, subspace estimation is imported from external prior work, and the lexicographic analysis is new.
full rationale
I walked the derivation chain of The paper's Theorems 1 and 2. The claimed regret bounds are obtained by (i) importing a subspace-estimation error bound from Kang et al. [2022, Theorem 4.1], which is external to this paper, (ii) converting that subspace error into a tail-coordinate bound via a standard singular-subspace perturbation argument (Lemmas 2 and 3), (iii) proving confidence intervals for the batch and online estimators using self-normalized martingale inequalities, and (iv) summing the resulting confidence widths with elliptical-potential arguments. No parameter is fitted to a subset of the data and then renamed as a prediction; the confidence radii are derived from stated noise, Lipschitz, and regularization constants. The lexicographic filtering analysis (Lemmas 10 and 11) is self-contained given Assumption 5 and the already-established confidence event, and the objective-wise regret is bounded by summing the actual reward gaps, not by assuming the rate. The self-citations to Xue et al. 2025a and 2025b are contextual or baseline references; the load-bearing subspace-estimation theorem is an external result, so there is no self-citation chain that forces the conclusion. The main caveat found is not circularity but an unstated correctness condition: the algorithm's parameter choices require knowing D_rr, the minimum r-th singular value of the objective parameter matrices, and the proof silently invokes the scaling condition sqrt(M(d1+d2)r)/D_rr less than or similar to (d1+d2)r to drop the 1/D_rr term. If D_rr is small, unknown, or misestimated, the displayed low-rank rate does not follow. This is a conditional-validity issue, not a case of the conclusion being equivalent to an input or to a fitted quantity.
Assumptions & free parameters
free parameters (4)
- T1 (exploration length)
- S_perp (tail-coordinate bound)
- lambda_perp (anisotropic regularization on complement)
- lambda0 (regularization on low-rank subspace)
assumptions (9)
- domain assumption Assumption 1: each objective parameter matrix Theta_i has rank at most r, with r much smaller than min(d1,d2).
- domain assumption Assumption 2: parameter matrices have Frobenius norm at most S, and arms have Frobenius norm at most 1.
- domain assumption Assumption 3: each inverse link function mu_i is continuously differentiable with derivatives bounded between c_mu and L_mu.
- domain assumption Assumption 4: rewards and noise are bounded by U and R respectively.
- domain assumption Assumption 5: there exists w >= 0 such that lower-priority objective gaps are bounded by w times the maximum higher-priority gap.
- domain assumption Assumption 6: there exists an exploration distribution D over X with bounded score-function second moments and with rows or columns pairwise independent.
- standard math Kang et al. 2022 Theorem 4.1: subspace estimation error bound for generalized low-rank matrix models.
- standard math Self-normalized martingale inequalities from Abbasi-Yadkori et al. 2011 and 2012.
- ad hoc to paper Standard scaling condition sqrt(M(d1+d2)r)/D_rr is at most (d1+d2)r.
Cite this review
Pith. "Pith review of Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits." pith.science (2026). https://pith.science/paper/DKJDF6SW
@misc{pith2026260804324,
author = {Pith},
title = {Pith review of: Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits},
year = {2026},
howpublished = {\url{https://pith.science/paper/DKJDF6SW}},
note = {Machine review of arXiv:2608.04324}
}
abstract
This paper studies generalized low-rank matrix bandits with multiple prioritized objectives. At each round, the learner selects a matrix-valued arm and observes a vector-valued reward, whose components correspond to multiple objectives with different priority levels. Each objective is governed by an objective-specific generalized low-rank matrix model, and the learner evaluates arms according to a lexicographic preference order, prioritizing higher-level objectives before lower-level ones. We propose \textsc{Lexi-LowGLM}, an efficient online algorithm that first estimates objective-specific low-rank subspaces and then performs lexicographic learning in the reduced feature spaces. Unlike existing single-objective algorithms that repeatedly solve a batch generalized linear estimator using all historical observations, \textsc{Lexi-LowGLM} updates each objective-specific estimator via an online Newton step, reducing the estimator-update complexity over $T$ rounds from $O(T^2)$ to $O(T)$. We establish a regret bound of $\widetilde O\left(W_i^{\rm lex}\sqrt{m}\,(d_1+d_2)r\sqrt{T}\right)$ for each objective $i\in[m]$, where $r$ is an upper bound on the ranks of the objective-specific parameter matrices and $W_i^{\rm lex}$ characterizes the lexicographic trade-off effect. This bound depends on the effective low-rank dimension $(d_1+d_2)r$ rather than the ambient dimension $d_1d_2$. Numerical experiments further validate the effectiveness and computational efficiency of the proposed method.
Figures
Reference graph
Works this paper leans on
-
[1]
Proceedings of the 29th International Joint Conference on Artificial Intelligence , pages =
Nearly Optimal Regret for Stochastic Linear Bandits with Heavy-Tailed Payoffs , author =. Proceedings of the 29th International Joint Conference on Artificial Intelligence , pages =
-
[2]
Efficient Algorithms for Generalized Linear Bandits with Heavy-tailed Rewards , year =
Xue, Bo and Wang, Yimu and Wan, Yuanyu and Yi, Jinfeng and Zhang, Lijun , booktitle =. Efficient Algorithms for Generalized Linear Bandits with Heavy-tailed Rewards , year =
-
[3]
arXiv preprint arXiv:2407.17466 , year=
Traversing pareto optimal policies: Provably efficient multi-objective reinforcement learning , author=. arXiv preprint arXiv:2407.17466 , year=
-
[4]
Multiobjective Lipschitz Bandits under Lexicographic Ordering , journal=
Xue, Bo and Cheng, Ji and Liu, Fei and Wang, Yimu and Zhang, Qingfu , year=. Multiobjective Lipschitz Bandits under Lexicographic Ordering , journal=
-
[5]
Proceedings of the 34th International Joint Conference on Artificial Intelligence , pages =
Problem-dependent Regret for Lexicographic Multi-Armed Bandits with Adversarial Corruptions , author =. Proceedings of the 34th International Joint Conference on Artificial Intelligence , pages =
-
[6]
Multiple Trade-offs: An Improved Approach for Lexicographic Linear Bandits , booktitle=
Xue, Bo and Lin, Xi and Zhang, Xiaoyuan and Zhang, Qingfu , year=. Multiple Trade-offs: An Improved Approach for Lexicographic Linear Bandits , booktitle=
-
[7]
Proceedings of the 42nd International Conference on Machine Learning , year=
Multi-objective Linear Reinforcement Learning with Lexicographic Rewards , author=. Proceedings of the 42nd International Conference on Machine Learning , year=
-
[8]
Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear Bandits , booktitle=
Cheng, Ji and Xue, Bo and Yi, Jiaxiang and Zhang, Qingfu , year=. Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear Bandits , booktitle=
Show all 276 references
-
[9]
Proceedings of the 34th International Joint Conference on Artificial Intelligence , pages =
Multi-Objective Neural Bandits with Random Scalarization , author =. Proceedings of the 34th International Joint Conference on Artificial Intelligence , pages =
-
[10]
Using Confidence Bounds for Exploitation-Exploration Trade-offs , journal =
Auer, Peter , year =. Using Confidence Bounds for Exploitation-Exploration Trade-offs , journal =
-
[11]
Bandits With Heavy Tail , year=
Bubeck, Sébastien and Cesa-Bianchi, Nicolò and Lugosi, Gábor , journal=. Bandits With Heavy Tail , year=
-
[12]
Weighted sums of certain dependent random variables
Azuma, Kazuoki. Weighted sums of certain dependent random variables. Tohoku Mathematical Journal. 1967
1967
-
[13]
Annales de l'I.H.P
Catoni, Olivier , title =. Annales de l'I.H.P. Probabilit\'es et statistiques , volume=. 2012 , pages =
2012
-
[14]
Some aspects of the sequential design of experiments
Robbins, Herbert. Some aspects of the sequential design of experiments. Bulletin of the American Mathematical Society. 1952
1952
-
[15]
and Long, Philip M
Abe, Naoki and Biermann, Alan W. and Long, Philip M. , title=. Algorithmica , year=
-
[16]
The heavy tail of the human brain
James A Roberts and Tjeerd W Boonstra and Michael Breakspear. The heavy tail of the human brain. Current Opinion in Neurobiology. 2015
2015
-
[17]
, title =
Li, Lihong and Chu, Wei and Langford, John and Schapire, Robert E. , title =. Proceedings of the 19th International Conference on World Wide Web , year =
-
[18]
Proceedings of the 33rd International Conference on International Conference on Machine Learning , pages =
Medina, Andres Munoz and Yang, Scott , title =. Proceedings of the 33rd International Conference on International Conference on Machine Learning , pages =
-
[19]
Proceedings of the 14th International Conference on Artificial Intelligence and Statistics , pages =
Contextual Bandits with Linear Payoff Functions , author =. Proceedings of the 14th International Conference on Artificial Intelligence and Statistics , pages =
-
[20]
Proceedings of the 31st International Conference on Machine Learning , pages =
Heavy-tailed regression with a generalized median-of-means , author =. Proceedings of the 31st International Conference on Machine Learning , pages =
-
[21]
Proceedings of the 31st Conference On Learning Theory , pages =
Efficient Contextual Bandits in Non-stationary Worlds , author =. Proceedings of the 31st Conference On Learning Theory , pages =
-
[22]
Proceedings of the 31st International Conference on Machine Learning , pages =
Taming the Monster: A Fast and Simple Algorithm for Contextual Bandits , author =. Proceedings of the 31st International Conference on Machine Learning , pages =
-
[23]
_1 -regression with Heavy-tailed Distributions , booktitle =
Lijun Zhang and Zhi. _1 -regression with Heavy-tailed Distributions , booktitle =
-
[24]
Online Stochastic Linear Optimization under One-bit Feedback , booktitle =
Lijun Zhang and Tianbao Yang and Rong Jin and Yichi Xiao and Zhi. Online Stochastic Linear Optimization under One-bit Feedback , booktitle =
-
[25]
, title =
Shao, Han and Yu, Xiaotian and King, Irwin and Lyu, Michael R. , title =. Advances in Neural Information Processing Systems 31 , pages =
-
[26]
Advances in Neural Information Processing Systems 24 , pages =
Improved Algorithms for Linear Stochastic Bandits , author =. Advances in Neural Information Processing Systems 24 , pages =
-
[27]
Proceedings of the 15th International Conference on Artificial Intelligence and Statistics , pages =
Online-to-Confidence-Set Conversions and Application to Sparse Stochastic Bandits , author =. Proceedings of the 15th International Conference on Artificial Intelligence and Statistics , pages =
-
[28]
Advances in Neural Information Processing Systems 24 , pages =
Linear Submodular Bandits and their Application to Diversified Retrieval , author =. Advances in Neural Information Processing Systems 24 , pages =
-
[29]
Sergey Foss and Dmitry Korshunov and Stan Zachary , TITLE =
-
[30]
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems , JOURNAL =
S\'. Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems , JOURNAL =. 2012 , volume =
2012
-
[31]
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins. Asymptotically efficient adaptive allocation rules. Advances in Applied Mathematics. 1985
1985
-
[32]
Proceedings of the 23rd Annual Conference on Learning Theory , pages =
Audibert, Jean-Yves and Bubeck, Sébastien , title =. Proceedings of the 23rd Annual Conference on Learning Theory , pages =
-
[33]
Best Arm Identification: A Unified Approach to Fixed Budget and Fixed Confidence , year =
Gabillon, Victor and Ghavamzadeh, Mohammad and Lazaric, Alessandro , booktitle =. Best Arm Identification: A Unified Approach to Fixed Budget and Fixed Confidence , year =
-
[34]
Bandit problems: sequential allocation of experiments , author=
-
[35]
and Kalai, Adam Tauman and McMahan, H
Flaxman, Abraham D. and Kalai, Adam Tauman and McMahan, H. Brendan , title =. Proceedings of the 16th Annual ACM-SIAM Symposium on Discrete Algorithms , year =
-
[36]
Proceedings of the 40th Annual ACM Symposium on Theory of Computing , year =
Kleinberg, Robert and Slivkins, Aleksandrs and Upfal, Eli , title =. Proceedings of the 40th Annual ACM Symposium on Theory of Computing , year =
-
[37]
Advances in Applied Probability , YEAR =
Rajeev Agrawal , TITLE =. Advances in Applied Probability , YEAR =
-
[38]
Finite-time Analysis of the Multiarmed Bandit Problem , journal =
Auer, Peter and Cesa-Bianchi, Nicol\`. Finite-time Analysis of the Multiarmed Bandit Problem , journal =
-
[39]
IEEE Journal of Selected Topics in Signal Processing , pages=
Deterministic sequencing of exploration and exploitation for multi-armed bandit problems , author=. IEEE Journal of Selected Topics in Signal Processing , pages=. 2013 , volume=
2013
-
[40]
Hayes and Sham M
Varsha Dani and Thomas P. Hayes and Sham M. Kakade , TITLE =. Proceedings of the 21st Annual Conference on Learning , YEAR =
-
[41]
The Nonstochastic Multiarmed Bandit Problem , journal =
Auer, Peter and Cesa-Bianchi, Nicol\`. The Nonstochastic Multiarmed Bandit Problem , journal =. 2002 , pages =
2002
-
[42]
and Thomas, Joy A
Cover, Thomas M. and Thomas, Joy A. , title =. 2006 , publisher =
2006
-
[43]
IEEE Transactions on Information Theory , volume=
PAC-Bayesian inequalities for martingales , author=. IEEE Transactions on Information Theory , volume=
-
[44]
Robust linear least squares regression
Audibert, Jean-Yves and Catoni, Olivier. Robust linear least squares regression. The Annals of Statistics. 2011
2011
-
[45]
Empirical risk minimization for heavy-tailed losses
Brownlees, Christian and Joly, Emilien and Lugosi, G \'a bor. Empirical risk minimization for heavy-tailed losses. The Annals of Statistics. 2015
2015
-
[46]
Journal of Machine Learning Research , year =
Daniel Hsu and Sivan Sabato , title =. Journal of Machine Learning Research , year =
-
[47]
Proceedings of the 36th International Conference on Machine Learning , pages =
Optimal Algorithms for Lipschitz Bandits with Heavy-tailed Rewards , author =. Proceedings of the 36th International Conference on Machine Learning , pages =
-
[48]
and Van Loan,, Charles F
Golub,, Gene H. and Van Loan,, Charles F. , title =. 1996 , publisher =
1996
-
[49]
and Hsu, Daniel and Kakade, Sham M
Agarwal, Alekh and Foster, Dean P. and Hsu, Daniel and Kakade, Sham M. and Rakhlin, Alexander , title =. SIAM Journal on Optimization , volume =
-
[50]
2014 , journal =
From Bandits to Monte-Carlo Tree Search: The Optimistic Principle Applied to Optimization and Planning , pages =. 2014 , journal =
2014
-
[51]
Lyu and Irwin King , title =
Xiaotian Yu and Han Shao and Michael R. Lyu and Irwin King , title =. 2018 , journal =
2018
-
[52]
Lipschitz Bandits Without the Lipschitz Constant , booktitle =
Bubeck, S. Lipschitz Bandits Without the Lipschitz Constant , booktitle =. 2011 , pages =
2011
-
[53]
SIAM Journal on Control and Optimization , volume =
Agrawal, Rajeev , title =. SIAM Journal on Control and Optimization , volume =
-
[54]
Bandit Convex Optimization:
Sébastien Bubeck and Ofer Dekel and Tomer Koren and Yuval Peres , booktitle =. Bandit Convex Optimization:
-
[55]
Bandit Algorithms , publisher=
Lattimore, Tor and Szepesvári, Csaba , year=. Bandit Algorithms , publisher=
-
[56]
Advances in Neural Information Processing Systems 23 , pages =
Parametric Bandits: The Generalized Linear Case , author =. Advances in Neural Information Processing Systems 23 , pages =
-
[57]
Advances in Neural Information Processing Systems 30 , pages =
Jun, Kwang-Sung and Bhargava, Aniruddha and Nowak, Robert and Willett, Rebecca , title =. Advances in Neural Information Processing Systems 30 , pages =
-
[58]
McCullagh, John A
P. McCullagh, John A. Nelder , title =
-
[59]
Hubert and Shi, Elaine and Song, Dawn , title =
Chan, T.-H. Hubert and Shi, Elaine and Song, Dawn , title =. Proceedings of the 37th International Colloquium Conference on Automata, Languages and Programming: Part II , year =
-
[60]
Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages =
Multi-Objective Generalized Linear Bandits , author =. Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages =
-
[61]
Logarithmic Regret Algorithms for Online Convex Optimization , journal =
Hazan, Elad and Agarwal, Amit and Kale, Satyen , year =. Logarithmic Regret Algorithms for Online Convex Optimization , journal =
-
[62]
Proceedings of the 34th International Conference on Machine Learning , pages =
Provably Optimal Algorithms for Generalized Linear Contextual Bandits , author =. Proceedings of the 34th International Conference on Machine Learning , pages =
-
[63]
Breaking the Moments Condition Barrier: No-Regret Algorithm for Bandits with Super Heavy-Tailed Payoffs , year =
Zhong, Han and Huang, Jiayi and Yang, Lin and Wang, Liwei , booktitle =. Breaking the Moments Condition Barrier: No-Regret Algorithm for Bandits with Super Heavy-Tailed Payoffs , year =
-
[64]
The Annals of Statistics , number =
G. The Annals of Statistics , number =
-
[65]
Bayesian Optimization under Heavy-tailed Payoffs , year =
Ray Chowdhury, Sayak and Gopalan, Aditya , booktitle =. Bayesian Optimization under Heavy-tailed Payoffs , year =
-
[66]
IEEE transactions on pattern analysis and machine intelligence , pages=
The apolloscape open dataset for autonomous driving and its application , author=. IEEE transactions on pattern analysis and machine intelligence , pages=
-
[67]
Proceedings of the 29th International Joint Conference on Artificial Intelligence , year =
Lijun Zhang , title =. Proceedings of the 29th International Joint Conference on Artificial Intelligence , year =
-
[68]
and Nowe, Ann , booktitle=
Drugan, Madalina M. and Nowe, Ann , booktitle=. Designing multi-objective multi-armed bandits algorithms: A study , year=
-
[69]
Proceedings of the 19th International Conference on Artificial Intelligence and Statistics , pages =
Pareto Front Identification from Stochastic Bandit Feedback , author =. Proceedings of the 19th International Conference on Artificial Intelligence and Statistics , pages =
-
[70]
2005 , publisher =
Ehrgott, Matthias , title =. 2005 , publisher =
2005
-
[71]
Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives , journal =
Alihan H. Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives , journal =
-
[72]
Proceedings of the 21st International Conference on Artificial Intelligence and Statistics , pages=
Multi-objective contextual bandit problem with similarity information , author=. Proceedings of the 21st International Conference on Artificial Intelligence and Statistics , pages=
-
[73]
Multi-objective
Van Moffaert, Kristof and Van Vaerenbergh, Kevin and Vrancx, Peter and Nowe, Ann , booktitle=. Multi-objective. 2014 , pages=
2014
-
[74]
Proceedings of the 39th International Conference on Machine Learning , pages =
Stochastic Contextual Dueling Bandits under Linear Stochastic Transitivity Models , author =. Proceedings of the 39th International Conference on Machine Learning , pages =
-
[75]
2016 , booktitle =
Joseph, Matthew and Kearns, Michael and Morgenstern, Jamie and Roth, Aaron , title =. 2016 , booktitle =
2016
-
[76]
2012 , booktitle =
Rodriguez, Mario and Posse, Christian and Zhang, Ethan , title =. 2012 , booktitle =
2012
-
[77]
2004 , publisher=
Convex optimization , author=. 2004 , publisher=
2004
-
[78]
Breaking the
Ghosh, Avishek and Sankararaman, Abishek , booktitle =. Breaking the
-
[79]
Proceedings of the 39th International Conference on Machine Learning , pages =
Adaptive Best-of-Both-Worlds Algorithm for Heavy-Tailed Multi-Armed Bandits , author =. Proceedings of the 39th International Conference on Machine Learning , pages =
-
[80]
Proceedings of the 39th International Conference on Machine Learning , pages =
A Simple Unified Framework for High Dimensional Bandit Problems , author =. Proceedings of the 39th International Conference on Machine Learning , pages =
-
[81]
Proceedings of the 39th International Conference on Machine Learning , pages =
Linear Bandit Algorithms with Sublinear Time Complexity , author =. Proceedings of the 39th International Conference on Machine Learning , pages =
-
[82]
Proceedings of the 39th International Conference on Machine Learning , pages =
Contextual Bandits with Smooth Regret: Efficient Learning in Continuous Action Spaces , author=. Proceedings of the 39th International Conference on Machine Learning , pages =
-
[83]
A Best-of-Both-Worlds Algorithm for Bandits with Delayed Feedback , year =
Masoudian, Saeed and Zimmert, Julian and Seldin, Yevgeny , booktitle =. A Best-of-Both-Worlds Algorithm for Bandits with Delayed Feedback , year =
-
[84]
Lipschitz Bandits with Batched Feedback , year =
Feng, Yasong and Huang, Zengfeng and Wang, Tianyu , booktitle =. Lipschitz Bandits with Batched Feedback , year =
-
[85]
Finite-Time Regret of Thompson Sampling Algorithms for Exponential Family Multi-Armed Bandits , year =
Jin, Tianyuan and Xu, Pan and Xiao, Xiaokui and Anandkumar, Anima , booktitle =. Finite-Time Regret of Thompson Sampling Algorithms for Exponential Family Multi-Armed Bandits , year =
-
[86]
Communication Efficient Federated Learning for Generalized Linear Bandits , year =
Li, Chuanhao and Wang, Hongning , booktitle =. Communication Efficient Federated Learning for Generalized Linear Bandits , year =
-
[87]
Outlier-Robust Sparse Mean Estimation for Heavy-Tailed Distributions , year =
Diakonikolas, Ilias and Kane, Daniel and Lee, Jasper and Pensia, Ankit , booktitle =. Outlier-Robust Sparse Mean Estimation for Heavy-Tailed Distributions , year =
-
[88]
Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed Noise , year =
Gorbunov, Eduard and Danilova, Marina and Dobre, David and Dvurechenskii, Pavel and Gasnikov, Alexander and Gidel, Gauthier , booktitle =. Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed Noise , year =
-
[89]
Proceedings of the 39th International Conference on Machine Learning , pages =
Improved Rates for Differentially Private Stochastic Convex Optimization with Heavy-Tailed Data , author =. Proceedings of the 39th International Conference on Machine Learning , pages =
-
[90]
Proceedings of the 39th International Conference on Machine Learning , pages =
High Probability Guarantees for Nonconvex Stochastic Gradient Descent with Heavy Tails , author =. Proceedings of the 39th International Conference on Machine Learning , pages =
-
[91]
Multi-armed Bandit Models for the Optimal Design of Clinical Trials: Benefits and Challenges , volume =
Sof. Multi-armed Bandit Models for the Optimal Design of Clinical Trials: Benefits and Challenges , volume =. Statistical Science , number =
-
[92]
MOEA/D: A Multiobjective Evolutionary Algorithm Based on Decomposition , year=
Zhang, Qingfu and Li, Hui , journal=. MOEA/D: A Multiobjective Evolutionary Algorithm Based on Decomposition , year=
-
[93]
Proceedings of the 38th International Conference on Machine Learning , pages =
Near-Optimal Representation Learning for Linear Bandits and Linear RL , author =. Proceedings of the 38th International Conference on Machine Learning , pages =
-
[94]
Proceedings of the 38th International Conference on Machine Learning , pages =
Robust Pure Exploration in Linear Bandits with Limited Budget , author =. Proceedings of the 38th International Conference on Machine Learning , pages =
-
[95]
2012 , volume =
Foundations and Trends in Machine Learning , title =. 2012 , volume =
2012
-
[96]
and Barto, Andrew G
Sutton, Richard S. and Barto, Andrew G. , edition =. Reinforcement Learning: An Introduction , year =
-
[97]
Proceedings of the 24th International Conference on Artificial Intelligence and Statistics , pages =
An Efficient Algorithm For Generalized Linear Bandit: Online Stochastic Gradient Descent and Thompson Sampling , author =. Proceedings of the 24th International Conference on Artificial Intelligence and Statistics , pages =
-
[98]
PAC-Bayesian Inequalities for Martingales , journal =
Yevgeny Seldin and Fran. PAC-Bayesian Inequalities for Martingales , journal =
-
[99]
Learning in Generalized Linear Contextual Bandits with Stochastic Delays , year =
Zhou, Zhengyuan and Xu, Renyuan and Blanchet, Jose , booktitle =. Learning in Generalized Linear Contextual Bandits with Stochastic Delays , year =
-
[100]
Advances in Neural Information Processing Systems 34 , pages =
Generalized Linear Bandits with Local Differential Privacy , author=. Advances in Neural Information Processing Systems 34 , pages =
-
[101]
John Myles White , TITLE =
-
[102]
Physics in Medicine & Biology , year=
Lexicographic ordering: intuitive multicriteria optimization for IMRT , author=. Physics in Medicine & Biology , year=
-
[103]
Integrated Assessment and Decision Support - Proceedings of the 1st Biennial Meeting of the International Environmental Modelling and Software Society , year=
Lexicographic Optimisation for Water Resources Planning: the Case of Lake Verbano, Italy , author=. Integrated Assessment and Decision Support - Proceedings of the 1st Biennial Meeting of the International Environmental Modelling and Software Society , year=
-
[104]
Proceedings of the 27th Conference on Learning Theory , pages =
Lipschitz Bandits: Regret Lower Bound and Optimal Algorithms , author =. Proceedings of the 27th Conference on Learning Theory , pages =
-
[105]
Advances in Neural Information Processing Systems 17 , year=
Nearly Tight Bounds for the Continuum-armed Bandit Problem , author=. Advances in Neural Information Processing Systems 17 , year=
-
[106]
X-Armed Bandits , journal =
S. X-Armed Bandits , journal =. 2011 , volume =
2011
-
[107]
Improved Rates for the Stochastic Continuum-Armed Bandit Problem
Auer, Peter and Ortner, Ronald and Szepesv \'a ri, Csaba. Improved Rates for the Stochastic Continuum-Armed Bandit Problem. Proceedings of the 20th Annual Conference on Learning Theory. 2007
2007
-
[108]
Online Optimization in X-Armed Bandits , year =
Bubeck, S\'. Online Optimization in X-Armed Bandits , year =. Advances in Neural Information Processing Systems 21 , pages =
-
[109]
2020 , booktitle =
Wang, Tianyu and Ye, Weicheng and Geng, Dawei and Rudin, Cynthia , title =. 2020 , booktitle =
2020
-
[110]
Yahyaa, Saba and M
Q. Yahyaa, Saba and M. Drugan, Madalina and Manderick, Bernard , title =. 2014 , booktitle =
2014
-
[111]
Multi-objective Contextual Multi-armed Bandit With a Dominant Objective , year=
Tekin, Cem and Turgay, Eralp , journal=. Multi-objective Contextual Multi-armed Bandit With a Dominant Objective , year=
-
[112]
2018 , booktitle =
Ma, Xiao and Zhao, Liqin and Huang, Guan and Wang, Zhi and Hu, Zelin and Zhu, Xiaoqiang and Gai, Kun , title =. 2018 , booktitle =
2018
-
[113]
Proceedings of 34th Conference on Learning Theory , pages =
Adaptive Discretization for Adversarial Lipschitz Bandits , author =. Proceedings of 34th Conference on Learning Theory , pages =
-
[114]
Proceedings of the 40th International Conference on International Conference on Machine Learning , pages=
Pareto Regret Analyses in Multi-objective Multi-armed Bandit , author=. Proceedings of the 40th International Conference on International Conference on Machine Learning , pages=
-
[115]
An Empirical Evaluation of Thompson Sampling , year =
Chapelle, Olivier and Li, Lihong , booktitle =. An Empirical Evaluation of Thompson Sampling , year =
-
[116]
Proceedings of the Workshop on On-line Trading of Exploration and Exploitation 2 , pages=
An unbiased offline evaluation of contextual bandit algorithms with generalized linear models , author=. Proceedings of the Workshop on On-line Trading of Exploration and Exploitation 2 , pages=
-
[117]
Proceedings of the 31st International Joint Conference on Artificial Intelligence , pages =
Lexicographic Multi-Objective Reinforcement Learning , author =. Proceedings of the 31st International Joint Conference on Artificial Intelligence , pages =
-
[118]
2015 , booktitle =
Wray, Kyle Hollins and Zilberstein, Shlomo , title =. 2015 , booktitle =
2015
-
[119]
Fair and Efficient Allocations under Lexicographic Preferences , booktitle=
Hosseini, Hadi and Sikdar, Sujoy and Vaish, Rohit and Xia, Lirong , year=. Fair and Efficient Allocations under Lexicographic Preferences , booktitle=
-
[120]
Multi-Objective MDPs with Conditional Lexicographic Reward Preferences , booktitle=
Wray, Kyle and Zilberstein, Shlomo and Mouaddib, Abdel-Illah , pages =. Multi-Objective MDPs with Conditional Lexicographic Reward Preferences , booktitle=
-
[121]
Stochastic Contextual Bandits with Long Horizon Rewards , booktitle=
Qin, Yuzhen and Li, Yingcong and Pasqualetti, Fabio and Fazel, Maryam and Oymak, Samet , year=. Stochastic Contextual Bandits with Long Horizon Rewards , booktitle=
-
[122]
Proceedings of the 34st Conference On Learning Theory , pages =
Chara Podimata and Alex Slivkins , title =. Proceedings of the 34st Conference On Learning Theory , pages =
-
[123]
Advances in Neural Information Processing Systems 32 , pages =
Nirandika Wanigasekara and Christina Lee Yu , title =. Advances in Neural Information Processing Systems 32 , pages =
-
[124]
Multi-Objective contextual bandits with a dominant objective , year=
Tekin, Cem and Turgay, Eralp , booktitle=. Multi-Objective contextual bandits with a dominant objective , year=
-
[125]
Customer Acquisition via Display Advertising Using Multi-Armed Bandit Experiments , booktitle =
Schwartz, Eric and Bradlow, Eric and Fader, Peter , year =. Customer Acquisition via Display Advertising Using Multi-Armed Bandit Experiments , booktitle =
-
[126]
Resource Allocation for Multi-source Multi-relay Wireless Networks
Khansa, Ali Al and Visoz, Raphael and Hayel, Yezekael and Lasaulce, Samson. Resource Allocation for Multi-source Multi-relay Wireless Networks. Ubiquitous Networking. 2021
2021
-
[127]
Multi-objective optimization for supply chain management problem: A literature review , volume =
Trisna, Trisna and Marimin, Marimin and Arkeman, Yandra and Sunarti, Titi , year =. Multi-objective optimization for supply chain management problem: A literature review , volume =
-
[128]
ACM Transactions on Information Systems , pages=
Wang, Yifan and Ma, Weizhi and Zhang, Min and Liu, Yiqun and Ma, Shaoping , title =. ACM Transactions on Information Systems , pages=. 2023 , volume =
2023
-
[129]
2022 , author =
Multiobjective optimization under uncertainty: A multiobjective robust (relative) regret approach , journal =. 2022 , author =
2022
-
[130]
2018 , booktitle =
Lykouris, Thodoris and Mirrokni, Vahab and Paes Leme, Renato , title =. 2018 , booktitle =
2018
-
[131]
2009 , journal =
Click Fraud , author =. 2009 , journal =
2009
-
[132]
Proceedings of the 32nd Conference on Learning Theory , pages =
Better Algorithms for Stochastic Bandits with Adversarial Corruptions , author =. Proceedings of the 32nd Conference on Learning Theory , pages =
-
[133]
Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics , pages =
Corruption-Tolerant Gaussian Process Bandit Optimization , author =. Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics , pages =
-
[134]
Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial Corruptions , year =
He, Jiafan and Zhou, Dongruo and Zhang, Tong and Gu, Quanquan , booktitle =. Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial Corruptions , year =
-
[135]
Advances in Neural Information Processing Systems 36 , year=
Robust Lipschitz Bandits to Adversarial Corruptions , author=. Advances in Neural Information Processing Systems 36 , year=
-
[136]
Stochastic Graphical Bandits with Adversarial Corruptions , booktitle=
Lu, Shiyin and Wang, Guanghui and Zhang, Lijun , year=. Stochastic Graphical Bandits with Adversarial Corruptions , booktitle=
-
[137]
Proceedings of the 23nd International Conference on Artificial Intelligence and Statistics , pages =
An Optimal Algorithm for Stochastic and Adversarial Bandits , author =. Proceedings of the 23nd International Conference on Artificial Intelligence and Statistics , pages =
-
[138]
Proceedings of the 25th International Conference on Artificial Intelligence and Statistics , pages =
Robust Stochastic Linear Contextual Bandits Under Adversarial Attacks , author =. Proceedings of the 25th International Conference on Artificial Intelligence and Statistics , pages =
-
[139]
and Hu, Timothy Y
Guan, Ziwei and Ji, Kaiyi and Bucci Jr., Donald J. and Hu, Timothy Y. and Palombo, Joseph and Liston, Michael and Liang, Yingbin , year=. Robust Stochastic Bandit Algorithms under Probabilistic Unbounded Adversarial Attack , booktitle=
-
[140]
Journal of Machine Learning Research , year =
Jason Altschuler and Victor-Emmanuel Brunel and Alan Malek , title =. Journal of Machine Learning Research , year =
-
[141]
Mean-based Best Arm Identification in Stochastic Bandits under Reward Contamination , year =
Mukherjee, Arpan and Tajer, Ali and Chen, Pin-Yu and Das, Payel , booktitle =. Mean-based Best Arm Identification in Stochastic Bandits under Reward Contamination , year =
-
[142]
2023 , eprint=
Adversarial Attacks on Combinatorial Multi-Armed Bandits , author=. 2023 , eprint=
2023
-
[143]
37th Conference on Neural Information Processing Systems , year=
Corruption-Robust Offline Reinforcement Learning with General Function Approximation , author=. 37th Conference on Neural Information Processing Systems , year=
-
[144]
Proceedings of the 25th International Conference on Artificial Intelligence and Statistics , pages =
Corruption-robust Offline Reinforcement Learning , author =. Proceedings of the 25th International Conference on Artificial Intelligence and Statistics , pages =
-
[145]
Proceedings of the 38th International Conference on Machine Learning , pages =
On Reinforcement Learning with Adversarial Corruption and Its Application to Block MDP , author =. Proceedings of the 38th International Conference on Machine Learning , pages =
-
[146]
Proceedings of the 26th International Conference on Artificial Intelligence and Statistics , pages =
Vector Optimization with Stochastic Bandit Feedback , author =. Proceedings of the 26th International Conference on Artificial Intelligence and Statistics , pages =
-
[147]
Advances in Neural Information Processing Systems 36 , year=
Adaptive Algorithms for Relaxed Pareto Set Identification , author=. Advances in Neural Information Processing Systems 36 , year=
-
[148]
No-regret Algorithms for Multi-task
Chowdhury, Sayak Ray and Gopalan, Aditya , booktitle =. No-regret Algorithms for Multi-task
-
[149]
Proceedings of the 36th International Conference on Machine Learning , pages =
Data Poisoning Attacks on Stochastic Bandits , author =. Proceedings of the 36th International Conference on Machine Learning , pages =
-
[150]
The End of Optimism?
Lattimore, Tor and Szepesvari, Csaba , booktitle =. The End of Optimism?
-
[151]
Proceedings of 34th Conference on Learning Theory , pages =
Fine-Grained Gap-Dependent Bounds for Tabular MDPs via Adaptive Multi-Step Bootstrap , author =. Proceedings of 34th Conference on Learning Theory , pages =
-
[152]
Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics , pages =
Adaptive Exploration in Linear Contextual Bandit , author =. Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics , pages =
-
[153]
Exploration in Structured Reinforcement Learning , year =
Ok, Jungseul and Proutiere, Alexandre and Tranos, Damianos , booktitle =. Exploration in Structured Reinforcement Learning , year =
-
[154]
Non-Asymptotic Gap-Dependent Regret Bounds for Tabular MDPs , year =
Simchowitz, Max and Jamieson, Kevin G , booktitle =. Non-Asymptotic Gap-Dependent Regret Bounds for Tabular MDPs , year =
-
[155]
2016 , booktitle =
Chen, Wei and Hu, Wei and Li, Fu and Li, Jian and Liu, Yu and Lu, Pinyan , title =. 2016 , booktitle =
2016
-
[156]
Interactive Submodular Bandit , year =
Chen, Lin and Krause, Andreas and Karbasi, Amin , booktitle =. Interactive Submodular Bandit , year =
-
[157]
Proceedings of the 36th International Conference on Machine Learning , pages =
Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds , author =. Proceedings of the 36th International Conference on Machine Learning , pages =
-
[158]
Exploration–exploitation Tradeoff Using Variance Estimates in Multi-armed Bandits , journal =
-
[159]
Improved Regret Analysis for Variance-Adaptive Linear Bandits and Horizon-Free Linear Mixture MDPs , year =
Kim, Yeoneung and Yang, Insoon and Jun, Kwang-Sung , booktitle =. Improved Regret Analysis for Variance-Adaptive Linear Bandits and Horizon-Free Linear Mixture MDPs , year =
-
[160]
Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP , year =
Zhang, Zihan and Yang, Jiaqi and Ji, Xiangyang and Du, Simon S , booktitle =. Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP , year =
-
[161]
Variance-Aware Off-Policy Evaluation with Linear Function Approximation , year =
Min, Yifei and Wang, Tianhao and Zhou, Dongruo and Gu, Quanquan , booktitle =. Variance-Aware Off-Policy Evaluation with Linear Function Approximation , year =
-
[162]
Llorens , title =
Xin-Qiang Cai and Pushi Zhang and Li Zhao and Bian Jiang and Masashi Sugiyama and Ashley J. Llorens , title =. Advances in Neural Information Processing Systems 36 , year =
-
[163]
1974 , author =
A two-armed bandit theory of market pricing , journal =. 1974 , author =
1974
-
[164]
2023 , author =
A multi-objective home healthcare delivery model and its solution using a branch-and-price algorithm and a two-stage meta-heuristic algorithm , journal =. 2023 , author =
2023
-
[165]
International Conference on Agents and Artificial Intelligence , year=
Thompson Sampling in the Adaptive Linear Scalarized Multi Objective Multi Armed Bandit , author=. International Conference on Agents and Artificial Intelligence , year=
-
[166]
Yahyaa, Saba and M
Q. Yahyaa, Saba and M. Drugan, Madalina and Manderick, Bernard , title =. 2014 , pages =
2014
-
[167]
and Zintgraf, Luisa M
Roijers, Diederik M. and Zintgraf, Luisa M. and Nowe, Ann , title =. 2017 , booktitle =
2017
-
[168]
and Zintgraf, Luisa M
Roijers, Diederik M. and Zintgraf, Luisa M. and Libin, Pieter and Reymond, Mathieu and Bargiacchi, Eugenio and Now\'. Interactive Multi-Objective Reinforcement Learning in Multi-Armed Bandits with Gaussian Process Utility Models , year =
-
[169]
Lexicographic Actor-Critic Deep Reinforcement Learning for Urban Autonomous Driving , year=
Zhang, Hengrui and Lin, Youfang and Han, Sheng and Lv, Kai , journal=. Lexicographic Actor-Critic Deep Reinforcement Learning for Urban Autonomous Driving , year=
-
[170]
Nonstochastic Multi-Armed Bandits with Graph-structured Feedback
Noga Alon and Nicolo Cesa-Bianchi and Claudio Gentile and Shie Mannor and Yishay Mansour and Ohad Shamir. Nonstochastic Multi-Armed Bandits with Graph-structured Feedback. SIAM Journal on Computing. 2017
2017
-
[171]
From Bandits to Experts: On the Value of Side-Observations , year =
Mannor, Shie and Shamir, Ohad , booktitle =. From Bandits to Experts: On the Value of Side-Observations , year =
-
[172]
, edition = 2, publisher =
West, Douglas B. , edition = 2, publisher =. Introduction to Graph Theory , year =
-
[173]
Leveraging Side Observations in Stochastic Bandits , year =
Caron, St\'. Leveraging Side Observations in Stochastic Bandits , year =. Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence , pages =
-
[174]
, title =
Buccapatnam, Swapna and Eryilmaz, Atilla and Shroff, Ness B. , title =. 2014 , booktitle =
2014
-
[175]
Shroff , title =
Swapna Buccapatnam and Fang Liu and Atilla Eryilmaz and Ness B. Shroff , title =. Journal of Machine Learning Research , year =
-
[176]
Proceedings of the 33rd International Conference on Machine Learning , pages =
Online Learning with Feedback Graphs Without the Graphs , author =. Proceedings of the 33rd International Conference on Machine Learning , pages =
-
[177]
Thompson Sampling for Stochastic Bandits with Graph Feedback , booktitle=
Tossou, Aristide and Dimitrakakis, Christos and Dubhashi, Devdatt , pages =. Thompson Sampling for Stochastic Bandits with Graph Feedback , booktitle=
-
[178]
Information Directed Sampling for Stochastic Bandits With Graph Feedback , booktitle=
Liu, Fang and Buccapatnam, Swapna and Shroff, Ness , pages =. Information Directed Sampling for Stochastic Bandits With Graph Feedback , booktitle=
-
[179]
Shroff , title =
Fang Liu and Zizhan Zheng and Ness B. Shroff , title =. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence , pages =
-
[180]
Proceedings of the 35th Uncertainty in Artificial Intelligence Conference , pages =
Problem-dependent Regret Bounds for Online Learning with Feedback Graphs , author =. Proceedings of the 35th Uncertainty in Artificial Intelligence Conference , pages =
-
[181]
Proceedings of the 31st International Conference on Algorithmic Learning Theory , pages =
Feedback graph regret bounds for Thompson Sampling and UCB , author =. Proceedings of the 31st International Conference on Algorithmic Learning Theory , pages =
-
[182]
Proceedings of the 39th Conference on Uncertainty in Artificial Intelligence , pages =
Stochastic Graphical Bandits with Heavy-Tailed Rewards , author =. Proceedings of the 39th Conference on Uncertainty in Artificial Intelligence , pages =
-
[183]
Journal of Machine Learning Research , year =
Eyal Even-Dar and Shie Mannor and Yishay Mansour , title =. Journal of Machine Learning Research , year =
-
[184]
Thompson , journal =
William R. Thompson , journal =. On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples , volume =
-
[185]
Efficient learning by implicit exploration in bandit problems with side observations , year =
Koc\'. Efficient learning by implicit exploration in bandit problems with side observations , year =. Advances in Neural Information Processing Systems 27 , pages =
-
[186]
Proceedings of the 28th Conference on Learning Theory , pages =
Online Learning with Feedback Graphs: Beyond Bandits , author =. Proceedings of the 28th Conference on Learning Theory , pages =
-
[187]
Proceedings of 33rd Conference on Learning Theory , pages =
A Closer Look at Small-loss Bounds for Bandits with Graph Feedback , author =. Proceedings of 33rd Conference on Learning Theory , pages =
-
[188]
and Mohri, Mehryar , title =
Arora, Raman and Marinov, Teodor V. and Mohri, Mehryar , title =. 2019 , booktitle =
2019
-
[189]
Proceedings of the 31st Conference On Learning Theory , pages =
Small-loss bounds for online learning with partial information , author =. Proceedings of the 31st Conference On Learning Theory , pages =
-
[190]
Proceedings of the 35th AAAI Conference on Artificial Intelligence , year =
Shiyin Lu and Yao Hu and Lijun Zhang , title =. Proceedings of the 35th AAAI Conference on Artificial Intelligence , year =
-
[191]
2000 , journal =
HERD BEHAVIOR AND AGGREGATE FLUCTUATIONS IN FINANCIAL MARKETS , author =. 2000 , journal =
2000
-
[192]
Nearly Optimal Best-of-Both-Worlds Algorithms for Online Learning with Feedback Graphs , year =
Ito, Shinji and Tsuchiya, Taira and Honda, Junya , booktitle =. Nearly Optimal Best-of-Both-Worlds Algorithms for Online Learning with Feedback Graphs , year =
-
[193]
Towards Best-of-All-Worlds Online Learning with Feedback Graphs , year =
Erez, Liad and Koren, Tomer , booktitle =. Towards Best-of-All-Worlds Online Learning with Feedback Graphs , year =
-
[194]
1999 , Address =
Nonlinear Multiobjective Optimization , Author =. 1999 , Address =
1999
-
[195]
2021 , author =
On the estimation of pareto front and dimensional similarity in many-objective evolutionary algorithm , journal =. 2021 , author =
2021
-
[196]
2018 , author =
Approximating the irregularly shaped Pareto front of multi-objective reservoir flood control operation problem , journal =. 2018 , author =
2018
-
[197]
An Evolutionary Many-Objective Optimization Algorithm Using Reference-Point-Based Nondominated Sorting Approach, Part I: Solving Problems With Box Constraints , year=
Deb, Kalyanmoy and Jain, Himanshu , journal=. An Evolutionary Many-Objective Optimization Algorithm Using Reference-Point-Based Nondominated Sorting Approach, Part I: Solving Problems With Box Constraints , year=
-
[198]
An Evolutionary Many-Objective Optimization Algorithm Based on Dominance and Decomposition , year=
Li, Ke and Deb, Kalyanmoy and Zhang, Qingfu and Kwong, Sam , journal=. An Evolutionary Many-Objective Optimization Algorithm Based on Dominance and Decomposition , year=
-
[199]
and Santana-Quintero, Luis V
Hernández-Díaz, Alfredo G. and Santana-Quintero, Luis V. and Coello Coello, Carlos A. and Molina, Julián , title = ". Evolutionary Computation , volume =
-
[200]
Multi-objective Bandits: Optimizing the Generalized
R. Multi-objective Bandits: Optimizing the Generalized. Proceedings of the 34th International Conference on Machine Learning , pages =
-
[201]
Annals of Operations Research , year=2022, volume=
Maciej Nowak and Tadeusz Trzaskalik , title=. Annals of Operations Research , year=2022, volume=
2022
-
[202]
Keeney , journal =
Ralph L. Keeney , journal =. Common Mistakes in Making Value Trade-Offs , volume =
-
[203]
A. D. Athanassopoulos and V. V. Podinovski , journal =. Dominance and Potential Optimality in Multiple Criteria Decision Analysis with Imprecise Information , volume =
-
[204]
Podinovski , number =
Victor V. Podinovski , number =. A DSS for multiple criteria decision analysis with imprecisely specified trade-offs , journal =
-
[205]
Ruiz and Francisco Ruiz and Kaisa Miettinen and Laura Delgado-Antequera and Vesa Ojalehto , title=
Ana B. Ruiz and Francisco Ruiz and Kaisa Miettinen and Laura Delgado-Antequera and Vesa Ojalehto , title=. Journal of Global Optimization , year=2019, volume=
2019
-
[206]
1999 , author =
Searching for psychologically stable solutions of multiple criteria decision problems , journal =. 1999 , author =
1999
-
[207]
2000 , author =
Using trade-off information in decision-making algorithms , journal =. 2000 , author =
2000
-
[208]
TopRank: A practical algorithm for online stochastic ranking , year =
Lattimore, Tor and Kveton, Branislav and Li, Shuai and Szepesvari, Csaba , booktitle =. TopRank: A practical algorithm for online stochastic ranking , year =
-
[209]
IEEE/ACM Transactions on Networking , pages =
Gai, Yi and Krishnamachari, Bhaskar and Jain, Rahul , title =. IEEE/ACM Transactions on Networking , pages =. 2012 , volume =
2012
-
[210]
Proceedings of the 30th International Conference on Machine Learning , pages =
Combinatorial Multi-Armed Bandit: General Framework and Applications , author =. Proceedings of the 30th International Conference on Machine Learning , pages =
-
[211]
Journal of Machine Learning Research , year =
Wei Chen and Yajun Wang and Yang Yuan and Qinshi Wang , title =. Journal of Machine Learning Research , year =
-
[212]
Proceedings of the 35th International Conference on Machine Learning , pages =
Thompson Sampling for Combinatorial Semi-Bandits , author =. Proceedings of the 35th International Conference on Machine Learning , pages =
-
[213]
Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence , pages =
Combinatorial semi-bandit in the non-stationary environment , author =. Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence , pages =
-
[214]
Hybrid Regret Bounds for Combinatorial Semi-Bandits and Adversarial Linear Bandits , year =
Ito, Shinji , booktitle =. Hybrid Regret Bounds for Combinatorial Semi-Bandits and Adversarial Linear Bandits , year =
-
[215]
Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =
Further Adaptive Best-of-Both-Worlds Algorithm for Combinatorial Semi-Bandits , author =. Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =
-
[216]
Proceedings of the 18th International Conference on Artificial Intelligence and Statistics , pages =
Tight Regret Bounds for Stochastic Combinatorial Semi-Bandits , author =. Proceedings of the 18th International Conference on Artificial Intelligence and Statistics , pages =
-
[217]
Proceedings of the 40th International Conference on Machine Learning , pages =
Probably Anytime-Safe Stochastic Combinatorial Semi-Bandits , author =. Proceedings of the 40th International Conference on Machine Learning , pages =
-
[218]
Neural Computation , volume =
Kuroki, Yuko and Xu, Liyuan and Miyauchi, Atsushi and Honda, Junya and Sugiyama, Masashi , title =. Neural Computation , volume =
-
[219]
Combinatorial Pure Exploration with Full-Bandit or Partial Linear Feedback , journal=
Du, Yihan and Kuroki, Yuko and Chen, Wei , year=. Combinatorial Pure Exploration with Full-Bandit or Partial Linear Feedback , journal=
-
[220]
Proceedings of the 38th Conference on Uncertainty in Artificial Intelligence , pages =
An explore-then-commit algorithm for submodular maximization under full-bandit feedback , author =. Proceedings of the 38th Conference on Uncertainty in Artificial Intelligence , pages =
-
[221]
Discover Artificial Intelligence , volume=
Multi-objective optimization for autonomous driving strategy based on Deep Q Network , author=. Discover Artificial Intelligence , volume=
-
[222]
Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages =
Learning Multi-Objective Rewards and User Utility Function in Contextual Bandits for Personalized Ranking , author =. Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages =
-
[223]
Proceedings of The 27th International Conference on Artificial Intelligence and Statistics , pages =
Sequential learning of the. Proceedings of The 27th International Conference on Artificial Intelligence and Statistics , pages =
-
[224]
Near-Optimal Regret Bounds for Contextual Combinatorial Semi-Bandits with Linear Payoff Functions , booktitle =
Kei Takemura and Shinji Ito and Daisuke Hatano and Hanna Sumita and Takuro Fukunaga and Naonori Kakimura and Ken. Near-Optimal Regret Bounds for Contextual Combinatorial Semi-Bandits with Linear Payoff Functions , booktitle =
-
[225]
Combinatorial Bandits with Linear Constraints: Beyond Knapsacks and Fairness , year =
Liu, Qingsong and Xu, Weihang and Wang, Siwei and Fang, Zhixuan , booktitle =. Combinatorial Bandits with Linear Constraints: Beyond Knapsacks and Fairness , year =
-
[226]
2022 , pages =
Hu, Xinyan and Ngo, Dung Daniel and Slivkins, Aleksandrs and Wu, Zhiwei Steven , title =. 2022 , pages =
2022
-
[227]
2016 , eprint=
Influence Maximization with Bandits , author=. 2016 , eprint=
2016
-
[228]
Proceedings of the 40th International Conference on Machine Learning , pages =
A Framework for Adapting Offline Algorithms to Solve Combinatorial Multi-Armed Bandit Problems with Bandit Feedback , author =. Proceedings of the 40th International Conference on Machine Learning , pages =
-
[229]
Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =
Randomized Greedy Learning for Non-monotone Stochastic Submodular Maximization Under Full-bandit Feedback , author =. Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =
-
[230]
2024 , booktitle =
Fourati, Fares and Alouini, Mohamed-Slim and Aggarwal, Vaneet , title =. 2024 , booktitle =
2024
-
[231]
Combinatorial Stochastic-Greedy Bandit , journal=
Fourati, Fares and Quinn, Christopher John and Alouini, Mohamed-Slim and Aggarwal, Vaneet , year=. Combinatorial Stochastic-Greedy Bandit , journal=
-
[232]
When Combinatorial Thompson Sampling meets Approximation Regret , year =
Perrault, Pierre , booktitle =. When Combinatorial Thompson Sampling meets Approximation Regret , year =
-
[233]
A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits , year =
Lee, Jungyhun and Yun, Se-Young and Jun, Kwang-Sung , booktitle =. A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits , year =
-
[234]
Neural Combinatorial Clustered Bandits for Recommendation Systems , journal=
Atalar, Baran and Joe-Wong, Carlee , year=. Neural Combinatorial Clustered Bandits for Recommendation Systems , journal=
-
[235]
Nemhauser, G. L. and Wolsey, L. A. and Fisher, M. L. , title =. Mathematical Programming , pages =. 1978 , volume =
1978
-
[236]
Journal of Machine Learning Research , year =
Jianqing Fan and Bai Jiang and Qiang Sun , title =. Journal of Machine Learning Research , year =
-
[237]
2013 , publisher =
Bach, Francis , title =. 2013 , publisher =
2013
-
[238]
Proceedings of the 31st International Conference on Algorithmic Learning Theory , pages =
Top- k Combinatorial Bandits with Full-Bandit Feedback , author =. Proceedings of the 31st International Conference on Algorithmic Learning Theory , pages =
-
[239]
Sadegh and Proutiere, Alexandre and Lelarge, Marc , title =
Combes, Richard and Talebi, M. Sadegh and Proutiere, Alexandre and Lelarge, Marc , title =. 2015 , booktitle =
2015
-
[240]
Combinatorial bandits , year =
Cesa-Bianchi, Nicol\`. Combinatorial bandits , year =. Journal of Computer and System Sciences , pages =
-
[241]
IEEE Transactions on Information Theory , pages =
Fang, Guanhua and Li, Ping and Samorodnitsky, Gennady , title =. IEEE Transactions on Information Theory , pages =. 2025 , volume =
2025
-
[242]
IEEE Transactions on Information Theory , pages =
Khaleghi, Azadeh , title =. IEEE Transactions on Information Theory , pages =. 2025 , volume =
2025
-
[243]
Karthik, P. N. and Reddy, Kota Srinivas and Tan, Vincent Y. F. , title =. IEEE Transactions on Information Theory , pages =. 2023 , volume =
2023
-
[244]
Multi-Armed Bandits With Correlated Arms , year =
Gupta, Samarth and Chaudhari, Shreyas and Joshi, Gauri and Ya. Multi-Armed Bandits With Correlated Arms , year =. IEEE Transactions on Information Theory , pages =
-
[245]
The Thirteenth International Conference on Learning Representations , year=
On Speeding Up Language Model Evaluation , author=. The Thirteenth International Conference on Learning Representations , year=
-
[246]
2024 , eprint=
Sample-Efficient Alignment for LLMs , author=. 2024 , eprint=
2024
-
[247]
Bandit-Based Prompt Design Strategy Selection Improves Prompt Optimizers
Ashizawa, Rin and Hirose, Yoichi and Yoshinari, Nozomu and Uchida, Kento and Shirakawa, Shinichi. Bandit-Based Prompt Design Strategy Selection Improves Prompt Optimizers. Findings of the Association for Computational Linguistics: ACL 2025. 2025
2025
-
[248]
2025 , eprint=
Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution , author=. 2025 , eprint=
2025
-
[249]
Efficient Prompt Optimization Through the Lens of Best Arm Identification , year =
Shi, Chengshuai and Yang, Kun and Chen, Zihan and Li, Jundong and Yang, Jing and Shen, Cong , booktitle =. Efficient Prompt Optimization Through the Lens of Best Arm Identification , year =
-
[250]
ACM Computing Surveys , numpages =
Liu, Pengfei and Yuan, Weizhe and Fu, Jinlan and Jiang, Zhengbao and Hayashi, Hiroaki and Neubig, Graham , title =. ACM Computing Surveys , numpages =. 2023 , volume =
2023
-
[251]
Advances in Neural Information Processing Systems 35 , pages =
Wei, Jason and Wang, Xuezhi and Schuurmans, Dale and Bosma, Maarten and ichter, brian and Xia, Fei and Chi, Ed and Le, Quoc V and Zhou, Denny , title=. Advances in Neural Information Processing Systems 35 , pages =
-
[252]
A Thorough Examination of Decoding Methods in the Era of LLM s
Shi, Chufan and Yang, Haoran and Cai, Deng and Zhang, Zhisong and Wang, Yifan and Yang, Yujiu and Lam, Wai. A Thorough Examination of Decoding Methods in the Era of LLM s. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024
2024
-
[253]
2024 , eprint=
GPT-4 Technical Report , author=. 2024 , eprint=
2024
-
[254]
2021 , booktitle =
Reynolds, Laria and McDonell, Kyle , title =. 2021 , booktitle =
2021
-
[255]
and Yang, Qiang and Xie, Xing , title =
Chang, Yupeng and Wang, Xu and Wang, Jindong and Wu, Yuan and Yang, Linyi and Zhu, Kaijie and Chen, Hao and Yi, Xiaoyuan and Wang, Cunxiang and Wang, Yidong and Ye, Wei and Zhang, Yue and Chang, Yi and Yu, Philip S. and Yang, Qiang and Xie, Xing , title =. 2024 , volume =
2024
-
[256]
From Generation to Judgment: Opportunities and Challenges of LLM -as-a-judge
Li, Dawei and Jiang, Bohan and Huang, Liangjie and Beigi, Alimohammad and Zhao, Chengshuai and Tan, Zhen and Bhattacharjee, Amrita and Jiang, Yuxuan and Chen, Canyu and Wu, Tianhao and Shu, Kai and Cheng, Lu and Liu, Huan. From Generation to Judgment: Opportunities and Challen...
2025
-
[257]
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena , year =
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric and Zhang, Hao and Gonzalez, Joseph E and Stoica, Ion , booktitle =. Judging LLM-as-a-Judge with MT-Bench and C...
-
[258]
Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition
Feng, Kehua and Ding, Keyan and Hongzhi, Tan and Ma, Kede and Wang, Zhihua and Guo, Shuangquan and Yuzhou, Cheng and Sun, Ge and Zheng, Guozhou and Zhang, Qiang and Chen, Huajun. Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition. Pr...
2025
-
[259]
2023 , eprint=
Evaluating Large Language Models: A Comprehensive Survey , author=. 2023 , eprint=
2023
-
[260]
Data Intelligence , volume =
Li, Linhan and Zhang, Huaping and Li, Chunjin and You, Haowen and Cui, Wenyao , title =. Data Intelligence , volume =
-
[261]
Proceedings of The 27th Conference on Learning Theory , pages =
lil' UCB : An Optimal Exploration Algorithm for Multi-Armed Bandits , author =. Proceedings of The 27th Conference on Learning Theory , pages =
-
[262]
2021 , eprint=
Training Verifiers to Solve Math Word Problems , author=. 2021 , eprint=
2021
-
[263]
Proceedings of the 34th AAAI Conference on Artificial Intelligence , year =
Yonatan Bisk and Rowan Zellers and Ronan Le Bras and Jianfeng Gao and Yejin Choi , title =. Proceedings of the 34th AAAI Conference on Artificial Intelligence , year =
-
[264]
Proceedings of The 33rd International Conference on Machine Learning , pages =
Anytime optimal algorithms in stochastic multi-armed bandits , author =. Proceedings of The 33rd International Conference on Machine Learning , pages =
-
[265]
Proceedings of the 36th International Conference on Machine Learning , pages =
Bilinear Bandits with Low-rank Structure , author =. Proceedings of the 36th International Conference on Machine Learning , pages =
-
[266]
Proceedings of the 38th International Conference on Machine Learning , pages =
Improved Regret Bounds of Bilinear Bandits using Action Space Analysis , author =. Proceedings of the 38th International Conference on Machine Learning , pages =
-
[267]
Proceedings of The 24th International Conference on Artificial Intelligence and Statistics , pages =
Low-Rank Generalized Linear Bandit Problems , author =. Proceedings of The 24th International Conference on Artificial Intelligence and Statistics , pages =
-
[268]
Efficient Frameworks for Generalized Low-Rank Matrix Bandit Problems , year =
Kang, Yue and Hsieh, Cho-Jui and Lee, Thomas Chun Man , booktitle =. Efficient Frameworks for Generalized Low-Rank Matrix Bandit Problems , year =
-
[269]
2025 , eprint =
Generalized Low-Rank Matrix Contextual Bandits with Graph Information , author=. 2025 , eprint =
2025
-
[270]
2026 , eprint =
Low-Rank Contextual Reinforcement Learning from Heterogeneous Human Feedback , author=. 2026 , eprint =
2026
-
[271]
Proceedings of the Fortieth Conference on Uncertainty in Artificial Intelligence , pages =
Low-rank Matrix Bandits with Heavy-tailed Rewards , author =. Proceedings of the Fortieth Conference on Uncertainty in Artificial Intelligence , pages =
-
[272]
Proceedings of the 41st International Conference on Machine Learning , pages =
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits , author =. Proceedings of the 41st International Conference on Machine Learning , pages =
-
[273]
Multi-task Representation Learning for Pure Exploration in Bilinear Bandits , year =
Mukherjee, Subhojyoti and Xie, Qiaomin and Hanna, Josiah and Nowak, Robert , booktitle =. Multi-task Representation Learning for Pure Exploration in Bilinear Bandits , year =
-
[274]
Advances in Neural Information Processing Systems 38 , year=
Thompson Sampling for Multi-Objective Linear Contextual Bandit , author=. Advances in Neural Information Processing Systems 38 , year=
-
[275]
Optimal Scalarizations for Sublinear Hypervolume Regret , year =
Zhang, Qiuyi (Richard) , booktitle =. Optimal Scalarizations for Sublinear Hypervolume Regret , year =
-
[276]
Prabhu , title =
Alperen Tercan, Vinayak S. Prabhu , title =. Proceedings of the 27th European Conference on Artificial Intelligence , pages =
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.