REVIEW 3 major objections 4 minor 98 references
Outlier-Robust Training of Machine Learning Models
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A concave per-sample loss kernel provably enlarges where stochastic training converges under arbitrary outliers.
desk verdict The paper's central theorem is unsupported: Lemma 27 drops a negative outlier term, so the claimed convergence-region enlargement over SGD does not follow; the survey and algorithm are still worth a look. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a modified Black-Rangarajan duality, stated as Corollary 4. For a concave kernel with derivative tending to one at zero and to zero at infinity, minimizing the sum of robust losses is equivalent to minimizing a sum of weighted original losses plus an outlier-process term, and the optimal weight is exactly the kernel's derivative evaluated at that sample's loss. This keeps the original problem structure intact, which is what makes it applicable to deep learning rather than only to weighted least squares. The Adaptive Alternation Algorithm alternates gradient descent on the weighted sum with a closed-form weight update, and adapts the kernel scale so that the average weight equals a target value interpreted as the expected inlier fraction. For a truncated kernel, the same update coincides with training on a conformal prediction set of the best-scoring samples.
What would settle it
Construct a linear regression problem with one outlier whose gradient perturbation points opposite the gradient of the outlier-free objective at every iterate, run the Adaptive Alternation Algorithm with a truncated kernel, and check whether the key descent inequality used in the proof holds along the trajectory. A single violation of that inequality, while the outlier remains fixed, would show that the enlarged convergence region does not follow from the stated assumptions.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that robust losses act through their derivative: replacing the stochastic update with a scaled version that multiplies each sample gradient by the robust kernel's derivative does not just heuristically dampen outliers; it provably enlarges the region from which gradient descent converges to the outlier-free optimum. The paper states this as Theorem 19: if the outlier-free loss components are L-smooth and Polyak-Lojasiewicz, the Adaptive Alternation Algorithm converges to an epsilon-neighborhood of the outlier-free optimal value whenever the iterates lie in a region defined by a weighted average of squared outlier gradient perturbations, whereas plain SGD requires the same bound without the squared derivative factor. Since this derivative lies in the unit interval, each outlier gradient contributes less to the variance that forces iterates out of the basin, so the robust region contains the SGD region. The claim is not that robust kernels make the objective globally easier; it is that, at the level of gradient variance and descent bounds, the per-sample weight is the mechanism that separates inlier signal from arbitrary outlier perturbation.
Load-bearing premise
The load-bearing premise is that the average alignment between outlier gradient perturbations and the outlier-free gradient cannot be so negative that it overpowers the descent signal; the paper relies on this to prove its enlarged region of convergence but does not state it among its assumptions.
Editorial extensions
If this is right
- Robust kernels developed for classification can be transplanted into robotics-style robust estimation and vice versa, because the modified duality preserves the original problem structure instead of squaring the loss.
- The Adaptive Alternation Algorithm turns the robust kernel's shape parameter from a hand-tuned hyperparameter into an adaptive variable controlled by one inlier-fraction target.
- The variance bound shows that outlier-contributed gradient variance is scaled by the squared derivative of the kernel, which predicts more stable descent trajectories under zero-mean outliers than plain SGD.
- The convergence theorems apply to arbitrary outlier gradients whose average squared perturbation is bounded, so the analysis covers adversarial outliers rather than only outliers drawn from a clean contamination model.
- For the truncated kernel, the algorithm's weight update is exactly a conformal prediction set over samples, giving a principled interpretation of iterative sample trimming.
Reading between the lines
- A practical robustness monitor could track the ratio of outlier-gradient energy before and after kernel weighting during training; a drop in that ratio would indicate the kernel is actively suppressing outlier influence.
- The same duality suggests a soft version of data cleaning: instead of hard-trimming samples, one could propagate the per-sample weights to downstream tasks as importance weights, an extension not pursued in the paper.
- Combining per-sample kernel weighting with norm-based methods such as gradient clipping might compound outlier suppression, since the former scales each gradient while the latter bounds the global step.
- If the inlier-fraction target is chosen by validation, the conformal-set interpretation suggests it could be calibrated to control the error rate of inlier selection, a directly testable extension of the parameter update rule.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a unified view of robust losses from M-estimation and risk minimization through a modified Black-Rangarajan duality, introduces a robust loss kernel, and presents the Adaptive Alternation Algorithm (AAA) for training models with outliers. The authors claim convergence of AAA to an epsilon-neighborhood of the outlier-free optimum under L-smoothness and Polyak-Lojasiewicz conditions, with a region of convergence W_AAA1 that is larger than the SGD region W_SGD because the outlier variance bound contains the factor sigma'_c(f_i(w))^2. The paper also reports experiments on linear regression, CIFAR-10 classification, and neural radiance field reconstruction, and releases implementation code.
Significance. The proposed unification of robust losses and the adaptive parameter-update rule are interesting and practically motivated, and the experimental results suggest the algorithm works well even at high outlier rates. If the convergence theorem were correct, the claim that robust loss kernels provably enlarge the convergence region under arbitrary outliers would be a significant contribution. However, the central proof contains load-bearing errors, so the theoretical contribution is not established. The algorithmic and experimental parts retain value, and the code release is a strength, but the stated convergence guarantees are currently unsupported.
major comments (3)
- [Appendix G, Lemma 27] The proof of Lemma 27 drops a negative term when moving from (57) to (50). The displayed inequality (57) is E_i[∇f_i(w)^T ∇f_I(w)] ≥ 2μ(f_I(w) − f*_I) − λ(1/nO) Σ_{i∈nO} ||h_i|| ||∇f_I(w)||, but the lemma concludes E_i[∇f_i(w)^T ∇f_I(w)] ≥ 2μ(f_I(w) − f*_I). This conclusion requires the subtracted term to be nonnegative or otherwise controlled, which is not assumed anywhere. A concrete violation is given by f_I(w) = 0.5w^2, one inlier with gradient 2w, and one outlier with h_i = −4w; for λ = 1/2 and μ = 1, E_i[∇f_i(w)^T ∇f_I(w)] = −w^2 whereas 2μ(f_I(w) − f*_I) = w^2, and Assumption 13 holds for |w| ≥ 1/4. Since Lemma 27 is used to derive the SGD descent inequality (65) and Lemma 28 repeats the same argument for Theorems 19 and 22, the descent inequalities (65), (71), and (74) do not follow, and the claimed comparison W_AAA1 versus W_SGD is not established.
- [Section 6.2, Lemma 14 and Appendix F] Lemma 14 states that the variance in the descent direction for SGD and AAA1 is given by (22) and (23), respectively. Both expressions vanish when λ = 0, but for batch-size-one stochastic gradients the variance around the full-batch gradient is generally nonzero even without outliers. The derivation in Appendix F actually produces an upper bound, not an equality, and the bound is obtained by discarding inlier variance terms such as E_i[||∇f_i,I||^2] − ||∇f_I||^2. As stated, the lemma is false; it should be rephrased as an upper bound that additionally retains the inlier variance contribution.
- [Theorem 19 and Lemma 28] The convergence region W_AAA1 in (26) imposes min_i σ_c(f_i(w)) ≥ β > 0, but the proof and Lemma 28 require a positive lower bound on the derivative σ'_c(f_i(w)), denoted φ in Lemma 28. These are distinct conditions: all kernels in Table 1 satisfy σ_c(0) = 0, so the min-σ condition fails for any iterate where an inlier loss is near zero, which is precisely the regime to which the theorem claims convergence. Moreover, Lemma 28 obtains its lower bound by the same invalid drop of a negative outlier-bias term identified in Lemma 27, so even replacing β with φ would not repair the proof without an additional sign or bounded-bias assumption on the outlier gradients.
minor comments (4)
- [Appendix I, proof of Theorem 19] The line 'We assume that w_t, w_{t+1} ∈ W_SGD' should instead refer to W_AAA1; as written, the proof invokes the wrong region.
- [Remark 9] Remark 9 defines the truncated kernel as σ_c(r) = c · max{r/c, 1}, but Table 1 and the subsequent update rule require c · min{r/c, 1}; with the max definition, the derivative does not yield the trimming indicator I{f_i(w_t) ≤ c}.
- [Theorem 22] The set H(w) in (28) uses (1/nO) Σ_{i∈nO} σ'_c(f_i(w')) = ζ, but the parameter update rule (18) imposes the same constraint over all measurements D, not over outliers only; the outlier-only version is inconsistent with the algorithm.
- [Theorem 18 statement] The theorem says SGD 'converges to the optimal value' and then states E[||f_I(w_t) − f*_I||] < ε; the wording should say 'converges to an ε-neighborhood of the optimal value' to match the displayed guarantee.
Circularity Check
Mild self-referential structure: Theorem 19 assumes exactly the zeta-constraint that the parameter-update rule (18) was constructed to enforce, while the main Lemma 27 has a separate non-circular proof gap.
-
self definitional
[Section 5.2, Eq. (18); Section 6.3, Theorem 19, W_AAA1 (Eq. (26)); Remark 20]
"ct = Find c∈[0,1] { 1/|D| ∑_{i∈D} σ′_c(fi(wt)) =ζ } ... The set WAAA1 has two more constraints 1/n ∑_{i=1}^n σ′_c(fi(w)) = ζ and mini σc(fi(w)) ≥ β > 0. The first comes from the step to update the parameter c in the algorithm and is always satisfied."
The algorithm's parameter update (18) solves for c by imposing exactly the equality (1/n)∑σ'_c(f_i(w_t)) = ζ. Theorem 19 then lists this same equality as one of the defining assumptions of the convergence region W_AAA1, and Remark 20 states that this constraint 'comes from the step to update the parameter c in the algorithm and is always satisfied.' Thus the theorem's premise is not an independent condition on the learning problem; it is manufactured by the algorithm's own update rule. This is a mild self-referential structure rather than a full fit-to-data circularity, because the theorem still requires the separate outlier-variance bound (27) to deliver convergence, but the 'increased region' claim is partially baked into the design of the parameter update.
full rationale
The core derivation chain—modified Black-Rangarajan duality (Corollary 4), the robust-kernel definition (Definition 6), the alternation update (17), the variance lemma (Lemma 14), and the descent theorems (18, 19, 22)—is not a case of fitting a parameter to data and then relabeling it as a prediction, nor does it depend on load-bearing self-citations. The enlarged-region comparison W_AAA1 versus W_SGD follows mathematically from σ'_c ≤ 1 and the construction of W_AAA1, so it is not circular in the input-output sense. The mild self-referential loop is the zeta-constraint: Eq. (18) defines c_t by the same equality that Theorem 19 assumes, so part of the theorem's hypothesis is guaranteed by the algorithm's design. Separately, Appendix G's Lemma 27 displays a lower bound containing a negative outlier-bias term, then drops that term without a sign condition; this is a genuine proof gap affecting Theorems 18, 19, and 22, but it is a correctness problem rather than a circularity. Overall the paper is not substantially circular, though it does contain a noticeable assumption-construction loop and a serious missing-support step in the main convergence proof.
Assumptions & free parameters
free parameters (1)
- zeta (inlier-fraction target) =
not reported in Section 7
assumptions (4)
- domain assumption Outlier gradients satisfy grad f_i(w) = grad f_i,I(w) + h_i(o_i,w), with h_i unknown; inlier gradient equals the outlier-free gradient.
- domain assumption Low signal-to-outlier ratio: for every outlier i, ||h_i|| >= 1 and ||h_i|| >= ||grad f_i,I||.
- ad hoc to paper The outlier-bias inner product term lambda (1/nO) sum_{i in nO} h_i^T grad f_I is assumed nonnegative or negligible in Lemma 27.
- domain assumption The robust kernel is strictly concave with invertible derivative, and min_i sigma_c(f_i(w)) >= beta > 0 throughout training.
Cite this review
Pith. "Pith review of Outlier-Robust Training of Machine Learning Models." pith.science (2026). https://pith.science/paper/C4ECWWTD
@misc{pith2026250100265,
author = {Pith},
title = {Pith review of: Outlier-Robust Training of Machine Learning Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/C4ECWWTD}},
note = {Machine review of arXiv:2501.00265}
}
abstract
Robust training of machine learning models in the presence of outliers has garnered attention across various domains. The use of robust losses is a popular approach and is known to mitigate the impact of outliers. We bring to light two literatures that have diverged in their ways of designing robust losses: one using M-estimation, which is popular in robotics and computer vision, and another using a risk-minimization framework, which is popular in deep learning. We first show that a simple modification of the Black-Rangarajan duality provides a unifying view. The modified duality brings out a definition of a robust loss kernel $\sigma$ that is satisfied by robust losses in both the literatures. Secondly, using the modified duality, we propose an Adaptive Alternation Algorithm (AAA) for training machine learning models with outliers. The algorithm iteratively trains the model by using a weighted version of the non-robust loss, while updating the weights at each iteration. The algorithm is augmented with a novel parameter update rule by interpreting the weights as inlier probabilities, and obviates the need for complex parameter tuning. Thirdly, we investigate convergence of the adaptive alternation algorithm to outlier-free optima. Considering arbitrary outliers (i.e., with no distributional assumption on the outliers), we show that the use of robust loss kernels {\sigma} increases the region of convergence. We experimentally show the efficacy of our algorithm on regression, classification, and neural scene reconstruction problems. We release our implementation code: https://github.com/MIT-SPARK/ORT.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Aftab and R
K. Aftab and R. Hartley. Convergence of Iteratively Re-weighted Least Squares to Robust M-Estimators . In IEEE Winter Conference on Applications of Computer Vision , pp.\ 480--487, Jan. 2015
2015
-
[2]
Aftab, R
K. Aftab, R. Hartley, and J. Trumpf. Generalized Weiszfeld Algorithms for Lq Optimization . IEEE Trans. Pattern Anal. Machine Intell. , 37 0 (4): 0 728--745, Apr. 2015
2015
-
[3]
Algan and I
G. Algan and I. Ulusoy. Image classification with deep learning in the presence of noisy labels: A survey. Knowledge-Based Systems, 215: 0 106771, Mar. 2021
2021
-
[4]
E. Amid, M. K. K. Warmuth, R. Anil, and T. Koren. Robust Bi-Tempered Logistic Loss Based on Bregman Divergences . In Advances in Neural Information Processing Systems (NIPS), volume 32, Dec. 2019
2019
-
[5]
Outlier-Robust Estimation: Hardness, Minimally Tuned Algorithms, and Applications
P. Antonante, V. Tzoumas, H. Yang, and L. Carlone. Outlier-robust estimation: Hardness, minimally tuned algorithms, and applications. IEEE Trans. Robotics , 38 0 (1): 0 281--301, 2021. https://arxiv.org/pdf/2007.15109.pdf
work page Pith review arXiv 2021
-
[6]
Armeni, O
I. Armeni, O. Sener, A. Zamir, H. Jiang, I. Brilakis, M. Fischer, and S. Savarese. 3d semantic parsing of large-scale indoor spaces. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp.\ 1534--1543, 2016
2016
-
[7]
Awasthi, M
P. Awasthi, M. F. Balcan, and P. M. Long. The power of localization for efficiently learning linear separators with noise. J. ACM, 63 0 (6), 2017
2017
-
[8]
J. T. Barron. A general and adaptive robust loss function. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 4331--4339, 2019
2019
Show all 98 references
-
[9]
Bhatia, P
K. Bhatia, P. Jain, and P. Kar. Robust regression via hard thresholding. In Advances in Neural Information Processing Systems (NIPS), pp.\ 721--729, 2015
2015
-
[10]
Bhatia, P
K. Bhatia, P. Jain, P. Kamalaruban, and P. Kar. Consistent robust regression. In Advances in Neural Information Processing Systems (NIPS), volume 30. Curran Associates, Inc., 2017
2017
-
[11]
M. J. Black and A. Rangarajan. On the unification of line processes, outlier rejection, and robust statistics with applications in early vision. Intl. J. of Computer Vision, 19 0 (1): 0 57--91, 1996
1996
-
[12]
Blake and A
A. Blake and A. Zisserman. Visual reconstruction. MIT Press, 1987
1987
-
[13]
G. E. P. Box and D. R. Cox. An Analysis of Transformations . Journal of the Royal Statistical Society: Series B (Methodological), 26 0 (2): 0 211--243, 1964
1964
-
[14]
Brimberg and R
J. Brimberg and R. F. Love. Global Convergence of a Generalized Iterative Procedure for the Minisum Location Problem with lp Distances . Operations Research, 41 0 (6): 0 1153--1163, 1993
1993
-
[15]
L. Carlone. Estimation contracts for outlier-robust geometric perception. Foundations and Trends (FnT) in Robotics, arXiv preprint: 2208.10521, 2023. https://arxiv.org/pdf/2208.10521.pdf
2023 arXiv
-
[16]
C. Chai, L. Cao, G. Li, J. Li, Y. Luo, and S. Madden. Human-in-the-loop Outlier Detection . In ACM SIGMOD International Conference on Management of Data , pp.\ 19--33, May 2020
2020
-
[17]
Chang, A
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang. Matterport3d: Learning from rgb-d data in indoor environments. International Conference on 3D Vision (3DV), 2017
2017
-
[18]
Charikar, J
M. Charikar, J. Steinhardt, and G. Valiant. Learning from untrusted data. In Proceedings of the 49th Annual ACM SIGACT Symposium on Theory of Computing, STOC 2017, pp.\ 47--60, 2017
2017
-
[19]
Chebrolu, T
N. Chebrolu, T. L \"a be, O. Vysotska, J. Behley, and C. Stachniss. Adaptive robust kernels for non-linear least squares problems. arXiv preprint arXiv:2004.14938, 2020
2004 arXiv
-
[20]
Chebrolu, T
N. Chebrolu, T. Läbe, O. Vysotska, J. Behley, and C. Stachniss. Adaptive robust kernels for non-linear least squares problems. IEEE Robotics and Automation Letters , 6 0 (2): 0 2240--2247, 2021
2021
-
[21]
Y. Chen, C. Caramanis, and S. Mannor. Robust sparse regression under adversarial corruption. In Intl. Conf. on Machine Learning (ICML), volume 28, pp.\ 774--782, 2013
2013
-
[22]
Chhabra, B
A. Chhabra, B. Li, J. Chen, P. Mohapatra, and H. Liu. Outlier Gradient Analysis : Efficiently Identifying Detrimental Training Samples for Deep Learning Models . arXiv: 2405.03869, Oct. 2024
2024
-
[23]
Demidovich, G
Y. Demidovich, G. Malinovsky, I. Sokolov, and P. Richtarik. A Guide Through the Zoo of Biased SGD . Advances in Neural Information Processing Systems (NIPS), 36: 0 23158--23171, Dec. 2023
2023
-
[24]
X. Deng, Y. Xiang, A. Mousavian, C. Eppner, T. Bretl, and D. Fox. Self-supervised 6D Object Pose Estimation for Robot Manipulation . In IEEE Intl. Conf. on Robotics and Automation (ICRA), pp.\ 3665--3671, May 2020
2020
-
[25]
Dereich and A
S. Dereich and A. Jentzen. Convergence rates for the Adam optimizer. arXiv: 2407.21078, Jul. 2024
2024 arXiv
-
[26]
Diakonikolas, G
I. Diakonikolas, G. Kamath, D. Kane, J. Li, A. Moitra, and A. Stewart. Robust estimators in high dimensions without the computational intractability. In IEEE 57th Annual Symposium on Foundations of Computer Science, pp.\ 655--664. IEEE, 2016
2016
-
[27]
Diakonikolas, G
I. Diakonikolas, G. Kamath, D. M. Kane, J. Li, A. Moitra, and A. Stewart. Robustly learning a gaussian: Getting optimal error, efficiently. In Proceedings of the Twenty-Ninth Annual ACM-SIAM Symposium on Discrete Algorithms, SODA '18, pp.\ 2683–2702, 2018 a
2018
-
[28]
Diakonikolas, D
I. Diakonikolas, D. M. Kane, and A. Stewart. Learning geometric concepts with nasty noise. In Proceedings of the 50th Annual ACM SIGACT Symposium on Theory of Computing, STOC 2018, pp.\ 1061--1073, 2018 b
2018
-
[29]
Diakonikolas, G
I. Diakonikolas, G. Kamath, D. Kane, J. Li, A. Moitra, and A. Stewart. Robust estimators in high-dimensions without the computational intractability. SIAM Journal on Computing, 48 0 (2): 0 742--864, 2019 a . doi:10.1137/17M1126680
2019 doi
-
[30]
Diakonikolas, G
I. Diakonikolas, G. Kamath, D. Kane, J. Li, J. Steinhardt, and A. Stewart. Sever: A robust meta-algorithm for stochastic optimization. In K. Chaudhuri and R. Salakhutdinov (eds.), Intl. Conf. on Machine Learning (ICML), volume 97 of Proceedings of Machine Learning Research, pp...
2019
-
[31]
Diakonikolas, W
I. Diakonikolas, W. Kong, and A. Stewart. Efficient algorithms and lower bounds for robust linear regression. In Proceedings of the Thirtieth Annual ACM-SIAM Symposium on Discrete Algorithms, SODA '19, pp.\ 2745--2754, 2019 c
2019
-
[32]
Elesedy and M
B. Elesedy and M. Hutter. U-clip: On-average unbiased stochastic gradient clipping. arXiv preprint arXiv:2302.02971, 2023
2023 arXiv
-
[33]
L. Feng, S. Shu, Z. Lin, F. Lv, L. Li, and B. An. Can Cross Entropy Loss Be Robust to Label Noise ? In Intl. Joint Conf. on AI (IJCAI), volume 3, pp.\ 2206--2212, Jul. 2020
2020
-
[34]
Ferrari and Y
D. Ferrari and Y. Yang. Maximum Lq-likelihood estimation. The Annals of Statistics, 38 0 (2): 0 753--783, Apr. 2010
2010
-
[35]
Fischler and R
M. Fischler and R. Bolles. Random sample consensus: a paradigm for model fitting with application to image analysis and automated cartography. Commun. ACM, 24: 0 381--395, 1981
1981
-
[36]
Foret, A
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur. Sharpness-aware Minimization for Efficiently Improving Generalization . In Intl. Conf. on Learning Representations (ICLR), Oct. 2020
2020
-
[37]
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, E. Orgad, R. Entezari, G. Daras, S. M. Pratt, V. Ramanujan, Y. Bitton, K. Marathe, S. Mussmann, R. Vencu, M. Cherti, R. Krishna, P. W. Koh, O. Saukh, A. Ratner, S. S...
2023
-
[38]
Garrigos and R
G. Garrigos and R. M. Gower. Handbook of Convergence Theorems for ( Stochastic ) Gradient Methods . arXiv preprint arXiv:2301.11235, Feb. 2023
2023 arXiv
-
[39]
Ghosh, N
A. Ghosh, N. Manwani, and P. S. Sastry. Making Risk Minimization Tolerant to Label Noise . Neurocomputing, 160: 0 93--107, Jul. 2015
2015
-
[40]
Ghosh, H
A. Ghosh, H. Kumar, and P. S. Sastry. Robust loss functions under label noise for deep neural networks. In Nat. Conf. on Artificial Intelligence (AAAI), pp.\ 1919–1925, Feb. 2017
1919
-
[41]
R. M. Gower, M. Schmidt, F. Bach, and P. Richt\'arik. Variance- Reduced Methods for Machine Learning . Proceedings of the IEEE , 108 0 (11): 0 1968--1983, Nov. 2020
1968
-
[42]
S. Hu, Z. Yang, X. Wang, Y. Ying, and S. Lyu. Outlier Robust Adversarial Training . In Proceedings of the 15th Asian Conference on Machine Learning , pp.\ 454--469. PMLR, Feb. 2024
2024
-
[43]
P. Huber. Robust Statistics. John Wiley & Sons, New York, NY, 1981
1981
-
[44]
Jawaid, R
M. Jawaid, R. Talak, Y. Latif, L. Carlone, and T.-J. Chin. Test-time certifiable self-supervision to bridge the sim2real gap in event-based satellite pose estimation. In IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), Oct. 2024
2024
-
[45]
Karmalkar and E
S. Karmalkar and E. Price. Compressed sensing with adversarial sparse noise via L1 regression. CoRR, abs/1809.08055, 2018. URL http://arxiv.org/abs/1809.08055
2018 arXiv
-
[46]
Karmalkar, A
S. Karmalkar, A. Klivans, and P. Kothari. List-decodable linear regression. In Advances in Neural Information Processing Systems (NIPS), volume 32, 2019
2019
-
[47]
A. R. Klivans, P. M. Long, and R. A. Servedio. Learning halfspaces with malicious noise. In S. Albers, A. Marchetti-Spaccamela, Y. Matias, S. Nikoletseas, and W. Thomas (eds.), Automata, Languages and Programming, pp.\ 609--621, 2009
2009
-
[48]
A. R. Klivans, P. K. Kothari, and R. Meka. Efficient algorithms for outlier-robust regression. CoRR, abs/1803.03241, 2018. URL http://arxiv.org/abs/1803.03241
2018 arXiv
-
[49]
Koloskova, H
A. Koloskova, H. Hendrikx, and S. U. Stich. Revisiting gradient clipping: Stochastic bias and tight convergence guarantees. In Intl. Conf. on Machine Learning (ICML), volume 202, pp.\ 17343--17363, Jul. 2023
2023
-
[50]
P. K. Kothari and J. Steinhardt. Better agnostic clustering via relaxed tensor norms. CoRR, abs/1711.07465, 2017. URL http://arxiv.org/abs/1711.07465
2017 arXiv
-
[51]
P. K. Kothari, J. Steinhardt, and D. Steurer. Robust moment estimation and improved clustering via sum of squares. In Proceedings of the 50th Annual ACM SIGACT Symposium on Theory of Computing, STOC 2018, pp.\ 1035--1046, 2018
2018
-
[52]
K. A. Lai, A. B. Rao, and S. Vempala. Agnostic estimation of mean and covariance. In 2016 IEEE 57th Annual Symposium on Foundations of Computer Science (FOCS), pp.\ 665--674. IEEE Computer Society, 2016. doi:10.1109/FOCS.2016.76. URL https://doi.ieeecomputersociety.org/10.1109...
2016 doi
-
[53]
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein. Visualizing the loss landscape of neural nets. In Advances in Neural Information Processing Systems (NIPS), volume 31, Dec. 2018. URL https://proceedings.neurips.cc/paper_files/paper/2018/file/a41b3bb3e6b050b6c9067c67f663b9...
2018
-
[54]
J. Li, R. Socher, and S. C. Hoi. Dividemix: Learning with noisy labels as semi-supervised learning. Intl. Conf. on Learning Representations (ICLR), 2020
2020
-
[55]
Liu and H
Y. Liu and H. Guo. Peer Loss Functions : Learning from Noisy Labels without Knowing Noise Rates . In Intl. Conf. on Machine Learning (ICML), pp.\ 6226--6236, Nov. 2020
2020
-
[56]
Z. Lu, Y. Zhang, K. Doherty, O. Severinsen, E. Yang, and J. Leonard. SLAM- supported self-training for 6d object pose estimation. In IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), pp.\ 2833--2840, Oct. 2022
2022
-
[57]
Lyu and I
Y. Lyu and I. W. Tsang. Curriculum loss: Robust learning and generalization against label corruption. In Intl. Conf. on Learning Representations (ICLR), Apr. 2020
2020
-
[58]
X. Ma, H. Huang, Y. Wang, S. Romano, S. Erfani, and J. Bailey. Normalized Loss Functions for Deep Learning with Noisy Labels . In Intl. Conf. on Machine Learning (ICML), pp.\ 6543--6553, Nov. 2020
2020
-
[59]
V. V. Mai and M. Johansson. Stability and convergence of stochastic gradient clipping: Beyond lipschitz continuity and smoothness. In Intl. Conf. on Machine Learning (ICML), pp.\ 7325--7335. PMLR, 2021
2021
-
[60]
A. K. Menon, A. S. Rawat, S. J. Reddi, and S. Kumar. Can gradient clipping mitigate label noise? In Intl. Conf. on Learning Representations (ICLR), 2020
2020
-
[61]
Merad and S
I. Merad and S. Ga\"iffas. Robust Stochastic Optimization via Gradient Quantile Clipping . Trans. on Machine Learning Research, May 2024
2024
-
[62]
Mildenhall, P
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. arXiv preprint arXiv:2003.08934, 2020
2003 arXiv
-
[63]
M\"uller, A
T. M\"uller, A. Evans, C. Schied, and A. Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph., 41 0 (4): 0 102:1--102:15, July 2022. doi:10.1145/3528223.3530127. URL https://doi.org/10.1145/3528223.3530127
2022
-
[64]
J. Naudts. Deformed exponentials and logarithms in generalized thermostatistics. Physica A: Statistical Mechanics and its Applications, 316 0 (1): 0 323--334, 2002
2002
-
[65]
Nguyen and T
N. Nguyen and T. Tran. Exact recoverability from dense corrupted observations via _ 1 -minimization. IEEE Trans. on Information Theory, 59 0 (4): 0 2017--2035, 2013
2017
-
[66]
L. Peng, C. K\"ummerle, and R. Vidal. On the Convergence of IRLS and Its Variants in Outlier-Robust Estimation . In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp.\ 17808--17818, Jun. 2023
2023
-
[67]
Prasad, A
A. Prasad, A. S. Suggala, S. Balakrishnan, and P. Ravikumar. Robust estimation via robust gradient estimation. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 82, 2020
2020
-
[68]
Raghavendra and M
P. Raghavendra and M. Yau. List decodable learning via sum of squares. In Proceedings of the Thirty-First Annual ACM-SIAM Symposium on Discrete Algorithms, SODA '20, pp.\ 161–180, 2020
2020
-
[69]
Reisizadeh, H
A. Reisizadeh, H. Li, S. Das, and A. Jadbabaie. Variance-reduced Clipping for Non-convex Optimization . arXiv: 2303.00883, Jun. 2023
2023 arXiv
-
[70]
M. Ren, W. Zeng, B. Yang, and R. Urtasun. Learning to Reweight Examples for Robust Deep Learning . In Intl. Conf. on Machine Learning (ICML), pp.\ 4334--4343, Jul. 2018
2018
-
[71]
Sabour, S
S. Sabour, S. Vora, D. Duckworth, I. Krasin, D. J. Fleet, and A. Tagliasacchi. Robustnerf: Ignoring distractors with robust losses. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp.\ 20626--20636, 2023
2023
-
[72]
Schmidt and D
T. Schmidt and D. Fox. Self-directed Lifelong Learning for Robot Vision . In Robotics Research , pp.\ 109--114. Springer International Publishing, 2020
2020
-
[73]
Shafer and V
G. Shafer and V. Vovk. A Tutorial on Conformal Prediction . J. of Machine Learning Research, pp.\ 51, 2008
2008
-
[74]
V. Shah, X. Wu, and S. Sanghavi. Choosing the Sample with Lowest Loss makes SGD Robust . In Twenty Third International Conference on Artificial Intelligence and Statistics , pp.\ 2120--2130, Jun. 2020
2020
-
[75]
Shen and S
Y. Shen and S. Sanghavi. Learning with Bad Training Data via Iterative Trimmed Loss Minimization . In Intl. Conf. on Machine Learning (ICML), pp.\ 5739--5748, May 2019
2019
-
[76]
J. Shi, R. Talak, D. Maggio, and L. Carlone. A correct-and-certify approach to self-supervise object pose estimators via ensemble self-training. In Robotics: Science and Systems (RSS), 2023. https://arxiv.org/pdf/2302.06019.pdf
2023 arXiv
-
[77]
H. Song, M. Kim, D. Park, Y. Shin, and J.-G. Lee. Learning From Noisy Labels With Deep Neural Networks : A Survey . IEEE Trans. Neural Netw. Learn. Syst., 34 0 (11): 0 8135--8153, Nov. 2023
2023
-
[78]
Sukhbaatar, J
S. Sukhbaatar, J. Bruna, M. Paluri, L. Bourdev, and R. Fergus. Training Convolutional Networks with Noisy Labels . In Intl. Conf. on Learning Representations (ICLR), May 2015
2015
-
[79]
Talak, L
R. Talak, L. Peng, and L. Carlone. Certifiable 3D object pose estimation: Foundations, learning models, and self-training. IEEE Trans. Robotics , 39 0 (4): 0 2805--2824, 2023. https://arxiv.org/pdf/2206.11215.pdf
2023 arXiv
-
[80]
Tancik, E
M. Tancik, E. Weber, E. Ng, R. Li, B. Yi, T. Wang, A. Kristoffersen, J. Austin, K. Salahi, A. Ahuja, et al. Nerfstudio: A modular framework for neural radiance field development. In SIGGRAPH, pp.\ 1--12, 2023
2023
-
[81]
K. M. Tavish and T. D. Barfoot. At all costs: A comparison of robust cost functions for camera correspondence outliers. In Conf. Computer and Robot Vision, pp.\ 62--69. IEEE, 2015
2015
-
[82]
Tkachenko, M
M. Tkachenko, M. Malyuk, A. Holmanyuk, and N. Liubimov. Label Studio : Data labeling software, 2020. URL https://github.com/HumanSignal/label-studio
2020
-
[83]
C. Wang, A. Wang, J. Li, A. Yuille, and C. Xie. Benchmarking Robustness in Neural Radiance Fields . In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp.\ 2926--2936, Jun. 2024 a
2024
-
[84]
Y. Wang, X. Ma, Z. Chen, Y. Luo, J. Yi, and J. Bailey. Symmetric Cross Entropy for Robust Learning With Noisy Labels . In Intl. Conf. on Computer Vision (ICCV), pp.\ 322--330, Oct. 2019
2019
-
[85]
Z. Wang, M. Chen, Y. Guo, Z. Li, and Q. Yu. Bridging the domain gap in satellite pose estimation: A self-training approach based on geometrical constraints. IEEE Trans. Aerosp. Electron. Syst., 60 0 (3): 0 2500--2514, 2024 b
2024
-
[86]
Wright and Y
J. Wright and Y. Ma. Dense error correction via ^1 -minimization. IEEE Trans. on Information Theory, 56 0 (7): 0 3540--3560, 2010
2010
-
[87]
Y. Xu, P. Cao, Y. Kong, and Y. Wang. L\_ DMI : A Novel Information-theoretic Loss Function for Training Deep Nets Robust to Label Noise . In Advances in Neural Information Processing Systems (NIPS), volume 32, Dec. 2019
2019
-
[88]
B. Yang, M. Bai, M. Liang, W. Zeng, and R. Urtasun. Auto4D : Learning to Label 4D Objects from Sequential Point Clouds . arXiv:2101.06586, Mar. 2021
2021 arXiv
-
[89]
Yang and L
H. Yang and L. Carlone. Certifiably optimal outlier-robust geometric perception: Semidefinite relaxations and scalable global optimization. IEEE Trans. Pattern Anal. Machine Intell. , 2022. https://arxiv.org/pdf/2109.03349.pdf
2022 arXiv
-
[90]
H. Yang, P. Antonante, V. Tzoumas, and L. Carlone. Graduated non-convexity for robust spatial perception: From non-minimal solvers to global outlier rejection. IEEE Robotics and Automation Letters ( RA-L ) , 5 0 (2): 0 1127--1134, 2020 a . arXiv preprint:1909.08605 (with suppl...
2020 arXiv
-
[91]
H. Yang, J. Shi, and L. Carlone. TEASER: Fast and Certifiable Point Cloud Registration . IEEE Trans. Robotics , 37 0 (2): 0 314--333, 2020 b . extended arXiv version 2001.07715 https://arxiv.org/pdf/2001.07715.pdf
2020 arXiv
-
[92]
F. Yu, D. Wang, E. Shelhamer, and T. Darrell. Deep layer aggregation. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp.\ 2403--2412, 2018
2018
-
[93]
Zhang, M
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz. mixup: Beyond empirical risk minimization. Intl. Conf. on Learning Representations (ICLR), 2018
2018
-
[94]
Zhang, T
J. Zhang, T. He, S. Sra, and A. Jadbabaie. Why Gradient Clipping Accelerates Training : A Theoretical Justification for Adaptivity . In Intl. Conf. on Learning Representations (ICLR), Mar. 2020 a
2020
-
[95]
Zhang, S
J. Zhang, S. P. Karimireddy, A. Veit, S. Kim, S. Reddi, S. Kumar, and S. Sra. Why are Adaptive Methods Good for Attention Models ? In Advances in Neural Information Processing Systems (NIPS), volume 33, pp.\ 15383--15393, Dec. 2020 b
2020
-
[96]
Zhang and M
Z. Zhang and M. Sabuncu. Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels . In Advances in Neural Information Processing Systems (NIPS), volume 31, Dec. 2018
2018
-
[97]
X. Zhou, X. Liu, D. Zhai, J. Jiang, and X. Ji. Asymmetric Loss Functions for Noise-Tolerant Learning : Theory and Applications . IEEE Trans. Pattern Anal. Machine Intell. , 45 0 (7): 0 8094--8109, Jul. 2023
2023
-
[98]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.