Pith. sign in

REVIEW 3 major objections 4 minor 98 references

Outlier-Robust Training of Machine Learning Models

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A concave per-sample loss kernel provably enlarges where stochastic training converges under arbitrary outliers.

desk verdict The paper's central theorem is unsupported: Lemma 27 drops a negative outlier term, so the claimed convergence-region enlargement over SGD does not follow; the survey and algorithm are still worth a look. read the letter →

arxiv 2501.00265 v1 pith:C4ECWWTD submitted 2024-12-31 cs.LG cs.CV

classification cs.LGcs.CV
keywords outlier-robusttrainingrobustlosskernelBlack-RangarajandualityAdaptiveAlternationAlgorithmconvergenceregionstochasticgradientdescentPolyak-Lojasiewiczneuralradiancefields
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Outliers in training data can derail stochastic gradient methods by injecting large, uncontrolled gradient perturbations. This paper tries to establish that a single object—a concave robust loss kernel applied to each sample's loss—gives a unified account of robust training and a provable cure. It modifies a classical duality so that minimizing the sum of robust losses is equivalent to minimizing a weighted version of the original loss, with the weight acting as an inlier probability, and it builds an Adaptive Alternation Algorithm that reweights samples and adapts its kernel parameter on the fly. Under smoothness and a Polyak-Lojasiewicz condition, the paper claims that this weighting multiplies outlier gradients by a factor no larger than one, so iterates converge to the outlier-free optimum from a strictly larger region than plain SGD. Experiments on linear regression, image classification, and neural scene reconstruction support the claim, including reconstructions with 80 percent of training pixels corrupted.

What carries the argument

The machinery is a modified Black-Rangarajan duality, stated as Corollary 4. For a concave kernel with derivative tending to one at zero and to zero at infinity, minimizing the sum of robust losses is equivalent to minimizing a sum of weighted original losses plus an outlier-process term, and the optimal weight is exactly the kernel's derivative evaluated at that sample's loss. This keeps the original problem structure intact, which is what makes it applicable to deep learning rather than only to weighted least squares. The Adaptive Alternation Algorithm alternates gradient descent on the weighted sum with a closed-form weight update, and adapts the kernel scale so that the average weight equals a target value interpreted as the expected inlier fraction. For a truncated kernel, the same update coincides with training on a conformal prediction set of the best-scoring samples.

What would settle it

Construct a linear regression problem with one outlier whose gradient perturbation points opposite the gradient of the outlier-free objective at every iterate, run the Adaptive Alternation Algorithm with a truncated kernel, and check whether the key descent inequality used in the proof holds along the trajectory. A single violation of that inequality, while the outlier remains fixed, would show that the enlarged convergence region does not follow from the stated assumptions.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that robust losses act through their derivative: replacing the stochastic update with a scaled version that multiplies each sample gradient by the robust kernel's derivative does not just heuristically dampen outliers; it provably enlarges the region from which gradient descent converges to the outlier-free optimum. The paper states this as Theorem 19: if the outlier-free loss components are L-smooth and Polyak-Lojasiewicz, the Adaptive Alternation Algorithm converges to an epsilon-neighborhood of the outlier-free optimal value whenever the iterates lie in a region defined by a weighted average of squared outlier gradient perturbations, whereas plain SGD requires the same bound without the squared derivative factor. Since this derivative lies in the unit interval, each outlier gradient contributes less to the variance that forces iterates out of the basin, so the robust region contains the SGD region. The claim is not that robust kernels make the objective globally easier; it is that, at the level of gradient variance and descent bounds, the per-sample weight is the mechanism that separates inlier signal from arbitrary outlier perturbation.

Load-bearing premise

The load-bearing premise is that the average alignment between outlier gradient perturbations and the outlier-free gradient cannot be so negative that it overpowers the descent signal; the paper relies on this to prove its enlarged region of convergence but does not state it among its assumptions.

Editorial extensions

If this is right

  • Robust kernels developed for classification can be transplanted into robotics-style robust estimation and vice versa, because the modified duality preserves the original problem structure instead of squaring the loss.
  • The Adaptive Alternation Algorithm turns the robust kernel's shape parameter from a hand-tuned hyperparameter into an adaptive variable controlled by one inlier-fraction target.
  • The variance bound shows that outlier-contributed gradient variance is scaled by the squared derivative of the kernel, which predicts more stable descent trajectories under zero-mean outliers than plain SGD.
  • The convergence theorems apply to arbitrary outlier gradients whose average squared perturbation is bounded, so the analysis covers adversarial outliers rather than only outliers drawn from a clean contamination model.
  • For the truncated kernel, the algorithm's weight update is exactly a conformal prediction set over samples, giving a principled interpretation of iterative sample trimming.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A practical robustness monitor could track the ratio of outlier-gradient energy before and after kernel weighting during training; a drop in that ratio would indicate the kernel is actively suppressing outlier influence.
  • The same duality suggests a soft version of data cleaning: instead of hard-trimming samples, one could propagate the per-sample weights to downstream tasks as importance weights, an extension not pursued in the paper.
  • Combining per-sample kernel weighting with norm-based methods such as gradient clipping might compound outlier suppression, since the former scales each gradient while the latter bounds the global step.
  • If the inlier-fraction target is chosen by validation, the conformal-set interpretation suggests it could be calibrated to control the error rate of inlier selection, a directly testable extension of the parameter update rule.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a unified view of robust losses from M-estimation and risk minimization through a modified Black-Rangarajan duality, introduces a robust loss kernel, and presents the Adaptive Alternation Algorithm (AAA) for training models with outliers. The authors claim convergence of AAA to an epsilon-neighborhood of the outlier-free optimum under L-smoothness and Polyak-Lojasiewicz conditions, with a region of convergence W_AAA1 that is larger than the SGD region W_SGD because the outlier variance bound contains the factor sigma'_c(f_i(w))^2. The paper also reports experiments on linear regression, CIFAR-10 classification, and neural radiance field reconstruction, and releases implementation code.

Significance. The proposed unification of robust losses and the adaptive parameter-update rule are interesting and practically motivated, and the experimental results suggest the algorithm works well even at high outlier rates. If the convergence theorem were correct, the claim that robust loss kernels provably enlarge the convergence region under arbitrary outliers would be a significant contribution. However, the central proof contains load-bearing errors, so the theoretical contribution is not established. The algorithmic and experimental parts retain value, and the code release is a strength, but the stated convergence guarantees are currently unsupported.

major comments (3)
  1. [Appendix G, Lemma 27] The proof of Lemma 27 drops a negative term when moving from (57) to (50). The displayed inequality (57) is E_i[∇f_i(w)^T ∇f_I(w)] ≥ 2μ(f_I(w) − f*_I) − λ(1/nO) Σ_{i∈nO} ||h_i|| ||∇f_I(w)||, but the lemma concludes E_i[∇f_i(w)^T ∇f_I(w)] ≥ 2μ(f_I(w) − f*_I). This conclusion requires the subtracted term to be nonnegative or otherwise controlled, which is not assumed anywhere. A concrete violation is given by f_I(w) = 0.5w^2, one inlier with gradient 2w, and one outlier with h_i = −4w; for λ = 1/2 and μ = 1, E_i[∇f_i(w)^T ∇f_I(w)] = −w^2 whereas 2μ(f_I(w) − f*_I) = w^2, and Assumption 13 holds for |w| ≥ 1/4. Since Lemma 27 is used to derive the SGD descent inequality (65) and Lemma 28 repeats the same argument for Theorems 19 and 22, the descent inequalities (65), (71), and (74) do not follow, and the claimed comparison W_AAA1 versus W_SGD is not established.
  2. [Section 6.2, Lemma 14 and Appendix F] Lemma 14 states that the variance in the descent direction for SGD and AAA1 is given by (22) and (23), respectively. Both expressions vanish when λ = 0, but for batch-size-one stochastic gradients the variance around the full-batch gradient is generally nonzero even without outliers. The derivation in Appendix F actually produces an upper bound, not an equality, and the bound is obtained by discarding inlier variance terms such as E_i[||∇f_i,I||^2] − ||∇f_I||^2. As stated, the lemma is false; it should be rephrased as an upper bound that additionally retains the inlier variance contribution.
  3. [Theorem 19 and Lemma 28] The convergence region W_AAA1 in (26) imposes min_i σ_c(f_i(w)) ≥ β > 0, but the proof and Lemma 28 require a positive lower bound on the derivative σ'_c(f_i(w)), denoted φ in Lemma 28. These are distinct conditions: all kernels in Table 1 satisfy σ_c(0) = 0, so the min-σ condition fails for any iterate where an inlier loss is near zero, which is precisely the regime to which the theorem claims convergence. Moreover, Lemma 28 obtains its lower bound by the same invalid drop of a negative outlier-bias term identified in Lemma 27, so even replacing β with φ would not repair the proof without an additional sign or bounded-bias assumption on the outlier gradients.
minor comments (4)
  1. [Appendix I, proof of Theorem 19] The line 'We assume that w_t, w_{t+1} ∈ W_SGD' should instead refer to W_AAA1; as written, the proof invokes the wrong region.
  2. [Remark 9] Remark 9 defines the truncated kernel as σ_c(r) = c · max{r/c, 1}, but Table 1 and the subsequent update rule require c · min{r/c, 1}; with the max definition, the derivative does not yield the trimming indicator I{f_i(w_t) ≤ c}.
  3. [Theorem 22] The set H(w) in (28) uses (1/nO) Σ_{i∈nO} σ'_c(f_i(w')) = ζ, but the parameter update rule (18) imposes the same constraint over all measurements D, not over outliers only; the outlier-only version is inconsistent with the algorithm.
  4. [Theorem 18 statement] The theorem says SGD 'converges to the optimal value' and then states E[||f_I(w_t) − f*_I||] < ε; the wording should say 'converges to an ε-neighborhood of the optimal value' to match the displayed guarantee.

Circularity Check

1 steps flagged · score 3.0 of 10

Mild self-referential structure: Theorem 19 assumes exactly the zeta-constraint that the parameter-update rule (18) was constructed to enforce, while the main Lemma 27 has a separate non-circular proof gap.

  1. self definitional [Section 5.2, Eq. (18); Section 6.3, Theorem 19, W_AAA1 (Eq. (26)); Remark 20]
    "ct = Find c∈[0,1] { 1/|D| ∑_{i∈D} σ′_c(fi(wt)) =ζ } ... The set WAAA1 has two more constraints 1/n ∑_{i=1}^n σ′_c(fi(w)) = ζ and mini σc(fi(w)) ≥ β > 0. The first comes from the step to update the parameter c in the algorithm and is always satisfied."

    The algorithm's parameter update (18) solves for c by imposing exactly the equality (1/n)∑σ'_c(f_i(w_t)) = ζ. Theorem 19 then lists this same equality as one of the defining assumptions of the convergence region W_AAA1, and Remark 20 states that this constraint 'comes from the step to update the parameter c in the algorithm and is always satisfied.' Thus the theorem's premise is not an independent condition on the learning problem; it is manufactured by the algorithm's own update rule. This is a mild self-referential structure rather than a full fit-to-data circularity, because the theorem still requires the separate outlier-variance bound (27) to deliver convergence, but the 'increased region' claim is partially baked into the design of the parameter update.

full rationale

The core derivation chain—modified Black-Rangarajan duality (Corollary 4), the robust-kernel definition (Definition 6), the alternation update (17), the variance lemma (Lemma 14), and the descent theorems (18, 19, 22)—is not a case of fitting a parameter to data and then relabeling it as a prediction, nor does it depend on load-bearing self-citations. The enlarged-region comparison W_AAA1 versus W_SGD follows mathematically from σ'_c ≤ 1 and the construction of W_AAA1, so it is not circular in the input-output sense. The mild self-referential loop is the zeta-constraint: Eq. (18) defines c_t by the same equality that Theorem 19 assumes, so part of the theorem's hypothesis is guaranteed by the algorithm's design. Separately, Appendix G's Lemma 27 displays a lower bound containing a negative outlier-bias term, then drops that term without a sign condition; this is a genuine proof gap affecting Theorems 18, 19, and 22, but it is a correctness problem rather than a circularity. Overall the paper is not substantially circular, though it does contain a noticeable assumption-construction loop and a serious missing-support step in the main convergence proof.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The theoretical claims rest on Assumptions 11 and 13, an unstated sign condition on the outlier bias in Lemma 27, and a beta > 0 strict-concavity condition that excludes the truncated kernel used in the best experiments. The only explicit user-set number central to the algorithm is zeta, which encodes expected inlier fraction. No new physical or model entities are introduced.

free parameters (1)
  • zeta (inlier-fraction target) = not reported in Section 7
    Equation (18) chooses c so the average robust weight equals zeta; by Eq. (19), zeta substitutes for the unknown inlier fraction n_I/n. Experiments vary outlier rate lambda without reporting zeta, so the algorithm's behavior may depend on an oracle-like setting of this parameter.
assumptions (4)
  • domain assumption Outlier gradients satisfy grad f_i(w) = grad f_i,I(w) + h_i(o_i,w), with h_i unknown; inlier gradient equals the outlier-free gradient.
    Assumption 11 (Section 6.1); verified for regression and classification in Appendix E, but for a generic model it is close to a definition of h_i.
  • domain assumption Low signal-to-outlier ratio: for every outlier i, ||h_i|| >= 1 and ||h_i|| >= ||grad f_i,I||.
    Assumption 13 (Section 6.1). It is not verified for classification and it excludes small but adversarial gradient perturbations, so it does not match the paper's arbitrary-outliers claim.
  • ad hoc to paper The outlier-bias inner product term lambda (1/nO) sum_{i in nO} h_i^T grad f_I is assumed nonnegative or negligible in Lemma 27.
    In Appendix G, Lemma 27, the proof obtains E_i[grad f_i^T grad f_I] >= 2 mu (f-f*) - lambda (1/nO) sum ||h_i|| ||grad f_I|| and then omits the negative term to conclude the desired inequality. Theorems 18-22 inherit this unstated sign condition.
  • domain assumption The robust kernel is strictly concave with invertible derivative, and min_i sigma_c(f_i(w)) >= beta > 0 throughout training.
    Corollary 4 and Theorem 19 require invertibility and a uniform positive lower bound. The truncated loss kernel used for Adaptive TL in Section 7.3 violates both.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Outlier-Robust Training of Machine Learning Models." pith.science (2026). https://pith.science/paper/C4ECWWTD

@misc{pith2026250100265,
  author       = {Pith},
  title        = {Pith review of: Outlier-Robust Training of Machine Learning Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/C4ECWWTD}},
  note         = {Machine review of arXiv:2501.00265}
}
abstract

Robust training of machine learning models in the presence of outliers has garnered attention across various domains. The use of robust losses is a popular approach and is known to mitigate the impact of outliers. We bring to light two literatures that have diverged in their ways of designing robust losses: one using M-estimation, which is popular in robotics and computer vision, and another using a risk-minimization framework, which is popular in deep learning. We first show that a simple modification of the Black-Rangarajan duality provides a unifying view. The modified duality brings out a definition of a robust loss kernel $\sigma$ that is satisfied by robust losses in both the literatures. Secondly, using the modified duality, we propose an Adaptive Alternation Algorithm (AAA) for training machine learning models with outliers. The algorithm iteratively trains the model by using a weighted version of the non-robust loss, while updating the weights at each iteration. The algorithm is augmented with a novel parameter update rule by interpreting the weights as inlier probabilities, and obviates the need for complex parameter tuning. Thirdly, we investigate convergence of the adaptive alternation algorithm to outlier-free optima. Considering arbitrary outliers (i.e., with no distributional assumption on the outliers), we show that the use of robust loss kernels {\sigma} increases the region of convergence. We experimentally show the efficacy of our algorithm on regression, classification, and neural scene reconstruction problems. We release our implementation code: https://github.com/MIT-SPARK/ORT.

Figures

Figures reproduced from arXiv: 2501.00265 by the authors.

Figure 1
Figure 1. Nerfacto (Tancik et al., 2023) reconstruction re [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Trajectory of (a) SGD (batch size = 1), (b) Adaptive Alternation Algorithm with Truncated Loss [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 5
Figure 5. Plot of the 1D training loss landscape as in [PITH_FULL_IMAGE:figures/full_fig_p015_5.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Test accuracy (PSNR ↑ and LPIPS ↓) of the trained model as a function of % outliers in the training data for various training algorithms: (i) Adam / SGD, the baseline approach proposed for training without outliers; (ii) Gradient Clipping, (iii) Normalized Gradient, (i…
Figure 6
Figure 6. Figure 6: Nerfacto reconstruction results after 80% of the training pixels have been perturbed by noise. The first row shows the result of training with our Adaptive Alternation Algorithm with Truncated Loss. The second row shows the result of running the original Adam optimizer…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

98 extracted references · 67 canonical work pages

  1. [1]

    Aftab and R

    K. Aftab and R. Hartley. Convergence of Iteratively Re-weighted Least Squares to Robust M-Estimators . In IEEE Winter Conference on Applications of Computer Vision , pp.\ 480--487, Jan. 2015

  2. [2]

    Aftab, R

    K. Aftab, R. Hartley, and J. Trumpf. Generalized Weiszfeld Algorithms for Lq Optimization . IEEE Trans. Pattern Anal. Machine Intell. , 37 0 (4): 0 728--745, Apr. 2015

  3. [3]

    Algan and I

    G. Algan and I. Ulusoy. Image classification with deep learning in the presence of noisy labels: A survey. Knowledge-Based Systems, 215: 0 106771, Mar. 2021

  4. [4]

    E. Amid, M. K. K. Warmuth, R. Anil, and T. Koren. Robust Bi-Tempered Logistic Loss Based on Bregman Divergences . In Advances in Neural Information Processing Systems (NIPS), volume 32, Dec. 2019

  5. [5]

    Outlier-Robust Estimation: Hardness, Minimally Tuned Algorithms, and Applications

    P. Antonante, V. Tzoumas, H. Yang, and L. Carlone. Outlier-robust estimation: Hardness, minimally tuned algorithms, and applications. IEEE Trans. Robotics , 38 0 (1): 0 281--301, 2021. https://arxiv.org/pdf/2007.15109.pdf

  6. [6]

    Armeni, O

    I. Armeni, O. Sener, A. Zamir, H. Jiang, I. Brilakis, M. Fischer, and S. Savarese. 3d semantic parsing of large-scale indoor spaces. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp.\ 1534--1543, 2016

  7. [7]

    Awasthi, M

    P. Awasthi, M. F. Balcan, and P. M. Long. The power of localization for efficiently learning linear separators with noise. J. ACM, 63 0 (6), 2017

  8. [8]

    J. T. Barron. A general and adaptive robust loss function. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 4331--4339, 2019

Show all 98 references
  1. [9]

    Bhatia, P

    K. Bhatia, P. Jain, and P. Kar. Robust regression via hard thresholding. In Advances in Neural Information Processing Systems (NIPS), pp.\ 721--729, 2015

  2. [10]

    Bhatia, P

    K. Bhatia, P. Jain, P. Kamalaruban, and P. Kar. Consistent robust regression. In Advances in Neural Information Processing Systems (NIPS), volume 30. Curran Associates, Inc., 2017

  3. [11]

    M. J. Black and A. Rangarajan. On the unification of line processes, outlier rejection, and robust statistics with applications in early vision. Intl. J. of Computer Vision, 19 0 (1): 0 57--91, 1996

  4. [12]

    Blake and A

    A. Blake and A. Zisserman. Visual reconstruction. MIT Press, 1987

  5. [13]

    G. E. P. Box and D. R. Cox. An Analysis of Transformations . Journal of the Royal Statistical Society: Series B (Methodological), 26 0 (2): 0 211--243, 1964

  6. [14]

    Brimberg and R

    J. Brimberg and R. F. Love. Global Convergence of a Generalized Iterative Procedure for the Minisum Location Problem with lp Distances . Operations Research, 41 0 (6): 0 1153--1163, 1993

  7. [15]

    L. Carlone. Estimation contracts for outlier-robust geometric perception. Foundations and Trends (FnT) in Robotics, arXiv preprint: 2208.10521, 2023. https://arxiv.org/pdf/2208.10521.pdf

  8. [16]

    C. Chai, L. Cao, G. Li, J. Li, Y. Luo, and S. Madden. Human-in-the-loop Outlier Detection . In ACM SIGMOD International Conference on Management of Data , pp.\ 19--33, May 2020

  9. [17]

    Chang, A

    A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang. Matterport3d: Learning from rgb-d data in indoor environments. International Conference on 3D Vision (3DV), 2017

  10. [18]

    Charikar, J

    M. Charikar, J. Steinhardt, and G. Valiant. Learning from untrusted data. In Proceedings of the 49th Annual ACM SIGACT Symposium on Theory of Computing, STOC 2017, pp.\ 47--60, 2017

  11. [19]

    Chebrolu, T

    N. Chebrolu, T. L \"a be, O. Vysotska, J. Behley, and C. Stachniss. Adaptive robust kernels for non-linear least squares problems. arXiv preprint arXiv:2004.14938, 2020

  12. [20]

    Chebrolu, T

    N. Chebrolu, T. Läbe, O. Vysotska, J. Behley, and C. Stachniss. Adaptive robust kernels for non-linear least squares problems. IEEE Robotics and Automation Letters , 6 0 (2): 0 2240--2247, 2021

  13. [21]

    Y. Chen, C. Caramanis, and S. Mannor. Robust sparse regression under adversarial corruption. In Intl. Conf. on Machine Learning (ICML), volume 28, pp.\ 774--782, 2013

  14. [22]

    Chhabra, B

    A. Chhabra, B. Li, J. Chen, P. Mohapatra, and H. Liu. Outlier Gradient Analysis : Efficiently Identifying Detrimental Training Samples for Deep Learning Models . arXiv: 2405.03869, Oct. 2024

  15. [23]

    Demidovich, G

    Y. Demidovich, G. Malinovsky, I. Sokolov, and P. Richtarik. A Guide Through the Zoo of Biased SGD . Advances in Neural Information Processing Systems (NIPS), 36: 0 23158--23171, Dec. 2023

  16. [24]

    X. Deng, Y. Xiang, A. Mousavian, C. Eppner, T. Bretl, and D. Fox. Self-supervised 6D Object Pose Estimation for Robot Manipulation . In IEEE Intl. Conf. on Robotics and Automation (ICRA), pp.\ 3665--3671, May 2020

  17. [25]

    Dereich and A

    S. Dereich and A. Jentzen. Convergence rates for the Adam optimizer. arXiv: 2407.21078, Jul. 2024

  18. [26]

    Diakonikolas, G

    I. Diakonikolas, G. Kamath, D. Kane, J. Li, A. Moitra, and A. Stewart. Robust estimators in high dimensions without the computational intractability. In IEEE 57th Annual Symposium on Foundations of Computer Science, pp.\ 655--664. IEEE, 2016

  19. [27]

    Diakonikolas, G

    I. Diakonikolas, G. Kamath, D. M. Kane, J. Li, A. Moitra, and A. Stewart. Robustly learning a gaussian: Getting optimal error, efficiently. In Proceedings of the Twenty-Ninth Annual ACM-SIAM Symposium on Discrete Algorithms, SODA '18, pp.\ 2683–2702, 2018 a

  20. [28]

    Diakonikolas, D

    I. Diakonikolas, D. M. Kane, and A. Stewart. Learning geometric concepts with nasty noise. In Proceedings of the 50th Annual ACM SIGACT Symposium on Theory of Computing, STOC 2018, pp.\ 1061--1073, 2018 b

  21. [29]

    Diakonikolas, G

    I. Diakonikolas, G. Kamath, D. Kane, J. Li, A. Moitra, and A. Stewart. Robust estimators in high-dimensions without the computational intractability. SIAM Journal on Computing, 48 0 (2): 0 742--864, 2019 a . doi:10.1137/17M1126680

  22. [30]

    Diakonikolas, G

    I. Diakonikolas, G. Kamath, D. Kane, J. Li, J. Steinhardt, and A. Stewart. Sever: A robust meta-algorithm for stochastic optimization. In K. Chaudhuri and R. Salakhutdinov (eds.), Intl. Conf. on Machine Learning (ICML), volume 97 of Proceedings of Machine Learning Research, pp...

  23. [31]

    Diakonikolas, W

    I. Diakonikolas, W. Kong, and A. Stewart. Efficient algorithms and lower bounds for robust linear regression. In Proceedings of the Thirtieth Annual ACM-SIAM Symposium on Discrete Algorithms, SODA '19, pp.\ 2745--2754, 2019 c

  24. [32]

    Elesedy and M

    B. Elesedy and M. Hutter. U-clip: On-average unbiased stochastic gradient clipping. arXiv preprint arXiv:2302.02971, 2023

  25. [33]

    L. Feng, S. Shu, Z. Lin, F. Lv, L. Li, and B. An. Can Cross Entropy Loss Be Robust to Label Noise ? In Intl. Joint Conf. on AI (IJCAI), volume 3, pp.\ 2206--2212, Jul. 2020

  26. [34]

    Ferrari and Y

    D. Ferrari and Y. Yang. Maximum Lq-likelihood estimation. The Annals of Statistics, 38 0 (2): 0 753--783, Apr. 2010

  27. [35]

    Fischler and R

    M. Fischler and R. Bolles. Random sample consensus: a paradigm for model fitting with application to image analysis and automated cartography. Commun. ACM, 24: 0 381--395, 1981

  28. [36]

    Foret, A

    P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur. Sharpness-aware Minimization for Efficiently Improving Generalization . In Intl. Conf. on Learning Representations (ICLR), Oct. 2020

  29. [37]

    S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, E. Orgad, R. Entezari, G. Daras, S. M. Pratt, V. Ramanujan, Y. Bitton, K. Marathe, S. Mussmann, R. Vencu, M. Cherti, R. Krishna, P. W. Koh, O. Saukh, A. Ratner, S. S...

  30. [38]

    Garrigos and R

    G. Garrigos and R. M. Gower. Handbook of Convergence Theorems for ( Stochastic ) Gradient Methods . arXiv preprint arXiv:2301.11235, Feb. 2023

  31. [39]

    Ghosh, N

    A. Ghosh, N. Manwani, and P. S. Sastry. Making Risk Minimization Tolerant to Label Noise . Neurocomputing, 160: 0 93--107, Jul. 2015

  32. [40]

    Ghosh, H

    A. Ghosh, H. Kumar, and P. S. Sastry. Robust loss functions under label noise for deep neural networks. In Nat. Conf. on Artificial Intelligence (AAAI), pp.\ 1919–1925, Feb. 2017

  33. [41]

    R. M. Gower, M. Schmidt, F. Bach, and P. Richt\'arik. Variance- Reduced Methods for Machine Learning . Proceedings of the IEEE , 108 0 (11): 0 1968--1983, Nov. 2020

  34. [42]

    S. Hu, Z. Yang, X. Wang, Y. Ying, and S. Lyu. Outlier Robust Adversarial Training . In Proceedings of the 15th Asian Conference on Machine Learning , pp.\ 454--469. PMLR, Feb. 2024

  35. [43]

    P. Huber. Robust Statistics. John Wiley & Sons, New York, NY, 1981

  36. [44]

    Jawaid, R

    M. Jawaid, R. Talak, Y. Latif, L. Carlone, and T.-J. Chin. Test-time certifiable self-supervision to bridge the sim2real gap in event-based satellite pose estimation. In IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), Oct. 2024

  37. [45]

    Karmalkar and E

    S. Karmalkar and E. Price. Compressed sensing with adversarial sparse noise via L1 regression. CoRR, abs/1809.08055, 2018. URL http://arxiv.org/abs/1809.08055

  38. [46]

    Karmalkar, A

    S. Karmalkar, A. Klivans, and P. Kothari. List-decodable linear regression. In Advances in Neural Information Processing Systems (NIPS), volume 32, 2019

  39. [47]

    A. R. Klivans, P. M. Long, and R. A. Servedio. Learning halfspaces with malicious noise. In S. Albers, A. Marchetti-Spaccamela, Y. Matias, S. Nikoletseas, and W. Thomas (eds.), Automata, Languages and Programming, pp.\ 609--621, 2009

  40. [48]

    A. R. Klivans, P. K. Kothari, and R. Meka. Efficient algorithms for outlier-robust regression. CoRR, abs/1803.03241, 2018. URL http://arxiv.org/abs/1803.03241

  41. [49]

    Koloskova, H

    A. Koloskova, H. Hendrikx, and S. U. Stich. Revisiting gradient clipping: Stochastic bias and tight convergence guarantees. In Intl. Conf. on Machine Learning (ICML), volume 202, pp.\ 17343--17363, Jul. 2023

  42. [50]

    P. K. Kothari and J. Steinhardt. Better agnostic clustering via relaxed tensor norms. CoRR, abs/1711.07465, 2017. URL http://arxiv.org/abs/1711.07465

  43. [51]

    P. K. Kothari, J. Steinhardt, and D. Steurer. Robust moment estimation and improved clustering via sum of squares. In Proceedings of the 50th Annual ACM SIGACT Symposium on Theory of Computing, STOC 2018, pp.\ 1035--1046, 2018

  44. [52]

    K. A. Lai, A. B. Rao, and S. Vempala. Agnostic estimation of mean and covariance. In 2016 IEEE 57th Annual Symposium on Foundations of Computer Science (FOCS), pp.\ 665--674. IEEE Computer Society, 2016. doi:10.1109/FOCS.2016.76. URL https://doi.ieeecomputersociety.org/10.1109...

  45. [53]

    H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein. Visualizing the loss landscape of neural nets. In Advances in Neural Information Processing Systems (NIPS), volume 31, Dec. 2018. URL https://proceedings.neurips.cc/paper_files/paper/2018/file/a41b3bb3e6b050b6c9067c67f663b9...

  46. [54]

    J. Li, R. Socher, and S. C. Hoi. Dividemix: Learning with noisy labels as semi-supervised learning. Intl. Conf. on Learning Representations (ICLR), 2020

  47. [55]

    Liu and H

    Y. Liu and H. Guo. Peer Loss Functions : Learning from Noisy Labels without Knowing Noise Rates . In Intl. Conf. on Machine Learning (ICML), pp.\ 6226--6236, Nov. 2020

  48. [56]

    Z. Lu, Y. Zhang, K. Doherty, O. Severinsen, E. Yang, and J. Leonard. SLAM- supported self-training for 6d object pose estimation. In IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), pp.\ 2833--2840, Oct. 2022

  49. [57]

    Lyu and I

    Y. Lyu and I. W. Tsang. Curriculum loss: Robust learning and generalization against label corruption. In Intl. Conf. on Learning Representations (ICLR), Apr. 2020

  50. [58]

    X. Ma, H. Huang, Y. Wang, S. Romano, S. Erfani, and J. Bailey. Normalized Loss Functions for Deep Learning with Noisy Labels . In Intl. Conf. on Machine Learning (ICML), pp.\ 6543--6553, Nov. 2020

  51. [59]

    V. V. Mai and M. Johansson. Stability and convergence of stochastic gradient clipping: Beyond lipschitz continuity and smoothness. In Intl. Conf. on Machine Learning (ICML), pp.\ 7325--7335. PMLR, 2021

  52. [60]

    A. K. Menon, A. S. Rawat, S. J. Reddi, and S. Kumar. Can gradient clipping mitigate label noise? In Intl. Conf. on Learning Representations (ICLR), 2020

  53. [61]

    Merad and S

    I. Merad and S. Ga\"iffas. Robust Stochastic Optimization via Gradient Quantile Clipping . Trans. on Machine Learning Research, May 2024

  54. [62]

    Mildenhall, P

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. arXiv preprint arXiv:2003.08934, 2020

  55. [63]

    M\"uller, A

    T. M\"uller, A. Evans, C. Schied, and A. Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph., 41 0 (4): 0 102:1--102:15, July 2022. doi:10.1145/3528223.3530127. URL https://doi.org/10.1145/3528223.3530127

  56. [64]

    J. Naudts. Deformed exponentials and logarithms in generalized thermostatistics. Physica A: Statistical Mechanics and its Applications, 316 0 (1): 0 323--334, 2002

  57. [65]

    Nguyen and T

    N. Nguyen and T. Tran. Exact recoverability from dense corrupted observations via _ 1 -minimization. IEEE Trans. on Information Theory, 59 0 (4): 0 2017--2035, 2013

  58. [66]

    L. Peng, C. K\"ummerle, and R. Vidal. On the Convergence of IRLS and Its Variants in Outlier-Robust Estimation . In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp.\ 17808--17818, Jun. 2023

  59. [67]

    Prasad, A

    A. Prasad, A. S. Suggala, S. Balakrishnan, and P. Ravikumar. Robust estimation via robust gradient estimation. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 82, 2020

  60. [68]

    Raghavendra and M

    P. Raghavendra and M. Yau. List decodable learning via sum of squares. In Proceedings of the Thirty-First Annual ACM-SIAM Symposium on Discrete Algorithms, SODA '20, pp.\ 161–180, 2020

  61. [69]

    Reisizadeh, H

    A. Reisizadeh, H. Li, S. Das, and A. Jadbabaie. Variance-reduced Clipping for Non-convex Optimization . arXiv: 2303.00883, Jun. 2023

  62. [70]

    M. Ren, W. Zeng, B. Yang, and R. Urtasun. Learning to Reweight Examples for Robust Deep Learning . In Intl. Conf. on Machine Learning (ICML), pp.\ 4334--4343, Jul. 2018

  63. [71]

    Sabour, S

    S. Sabour, S. Vora, D. Duckworth, I. Krasin, D. J. Fleet, and A. Tagliasacchi. Robustnerf: Ignoring distractors with robust losses. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp.\ 20626--20636, 2023

  64. [72]

    Schmidt and D

    T. Schmidt and D. Fox. Self-directed Lifelong Learning for Robot Vision . In Robotics Research , pp.\ 109--114. Springer International Publishing, 2020

  65. [73]

    Shafer and V

    G. Shafer and V. Vovk. A Tutorial on Conformal Prediction . J. of Machine Learning Research, pp.\ 51, 2008

  66. [74]

    V. Shah, X. Wu, and S. Sanghavi. Choosing the Sample with Lowest Loss makes SGD Robust . In Twenty Third International Conference on Artificial Intelligence and Statistics , pp.\ 2120--2130, Jun. 2020

  67. [75]

    Shen and S

    Y. Shen and S. Sanghavi. Learning with Bad Training Data via Iterative Trimmed Loss Minimization . In Intl. Conf. on Machine Learning (ICML), pp.\ 5739--5748, May 2019

  68. [76]

    J. Shi, R. Talak, D. Maggio, and L. Carlone. A correct-and-certify approach to self-supervise object pose estimators via ensemble self-training. In Robotics: Science and Systems (RSS), 2023. https://arxiv.org/pdf/2302.06019.pdf

  69. [77]

    H. Song, M. Kim, D. Park, Y. Shin, and J.-G. Lee. Learning From Noisy Labels With Deep Neural Networks : A Survey . IEEE Trans. Neural Netw. Learn. Syst., 34 0 (11): 0 8135--8153, Nov. 2023

  70. [78]

    Sukhbaatar, J

    S. Sukhbaatar, J. Bruna, M. Paluri, L. Bourdev, and R. Fergus. Training Convolutional Networks with Noisy Labels . In Intl. Conf. on Learning Representations (ICLR), May 2015

  71. [79]

    Talak, L

    R. Talak, L. Peng, and L. Carlone. Certifiable 3D object pose estimation: Foundations, learning models, and self-training. IEEE Trans. Robotics , 39 0 (4): 0 2805--2824, 2023. https://arxiv.org/pdf/2206.11215.pdf

  72. [80]

    Tancik, E

    M. Tancik, E. Weber, E. Ng, R. Li, B. Yi, T. Wang, A. Kristoffersen, J. Austin, K. Salahi, A. Ahuja, et al. Nerfstudio: A modular framework for neural radiance field development. In SIGGRAPH, pp.\ 1--12, 2023

  73. [81]

    K. M. Tavish and T. D. Barfoot. At all costs: A comparison of robust cost functions for camera correspondence outliers. In Conf. Computer and Robot Vision, pp.\ 62--69. IEEE, 2015

  74. [82]

    Tkachenko, M

    M. Tkachenko, M. Malyuk, A. Holmanyuk, and N. Liubimov. Label Studio : Data labeling software, 2020. URL https://github.com/HumanSignal/label-studio

  75. [83]

    C. Wang, A. Wang, J. Li, A. Yuille, and C. Xie. Benchmarking Robustness in Neural Radiance Fields . In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp.\ 2926--2936, Jun. 2024 a

  76. [84]

    Y. Wang, X. Ma, Z. Chen, Y. Luo, J. Yi, and J. Bailey. Symmetric Cross Entropy for Robust Learning With Noisy Labels . In Intl. Conf. on Computer Vision (ICCV), pp.\ 322--330, Oct. 2019

  77. [85]

    Z. Wang, M. Chen, Y. Guo, Z. Li, and Q. Yu. Bridging the domain gap in satellite pose estimation: A self-training approach based on geometrical constraints. IEEE Trans. Aerosp. Electron. Syst., 60 0 (3): 0 2500--2514, 2024 b

  78. [86]

    Wright and Y

    J. Wright and Y. Ma. Dense error correction via ^1 -minimization. IEEE Trans. on Information Theory, 56 0 (7): 0 3540--3560, 2010

  79. [87]

    Y. Xu, P. Cao, Y. Kong, and Y. Wang. L\_ DMI : A Novel Information-theoretic Loss Function for Training Deep Nets Robust to Label Noise . In Advances in Neural Information Processing Systems (NIPS), volume 32, Dec. 2019

  80. [88]

    B. Yang, M. Bai, M. Liang, W. Zeng, and R. Urtasun. Auto4D : Learning to Label 4D Objects from Sequential Point Clouds . arXiv:2101.06586, Mar. 2021

  81. [89]

    Yang and L

    H. Yang and L. Carlone. Certifiably optimal outlier-robust geometric perception: Semidefinite relaxations and scalable global optimization. IEEE Trans. Pattern Anal. Machine Intell. , 2022. https://arxiv.org/pdf/2109.03349.pdf

  82. [90]

    H. Yang, P. Antonante, V. Tzoumas, and L. Carlone. Graduated non-convexity for robust spatial perception: From non-minimal solvers to global outlier rejection. IEEE Robotics and Automation Letters ( RA-L ) , 5 0 (2): 0 1127--1134, 2020 a . arXiv preprint:1909.08605 (with suppl...

  83. [91]

    H. Yang, J. Shi, and L. Carlone. TEASER: Fast and Certifiable Point Cloud Registration . IEEE Trans. Robotics , 37 0 (2): 0 314--333, 2020 b . extended arXiv version 2001.07715 https://arxiv.org/pdf/2001.07715.pdf

  84. [92]

    F. Yu, D. Wang, E. Shelhamer, and T. Darrell. Deep layer aggregation. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp.\ 2403--2412, 2018

  85. [93]

    Zhang, M

    H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz. mixup: Beyond empirical risk minimization. Intl. Conf. on Learning Representations (ICLR), 2018

  86. [94]

    Zhang, T

    J. Zhang, T. He, S. Sra, and A. Jadbabaie. Why Gradient Clipping Accelerates Training : A Theoretical Justification for Adaptivity . In Intl. Conf. on Learning Representations (ICLR), Mar. 2020 a

  87. [95]

    Zhang, S

    J. Zhang, S. P. Karimireddy, A. Veit, S. Kim, S. Reddi, S. Kumar, and S. Sra. Why are Adaptive Methods Good for Attention Models ? In Advances in Neural Information Processing Systems (NIPS), volume 33, pp.\ 15383--15393, Dec. 2020 b

  88. [96]

    Zhang and M

    Z. Zhang and M. Sabuncu. Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels . In Advances in Neural Information Processing Systems (NIPS), volume 31, Dec. 2018

  89. [97]

    X. Zhou, X. Liu, D. Zhai, J. Jiang, and X. Ji. Asymmetric Loss Functions for Noise-Tolerant Learning : Theory and Applications . IEEE Trans. Pattern Anal. Machine Intell. , 45 0 (7): 0 8094--8109, Jul. 2023

  90. [98]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.