Pith. sign in

REVIEW 1 cited by

Understanding Square Loss in Training Overparametrized Neural Network Classifiers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.03657 v1 pith:HQM7BTPJ submitted 2021-12-07 stat.ML cs.LG

classification stat.MLcs.LG
keywords losssquareerrorneuralcalibrationdataraterobustness
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning has achieved many breakthroughs in modern classification tasks. Numerous architectures have been proposed for different data structures but when it comes to the loss function, the cross-entropy loss is the predominant choice. Recently, several alternative losses have seen revived interests for deep classifiers. In particular, empirical evidence seems to promote square loss but a theoretical justification is still lacking. In this work, we contribute to the theoretical understanding of square loss in classification by systematically investigating how it performs for overparametrized neural networks in the neural tangent kernel (NTK) regime. Interesting properties regarding the generalization error, robustness, and calibration error are revealed. We consider two cases, according to whether classes are separable or not. In the general non-separable case, fast convergence rate is established for both misclassification rate and calibration error. When classes are separable, the misclassification rate improves to be exponentially fast. Further, the resulting margin is proven to be lower bounded away from zero, providing theoretical guarantees for robustness. We expect our findings to hold beyond the NTK regime and translate to practical settings. To this end, we conduct extensive empirical studies on practical neural networks, demonstrating the effectiveness of square loss in both synthetic low-dimensional data and real image data. Comparing to cross-entropy, square loss has comparable generalization error but noticeable advantages in robustness and model calibration.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Super-fast Rates of Convergence for Neural Network Classifiers under the Hard Margin Condition

    cs.LG 2025-05 unverdicted novelty 6.0 of 10

    DNN classifiers achieve excess risk O(n^{-α}) with α arbitrarily large under hard margin (q=∞) and distribution-adapted smoothness of the regression function, with matching minimax lower bounds for q≥2.

Pith tools