REVIEW 2 cited by
The Computational Complexity of Training ReLU(s)
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We consider the computational complexity of training depth-2 neural networks composed of rectified linear units (ReLUs). We show that, even for the case of a single ReLU, finding a set of weights that minimizes the squared error (even approximately) for a given training set is NP-hard. We also show that for a simple network consisting of two ReLUs, the error minimization problem is NP-hard, even in the realizable case. We complement these hardness results by showing that, when the weights and samples belong to the unit ball, one can (agnostically) properly and reliably learn depth-2 ReLUs with $k$ units and error at most $\epsilon$ in time $2^{(k/\epsilon)^{O(1)}}n^{O(1)}$; this extends upon a previous work of Goel, Kanade, Klivans and Thaler (2017) which provided efficient improper learning algorithms for ReLUs.
Forward citations
Cited by 2 Pith papers
-
Equivalence of Coarse and Fine-Grained Models for Learning with Distribution Shift
PQ and TDS learning are equivalent in the distribution-free setting for Boolean classes, implying hardness for TDS halfspace learning but efficient algorithms with membership queries.
-
Omnipredicting Single-Index Models with Multi-Index Models
A new analysis of the Isotron algorithm yields omnipredictors for single-index models with about ε^-4 samples (ε^-2 for bi-Lipschitz links), improving the previous ε^-10 construction.
Discussion (0). Continue with ORCID to comment.