REVIEW 14 cited by
Understanding Deep Neural Networks with Rectified Linear Units
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
In this paper we investigate the family of functions representable by deep neural networks (DNN) with rectified linear units (ReLU). We give an algorithm to train a ReLU DNN with one hidden layer to *global optimality* with runtime polynomial in the data size albeit exponential in the input dimension. Further, we improve on the known lower bounds on size (from exponential to super exponential) for approximating a ReLU deep net function by a shallower ReLU net. Our gap theorems hold for smoothly parametrized families of "hard" functions, contrary to countable, discrete families known in the literature. An example consequence of our gap theorems is the following: for every natural number $k$ there exists a function representable by a ReLU DNN with $k^2$ hidden layers and total size $k^3$, such that any ReLU DNN with at most $k$ hidden layers will require at least $\frac{1}{2}k^{k+1}-1$ total nodes. Finally, for the family of $\mathbb{R}^n\to \mathbb{R}$ DNNs with ReLU activations, we show a new lowerbound on the number of affine pieces, which is larger than previous constructions in certain regimes of the network architecture and most distinctively our lowerbound is demonstrated by an explicit construction of a *smoothly parameterized* family of functions attaining this scaling. Our construction utilizes the theory of zonotopes from polyhedral theory.
Forward citations
Cited by 14 Pith papers
-
A law of robustness for two-layer neural networks with arbitrary weights
Any width-m two-layer piecewise-linear network with arbitrary weights that fits n noisy labels below the noise floor has Lip ≳ ε sqrt(n/(m log(m n d/ε))) with high probability on the sphere or Gaussian.
-
Reconstructing the Stripping History of the Sagittarius Stream with Neural Networks
A neural network trained on simulations infers stripping times for Sagittarius stream stars from phase-space data, measuring a 0.3 dex/Gyr metallicity gradient and estimating ages for globular clusters such as Pal 12 ...
-
Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
One of the Q, K or V weights in transformer self-attention is redundant and replaceable by the identity matrix under mild assumptions, reducing parameters by 25 percent with no loss in small-model performance.
-
Laguerre Geometry for Interpreting Large Language Models
LLM concepts are Laguerre–Voronoi cells; Geometric Lens reads the exact cell of any hidden vector by isolating residual piecewise-linear flow from cross-token attention transport.
-
Estimation of High Dimensional Bounded Discrete Graphical Models via Regularized Generalized Score Matching
Introduces bounded discrete graphical models and the BRIDGE regularized score matching estimator with nonasymptotic error bounds and exact support recovery for high-dimensional discrete data.
-
A Theory on Flow Matching with Neural Networks
Establishes convergence guarantees for overparameterized 2-layer ReLU networks in flow matching, generalization bounds for the velocity-field objective, and Wasserstein guarantees for generated samples, using multi-ta...
-
"Show Me You Comply... Without Showing Me Anything": Zero-Knowledge Software Auditing for AI-Enabled Systems
ZKMLOps is an MLOps framework that uses zero-knowledge proofs to generate verifiable cryptographic evidence of AI model compliance without revealing confidential information.
-
Exploring substructures in the Milky Way halo Neural networks applied to Gaia and APOGEE DR 17
A chemo-dynamical graph neural network pipeline recovers most known globular clusters in APOGEE halo data, re-identifies known streams, and proposes a new stream candidate, Arnus I.
-
On the inductive bias of infinite-depth ResNets and the bottleneck rank
The minimum-weight cost of deep linear ResNets interpolates between nuclear norm and rank as the weight-decay ratio varies, implying a low bottleneck-rank bias for deep nonlinear ResNets.
-
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
SPIN lets weak LLMs become strong by self-generating training data from previous model versions and training to prefer human-annotated responses over its own outputs, outperforming DPO even with extra GPT-4 data on be...
-
On Symmetry and Initialization for Neural Networks
For symmetric target functions, chosen initial conditions in one-hidden-layer networks enable SGD to produce generalization guarantees, unlike random initialization.
-
Bayesian meta-learning for modeling Alzheimer's disease progression
Bayesian meta-learner predicts individualized Alzheimer's disease progression distributions from MRI and trajectories, competitive on ADNI data and less overconfident for long-term scores than deterministic versions.
-
Engineering Trustworthy Machine-Learning Operations with Zero-Knowledge Proofs
A systematic review of 57 ZKP-for-ML papers concludes that inference verification dominates the field and that research is converging toward a unified ZKMLOps framework for trustworthy, auditable AI.
-
Deep learning applied to computational mechanics: A comprehensive review, state of the art, and the classics
A comprehensive review of deep learning techniques for computational mechanics, including LSTM for constitutive modeling, PINNs for PDE solving, optimizers, and kernel methods.
Discussion (0). Continue with ORCID to comment.