Pith. sign in

REVIEW 14 cited by

Understanding Deep Neural Networks with Rectified Linear Units

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1611.01491 v6 pith:A23CZDWC submitted 2016-11-04 cs.LG cond-mat.dis-nncs.AIcs.CCstat.ML

classification cs.LGcond-mat.dis-nncs.AIcs.CCstat.ML
keywords reludeepexponentialfamilyfunctionshiddensizeconstruction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this paper we investigate the family of functions representable by deep neural networks (DNN) with rectified linear units (ReLU). We give an algorithm to train a ReLU DNN with one hidden layer to *global optimality* with runtime polynomial in the data size albeit exponential in the input dimension. Further, we improve on the known lower bounds on size (from exponential to super exponential) for approximating a ReLU deep net function by a shallower ReLU net. Our gap theorems hold for smoothly parametrized families of "hard" functions, contrary to countable, discrete families known in the literature. An example consequence of our gap theorems is the following: for every natural number $k$ there exists a function representable by a ReLU DNN with $k^2$ hidden layers and total size $k^3$, such that any ReLU DNN with at most $k$ hidden layers will require at least $\frac{1}{2}k^{k+1}-1$ total nodes. Finally, for the family of $\mathbb{R}^n\to \mathbb{R}$ DNNs with ReLU activations, we show a new lowerbound on the number of affine pieces, which is larger than previous constructions in certain regimes of the network architecture and most distinctively our lowerbound is demonstrated by an explicit construction of a *smoothly parameterized* family of functions attaining this scaling. Our construction utilizes the theory of zonotopes from polyhedral theory.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 127 citations worldwide. Full citation record

  1. A law of robustness for two-layer neural networks with arbitrary weights

    cs.LG 2026-07 accept novelty 8.0 of 10

    Any width-m two-layer piecewise-linear network with arbitrary weights that fits n noisy labels below the noise floor has Lip ≳ ε sqrt(n/(m log(m n d/ε))) with high probability on the sphere or Gaussian.

  2. Reconstructing the Stripping History of the Sagittarius Stream with Neural Networks

    astro-ph.GA 2026-05 unverdicted novelty 7.0 of 10

    A neural network trained on simulations infers stripping times for Sagittarius stream stars from phase-space data, measuring a 0.3 dex/Gyr metallicity gradient and estimating ages for globular clusters such as Pal 12 ...

  3. Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers

    cs.LG 2025-10 unverdicted novelty 7.0 of 10

    One of the Q, K or V weights in transformer self-attention is redundant and replaceable by the identity matrix under mild assumptions, reducing parameters by 25 percent with no loss in small-model performance.

  4. Laguerre Geometry for Interpreting Large Language Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    LLM concepts are Laguerre–Voronoi cells; Geometric Lens reads the exact cell of any hidden vector by isolating residual piecewise-linear flow from cross-token attention transport.

  5. Estimation of High Dimensional Bounded Discrete Graphical Models via Regularized Generalized Score Matching

    stat.ME 2026-06 unverdicted novelty 6.0 of 10

    Introduces bounded discrete graphical models and the BRIDGE regularized score matching estimator with nonasymptotic error bounds and exact support recovery for high-dimensional discrete data.

  6. A Theory on Flow Matching with Neural Networks

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Establishes convergence guarantees for overparameterized 2-layer ReLU networks in flow matching, generalization bounds for the velocity-field objective, and Wasserstein guarantees for generated samples, using multi-ta...

  7. "Show Me You Comply... Without Showing Me Anything": Zero-Knowledge Software Auditing for AI-Enabled Systems

    cs.SE 2025-10 unverdicted novelty 6.0 of 10

    ZKMLOps is an MLOps framework that uses zero-knowledge proofs to generate verifiable cryptographic evidence of AI model compliance without revealing confidential information.

  8. Exploring substructures in the Milky Way halo Neural networks applied to Gaia and APOGEE DR 17

    astro-ph.GA 2025-07 conditional novelty 6.0 of 10

    A chemo-dynamical graph neural network pipeline recovers most known globular clusters in APOGEE halo data, re-identifies known streams, and proposes a new stream candidate, Arnus I.

  9. On the inductive bias of infinite-depth ResNets and the bottleneck rank

    cs.LG 2025-01 conditional novelty 6.0 of 10

    The minimum-weight cost of deep linear ResNets interpolates between nuclear norm and rank as the weight-decay ratio varies, implying a low bottleneck-rank bias for deep nonlinear ResNets.

  10. Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

    cs.LG 2024-01 unverdicted novelty 6.0 of 10

    SPIN lets weak LLMs become strong by self-generating training data from previous model versions and training to prefer human-annotated responses over its own outputs, outperforming DPO even with extra GPT-4 data on be...

  11. On Symmetry and Initialization for Neural Networks

    cs.LG 2019-07 unverdicted novelty 5.0 of 10

    For symmetric target functions, chosen initial conditions in one-hidden-layer networks enable SGD to produce generalization guarantees, unlike random initialization.

  12. Bayesian meta-learning for modeling Alzheimer's disease progression

    stat.ML 2026-06 unverdicted novelty 4.0 of 10

    Bayesian meta-learner predicts individualized Alzheimer's disease progression distributions from MRI and trajectories, competitive on ADNI data and less overconfident for long-term scores than deterministic versions.

  13. Engineering Trustworthy Machine-Learning Operations with Zero-Knowledge Proofs

    cs.SE 2025-05 conditional novelty 4.0 of 10

    A systematic review of 57 ZKP-for-ML papers concludes that inference verification dominates the field and that research is converging toward a unified ZKMLOps framework for trustworthy, auditable AI.

  14. Deep learning applied to computational mechanics: A comprehensive review, state of the art, and the classics

    cs.LG 2022-12 unverdicted novelty 2.0 of 10

    A comprehensive review of deep learning techniques for computational mechanics, including LSTM for constitutive modeling, PINNs for PDE solving, optimizers, and kernel methods.

Pith tools