REVIEW 5 major objections 5 minor 1 cited by
Deep ReLU networks -- injectivity capacity upper bounds
T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Deep ReLU networks become injective once total width reaches about 10 times the input, and depth beyond four layers adds almost nothing.
desk verdict The recursive RDT program and saturation numbers are interesting, but Lemma 1's proof breaks on a false support-size inequality, so the central equivalence and all Table 1 bounds are unproven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the random-dual function $\varphi_0$ (and its lifted counterpart $\bar{\varphi}_0$), a scalar quantity built from independent Gaussian vectors that lower-bounds the minimax objective of the feasibility problem. Its sign decides the issue: positivity of $\varphi_0$ means the sparse-output feasibility problem is infeasible, which in turn means the ReLU network is typically injective. The machinery around it consists of a Gaussian comparison theorem that replaces the random matrices $A^{(i)}$ by Gaussian vectors, a Lagrangian reformulation of the sparse cardinality constraint, and the square-root trick that turns the resulting quadratic terms into one-dimensional Gaussian integrals. The optimized scalar integrals $f_{q,1}$ and $f_{q,2}$ are what the numerical tables actually evaluate to locate the zero of $\varphi_0$.
What would settle it
Simulate a two-layer Gaussian ReLU network near the claimed threshold, for example with $\alpha_1=6.7004$ and $\alpha_2=8.0$, and check whether two distinct finite-dimensional inputs produce the same output with non-negligible probability; finding such collisions would falsify the claimed equivalence. A more direct check is to compute the rank of the matrix in equation (14) for overlapping support sets with $|S_0|=2n$: a positive-probability configuration with rank below $2n$ would break the proof's key step.
Extended reading notes
Core claim
The paper's central claim is that injectivity of an $l$-layer ReLU network is equivalent to infeasibility of an $l$-extended $\ell_0$ spherical perceptron. For two layers, the feasibility problem asks for a unit-norm $x$ satisfying $A^{(1)}x=z$, $A^{(2)}\max(z,0)=t$, and $\|\max(t,0)\|_0<2n$ (or $<n$ for weak injectivity); if no such $x$ exists, the network is typically injective. The paper converts this feasibility question into a minimax random optimization, then uses Gaussian comparison and random duality theory to construct a scalar random-dual value $\varphi_0$ whose positivity implies infeasibility with probability tending to one. Setting $\varphi_0=0$ selects the capacity threshold, and numerical evaluation gives $\alpha_{\mathrm{ReLU}}^{(\mathrm{inj})}=6.7004$ for one layer, $8.267$ for two layers, $9.49$ for three layers, and $10.124$ for four layers, with per-layer expansions $6.70$, $1.23$, $1.15$, and $1.07$. These are upper bounds, because the reversal step needed for exactness is not available.
Load-bearing premise
The entire calculation rests on the assumption that infeasibility of the sparse-output feasibility problem forces injectivity of the multi-layer map, a step that depends on a rank lower bound for the combined matrix which the paper asserts but does not prove for overlapping intermediate support sets.
Editorial extensions
If this is right
- For Gaussian iid weights, a two-layer ReLU network is typically weakly injective at total expansion $\alpha_2=8.267$, meaning the second layer only needs a relative expansion of about $1.234$ beyond the single-layer capacity.
- Adding a third and fourth layer lowers the per-layer expansion to about $1.148$ and $1.067$, so the expansion requirement saturates quickly with depth.
- The equivalence between deep ReLU injectivity and the $l$-extended $\ell_0$ spherical perceptron transfers any future improvement in perceptron capacity bounds directly into improved injectivity bounds for deep networks.
- In the deep generative compressed sensing setting, the reciprocals of these expansion ratios correspond to undersampling ratios that classical non-network methods do not reach, provided the networks generalize and the recovery algorithm runs fast.
Reading between the lines
- The paper leaves implicit that the observed saturation suggests a finite limiting total expansion as $l\to\infty$; if the upper bounds are anywhere near tight, depth alone cannot push the required width much below roughly ten outputs per input.
- A testable extension is to run the equivalence in reverse: search for feasible points of the $l$-extended $\ell_0$ spherical perceptron just below the claimed thresholds, since any such point would directly produce two colliding inputs in the corresponding ReLU network.
- The gap between weak and strong injectivity (for two layers, $8.267$ versus $12.35$) indicates that recovering a fixed generative input is substantially cheaper than guaranteeing worst-case uniqueness, which is the regime most recovery algorithms actually operate in.
- If the unproved rank condition on overlapping intermediate supports can be established, these conditional upper bounds would become certified injectivity guarantees; until then the numerical thresholds rest on that missing step.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies injectivity of deep ReLU networks in the proportional high-dimensional regime and claims an equivalence between l-layer ReLU injectivity and l-extended ℓ0 spherical perceptron feasibility (Eq. (9)). Building on this equivalence, the author develops a random-duality-theory (RDT) framework and reports weak injectivity capacity upper bounds in Table 1: 6.7004, 8.267, 9.49, and 10.124 for 1, 2, 3, and 4 layers, respectively, with per-layer expansions decreasing to about 1.07 by the fourth layer. The paper also introduces a partially lifted RDT variant and reports lowered 2-layer bounds in Tables 8 and 9.
Significance. If the central equivalence and the subsequent random-duality steps were rigorous, the quantitative capacity bounds and the observed expansion-saturation effect would constitute a substantial advance over trivial depth-multiplicative bounds such as 6.7^l. The paper is explicit about the numerical values of all RDT parameters and honestly distinguishes upper bounds from exact capacities; the lifted RDT computation is a useful methodological contribution. However, the load-bearing Lemma 1 is not proven, and the numerical tables inherit this gap. The significance of the work is therefore strictly conditional on repairing the proof of the injectivity-to-feasibility equivalence.
major comments (5)
- [Section 3, Eq. (15)] The support-size inequality |S0| ≤ min(|S1|,|S2|) in Eq. (15) is false in general. S0 indexes nonzero coordinates of the second-layer output, while S1 and S2 index nonzero coordinates of the first-layer output; ReLU can create more nonzero output coordinates than its input has. For example, with n=1, m1=2, A(1)=[1; -1] and x=1, one has S1={1}, but for a generic iid Gaussian A(2) with m2=20, |S0| is approximately 10, far exceeding |S1|=1. Thus Eq. (15) does not follow from the non-degenerative assumption, and the subsequent rank lower bound in Eq. (16) is not established.
- [Section 3, Eqs. (14)–(16)] Even when |S0| ≥ 2n, the rank bound in Eq. (16) is not justified. The matrix in Eq. (14) contains the blocks A(2)_{S0,So} A(1)_{So,:} and -A(2)_{S0,So} A(1)_{So,:}, which are negatives of each other when S1 and S2 overlap, and no genericity argument is supplied to show that the full 2n-column concatenation has rank 2n. Since the existence of a nonzero collision vector [xbar; x] in Eq. (14) is excluded only by this rank bound, the claimed implication from infeasibility of (11) to injectivity is not proven.
- [Section 3, weak injectivity definition after Eq. (17)] For the weak injectivity case actually used in Table 1, infeasibility of (11) with finj=f(w) only guarantees |S0| ≥ n, not |S0| ≥ 2n. The rank bound in Eq. (16) requires the factor 2n, so the proof cannot establish weak injectivity from the stated feasibility problem. The capacity values 8.267, 9.49, and 10.124 in Tables 1, 2, 4, and 6 are therefore unsupported by the argument as written.
- [Section 3, Theorem 1 and Eq. (27)] The proof of Theorem 1 is only a one-line reference to a two-fold application of Gordon's probabilistic comparison theorem. The functional in Eq. (21) contains max(z,0) inside the term y(2)^T A(2) max(z,0), so the Gaussian process is not linear in the optimization variable z; moreover z is coupled to A(1) through the constraint A(1)x=z. The proof does not verify that the comparison theorem applies to the resulting constrained, nonconvex min-max problem, nor does it identify the required Lipschitz or index-set conditions. The implication (φ0>0) ⇒ typical injectivity is therefore not established by the cited theorem.
- [Section 2.2 and Table 1; Eq. (8)] The deep capacity computation takes α1 = 6.7004 as an input, which the author's own reference [82] provides as an RDT-based upper bound (or a statistical-physics prediction). If α1 is not the exact minimally admissible value, then the recursively defined minimally admissible sequence in Eq. (8) is not being used; using a larger-than-minimal first-layer expansion can only make the second-layer injectivity easier, so the resulting values in Table 1 cannot be claimed as upper bounds for the minimally admissible sequence. The paper should state explicitly that the reported numbers are conditional on the exactness of the single-layer value α1=6.7004.
minor comments (5)
- [Section 4.1, Lemma 2 and Eq. (55)] The notation frp(A1:2) in Eq. (55) should be frp(A1:3) to match the three-layer setup.
- [Section 4.1, proof of Theorem 2] In the proof of Theorem 2, the term described as corresponding to A(3) repeats the A(2) expression with max(z,0); it should involve max(t,0) and the appropriate h(3), g(3) variables.
- [Section 3, Eq. (22)] The chain in Eq. (22) equates injectivity with P(F is feasible) → 1; for injectivity one expects infeasibility of the collision problem. The displayed equality appears to have the feasibility direction reversed and should be corrected.
- [Throughout] There are numerous typographical errors, including 'evem', 'od', 'Lipshitzian', 'extened', 'forth' for fourth, and the comma-decimal confusion '3, 68' in Table 6. A careful proofreading pass is needed.
- [Section 3, Eqs. (21) and (26)] The norm notation in the constraint on y(2) is typeset as '|y(2)‖_2 = 1/√n', with a stray pipe; this should be ‖y(2)‖_2 = 1/√n.
Circularity Check
No significant circularity: the deep-layer capacity values are outputs of the RDT optimization, not fits to the single-layer input; the main weakness is a correctness gap in Lemma 1, not circular reasoning.
full rationale
The deep-layer computation legitimately takes the single-layer value alpha1=6.7004 from the author's prior work [82] as an input parameter. This is a normal dependence on a separate result with stated assumptions, not a fitted-input-called-prediction: the RDT parameters (r, gamma-bar, nu, gamma) are optimized within the bound, the multi-layer capacities are solutions of those optimization problems, and the expansion saturation is an emergent output rather than an assumed conclusion. The 'l-extended l0 spherical perceptron' is introduced as an explicit optimization reformulation of the injectivity feasibility problem (eq. (11) to eq. (18)), so the stated equivalence (9) is partly a naming/representation; however, the numerical upper bounds are obtained by solving that same optimization via Gordon comparison and RDT, not by assuming injectivity from an independent perceptron model. The paper also explicitly states that strong random duality is not in place, so the results are honestly labeled as upper bounds, and no uniqueness theorem is imported from the author's earlier work. The most serious issue is a correctness gap in Lemma 1: the inequality |S0| <= min(|S1|,|S2|) in eq. (15) is not established, and the rank bound in eq. (16) therefore does not follow. That is a mathematical error in the proof of the injectivity-to-feasibility equivalence, not a circular reduction of the claimed results to their own inputs. Since every explicit circularity pattern requires a specific reduction by construction or a fitted parameter renamed as a prediction, and none of the paper's core derivations exhibit that, the circularity score is 0.
Assumptions & free parameters
free parameters (2)
- single-layer injectivity capacity alpha1 =
6.7004
- RDT variational parameters (r, gamma-bar1, gamma-bar2, nu, gamma; c3 for lifted RDT) =
e.g., 2-layer weak: r=1.7697, gamma-bar1=0.8935, gamma-bar2=0.9642, nu=0.5560, gamma=0.3078
assumptions (5)
- standard math Gordon's probabilistic comparison theorem
- standard math Concentration of measure for Lipschitz functions of iid Gaussian matrices
- domain assumption Non-degenerative assumption that every subset of rows and columns of A(i) is full rank
- ad hoc to paper Equivalence between infeasibility of (11) and injectivity of the ReLU network
- domain assumption iid standard normal entries of A(i)
Cite this review
Pith. "Pith review of Deep ReLU networks -- injectivity capacity upper bounds." pith.science (2026). https://pith.science/paper/TDL62HZP
@misc{pith2026241219677,
author = {Pith},
title = {Pith review of: Deep ReLU networks -- injectivity capacity upper bounds},
year = {2026},
howpublished = {\url{https://pith.science/paper/TDL62HZP}},
note = {Machine review of arXiv:2412.19677}
}
abstract
We study deep ReLU feed forward neural networks (NN) and their injectivity abilities. The main focus is on \emph{precisely} determining the so-called injectivity capacity. For any given hidden layers architecture, it is defined as the minimal ratio between number of network's outputs and inputs which ensures unique recoverability of the input from a realizable output. A strong recent progress in precisely studying single ReLU layer injectivity properties is here moved to a deep network level. In particular, we develop a program that connects deep $l$-layer net injectivity to an $l$-extension of the $\ell_0$ spherical perceptrons, thereby massively generalizing an isomorphism between studying single layer injectivity and the capacity of the so-called (1-extension) $\ell_0$ spherical perceptrons discussed in [82]. \emph{Random duality theory} (RDT) based machinery is then created and utilized to statistically handle properties of the extended $\ell_0$ spherical perceptrons and implicitly of the deep ReLU NNs. A sizeable set of numerical evaluations is conducted as well to put the entire RDT machinery in practical use. From these we observe a rapidly decreasing tendency in needed layers' expansions, i.e., we observe a rapid \emph{expansion saturation effect}. Only $4$ layers of depth are sufficient to closely approach level of no needed expansion -- a result that fairly closely resembles observations made in practical experiments and that has so far remained completely untouchable by any of the existing mathematical methodologies.
Forward citations
Cited by 1 Pith paper
Reference graph
Works this paper leans on
-
[82]
M. Stojnic. Injectivity capacity of relu gates. 2024. a vailable online at http://arxiv.org/abs/2410. 20646
work page 2024
-
[1]
E. Abbe, S. Li, and A. Sly. Proof of the contiguity conject ure and lognormal limit for the symmetric perceptron. In 62nd IEEE Annual Symposium on Foundations of Computer Scienc e, FOCS 2021, Denver, CO, USA, February 7-10, 2022 , pages 327–338. IEEE, 2021
2021
-
[2]
E. Abbe, S. Li, and A. Sly. Binary perceptron: efficient alg orithms can find solutions in a rare well- connected cluster. In STOC ’22: 54th Annual ACM SIGACT Symposium on Theory of Computi ng, Rome, Italy, June 20 - 24, 2022 , pages 860–873. ACM, 2022
2022
-
[3]
Alweiss, Y
R. Alweiss, Y. P. Liu, and M. Sawhney. Discrepancy minimi zation via a self-balancing walk. In Proc. 53rd STOC, ACM , pages 14–20, 2021
2021
-
[4]
Amelunxen, M
D. Amelunxen, M. Lotz, M. McCoy, and J. Tropp. Living on th e edge: phase transitions in convex programs with random data. Information and Inference: A Journal of the IMA , 3(3):224–294, 2014
2014
-
[5]
B. L. Annesi, E. M. Malatesta, and F. Zamponi. Exact full- RSB SAT/UNSAT transition in infinitely wide two-layer neural networks. 2023. available online at http://arxiv.org/abs/2410.06717
work page Pith review arXiv 2023
-
[6]
S. R. Arridge, P. Maass, O. Öktem, and C. B. Schönlieb. Sol ving inverse problems using data-driven models. Acta Numer. , 28:1–174, 2019
2019
-
[7]
Aubin, W
B. Aubin, W. Perkins, and L. Zdeborova. Storage capacity in symmetric binary perceptrons. J. Phys. A, 52(29):294003, 2019
2019
Show all 91 references
-
[8]
Baldassi, E
C. Baldassi, E. M. Malatesta, and R. Zecchina. Propertie s of the geometry of solutions and capacity of multilayer neural networks with rectified linear unit activ ations. Phys. Rev. Lett. , 123:170602, October 2019
2019
-
[9]
Baldi and S
P. Baldi and S. Venkatesh. Number od stable points for spi n-glasses and neural networks of higher orders. Phys. Rev. Letters , 58(9):913–916, Mar. 1987
1987
-
[10]
Barkai, D
E. Barkai, D. Hansel, and H. Sompolinsky. Broken symmet ries in multilayered perceptrons. Phys. Rev. A, 45(6):4146, March 1992
1992
-
[11]
A. Bora, A. Jalal, E. Price, and A. G. Dimakis. Compresse d sensing using generative models. In Proceedings of the 34th International Conference on Machine L earning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 , volume 70 of Proceedings of Machine Learning Research , ...
2017
-
[12]
S. H. Cameron. Tech-report 60-600. Proceedings of the bionics symposium , pages 197–212, 1960. Wright air development division, Dayton, Ohio
1960
-
[13]
C. Clum. Topics in the Mathematics of Data Science . PhD thesis, The Ohio State University, 2022
2022
-
[14]
T. Cover. Geomretrical and statistical properties of s ystems of linear inequalities with applications in pattern recognition. IEEE Transactions on Electronic Computers , (EC-14):326–334, 1965
1965
-
[15]
Daskalakis, D
C. Daskalakis, D. Rohatgi, and E. Zampetakis. Constant -expansion suffices for compressed sensing with generative priors. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, Dece mber 6-12, 2020,...
2020
-
[16]
M. Dhar, A. Grover, and S. Ermon. Modeling sparse deviat ions for compressed sensing using gener- ative models. In Proceedings of the 35th International Conference on Machine L earning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018 , volume 80 of Proceedings...
2018
-
[17]
D. Donoho. High-dimensional centrally symmetric poly topes with neighborlines proportional to dimen- sion. Disc. Comput. Geometry , 35(4):617–652, 2006. 24
2006
-
[18]
Engel, H
A. Engel, H. M. Kohler, F. Tschepke, H. Vollmayr, and A. Z ippelius. Storage capacity and learning algorithms for two-layer neural networks. Phys. Rev. A , 45(10):7590, May 1992
1992
-
[19]
J. B. Estrach, A. Szlam, and Y. LeCun. Signal recovery fr om pooling representations. In Proceedings of the 31th International Conference on Machine Learning, ICML 201 4, Beijing, China, 21-26 June 2014 , volume 32 of JMLR Workshop and Conference Proceedings , pages 307–315. J...
2014
-
[20]
Fazlyab, A
M. Fazlyab, A. Robey, H. Hassani, M. Morari, and G. J. Pap pas. Efficient and accurate estimation of lipschitz constants for deep neural networks. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2 019, NeurIPS 2...
2019
-
[21]
Fletcher, Sundeep Rangan, and Philip Schnite r
Alyson K. Fletcher, Sundeep Rangan, and Philip Schnite r. Inference in deep networks in high dimen- sions. In 2018 IEEE International Symposium on Information Theory, ISIT 2018, Vail, CO, USA, June 17-22, 2018 , pages 1884–1888. IEEE, 2018
2018
-
[22]
Franz, G
S. Franz, G. Parisi, M. Sevelev, P. Urbani, and F. Zampon i. Universality of the SAT-UNSAT (jamming) threshold in non-convex continuous constraint satisfacti on problems. SciPost Physics , 2:019, 2017
2017
-
[23]
Furuya, M
T. Furuya, M. Puthawala, M. Lassas, and M. V. de Hoop. Glo bally injective and bijective neural operators. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 ...
2023
-
[24]
R. M. Durbin G. J. Mitchison. Bounds on the learning capa city of some multi-layer networks. Biological Cybernetics, 60:345–365, 1989
1989
-
[25]
Kizildag, Will Perkins, and Cha ngji Xu
David Gamarnik, Eren C. Kizildag, Will Perkins, and Cha ngji Xu. Algorithms and barriers in the symmetric binary perceptron model. In 63rd IEEE Annual Symposium on Foundations of Computer Science, FOCS 2022, Denver, CO, USA, October 31 - November 3, 2 022, pages 576–587. IEEE, 2022
2022
-
[26]
E. Gardner. The space of interactions in neural network s models. J. Phys. A: Math. Gen. , 21:257–270, 1988
1988
-
[27]
Gardner and B
E. Gardner and B. Derrida. Optimal storage properties o f neural networks models. J. Phys. A: Math. Gen., 21:271–284, 1988
1988
-
[28]
Y. Gordon. Some inequalities for Gaussian processes an d applications. Israel Journal of Mathematics , 50(4):265–289, 1985
1985
-
[29]
Y. Gordon. On Milman’s inequality and random subspaces which escape through a mesh in Rn. Geo- metric Aspect of of functional analysis, Isr. Semin. 1986-87, L ect. Notes Math , 1317, 1988
1986
-
[30]
H. Gouk, E. Frank, B. Pfahringer, and M. J. Cree. Regular isation of neural networks by enforcing lipschitz continuity. Mach. Learn., 110(2):393–416, 2021
2021
-
[31]
Gutfreund and Y
H. Gutfreund and Y. Stein. Capacity of neural networks w ith discrete synaptic couplings. J. Physics A: Math. Gen , 23:2613, 1990
1990
-
[32]
P. Hand, O. Leong, and V. Voroninski. Phase retrieval un der a generative prior. In Advances in Neural Information Processing Systems 31: Annual Conference on Neur al Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada , pages 9154–9164, 2018
2018
-
[33]
He, C.-K
H. He, C.-K. Wen, and S. Jin. Generalized expectation co nsistent signal recovery for nonlinear measure- ments. In 2017 IEEE International Symposium on Information Theory, ISIT 2017, Aachen, Germany, June 25-30, 2017 , pages 2333–2337. IEEE, 2017
2017
-
[34]
Heckel, W
R. Heckel, W. Huang, P. Hand, and V. Voroninski. Deep den oising: Rate-optimal recovery of structured signals with a deep prior. 2018. available online at http://arxiv.org/abs/1805.08855. 25
2018 arXiv
-
[35]
Hegde, M
C. Hegde, M. B. Wakin, and R. G. Baraniuk. Random project ions for manifold learning. In Advances in Neural Information Processing Systems 20, Proceedings of t he Twenty-First Annual Conference on Neural Information Processing Systems, Vancouver, British Col umbia, Canada, Dec...
2007
-
[36]
Huang, P
W. Huang, P. Hand, R. Heckel, and V. Voroninski. A provab ly convergent scheme for compressive sensing under random generative priors. 2018. available online at http://arxiv.org/abs/1812.04176
2018 arXiv
-
[37]
Jordan and A
M. Jordan and A. G. Dimakis. Exactly computing the local lipschitz constant of relu networks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, v irtual, 2020
2020
-
[38]
R. D. Joseph. The number of orthants in n-space instersected by an s-dimensional subspace. Tech. memo 8, project PARA , 1960. Cornel aeronautical lab., Buffalo, N.Y
1960
-
[39]
Kothari, A
K. Kothari, A. Khorashadizadeh, M. V. de Hoop, and I. Dok manic. Trumpets: Injective flows for inference and inverse problems. In Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence, UAI 2021, Virtual Event, 27-30 July 20 21, volume 161 of Proc...
2021
-
[40]
Krauth and M
W. Krauth and M. Mezard. Storage capacity of memory netw orks with binary couplings. J. Phys. France, 50:3057–3066, 1989
1989
-
[41]
Q. Lei, A. Jalal, I. S. Dhillon, and A. G. Dimakis. Invert ing deep generative models, one layer at a time. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, V ancouver, ...
2019
-
[42]
Louart, Z
C. Louart, Z. Liao, and R. Couillet. A random matrix appr oach to neural networks. CoRR, abs/1702.05419, 2017
2017 arXiv
-
[43]
Maillard, A
A. Maillard, A. S. Bandeira, D. Belius, I. Dokmanic, and S. Nakajima. Injectivity of relu networks: perspectives from statistical physics. 2023. available on line at http://arxiv.org/abs/2302.14112
2023 arXiv
-
[44]
Mardani, Q
M. Mardani, Q. Sun, D. L. Donoho, V. Papyan, H. Monajemi, S. Vasanawala, and J. M. Pauly. Neural proximal gradient descent for compressive imaging. In Advances in Neural Information Processing Sys- tems 31: Annual Conference on Neural Information Processing S ystems 2018, Neur...
2018
-
[45]
D. Paleka. Injectivity of ReLU neural networks at initialization . Master thesis, ETH Zurich, 2021
2021
-
[46]
Pennington and P
J. Pennington and P. Worah. Nonlinear random matrix the ory for deep learning. In Advances in Neural Information Processing Systems 30: Annual Conference on Neur al Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA , pages 2637–2646, 2017
2017
-
[47]
Perkins and C
W. Perkins and C. Xu. Frozen 1-RSB structure of the symme tric Ising perceptron. STOC 2021: Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory o f Computing , pages 1579–1588, 2021
2021
-
[48]
Puthawala, K
M. Puthawala, K. Kothari, M. Lassas, I. Dokmanic, and M. V. de Hoop. Globally injective relu networks. J. Mach. Learn. Res. , 23:105:1–105:55, 2022
2022
-
[49]
Puthawala, M
M. Puthawala, M. Lassas, I. Dokmanic, and M. V. de Hoop. U niversal joint approximation of manifolds and densities by simple injective flows. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA , volume 162 of Proceedings of Mac...
2022
-
[50]
Rangan, P
S. Rangan, P. Schniter, and A. K. Fletcher. Vector appro ximate message passing. In 2017 IEEE International Symposium on Information Theory, ISIT 2017, Aachen, Germany, June 25-30, 2017 , pages 1588–1592. IEEE, 2017. 26
2017
-
[51]
Romano, M
Y. Romano, M. Elad, and P. Milanfar. The little engine th at could: Regularization by denoising (RED). SIAM J. Imaging Sci. , 10(4):1804–1844, 2017
2017
-
[52]
Leigh Ross and J
B. Leigh Ross and J. C. Cresswell. Tractable density est imation on learned manifolds with conformal embedding flows. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, Dece mber 6-14, 2021, vi...
2021
-
[53]
L. Schlafli. Gesammelte Mathematische AbhandLungen I . Basel, Switzerland: Verlag Birkhauser, 1950
1950
-
[54]
Schneider and W
R. Schneider and W. Weil. Stochastic and Integral Geometry . Springer-Verlag Berlin Heidelberg, 2008
2008
-
[55]
Schniter, S
P. Schniter, S. Rangan, and A. K. Fletcher. Vector appro ximate message passing for the generalized linear model. In 50th Asilomar Conference on Signals, Systems and Computers, ACSSC 2016, Pacific Grove, CA, USA, November 6-9, 2016 , pages 1525–1529. IEEE, 2016
2016
-
[56]
Shah and C
V. Shah and C. Hegde. Solving linear inverse problems us ing gan priors: An algorithm with provable guarantees. In 2018 IEEE International Conference on Acoustics, Speech and Si gnal Processing, ICASSP 2018, Calgary, AB, Canada, April 15-20, 2018 , pages 4609–4613. IEEE, 2018
2018
-
[57]
Shcherbina and B
M. Shcherbina and B. Tirozzi. On the volume of the intrer section of a sphere with random half spaces. C. R. Acad. Sci. Paris. Ser I , (334):803–806, 2002
2002
-
[58]
Shcherbina and B
M. Shcherbina and B. Tirozzi. Rigorous solution of the G ardner problem. Comm. on Math. Physics , (234):383–422, 2003
2003
-
[59]
M. Stojnic. Block-length dependent thresholds in bloc k-sparse compressed sensing. available online at http://arxiv.org/abs/0907.3679
-
[60]
M. Stojnic. Various thresholds for ℓ1-optimization in compressed sensing. available online at http:// arxiv.org/abs/0907.3666
-
[61]
M. Stojnic. Block-length dependent thresholds for ℓ2/ℓ1-optimization in block-sparse compressed sens- ing. ICASSP, IEEE International Conference on Acoustics, Signal and Speech Processing, pages 3918– 3921, 14-19 March 2010. Dallas, TX
2010
-
[62]
M. Stojnic. ℓ1 optimization and its various thresholds in compressed sens ing. ICASSP, IEEE Inter- national Conference on Acoustics, Signal and Speech Proces sing, pages 3910–3913, 14-19 March 2010. Dallas, TX
2010
-
[63]
M. Stojnic. Another look at the Gardner problem. 2013. a vailable online at http://arxiv.org/abs/ 1306.3979
2013 arXiv
-
[64]
M. Stojnic. Discrete perceptrons. 2013. available onl ine at http://arxiv.org/abs/1303.4375
2013 arXiv
-
[65]
M. Stojnic. Lifting ℓ1-optimization strong and sectional thresholds. 2013. avai lable online at http:// arxiv.org/abs/1306.3770
2013 arXiv
-
[66]
M. Stojnic. Lifting/lowering Hopfield models ground st ate energies. 2013. available online at http:// arxiv.org/abs/1306.3975
2013 arXiv
-
[67]
M. Stojnic. Meshes that trap random subspaces. 2013. av ailable online at http://arxiv.org/abs/ 1304.0003
2013 arXiv
-
[68]
M. Stojnic. Negative spherical perceptron. 2013. avai lable online at http://arxiv.org/abs/1306. 3980
2013
-
[69]
M. Stojnic. Regularly random duality. 2013. available online at http://arxiv.org/abs/1303.7295
2013 arXiv
-
[70]
M. Stojnic. Spherical perceptron as a storage memory wi th limited errors. 2013. available online at http://arxiv.org/abs/1306.3809. 27
2013 arXiv
-
[71]
M. Stojnic. Fully bilinear generic and lifted random pr ocesses comparisons. 2016. available online at http://arxiv.org/abs/1612.08516
2016 arXiv
-
[72]
M. Stojnic. Generic and lifted probabilistic comparis ons – max replaces minmax. 2016. available online at http://arxiv.org/abs/1612.08506
2016 arXiv
-
[73]
M. Stojnic. Bilinearly indexed random processes – stationarization of fully lifted interpolation. 2023. available online at http://arxiv.org/abs/2311.18097
2023 arXiv
-
[74]
M. Stojnic. Binary perceptrons capacity via fully lift ed random duality theory. 2023. available online at http://arxiv.org/abs/2312.00073
2023 arXiv
-
[75]
M. Stojnic. Capacity of the treelike sign perceptrons n eural networks with one hidden layer – rdt based upper bounds. 2023. available online at http://arxiv.org/abs/2312.08244
2023 arXiv
-
[76]
M. Stojnic. Fl rdt based ultimate lowering of the negati ve spherical perceptron capacity. 2023. available online at http://arxiv.org/abs/2312.16531
2023 arXiv
-
[77]
M. Stojnic. Fully lifted interpolating comparisons of bilinearly indexed random processes. 2023. available online at http://arxiv.org/abs/2311.18092
2023 arXiv
-
[78]
M. Stojnic. Fully lifted random duality theory. 2023. a vailable online at http://arxiv.org/abs/2312. 00070
2023
-
[79]
M. Stojnic. Lifted rdt based capacity analysis of the 1-hidden layer treelike sign perceptrons neural networks. 2023. available online at http://arxiv.org/abs/2312.08257
2023 arXiv
-
[80]
M. Stojnic. Exact capacity of the wide hidden layer treelike neural networks with generic activat ions
-
[81]
M. Stojnic. Fixed width treelike neural networks capac ity analysis – generic activations. 2024. available online at http://arxiv.org/abs/2402.05696
2024 arXiv
-
[83]
Talagrand
M. Talagrand. The Parisi formula. Annals of mathematics , 163(2):221–263, 2006
2006
-
[84]
Talagrand
M. Talagrand. Mean field models and spin glasses: Volume I . A series of modern surveys in mathematics 54, Springer-Verlag, Berlin Heidelberg, 2011
2011
-
[85]
Venkatesh
S. Venkatesh. Epsilon capacity of neural networks. Proc. Conf. on Neural Networks for Computing, Snowbird, UT, 1986
1986
-
[86]
J. G. Wendel. A problem in geometric probablity. Mathematics Scandinavia, 11:109–111, 1962
1962
-
[87]
R. O. Winder. Single stage threshold logic. Switching circuit theory and logical design , pages 321–332, Sep. 1961. AIEE Special publications S-134
1961
-
[88]
R. O. Winder. Threshold logic. Ph. D. dissertation, Princetoin University, 1962
1962
-
[89]
Y. Wu, M. Rosca, and T. P. Lillicrap. Deep compressed sen sing. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Lon g Beach, California, USA , volume 97 of Proceedings of Machine Learning Research , pages 6850–6860. PMLR, 2019
2019
-
[90]
J. A. Zavatone-Veth and C. Pehlevan. Activation functi on dependence of the storage capacity of treelike neural networks. Phys. Rev. E , 103:L020301, February 2021. 28
2021
-
[2024]
available online at http://arxiv.org/abs/2402.05719
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.