REVIEW 2 major objections 6 minor 1 cited by
The Capacity of Quantum Neural Networks
T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A universal bound caps what quantum neural networks can memorize: C ≤ W, with exponential headroom only when parameters are quantum states.
desk verdict The C≤W pigeonhole bound is correct and useful, but the claimed exponential capacity advantage for quantum-parameterized QNNs rests on an unproven trainability assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key object is the pair $(C, W)$ and the inequality $C \leq W$. Capacity $C$ is defined as the largest product $Nm$ for which a learning machine can be trained to reproduce any one of all possible $N$, $m$-bit labellings of $N$ inputs in general position—essentially, the largest random-access memory the machine can emulate. $W$ is the trainable-parameter information content: for classical parameters $W = \sum_i \log_2(M_i)$, while for a quantum parameter state described by a density operator of dimension $D$, $W = m(D^2 - 1)$ at $m$-bit precision. The inequality is proven by the pigeonhole principle (a model requiring $T$ bits to specify needs at least $T$ bits of trainable storage), and the paper uses it to predict capacity scalings such as $W \propto N_w \log_2(N_s)$ for the Gaussian Boson Sampler reservoir, where $N_s$ is the number of measurement samples per expectation value.
What would settle it
Train a small quantum-parameterized QNN (e.g., with a two-qubit parameter state, $D=4$) on random $m$-bit mappings and measure its capacity; if $C$ remains well below $m(D^2-1) = 15m$ bits even with ideal noiseless simulation and a thorough search over training strategies, the exponential-capacity interpretation of $C \leq W$ fails in practice.
Extended reading notes
Core claim
The central discovery is an information-theoretic capacity bound: for any learning machine with trainable parameters, $C \leq W$, where $C$ is the memory capacity (the largest $Nm$ bits of data it can be trained to memorize for $N$ general-position inputs) and $W$ is the amount of information that can be stored in its trainable parameters through training. Classically parameterized QNNs therefore have $C$ bounded by the same formula as classical NNs with equal parameter count and precision. If the trainable parameters are quantum states, however, a density operator on a Hilbert space of dimension $D$ carries $W = m(D^2 - 1)$ bits at $m$-bit precision, so the capacity can grow exponentially with the number of quantum systems (modes or qubits) used for parameterization—potentially on both classical and quantum tasks. The paper proves the inequality by the pigeonhole principle and verifies the scaling behavior numerically with a 5-mode Gaussian Boson Sampler reservoir computer, showing that capacity saturates with measurement samples and that generalization occurs when training data exceeds $C$.
Load-bearing premise
The exponential capacity advantage of quantum-parameterized QNNs depends on the assumption that each independent matrix element of the parameter density operator can really be trained to $m$-bit precision; the paper does not exhibit a training procedure that achieves such precision, and if only limited precision is reachable the advantage shrinks.
Editorial extensions
If this is right
- Classically parameterized QNNs (parameters are classical numbers adjusted by classical training) have $C$ bounded exactly like classical NNs of the same size and precision; the quantum circuit itself confers no capacity advantage.
- Quantum parameterization—e.g., a trainable density operator $\hat{\rho}_w$—can in principle give exponentially larger $C$, since $W = m(D^2 - 1)$ grows quadratically with Hilbert-space dimension $D$.
- The bound gives a generalization criterion: when training data exceeds $C$, memorization is impossible and the network must learn a compressed model, which is when generalization begins.
- Measurement and sampling noise reduce capacity logarithmically: in the GBS-fQRC, $C$ increases linearly with parameter count but only as $\log_2(N_s)$ with measurement samples, so suppressing readout noise is a direct way to raise capacity.
- The capacity bound is universal across classical or quantum data, training methods, and hybrid schemes, so it can serve as a common metric for comparing QNN architectures.
Reading between the lines
- The paper's own examples leave open whether the exponential capacity promised by quantum parameter states is actually trainable; a natural next step is to construct and train a small quantum-parameterized QNN and measure whether $C$ approaches $m(D^2 - 1)$ in practice.
- If $C \leq W$ transfers to other complexity notions such as VC dimension or Rademacher complexity, quantum-parameterized networks may require correspondingly more training data to generalize, offsetting part of the capacity advantage.
- The fQRC demonstration suggests that even without a capacity advantage, quantum reservoir computers could still win on speed or on handling quantum data directly, since they avoid the need to perform tomography on inputs.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces an information-theoretic memory-capacity measure C for quantum neural networks (QNNs), defined as the largest number N of m-bit labels that the network can learn to assign to N general-position inputs, and a companion quantity W, the number of bits that can be stored in the trainable parameters. The central result is the inequality C≤W, proved by a pigeonhole/counting argument. The paper applies this bound to classically parameterized QNNs, concluding that they offer no capacity advantage over classical networks with equally precise parameters, and to quantum-parameterized QNNs, where it suggests W=m(D^2−1) and hence a possible exponential capacity advantage. The theoretical part is illustrated with numerical simulations of a feed-forward reservoir computer based on a Gaussian Boson Sampler, including capacity estimates, function-approximation tasks, and a study of the effect of finite sampling of expectation values on capacity and generalization.
Significance. If the central bound and its interpretation hold, the paper provides a simple and general organizing principle for QNN expressiveness: capacity is limited by the information that training can write into the parameters, independent of whether the computation is quantum or classical. The proof of C≤W is elementary and correct, and the paper deserves credit for stating it cleanly and for testing its consequences numerically on a concrete GBS-based architecture, with simulations in Strawberry Fields and QuTiP. The numerical demonstrations that capacity grows with the number of trained parameters and with sampling precision, and the generalization threshold near T≈C, are useful illustrations. However, the headline claim that quantum-parameterized QNNs could have exponentially larger capacities rests on an additional premise about trainability of the full density-matrix parameter space; this premise is not proven and is not supported by the fQRC example, which only trains classical output weights. The paper would be an important contribution if that gap is closed or if the claim is explicitly reframed as a conjecture.
major comments (2)
- [Main text, Eq. (1) and the paragraph 'The interpretation of Eqn. 1 in the case of quantum parameterization...'] The step in the paragraph following Eq. (1) that sets W=m(D^2−1) for a trainable quantum parameter state is not justified. Equation (1) alone only says C is bounded by the information that the training procedure can store in the parameters; it does not imply that every independent real matrix element of a general D-dimensional density operator can be trained to m-bit precision. Trace and positivity constraints restrict the reachable set, and a generic training rule (e.g., parameterized unitary gates acting on an ancilla) reaches only a submanifold of state space. The fQRC example trains only the classical matrix Wout, so it provides no demonstration that a quantum-parameterized QNN can saturate Eq. (1) with W=m(D^2−1). For the advertised exponential advantage to follow, the authors should either exhibit a concrete training protocol whose reachable set is an m-bit-resolution net over the D^2−1-dimensional parameter space, or state the exponential-advantage claim as a conjecture contingent on such trainability.
- [Supplemental Material, 'Achieving the capacity upper bound with fQRC'] The assertion that the fQRC 'can theoretically achieve C=W' is not proven. The argument requires that for every set of N≤Nw distinct general-position inputs, the N×Nw feature matrix Rout has rank N. The text justifies this by saying the chosen expectation values are linearly independent transformations because the operators cannot be written as linear combinations of one another; this is insufficient. Linear independence of functions on the whole input space does not imply linear independence of their evaluations at a finite set of points, which is the condition needed for exact inversion for arbitrary labels. Please provide a rigorous proof, or a precise genericity statement with proof, that the GBS feature vectors are in general position, or restrict the claim accordingly.
minor comments (6)
- [Supplemental Material, 'Estimation of capacity'] Equation numbering is reused: the Supplemental Material's capacity estimator is also labeled Eq. (1), which conflicts with Eq. (1) of the main text; please renumber the SM equations.
- [Conclusion] The phrase 'exponential decrease in the capacity' is not consistent with the stated scaling W∝Nw log2(Ns); the capacity decreases logarithmically with Ns, or equivalently the number of samples needed to maintain a given capacity grows exponentially with the capacity. Please rephrase.
- [Main text, 'Definitions' and quantum-parameterization paragraph] The notion of m-bit precision for density matrices and unitaries ('each matrix element correct to within m-bit precision') is informal; please specify a norm or a formal elementwise error model and state how trace and positivity constraints are handled.
- [Main text, sampling discussion around Fig. 2(c)-(d)] For the stochastic version of the QNN with Ns simulated measurements, the meaning of 'learning a labeling' is not made precise; a probabilistic model should specify the allowed error probability or confidence threshold.
- [Fig. 2 and related SM discussion] The capacity estimates are based on only 10-100 random labellings, yet the plotted curves show no error bars or confidence intervals; adding variability estimates would make the quantitative claims easier to assess.
- [Main text, footnote 72] The footnote argues that capacity-limited QNNs offer a natural advantage on quantum data because tomography is avoided; this is a sample-complexity argument rather than a capacity argument, and the distinction should be stated explicitly.
Circularity Check
No significant circularity: the central capacity bound is a pigeonhole/counting inequality, and the numerical simulations are illustrative rather than fitted inputs masquerading as predictions.
full rationale
The paper's central claim C≤W is derived as a counting/pigeonhole bound, not as a restatement of its own conclusion. C is defined as the largest Nm-bit random labeling task a learning machine can memorize, while W is given an operational definition as the number of distinct training-settable parameter values, e.g. W=Σ log2(Mi) for classical parameters. The proof argues that exactly learning a T-bit random map requires the trainable degrees of freedom to store at least T bits, so C cannot exceed W; this is an independent information-theoretic argument rather than a definitional equality. The quantum extension W=m(D^2−1) is presented as an interpretive assumption about counting the information in a trainable density operator, and the paper explicitly hedges the resulting advantage as one that 'could' occur; even if this assumption is unproven or optimistic, it is not circular because it is not obtained by substituting the target result into itself. The GBS-fQRC simulations estimate capacity from training error and compare it with W computed from parameter precision; this is a consistency check and illustration, not a fitted parameter being renamed as a prediction. The supplemental proof that fQRC can reach C=W rests on a linear-algebra rank condition over general-position inputs; the rank assertion is stated rather than fully demonstrated, but that is a rigor gap, not circular reasoning. There are no load-bearing self-citations or imported uniqueness theorems: the cited classical capacity literature provides background, and the related preprints mentioned in the note are not used to justify the main inequality. The derivation chain is therefore self-contained with respect to circularity, while the realizability of an exponential quantum-parameter capacity advantage depends on an unproven trainability assumption that belongs to correctness risk, not to circularity.
Assumptions & free parameters
free parameters (1)
- numerical precision cutoff epsilon_p =
~45 bits (log2(1/1e-14))
assumptions (4)
- standard math Pigeonhole principle: a machine implementing T bits of distinguishable behaviors must have at least T bits of configurable state.
- domain assumption The trained behavior of a learning machine is determined entirely by its trainable parameters and the input.
- domain assumption General position of inputs: N distinct random inputs drawn uniformly.
- standard math Density matrices of dimension D require D^2−1 real parameters.
Cite this review
Pith. "Pith review of The Capacity of Quantum Neural Networks." pith.science (2026). https://pith.science/paper/UC5W6BMG
@misc{pith2026190801364,
author = {Pith},
title = {Pith review of: The Capacity of Quantum Neural Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/UC5W6BMG}},
note = {Machine review of arXiv:1908.01364}
}
abstract
A key open question in quantum computation is what advantages quantum neural networks (QNNs) may have over classical neural networks (NNs), and in what situations these advantages may transpire. Here we address this question by studying the memory capacity $C$ of QNNs, which is a metric of the expressive power of a QNN that we have adapted from classical NN theory. We present a capacity inequality showing that the capacity of a QNN is bounded by the information $W$ that can be trained into its parameters: $C \leq W$. One consequence of this bound is that QNNs that are parameterized classically do not show an advantage in capacity over classical NNs having an equal number of parameters. However, QNNs that are parametrized with quantum states could have exponentially larger capacities. We illustrate our theoretical results with numerical experiments by simulating a particular QNN based on a Gaussian Boson Sampler. We also study the influence of sampling due to wavefunction collapse during operation of the QNN, and provide an analytical expression connecting the capacity to the number of times the quantum system is measured.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Quantum Computational-Sensing Advantage
A perspective defines quantum computational sensing (QCS) and its advantage (QCSA), and organizes many recent sensing-plus-computing protocols into a single taxonomy.
Reference graph
Works this paper leans on
-
[1]
School of Applied and Engineering Physics, Cornell University, Ithaca, NY 14853, USA and
-
[2]
E. L. Ginzton Laboratory, Stanford University, Stanford, CA 94305, USA A key open question in quantum computation is what advantages quantum neural networks (QNNs) may have over classical neural networks (NNs), and in what situations these advantages may transpire. Here we address this question by studying the memory capacity C of QNNs, which is a metric ...
arXiv 1908
-
[3]
E. Behrman, J. Niemel, J. Steck, and S. Skinner, in Pro- ceedings of the 4th Workshop on Physics of Computation (1996)
work page 1996
-
[4]
T. Menneer and A. Narayanan, Department of Computer Science, University of Exeter, Exeter, United Kingdom, Technical Report 329 (1995)
work page 1995
-
[5]
S. C. Kak, Advances in Imaging and Electron Physics (1995)
work page 1995
-
[6]
Peruzzo, J
A. Peruzzo, J. McClean, P. Shadbolt, M.-H. Yung, X.-Q. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’Brien, Nature Communications 5, 4213 (2014)
2014
- [7]
- [8]
Show all 77 references
-
[9]
Schuld, I
M. Schuld, I. Sinayskiy, and F. Petruccione, Quantum Information Processing 13, 2567 (2016)
2016
-
[10]
J. R. McClean, J. Romero, R. Babbush, and A. Aspuru- Guzik, New Journal of Physics 18, 023023 (2016)
2016
- [11]
-
[12]
Killoran, T
N. Killoran, T. R. Bromley, J. M. Arrazola, M. Schuld, N. Quesada, and S. Lloyd, arXiv:1806.06871
-
[13]
K. H. Wan, O. Dahlsten, H. Kristj´ ansson, R. Gardner, and M. S. Kim, npj Quantum Information 3, 36 (2017)
2017
-
[14]
Rebentrost, T
P. Rebentrost, T. R. Bromley, C. Weedbrook, and S. Lloyd, Physical Review A 98, 042308 (2018)
2018
-
[15]
Tacchino, C
F. Tacchino, C. Macchiavello, D. Gerace, and D. Bajoni, npj Quantum Information 5, 26 (2019)
2019
-
[16]
G. R. Steinbrecher, J. P. Olson, D. Englund, and J. Car- olan, npj Quantum Information 5, 60 (2019)
2019
-
[17]
K. Beer, D. Bondarenko, T. Farrelly, T. J. Osborne, R. Salzmann, and R. Wolf, arXiv:1902.10445
1902 arXiv
- [18]
-
[19]
I. Cong, S. Choi, and M. D. Lukin, arXiv:1810.03787
-
[20]
Carolan, M
J. Carolan, M. Mosheni, J. P. Olson, M. Prabhu, C. Chen, D. Bunandar, N. C. Harris, F. N. C. Wong, M. Hochberg, S. Lloyd, and D. Englund, arXiv:1904.10463
1904 arXiv
-
[21]
S. H. Adachi and M. P. Henderson, arXiv:1510.06356
-
[22]
Benedetti, J
M. Benedetti, J. Realpe-G´ omez, R. Biswas, and A. Perdomo-Ortiz, Physical Review A94, 022308 (2016)
2016
- [23]
-
[24]
C. M. Wilson, J. S. Otterbach, N. Tezak, R. S. Smith, G. E. Crooks, and M. P. da Silva, arXiv:1806.08321
-
[25]
M. H. Amin, E. Andriyash, J. Rolfe, B. Kulchytskyy, and R. Melko, Physical Review X 8, 021050 (2018)
2018
-
[26]
Mitarai, M
K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii, Phys- ical Review A 98, 032309 (2018)
2018
-
[27]
Havl´ ıˇ cek, A
V. Havl´ ıˇ cek, A. D. C´ orcoles, K. Temme, A. W. Harrow, A. Kandala, J. M. Chow, and J. M. Gambetta, Nature 567, 209 (2019)
2019
- [28]
-
[29]
Schuld and N
M. Schuld and N. Killoran, Physical Review Letters 122, 6 040504 (2019)
2019
- [30]
-
[31]
Ventura and T
D. Ventura and T. Martinez, arXiv:9807053
- [32]
- [33]
- [34]
-
[35]
J. D. Biamonte and P. J. Love, Physical Review A 78, 012352 (2008)
2008
-
[36]
Bergholm, J
V. Bergholm, J. Izaac, M. Schuld, C. Gogolin, C. Blank, K. McKiernan, and N. Killoran, arXiv:1811.04968
-
[37]
Preskill, Quantum 2, 79 (2018)
J. Preskill, Quantum 2, 79 (2018)
2018
-
[38]
K. M. Nakanishi, K. Fujii, and S. Todo, arXiv:1903.12166
1903 arXiv
-
[39]
R. M. Parrish, E. G. Hohenstein, P. L. McMahon, and T. J. Martinez, arXiv:1906.08728
1906 arXiv
-
[40]
Schuld, V
M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, and N. Killoran, Physical Review A 99, 032331 (2019)
2019
-
[41]
J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, Nature Communications 9, 4812 (2018)
2018
- [42]
-
[43]
V. N. Vapnik and A. Y. Chervonenkis, in Measures of Complexity (Springer International Publishing, Cham,
-
[44]
D. J. C. Mackay, Information Theory, Inference, and Learning Algorithms (Learning, 2005)
2005
-
[45]
Baldi and R
P. Baldi and R. Vershynin, Advances in Neural Informa- tion Processing Systems, 7729-7738 (2018)
2018
- [46]
- [47]
-
[48]
T. M. Cover, IEEE Transactions on Electronic Comput- ers 3, 326 (1965)
1965
- [49]
- [50]
-
[51]
Shalev-Shwartz and S
S. Shalev-Shwartz and S. Ben-David, Understanding ma- chine learning: From theory to algorithms (Cambridge University Press, 2013)
2013
-
[52]
Gardner, Europhysics Letters 4, 481 (1987)
E. Gardner, Europhysics Letters 4, 481 (1987)
1987
-
[53]
Kinzel, Philosophical Magazine B 77, 1455 (1998)
W. Kinzel, Philosophical Magazine B 77, 1455 (1998)
1998
-
[54]
Recent works have ex- plored quantum reservoir computers (QRC) [55–60] that adapt the concept of reservoir computing [61, 62] to quan- tum dynamical systems
in the classical literature. Recent works have ex- plored quantum reservoir computers (QRC) [55–60] that adapt the concept of reservoir computing [61, 62] to quan- tum dynamical systems. The GBS-fQRC simplifies this approach by considering a feed-forward device, imple- mented w...
-
[55]
V. N. Vapnik, The nature of statistical learning theory (Springer, 2000)
2000
-
[56]
See the attached Supplemental Material
-
[57]
G. B. Huang, Q. Y. Zhu, and C. K. Siew, in IEEE In- ternational Conference on Neural Networks - Conference Proceedings (2004)
2004
-
[58]
Fujii and K
K. Fujii and K. Nakajima, Physical Review Applied 8, 024030 (2017)
2017
- [59]
- [60]
- [61]
-
[62]
Ghosh, A
S. Ghosh, A. Opala, M. Matuszewski, T. Paterek, and T. C. H. Liew, npj Quantum Information 5, 35 (2019)
2019
- [63]
-
[64]
Jaeger, in GMD-Forschungszentrum Information- stechnik Report 148 (2001)
H. Jaeger, in GMD-Forschungszentrum Information- stechnik Report 148 (2001)
2001
-
[65]
Maass, T
W. Maass, T. Natschl¨ ager, and H. Markram, Neural Computation 14, 2531 (2002)
2002
-
[66]
Aaronson and A
S. Aaronson and A. Arkhipov, in Proceedings of the 43rd annual ACM symposium on Theory of computing - STOC ’11 (ACM Press, New York, New York, USA, 2011) p. 333
2011
-
[67]
A. Lund, A. Laing, S. Rahimi-Keshari, T. Rudolph, J. O’Brien, and T. Ralph, Physical Review Letters 113, 100502 (2014)
2014
-
[68]
C. S. Hamilton, R. Kruse, L. Sansoni, S. Barkhofen, C. Silberhorn, and I. Jex, Physical Review Letters 119, 170501 (2017)
2017
-
[69]
J. Huh, G. G. Guerreschi, B. Peropadre, J. R. McClean, and A. Aspuru-Guzik, Nature Photonics 9, 615 (2015)
2015
-
[70]
Br´ adler, P.-L
K. Br´ adler, P.-L. Dallaire-Demers, P. Rebentrost, D. Su, and C. Weedbrook, Physical Review A98, 032310 (2018)
2018
-
[71]
J. M. Arrazola, T. R. Bromley, and P. Rebentrost, Phys- ical Review A 98, 012322 (2018)
2018
- [72]
-
[73]
Killoran, J
N. Killoran, J. Izaac, N. Quesada, V. Bergholm, M. Amy, and C. Weedbrook, arXiv:1804.03159
-
[74]
Johansson, P
J. Johansson, P. Nation, and F. Nori, Computer Physics Communications 184, 1234 (2013)
2013
-
[75]
For example, even capacity-limited QNNs offer a natural advantage over classical NNs for operations on quantum data, since using a QNN avoids needing to first perform tomography to determine the classical representation of the input quantum state
-
[76]
Banchi and S
L. Banchi and S. Pirandola, in Quantum Information and Measurement (QIM) V: Quantum Technologies (OSA, Washington, D.C., 2019) p. S4B.5
2019
-
[77]
Dolzhkov, B
E. Dolzhkov, B. Ghandchi, and D. O. Theis, arXiv:1903.12611. The Capacity of Quantum Neural Networks: Supplemental Material Detailed description of the GBS-fQRC Our GBS-fQRC reservoir circuit consists of M input squeezers, followed by a random M-mode interferometer and finally,...
1903 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.