REVIEW 2 major objections 5 minor 66 references
A weighted mixture of ancilla-free IQP circuits, the mixed IQP-QCBM, is locally trainable with polynomially many branches and can beat any single ancilla-free circuit only when its branches generate distinct output distributions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-07-31 23:41 UTC pith:IJGLTQXY
load-bearing objection Clean unconditional trainability result for mixed IQP-QCBMs plus a conditional extension that rests on an assumption the paper's own diagnostics do not validate at scale; honest numerics, worth engaging seriously. the 2 major comments →
Trainability and Mode Separation of Mixed IQP-QCBMs
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central object is the Walsh-Hadamard branch decomposition: a compiled IQP circuit with a ancillas has a data marginal that is exactly a uniform mixture of L=2^a ancilla-free IQP circuits, generalized to weights pi_l. For this mixed IQP-QCBM the paper proves local barren-plateau avoidance from the data-agnostic center (Theorem 1): the loss curvature for every one-body branch angle is c = 8w_j/L^2 = Omega(1/poly(n)), so the loss variance stays inverse-polynomial on an inverse-polynomial patch. Under Assumptions 1 and 2, approximately factorizable groups and a noncollapsed well-weighted group, the global and cluster centers enjoy the same guarantee (Theorem 2), with curvature at least 8w_{j
What carries the argument
The machinery is the branch decomposition and the curvature identity. Tracing out the ancillas of an (n+a)-qubit compiled IQP circuit yields a uniform mixture of L=2^a order-2 IQP circuits, each with its own angle vector; the Walsh-Hadamard transform (theta_G = H theta-tilde_G, with H^2 = L I) converts between branch angles and compiled ancilla-coupling angles with condition number 1. The loss is the low-body MMD, a weighted sum over Pauli-Z correlator mismatches L(Theta) = sum_A w_A (z_A(Theta) - t_A)^2, and its second derivative with respect to one branch angle splits into a nonnegative model-sensitivity term and a signed data-mismatch term. A curvature-to-variance lemma converts a center
Load-bearing premise
The data-dependent trainability guarantee hinges on Assumption 1 — that within each assigned cluster the higher-body correlators decay like (C/n)^{|A|/2} for an n-independent constant C — which the paper's finite-size diagnostics (C_n = 11.9-14.3 at n=16) do not establish for growing n; if real grouped data violates this, the data-dependent half of the headline result has no proven instances.
What would settle it
Take a clustered dataset family with growing n (e.g., binarized MNIST or an Ising model at increasing lattice size) and compute the effective constant C_n = n * max_{l,A} |t_A^{(l)} - prod_{j in A} t_j^{(l)}|^{2/|A|}. If C_n is unbounded, Assumption 1 fails and Theorem 2 has no guarantee; if instead C_n stays bounded while the data-agnostic curvature still scales as 8w_j/L^2 across n, the trainability claim is supported. Alternatively, test Theorem 3: measure the branch-separating gradient at finite spread delta; if it grows faster than linearly, or fails to vanish at coincidence for some comp
If this is right
- With L=O(poly(n)) branches, the data-agnostic start is locally free of barren plateaus for any target; the exact curvature 8w_j/L^2 quantifies the 1/L^2 cost of adding branches.
- Under Assumptions 1 and 2, global and cluster starts inherit a local trainability certificate, so data-dependent initialization remains viable for mixtures.
- A mixture cannot outperform the best ancilla-free IQP circuit on the same graph unless D_branch >= Delta_anc-free - sqrt(L(Theta)) > 0; coincident branches are provably stuck at the ancilla-free floor.
- Because branch-separating gradients vanish at coincidence, deterministic gradient descent from a coincident start cannot create branch diversity; a perturbation or a cluster-based start is required.
- Cluster initialization gave the lowest or tied-lowest test MMD^2 on all four benchmarks, with branches specializing to blob patterns, magnetization sectors, and digit shapes.
Where Pith is reading between the lines
- The 1/L^2 curvature suppression suggests a trainability-expressivity trade-off: if L grows faster than polynomial, the data-agnostic patch shrinks and the local guarantee disappears; one could test whether adaptive initialization could restore the curvature.
- Assumption 1 is only diagnosed at fixed sizes; if C_n grows with n on realistic clustered data, the data-dependent theorem loses its instances, so a scaling test on larger MNIST or spin-glass clusters would directly bound the theorem's domain.
- The Walsh-Hadamard mixture structure is generic, so the same branch-decomposition and diversity-floor argument likely carries to other classically trainable quantum generative families, making mode separation a design principle rather than a special feature.
- The MNIST result that global initialization can also specialize suggests that a dynamic diversity-promoting objective, which the paper floats, might be a more robust route than a fixed partition.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the mixed IQP-QCBM, a weighted mixture of ancilla-free IQP circuits realized as an ancilla-extended compiled IQP circuit, and analyzes its local trainability and expressivity. The main theoretical results are: (i) Theorem 1, an exact closed-form curvature c = 8w_{j}/L^2 at the data-agnostic center, giving an inverse-polynomial loss-variance lower bound and hence local barren-plateau avoidance for L = poly(n) branches; (ii) Theorem 2, a conditional lower bound on the curvature at the global and cluster data-dependent centers under Assumptions 1 and 2, proved in Appendix B; (iii) Theorem 3, showing that branch-distinguishing (ancilla-coupling) gradients vanish at branch coincidence and grow at most linearly with branch spread; and (iv) Proposition 2, a diversity floor stating that improving below the best ancilla-free MMD requires positive branch diversity. The paper also reports exact n=16 checks of the curvature and gradient predictions, and numerical benchmarks on four datasets showing that cluster initialization converges fastest and reaches the lowest or comparable test MMD^2, with branches specializing to distinct data modes.
Significance. If the results hold, the paper makes a substantial contribution to the trainability analysis of quantum generative models. Theorem 1 is a clean, unconditional, parameter-free closed-form prediction that is verified by exact evaluation at n=16. Theorem 3 identifies an explicit Walsh-orthogonality mechanism that suppresses branch-separating gradients at coincidence, and the numerical study provides a concrete demonstration of mode separation, connecting branch diversity to test performance. The paper is honest about the conditional status of Theorem 2 and about the limitations of the low-body MMD. The main weakness is that the data-dependent trainability guarantee rests on an unverified assumption about the grouped data, which limits the breadth of the central claim.
major comments (2)
- [Sec. III C / Appendix B5 / Eq. (B20)] Assumption 1 is load-bearing for Theorem 2 and is not established. The proof bounds every higher-body mismatch term by sum_ell pi_ell (C/n)^{|A|/2}; without an n-independent C, the O(1/n^3) remainder in Eqs. (24) and (B15) is uncontrolled. The paper's own finite-size diagnostics at n=16 give C_n = 11.9-14.3 for the blobs and Ising, and state that these 'do not establish the n-independent constant required by Assumption 1.' For MNIST and D-Wave, only median/95th-percentile low-body residuals are measured, not the uniform maximum over all subsets and groups that the assumption requires. Consequently, the data-dependent (global and cluster) trainability guarantee, one of the two headline theorems, is unsupported for the actual benchmarks as n grows. Please provide scaling tests that estimate C_n at several n, prove Assumption 1 for the data families used, or explicitly restrict Theorem 2 to
- [Sec. II E, Sec. V A / Corollary 1] The coincidence-breaking perturbation used in the experiments is not certified to lie inside the inverse-polynomial patch on which Theorem 1's variance bound holds. The paper states in Sec. III B and Corollary 1 that the admissible radius 'does not certify that the finite run perturbation lies inside it.' Since the reported training runs start from these perturbed angles, the numerical demonstrations do not directly validate the local barren-plateau guarantee at the exact starting points used. Please either choose the perturbation radius within the certified value, or state clearly that the theorem applies to the exact center and that the experiments test the center curvature and gradient behavior only.
minor comments (5)
- [Abstract] Typo: 'we callmode separation' should read 'we call mode separation'.
- [Fig. 2(b) / Appendix D5] The quantity c_lead is used in the main text but defined only in Appendix D5. Define it in the main text or in the figure caption for readability.
- [Sec. III and Appendix B] Theorems 2 and 3 are stated as informal in the main text with formal versions deferred to the appendices. For a journal, consider stating the formal versions (or at least the exact assumptions) in the main text, since the informal statements omit the partition-identity condition and the precise regularity hypotheses.
- [Appendix B5] The finite-size diagnostics would be considerably more informative if they reported C_n at multiple system sizes (e.g., n=16, 32, 64) rather than at a single n. As written, a single size cannot provide evidence about the asymptotic behavior required by Assumption 1.
- [Fig. 3(e)-(h)] The RBM reference is labeled 'best classical RBM'; since the comparison is explicitly not a controlled head-to-head, consider labeling it 'RBM (validation-selected)' to avoid implying a state-of-the-art baseline.
Circularity Check
No significant circularity: the main trainability and expressivity results are exact closed-form derivations or honestly conditional theorems, with no fitted parameter renamed as a prediction and no load-bearing self-citation.
full rationale
The derivation chain is self-contained in the relevant sense. Theorem 1's curvature c = 8w_{j}/L^2 (Eq. 23) is an exact evaluation of the general curvature identity (Eq. 22/B3) at the data-agnostic center: only A={j} survives because all nontrivial branch correlators vanish there, and the n=16 exact evaluation is a genuine verification rather than a fit. Theorem 2 is explicitly conditional on Assumptions 1 and 2; the proof (Appendix B 4, Eq. B20) uses Assumption 1 to bound every higher-body mismatch term by (C/n)^{|A|/2}, and Assumption 2 to make the one-body term dominate. The paper's own Appendix B 5 states that the measured C_n = 11.9-14.3 at n=16 'do not establish the n-independent constant required by Assumption 1.' This is an honest, clearly located limitation of the data-dependent guarantee, not a circular reduction: Theorem 2 remains a conditional theorem whose premise is unverified for growing n. The curvature-to-variance bridge, Lemma 1, is 'Theorem 2 of Ref. [13], invoked here verbatim rather than re-derived'; that source is external (Lerch et al.) and does not overlap with the present authors, and the paper's mixture-specific contribution is checking that the regularity constants remain polynomially bounded. Proposition 2 is a triangle-inequality bound following from the definition of D_branch, not a fit masquerading as a prediction. Theorem 3's gradient suppression follows from Walsh orthogonality of the compiled angles. I find no step where a claimed prediction is equivalent by construction to its input, and no self-citation chain is load-bearing. The correct concern is empirical/correctness risk about Assumption 1's scaling, which belongs outside the circularity score.
Axiom & Free-Parameter Ledger
free parameters (4)
- MMD kernel bandwidth sigma (m_bar = 2, 6) =
sigma = 1.3183, 0.6006 (n=16); 9.8868, 5.6935 (MNIST); 7.7621, 4.4627 (D-Wave)
- Spectral clustering bandwidth sigma_c =
sqrt(d_med/2): 1.732 (blobs), 2.000 (Ising), 8.185 (MNIST), 10.909 (D-Wave)
- Coincidence-breaking perturbation radius r_0 =
pi/(2*sqrt(m)) per angle
- Branch weights pi_ell =
cluster masses |C_ell|/N (Table VI); 1/L for global/agnostic
axioms (5)
- domain assumption Lemma 1 = Theorem 2 of Ref. [13]: a center curvature c=Omega(1/poly(n)) implies inverse-polynomial loss variance on an inverse-polynomial patch
- domain assumption Loss-variance concentration is taken as the barren-plateau criterion, via equivalence to gradient variance (Refs. [12, 34])
- ad hoc to paper Assumption 1: approximately factorizable groups, |t_A^ell - Pi_{j in A} t_j^ell| <= (C/n)^{|A|/2}
- ad hoc to paper Assumption 2: a noncollapsed, well-weighted group (pi_ell* = omega(1/n), 1-(t_j*^ell*)^2 = Theta(1))
- standard math IQP correlator identity d^2 z_A / d theta_G^2 = -4 z_A for G.A=1, and the randomized correlator estimator (Eqs. A5-A7)
read the original abstract
Quantum circuit Born machines (QCBMs) based on instantaneous quantum polynomial-time (IQP) circuits are promising quantum generative models for their classical trainability. It is known that their ancilla-free form avoids barren plateaus under certain initializations, but remains non-universal. Although adding ancilla qubits raises the expressivity, whether the ancilla-extended model retains local trainability remains unknown. We propose the mixed IQP-QCBM, which generalizes the ancilla-extended circuit as a weighted mixture of ancilla-free IQP circuits, called branches. For a polynomial number of branches, we prove local barren-plateau avoidance from data-agnostic and, under certain assumptions, data-dependent initializations. We further show that the mixed IQP-QCBM can surpass the best ancilla-free IQP circuit only if its branches generate a number of distinct distributions. In particular, we focus on a behavior we call \emph{mode separation}, in which each branch captures a particular feature of the target. Mode separation is hard to attain from an initialization whose branches generate the same distribution: the gradients that would separate them are suppressed while the distributions they generate remain close. This motivates \emph{cluster initialization}, which assigns a different unsupervised data cluster to each branch and provides an initial degree of mode separation. Exact calculations on two 16-bit datasets support the barren-plateau and gradient-suppression claims. On four benchmarks, binary clusters, a two-dimensional Ising model, binarized MNIST, and a 484-spin glass, cluster initialization converges fastest and reaches the lowest mean test $\mathrm{MMD}^2$. We observe that, when achieving the lowest test $\mathrm{MMD}^2$, the mixed IQP-QCBM contains branches specialized to distinguishable data features such as blob patterns, magnetization sectors, or digit shapes.
Figures
Reference graph
Works this paper leans on
-
[1]
Benedetti, D
M. Benedetti, D. Garcia-Pintos, O. Perdomo, V. Leyton- Ortega, Y. Nam, and A. Perdomo-Ortiz, A generative modeling approach for benchmarking and training shal- low quantum circuits, npj Quantum Information5, 45 (2019)
2019
-
[2]
Liu and L
J.-G. Liu and L. Wang, Differentiable learning of quan- tum circuit born machines, Phys. Rev. A98, 062324 (2018)
2018
-
[3]
Coyle, D
B. Coyle, D. Mills, V. Danos, and E. Kashefi, The Born supremacy: quantum advantage and training of an Ising Born machine, npj Quantum Information6, 60 (2020)
2020
-
[4]
M. J. Bremner, R. Jozsa, and D. J. Shepherd, Classical simulation of commuting quantum computations implies collapse of the polynomial hierarchy, Proceedings of the Royal Society A: Mathematical, Physical and Engineer- ing Sciences467, 459 (2011), arXiv:1005.1407 [quant-ph]
Pith/arXiv arXiv 2011
-
[5]
Y. Du, Z. Tu, B. Wu, X. Yuan, and D. Tao, Power of quantum generative learning (2022), arXiv:2205.04730 [quant-ph]
Pith/arXiv arXiv 2022
-
[6]
J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Bab- bush, and H. Neven, Barren plateaus in quantum neural network training landscapes, Nature Communications9, 4812 (2018)
2018
-
[7]
Cerezo, A
M. Cerezo, A. Sone, T. Volkoff, L. Cincio, and P. J. Coles, Cost function dependent barren plateaus in shal- low parametrized quantum circuits, Nature Communica- tions12, 1791 (2021)
2021
-
[8]
Van den Nest, Simulating quantum computers with probabilistic methods, Quantum Information and Com- putation11, 784 (2011)
M. Van den Nest, Simulating quantum computers with probabilistic methods, Quantum Information and Com- putation11, 784 (2011)
2011
-
[9]
E. Recio-Armengol, S. Ahmed, and J. Bowles, Train on classical, deploy on quantum: scaling generative quantum machine learning to a thousand qubits (2026), arXiv:2503.02934 [quant-ph]
arXiv 2026
-
[10]
M. J. Bremner, A. Montanaro, and D. J. Shepherd, Average-case complexity versus approximate simulation of commuting quantum computations, Phys. Rev. Lett. 117, 080501 (2016)
2016
-
[11]
M. J. Bremner, B. Cheng, and Z. Ji, Instantaneous quan- tum polynomial-time sampling and verifiable quantum advantage: Stabilizer scheme and classical security, PRX Quantum6, 020315 (2025). 14
2025
-
[12]
M. S. Rudolph, S. Lerch, S. Thanasilp, O. Kiss, O. Shaya, S. Vallecorsa, M. Grossi, and Z. Holmes, Trainability bar- riers and opportunities in quantum generative modeling, npj Quantum Information10, 116 (2024)
2024
- [13]
-
[14]
K. Shen, S. Pielawa, V. Dunjko, and H. Wang, Character- izing trainability of instantaneous quantum polynomial circuit born machines (2026), arXiv:2602.11042 [quant- ph]
arXiv 2026
-
[15]
G. De Luca, Trainability of IQP quantum circuit born machines under gaussian initialization (2026), arXiv:2606.10179 [quant-ph]
Pith/arXiv arXiv 2026
- [16]
-
[17]
A. Kurkin, K. Shen, S. Pielawa, H. Wang, and V. Dunjko, Note on the universality of parameterized iqp circuits with hidden units for generating probability distributions (2025), arXiv:2504.05997 [quant-ph]
Pith/arXiv arXiv 2025
-
[18]
Zhong, X
W. Zhong, X. Gao, S. F. Yelin, and K. Najafi, Many-body localized hidden generative models, Phys. Rev. Research 6, 043041 (2024)
2024
-
[19]
N. Wiebe and L. Wossnig, Generative training of quan- tum boltzmann machines with hidden units (2019), arXiv:1905.09902 [quant-ph]
Pith/arXiv arXiv 2019
-
[20]
J. Slim, S. Monaco, F. Rehm, D. Kr¨ ucker, and K. Bor- ras, An iqp born machine for calorimeter image genera- tion at 64 qubits with compiled-iqp deployment (2026), arXiv:2605.27735 [quant-ph]
Pith/arXiv arXiv 2026
-
[21]
Cheng, J
S. Cheng, J. Chen, and L. Wang, Information perspective to probabilistic modeling: Boltzmann machines versus born machines, Entropy20, 583 (2018)
2018
-
[22]
Z.-Y. Han, J. Wang, H. Fan, L. Wang, and P. Zhang, Unsupervised generative modeling using matrix product states, Physical Review X8, 10.1103/physrevx.8.031012 (2018)
-
[23]
Shepherd and M
D. Shepherd and M. J. Bremner, Temporally unstruc- tured quantum computation, Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sci- ences465, 1413–1439 (2009)
2009
-
[24]
Gretton, K
A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Sch¨ olkopf, and A. Smola, A kernel two-sample test, Journal of Ma- chine Learning Research13, 723 (2012)
2012
-
[25]
Muandet, K
K. Muandet, K. Fukumizu, B. Sriperumbudur, and B. Sch¨ olkopf, Kernel mean embedding of distributions: A review and beyond, Found. Trends Mach. Learn.10, 1 (2017)
2017
-
[26]
B. K. Sriperumbudur, A. Gretton, K. Fukumizu, B. Sch¨ olkopf, and G. R. Lanckriet, Hilbert space em- beddings and metrics on probability measures, J. Mach. Learn. Res.11, 1517 (2010)
2010
-
[27]
Li, W.-C
C.-L. Li, W.-C. Chang, Y. Cheng, Y. Yang, and B. P´ oczos, MMD GAN: Towards deeper understanding of moment matching network, inAdvances in Neural Infor- mation Processing Systems 30(Curran Associates, Inc.,
-
[28]
Y. Li, K. Swersky, and R. Zemel, Generative moment matching networks, inProceedings of the 32nd Interna- tional Conference on Machine Learning - Volume 37, ICML’15 (JMLR.org, 2015) p. 1718–1727
2015
-
[29]
Grant, L
E. Grant, L. Wossnig, M. Ostaszewski, and M. Benedetti, An initialization strategy for addressing barren plateaus in parametrized quantum circuits, Quantum3, 214 (2019)
2019
-
[30]
Zhang, L
K. Zhang, L. Liu, M.-H. Hsieh, and D. Tao, Escap- ing from the barren plateau via gaussian initializations in deep variational quantum circuits, inProceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22 (Curran Associates Inc., Red Hook, NY, USA, 2022)
2022
-
[31]
G. Verdon, M. Broughton, J. R. McClean, K. J. Sung, R. Babbush, Z. Jiang, H. Neven, and M. Mohseni, Learn- ing to learn with quantum neural networks via classical neural networks (2019), arXiv:1907.05415 [quant-ph]
Pith/arXiv arXiv 2019
-
[32]
A. Y. Ng, M. I. Jordan, and Y. Weiss, On spectral clus- tering: Analysis and an algorithm, inAdvances in Neural Information Processing Systems, Vol. 14 (2001) pp. 849– 856
2001
-
[33]
von Luxburg, A tutorial on spectral clustering, Statis- tics and Computing17, 395 (2007)
U. von Luxburg, A tutorial on spectral clustering, Statis- tics and Computing17, 395 (2007)
2007
-
[34]
Arrasmith, Z
A. Arrasmith, Z. Holmes, M. Cerezo, and P. J. Coles, Equivalence of quantum barren plateaus to cost concen- tration and narrow gorges, Quantum Science and Tech- nology7, 045015 (2022)
2022
-
[35]
Ghosh, V
A. Ghosh, V. Kulharia, V. P. Namboodiri, P. H. S. Torr, and P. K. Dokania, Multi-agent diverse generative adver- sarial networks, inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018) pp. 8513–8521
2018
-
[36]
Lecun, L
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, Gradient-based learning applied to document recogni- tion, Proceedings of the IEEE86, 2278 (1998)
1998
-
[37]
D. P. Kingma and J. Ba, Adam: A method for stochas- tic optimization, inProceedings of the 3rd International Conference on Learning Representations (ICLR)(2015)
2015
-
[38]
E. Armengol and J. Bowles, IQPopt: Fast optimization of instantaneous quantum polynomial circuits in JAX (2025), arXiv:2501.04776 [quant-ph]
Pith/arXiv arXiv 2025
-
[39]
J. M. K¨ ubler, S. Buchholz, and B. Sch¨ olkopf, The in- ductive bias of quantum kernels, inAdvances in Neural Information Processing Systems 34(Curran Associates, Inc., 2021) arXiv:2106.03747 [quant-ph]
Pith/arXiv arXiv 2021
-
[40]
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, Outrageously large neural net- works: The sparsely-gated mixture-of-experts layer, in International Conference on Learning Representations (2017) arXiv:1701.06538 [cs.LG]
Pith/arXiv arXiv 2017
-
[41]
A. Huang, W. Maxwell, V. Belis, E. Peters, J. Pye, S. Ja- hangiri, and J. Bowles, Spectral born machines: Classi- cally trainable quantum generative models for discrete data (2026), arXiv:2607.06675 [quant-ph]
Pith/arXiv arXiv 2026
-
[42]
D. Maslov, S. Bravyi, F. Tripier, A. Maksymov, and J. Latone, Fast classical simulation of Harvard/QuEra IQP circuits (2024), arXiv:2402.03211 [quant-ph]
Pith/arXiv arXiv 2024
-
[43]
Scriva, E
G. Scriva, E. Costa, B. McNaughton, and S. Pilati, Accel- erating equilibrium spin-glass simulations using quantum annealers via generative deep learning, SciPost Phys.15, 018 (2023)
2023
-
[44]
Demidik, C
M. Demidik, C. T¨ uys¨ uz, N. Piatkowski, M. Grossi, and K. Jansen, Expressive equivalence of classical and quan- tum restricted boltzmann machines, Communications Physics8, 413 (2025)
2025
-
[45]
T. Tieleman, Training restricted Boltzmann machines us- ing approximations to the likelihood gradient, inPro- 15 ceedings of the 25th International Conference on Machine Learning(2008) pp. 1064–1071
2008
-
[46]
G. E. Hinton, A practical guide to training restricted boltzmann machines, inNeural Networks: Tricks of the Trade: Second Edition, edited by G. Montavon, G. B. Orr, and K.-R. M¨ uller (Springer Berlin Heidelberg, Berlin, Heidelberg, 2012) pp. 599–619. 16 CONTENTS OF THE APPENDICES Appendix A: Model, estimator, and deployment. . . . . . . . . . . . . . . ....
2012
-
[47]
[8] in the main text
Randomized estimation of IQP correlators This subsection records the classical randomized estimator that evaluates one branch correlatorz A(θ(ℓ)), the esti- mator referred to as Ref. [8] in the main text. To lighten notation we drop the branch label within this subsection and writeθfor a single branch’s angles; the result applies to eachθ (ℓ) in turn. Red...
-
[48]
(12), so that one IQP-form circuit estimates the uniform mixture correlatorz A(Θ)
Compilation to the ancilla circuit This subsection compiles theL= 2 a branches into the uniform-weight (n+a)-qubit IQP circuit of Eq. (12), so that one IQP-form circuit estimates the uniform mixture correlatorz A(Θ). The derivation reuses theX-basis correlator reduction of the preamble [Eqs. (A2)–(A6)] undern→n+a,s→(s,u),G→G⊗S,A→(A,0), whereu∈ {0,1} a are...
-
[49]
It fixes the branch weights and proves the estimator equivalence
Branch weights and deployment equivalence This subsection assembles the two estimators into the deployment statement used in the main text: the trained mixed IQP marginal, the weighted mixture, and a decomposed randomized estimator that never builds the joint register all share one estimand. It fixes the branch weights and proves the estimator equivalence...
-
[50]
choose integer countsn ℓ for the active branches with P ℓ:πℓ>0 nℓ =Mandn ℓ ≥1 wheneverπ ℓ >0 (in practice largest-remainder rounding ofπ ℓM)
-
[51]
for each active branchℓ, drawn ℓ auxiliary bit stringss∼Unif({0,1} n) and form the single-circuit unbiased estimatorbzA(θ(ℓ)) of Eq. (A7)
-
[52]
Theπ ℓ weighting enters exactly once, in step 3; the allocation of step 1 only distributes the budget and applies no second weighting
combine as the weighted meanbz dec A = P ℓ πℓ bzA(θ(ℓ)). Theπ ℓ weighting enters exactly once, in step 3; the allocation of step 1 only distributes the budget and applies no second weighting. Equivalently one may draw a branch index with probabilityπ ℓ per sample and average the resulting single-branch cosine samples directly, using sampling frequency rat...
-
[53]
(A21) (Appendix A 3), each branch correlator obeysz A(θ(ℓ))∈[−1,1] as the expectation of a±1-eigenvalue Pauli word, andw A = Θ(n−|A|) is the low-body weight atσ= Θ( √n)
Branch-wise view of the loss and its curvature We work throughout with the low-body MMD loss L(Θ) = X A⊆[n] wA zA(Θ)−t A 2 , z A(Θ) = L−1X ℓ=0 πℓ zA(θ(ℓ)),(B1) with weightsπ ℓ on the simplex (π ℓ = 1/Lfor the standard all-zero ancilla input), where the second equality is the expectation-value linearity of Eq. (A21) (Appendix A 3), each branch correlator o...
-
[54]
Its curvature- to-variance step is Theorem 2 of Ref
T rainability lemma: center curvature implies inverse-polynomial loss variance The following lemma is the single trainability statement on which both barren-plateau theorems rest. Its curvature- to-variance step is Theorem 2 of Ref. [13], invoked here verbatim with its explicit constants rather than re-derived; the only mixture-specific content is checkin...
-
[55]
We evaluate the curvature Eq
Data-agnostic center: inverse-polynomial curvature At the unbiased, data-agnostic center the single-parameter curvature of the mixture loss isc= 8w {j}/L2 = Ω(1/poly(n)), which by the trainability lemma rules out a barren plateau. We evaluate the curvature Eq. (B3) at this center and then invoke Lemma 1. Evaluation at the unbiased center.Set, in every bra...
-
[56]
III C and feeds it to the Master Lemma 1 to obtain the data-dependent counterpart of Theorem 1
Data-dependent center: curvature This subsection computes the loss curvature at the data-dependent center of Sec. III C and feeds it to the Master Lemma 1 to obtain the data-dependent counterpart of Theorem 1. The data-agnostic analysis of the preceding subsection anchors the curvature at the unbiased centerθ (ℓ) j =π/4,θ (ℓ) jk = 0, at which the data-mis...
-
[57]
Its body-order-rescaled finite-size constant is Ck,eff :=nϵ 2/k k , q k,eff :=C k,eff /n.(B32) We compute the same quantities from the 95th percentile to expose the tail
Finite-size diagnostics of grouped data We first measure a typical residual within each group, ϵ(ℓ) k = median |A|=k r(ℓ) A ,(B31) and aggregate it asϵ k = P ℓ πℓϵ(ℓ) k . Its body-order-rescaled finite-size constant is Ck,eff :=nϵ 2/k k , q k,eff :=C k,eff /n.(B32) We compute the same quantities from the 95th percentile to expose the tail. Fork= 2, . . . ...
-
[58]
II C) that an ancilla-free IQP circuit is not universal [16]; we give a constructive version through the mixture
Universality via a trivial mixture Recall (Sec. II C) that an ancilla-free IQP circuit is not universal [16]; we give a constructive version through the mixture. A deliberately trivial mixture, one branch per support point of the target, already representsany distribution, so the mixed IQP family is universal. The construction has no generative value on i...
-
[59]
II E 1 is branch-coincident, and a mixture whose branches coincide reduces to a single ancilla-free circuit
Branch diversity is necessary to surpass the best ancilla-free circuit The data-agnostic center of Sec. II E 1 is branch-coincident, and a mixture whose branches coincide reduces to a single ancilla-free circuit. Here we make this quantitative: the MMD loss attainable by the mixture is controlled from below by how much its branches differ, so that branch ...
-
[60]
Dataset details Binary blobs (n= 16).This dataset is a binary analog of Gaussian blobs [9]. A sample is generated by choosing one of eight fixed 16-bit patterns uniformly and flipping each bit independently with probabilityη= 0.05: pdata(z) = 1 8 8X m=1 η dH (z,b(m)) (1−η) n−dH (z,b(m)),(D1) whered H is the Hamming distance. The resulting distribution has...
2000
-
[61]
The MNIST data-agnostic runs are listed separately because they use a smaller learning rate,η= 5×10 −4, and a longer optimization budget
IQP-QCBM hyperparameters Table IV gives the complete circuit, optimization, and estimator settings used for the numerical benchmarks. The MNIST data-agnostic runs are listed separately because they use a smaller learning rate,η= 5×10 −4, and a longer optimization budget. These settings were used because the shared setting did not converge consistently acr...
2000
-
[62]
For the global and data-agnostic initializations, the ancilla branches are sampled uniformly, πℓ = 1 L , ℓ= 0,
Branch weights and clustering settings The reported experiments use two branch-weight rules, fixed before joint training. For the global and data-agnostic initializations, the ancilla branches are sampled uniformly, πℓ = 1 L , ℓ= 0, . . . , L−1,(D3) for every dataset, ancilla count, and optimization seed (with the single-circuit casea= 0 understood asπ 0 ...
1906
-
[63]
The RBM family also admits quantum extensions with quantified expressivity relations to its classical form [44]; here it serves purely as a classical reference
Classical RBM baseline The classical reference is a Bernoulli–Bernoulli restricted Boltzmann machine (RBM) withnvisible units and one layer ofhhidden units, givingnh+n+htrainable parameters. The RBM family also admits quantum extensions with quantified expressivity relations to its classical form [44]; here it serves purely as a classical reference. We tr...
2000
-
[64]
Protocol for the numerical tests of the theorems This appendix records the protocol and the per-configuration values behind Sec. V B. Each center is rebuilt through the training code path with the coincidence-breaking radius set to zero, so the point evaluated is the analyzed center itself. Exact loss.Forn= 16 the loss is evaluated in closed form. The pha...
-
[65]
Each ablated branch is therefore a product distribution
Ablation of the trained two-body angles To isolate the trained non-product structure, we set every two-body angle of the headline cluster-initialized MNIST (a= 4) and D-Wave (a= 3) models to zero while retaining their one-body angles and branch weights, then repeat the test-MMD2 evaluation with the same estimator seeds and bandwidths. Each ablated branch ...
-
[66]
3(e)–(h) by kernel bandwidth at each benchmark’s headline mixture size
MMD bandwidth dependence Figure 9 resolves the sweep averages of Fig. 3(e)–(h) by kernel bandwidth at each benchmark’s headline mixture size. 1 2 3 4 5 6 mean Pauli weight ̄m of the kernel 10−4 10−3 10−2 test MMD2 (a) blobs, a = 3 1 2 3 4 5 6 mean Pauli weight ̄m of the kernel 10−3 2 × 10−4 3 × 10−4 4 × 10−4 6 × 10−4 (b) Ising, a = 3 1 2 3 4 5 6 mean Paul...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.