REVIEW 3 major objections 6 minor 94 references
DAGAF: A directed acyclic generative adversarial framework for joint structure learning and tabular data synthesis
T0 review · 3 major / 6 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read An architecture-only matrix of joint Fourier coefficients predicts how parametrised quantum circuits couple frequencies and shape training kernels.
desk verdict The supplied manuscript is a clean spectral framework for PQCs (joint harmonic matrix C), not the DAGAF abstract in the wrapper; the algebra is solid inside a well-flagged circuit class. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The circuit harmonic matrix C: the matrix of joint Fourier coefficients in f(x;θ) = Σ_ω Σ_k C_ωk e^{iω·x} e^{ik·θ}, obtained from encoder path amplitudes or Pauli-propagation branches. All second-order coefficient statistics and both harmonic and data-space QNTKs factor as quadratic forms through C.
What would settle it
On a fixed re-uploading circuit, estimate C by Fourier analysis over parameters, form CPC† and C diag(‖k‖₂²) C†, and check whether they match Monte-Carlo coefficient covariances and parameter-averaged harmonic QNTKs within sampling and truncation error; systematic mismatch under single-use Pauli rotations would refute the identities.
Extended reading notes
Core claim
For re-uploading quantum Fourier models with commuting phase encoders and single-use Pauli-rotation trainable blocks, the model admits a joint harmonic expansion whose coefficients form an architecture-level circuit harmonic matrix C. Under uniform parameter sampling, centred coefficient covariances equal the row Gram matrix C P C† (P removing the constant mode), and the parameter-averaged harmonic quantum neural tangent kernel equals C diag(‖k‖₂²) C†; the usual data-space kernel is recovered by projecting with the design matrix V.
Load-bearing premise
The closed forms need each trainable angle to appear in only one rotation and parameters to be averaged with a product uniform measure; sharing angles or non-product priors break the character orthogonality that collapses the formulas.
Editorial extensions
If this is right
- Circuit architectures can be compared by their C-induced coefficient correlation matrices and averaged QNTKs before any dataset is chosen.
- Encoder redundancy and trainer branching control spectral bias through row energies and shared k-support of C, giving design handles on which frequencies couple.
- The column space of the data-space QNTK is constrained by V C at every parameter point, so architecture restricts reachable function-space directions throughout training.
- The same C construction extends in principle to correlators of quantum circuit Born machines, linking generative trainability to row energies of C.
Reading between the lines
- If C is cheap to approximate for a circuit family, architecture search could optimise for desired correlation or kernel spectra rather than only expressivity heuristics.
- Truncation of C to low Hamming-weight harmonics may already be enough for relative ranking of circuits even when absolute kernel scale is wrong.
- Negative off-diagonal blocks in Corr derived from C would predict competitive rather than cooperative learning of those frequency pairs, a testable training-dynamics signature.
- Relaxing single-use parameters while keeping a Pauli-propagation view would require a non-product character calculus; that is the natural next algebraic object.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces a data-agnostic spectral framework for a broad class of re-uploading parametrised quantum circuits that admit a quantum Fourier model description. For commuting phase encoders and single-use Pauli-rotation trainable blocks with Clifford interleavings, the model is written as a joint input–parameter Fourier expansion whose coefficients form an architecture-level circuit harmonic matrix C. Under uniform parameter sampling on the torus, centred coefficient covariances equal the row Gram matrix CPC† (P removing the constant mode), and the parameter-averaged harmonic QNTK equals C diag(‖k‖₂²) C†; the usual data-space QNTK is recovered by projection with the design matrix V. Three equivalent constructions of C (joint Fourier coefficients, encoder-path aggregation, and Pauli-propagation node expansions) are given, with supporting appendices and numerical checks that compare truncated-C estimators to split-sample Monte Carlo on variances, correlation matrices, and averaged harmonic QNTKs for several standard ansätze.
Significance. If the identities hold in the stated class, the work supplies a clean architectural prior that separates design-level structure from data-dependent effects: coefficient second-order statistics and local training kernels factor algebraically through a single circuit-defined matrix C, without reference to a dataset or optimisation trajectory. That link is useful for comparing ansätze, diagnosing spectral bias and frequency coupling, and relating Fourier fingerprints to QNTK geometry. Strengths include explicit factorisations grounded in standard torus harmonic analysis and Heisenberg conjugation, equivalent constructive definitions of C, a careful limitations section, and reproducible-style numerical protocol (split samples, normalised Frobenius/cosine diagnostics). The framework is scoped rather than universal, but within that scope it organises several previously separate diagnostics under one object.
major comments (3)
- The load-bearing closed forms (Eqs. (42), (70), and the averaged kernel identities in §V) rest on single-use parameters and product Haar measure yielding character orthogonality E_θ[e^{i(k−l)·θ}]=δ_kl (§II.C–II.D, VII.A.2). The manuscript already flags that parameter tying, non-product priors, or non-Pauli trainables enlarge K and break orthogonality. For the central claim to be usable beyond the stated class, either (i) state the theorem statements with these assumptions as explicit hypotheses in the main text (not only in limitations), or (ii) give at least one worked extension (e.g. degree-r reuse with ka∈{−r,…,r}) showing how CPC† and the weighted Gram form deform. Without that, the abstract-level claim of a general architectural matrix risks over-reading.
- Numerical validation (§VI) truncates K to low Hamming weight (and a hard |K| cap), then compares normalised variance profiles, correlations, and ‖k‖₂²-weighted kernels to Monte Carlo (§VI.A.4, VI.D). Agreement is structural and often good, but absolute scale—especially for the QNTK, where truncation compounds with ‖k‖₂² weights—is not controlled. Because the paper’s claim is that objects “factor through C” and can be reconstructed from design, the experiments should either (a) report absolute (not only normalised) errors on a small fully enumerable instance where K is complete, or (b) quantify residual mass outside the truncated sector as a function of depth. Otherwise the numerics support pattern similarity more than reconstruction fidelity.
- The manuscript title/abstract in the submission metadata (DAGAF / causal tabular synthesis) do not match the supplied full text (Circuit Harmonic Matrices / QML). This is a cataloguing issue rather than a mathematical error, but it must be corrected before any decision; the report below assesses only the quantum Fourier / C-matrix manuscript actually provided.
minor comments (6)
- Notation: aω(θ) is used for trainable Fourier coefficients while much of the QFM literature uses cω; a short “notation vs prior work” sentence in §II would reduce confusion (already noted briefly).
- Figures 1 and 4–6 are dense composite panels; ensure captions fully define colour scales, grey diagonals, and which encoder (RX/RY) is used in each panel without relying on appendix cross-references alone.
- Eq. (1) and later displays mix ω·x and k·θ; consistently mark complex conjugation / reality conditions for real-valued f when reporting complex correlation matrices.
- Code availability URL appears as placeholder underscores in the manuscript; replace with a working repository link and pin the commit used for the reported figures.
- Related work on Fourier fingerprints and spectral bias [21–23] is cited; a short table mapping each prior diagnostic to the corresponding C-derived object would clarify complementarity.
- Appendix D proofs are sketched; for journal archival value, expand the covariance identity to a fully self-contained proof with the projector P made explicit in components.
Circularity Check
No significant circularity: CPC† and C diag(‖k‖₂²)C† are standard torus Fourier identities once a(θ)=Cψ(θ) is granted; Monte Carlo checks are independent, not tautological.
full rationale
The manuscript defines the circuit harmonic matrix C from joint input–parameter Fourier coefficients (equivalently path amplitudes or Pauli-propagation nodes), then derives second-order coefficient statistics and harmonic QNTKs as quadratic forms in C under uniform sampling on the parameter torus. Those identities follow from character orthogonality E_θ[e^{i(k−l)·θ}]=δ_kl and differentiation of characters—standard harmonic analysis on T^m—once the factorisation a(θ)=Cψ(θ) is granted by the single-use Pauli-rotation + commuting phase-encoder assumptions. C is fixed by encoder, ansatz, observable, and input state; it is not fitted to a supervised loss or to the Monte Carlo targets. Split-sample Monte Carlo estimators of the same population objects (variances, Pearson correlations, parameter-averaged H) are independent numerical checks of the algebraic identities under truncation, not predictions forced by construction. Related-work citations (QFM spectra, QNTK, Pauli propagation, Fourier fingerprints) supply context and experimental circuit families; none is a load-bearing uniqueness theorem that forbids alternatives or smuggles the central Gram forms. Uniform-sampling statistics are correctly scoped as initialisation-time architectural priors (Section VII.A.4), not trained-model predictions. No self-definitional loop, fitted-input-as-prediction, or renaming-only step appears in the derivation chain.
Assumptions & free parameters
free parameters (2)
- Hamming-weight truncation radius h and hard cap on |K| =
h=3 for d>1 (with additional |K| cap); exact cap not fully tabulated
- Monte Carlo sample size S and DFT grid size n_x =
S=100096; n_x=128 (variances) / 126 (correlations, kernels)
assumptions (5)
- domain assumption Commuting phase encoders yield a finite encoder-accessible frequency set Ω via difference sets and Minkowski sums under re-uploading.
- domain assumption Trainable blocks are single-use Pauli rotations with Clifford interleavings, so each a_ω(θ) is a degree-≤1 trigonometric polynomial per coordinate and K ⊆ {−1,0,1}^m.
- standard math Character orthogonality on the m-torus under Haar measure: E_θ[e^{i(k−l)·θ}] = δ_kl.
- standard math Heisenberg conjugation of Pauli strings through Pauli rotations yields the two-branch cos/sin rule (Eq. 10), generating the parameter harmonics.
- domain assumption Population coefficient statistics and averaged kernels are taken under θ ~ Unif(T^m), interpreted as initialisation-time architectural priors.
invented entities (2)
-
Circuit harmonic matrix C (joint input–parameter Fourier coefficient matrix)
independent evidence
-
Coefficient-space (harmonic) QNTK H(θ) = C M(θ) C†
independent evidence
Cite this review
Pith. "Pith review of DAGAF: A directed acyclic generative adversarial framework for joint structure learning and tabular data synthesis." pith.science (2026). https://pith.science/paper/2604.04290
@misc{pith2026260404290,
author = {Pith},
title = {Pith review of: DAGAF: A directed acyclic generative adversarial framework for joint structure learning and tabular data synthesis},
year = {2026},
howpublished = {\url{https://pith.science/paper/2604.04290}},
note = {Machine review of arXiv:2604.04290}
}
read the original abstract
Understanding the causal relationships between data variables can provide crucial insights into the construction of tabular datasets. Most existing causality learning methods typically focus on applying a single identifiable causal model, such as the Additive Noise Model (ANM) or the Linear non-Gaussian Acyclic Model (LiNGAM), to discover the dependencies exhibited in observational data. We improve on this approach by introducing a novel dual-step framework capable of performing both causal structure learning and tabular data synthesis under multiple causal model assumptions. Our approach uses Directed Acyclic Graphs (DAG) to represent causal relationships among data variables. By applying various functional causal models including ANM, LiNGAM and the Post-Nonlinear model (PNL), we implicitly learn the contents of DAG to simulate the generative process of observational data, effectively replicating the real data distribution. This is supported by a theoretical analysis to explain the multiple loss terms comprising the objective function of the framework. Experimental results demonstrate that DAGAF outperforms many existing methods in structure learning, achieving significantly lower Structural Hamming Distance (SHD) scores across both real-world and benchmark datasets (Sachs: 47%, Child: 11%, Hailfinder: 5%, Pathfinder: 7% improvement compared to state-of-the-art), while being able to produce diverse, high-quality samples.
Reference graph
Works this paper leans on
-
[1]
Path-amplitude harmonics and aggregation Under single-use Pauli rotations, each path amplitude is itself a finite Fourier series on� m, Ap(θ) = � k∈K cp,keik·θ,(31) whereK� ��1,0,1� m arises from selecting the constant, positive-frequency, or negative-frequency character term generated by each non-commuting rotation encountered along the branch (Section I...
-
[2]
Node expansion and separability We write the resulting expansion in the separable form f(x;θ) = � ν∈N dνNν(x)M ν(θ),(33) whereνindexes nodes (branches),d ν��is anx- andθ-independent scalar determined by operator algebra along the branch,N ν(x) is a product of encoder-induced trigonometric factors, andMν(θ) is a product of trainable trigonometric factors. ...
-
[3]
Hence� θ[a(θ)] =C•0 and centring removes exactly thek= 0 contribution
Character orthogonality and the�= 0mode When 0�K, character orthogonality on� m implies �θ[ψk] = 0 for allk�= 0 and� θ[ψ0] = 1 which is standard harmonic analysis on compact abelian groups (since� m �= U(1)m) [25, 26]. Hence� θ[a(θ)] =C•0 and centring removes exactly thek= 0 contribution. Here�indicates theωindex is vectorised over the whole spectrum Ω, s...
-
[4]
Parseval/Plancherel and row energies More generally, Parseval’s identity (equivalently, Plancherel onL 2(�m)) yields the squared-magnitude second moment [25, 26] �θ ��aω(θ)�2� = � k∈K �Cωk�2 =�Cω0�2 + � k̸=0 �Cωk�2.(45) This motivates therow energyofC, Eω:= � k∈K �Cωk�2 =�� θ[aω(θ)]�2 + Var[aω(θ)],(46) so that, after centring (or when 0/�K),E ωcoincides w...
-
[5]
Pearson normalisation and correlation as a Gram matrix To isolatefrequency–frequency coupling structure independent of marginal scale, define the population Pearson correlation matrix by Corr�a(θ),a(θ)†� =D−1/2Cov�a(θ),a(θ)†�D−1/2 =D−1/2(CPC†)D−1/2,(47) whereD�� |Ω|×|Ω|is the diagonal matrix of variances, D:= diag (Var[aω(θ)])ω∈Ω.(48) 9 Equivalently, with...
-
[6]
The accessible spectrum and its construction via difference sets and Minkowski sums is standard in the QFM literature [14, 15, 19, 20]
Encoder-side redundancy and variance scaling For a fixed encoder, eachω�Ω can arise from multiple layer-wise spectral choices; letR(ω) denote the set of encoder-induced paths producingω(Section III B and Appendix C). The accessible spectrum and its construction via difference sets and Minkowski sums is standard in the QFM literature [14, 15, 19, 20]. Usin...
-
[7]
From (43), Cov[a] ωµis large when the rowsC ω,•andC µ,•have aligned phases on a shared set of centred harmonics
Trainer-side coupling Similarly, coefficient correlations are controlled by overlap ofk-support. From (43), Cov[a] ωµis large when the rowsC ω,•andC µ,•have aligned phases on a shared set of centred harmonics. V. TRAINING KERNELS This section formalises learning kernels as geometric objects derived from the circuit harmonic matrix. We introduce a generali...
-
[8]
(67) This isolates a universal torus objectM(θ) (depending only on the character map and the parameter space manifold structure) from the architecture-dependent mapping encoded byC
Harmonic factorisation through� InC-matrix notation the harmonic QNTK factors as H(θ) =CM(θ)C†, M kl(θ) = m� a=1 ∂ψk(θ) ∂θa ∂ψl(θ) ∂θa . (67) This isolates a universal torus objectM(θ) (depending only on the character map and the parameter space manifold structure) from the architecture-dependent mapping encoded byC
Show all 94 references
-
[9]
Parameter averaging Using (61) and character orthogonality� θ[ei(k−l)·θ] = δkl (Section II D; [25]), parameter averaging diagonalises the (k,l) sum and yields �θ[H(θ)] =C�θ[M(θ)]C†,(68) �θ[M(θ)]kk = m� a=1 k2 a =�k� 2 2.(69)
-
[10]
Interpretation of the� 2 � weights In the single-use Pauli-rotation regimeK� ��1,0,1� m, sok 2 a � �0,1�acts as an indicator: only 11 those parameter-harmonic patternskin which parameter θa appears (i.e.k a =�1) contribute to∂ θ�a(θ), while harmonics withk a = 0 give a vanishi...
-
[11]
Throughout, the trainable blockW ℓis taken to be a depth-drepetition of a fixed ans¨ atz pattern as defined in [10] and also used in [21]
Circuit family We study re-uploading quantum Fourier models of the form U(θ,x) = L� ℓ=1 Wℓ(θ(ℓ))Sℓ(x),(77) whereS ℓ(x) are Pauli-encoding blocks andW ℓ(θ(ℓ)) are single-use parametrised Pauli-rotation blocks interleaved with Clifford entangling gates. Throughout, the trainable...
-
[12]
For our circuits where we re-encode on each qubit and each layer,ω max =nL
Coefficient estimation via discrete input Fourier transform For each sampled parameter vectorθ, coefficientsaω(θ) are estimated from the circuit outputf(x;θ) =�O�U(θ,x) by a discrete Fourier transform over a uniform gridx j � [0,2π),j= 1,...,n x: aω(θ)� 2π nx n�� j=1 f(xj;θ)e−...
-
[13]
Sampling and split-sample estimators Parameter-dependent quantities are estimated either (i) from joint Fourier coefficientsC ωkvia constructions made in Sections IV and V or (ii) by usual Monte Carlo samplingθ(s) �Unif(� m) and forming empirical averages. To avoid spuriously ...
-
[14]
Matrix similarity measures To quantify agreement between matrices, we use scale-sensitive and scale-insensitive diagnostics. The relative Frobenius error is εF (A,B) = �A�B� F �B�F ,(81) and a scale-insensitive off-diagonal alignment (Frobenius cosine similarity) is �(A,B) = R...
-
[15]
Setup We first test the variance identity Varθ[aω(θ)] = � k̸=0 �Cωk�2 = [CPC†]ωω,(83) which is the diagonal of the centred Gram matrixCPC † derived in Section IV. We setn= 6 qubits andL= 1 layer and vary the training-block depthd� �1,...,5�for circuits YZY (no entangling), YZY...
-
[16]
Results Subfigures (b) in Figures 4–6 compare the normalised empirical variance profile ¯Vωto the normalised truncated row-energy profile ¯Eωacross depthsd= 1,...,5. Across all depths tested, the two curves track one another closely as functions ofω, and the per-depth Pearson ...
-
[17]
Setup Section IV shows that, under uniform parameter samplingθ�Unif(� m), the centred coefficient covariance is the row Gram matrix Covθ[˜a(θ)] =CPC†,(88) and the corresponding Pearson correlation matrix is obtained by normalising by the marginal variances, Corrθ[a(θ)] =D−1/2(...
-
[18]
Results Subfigures (c) in Figures 4–6 compare the Hermitian Complex Pearson correlation matrices obtained via our Cmatrix construction, and usual normalised Monte Carlo covariances. Note that the colour-bar is scaled to the magnitude of the off-diagonal values to illustrate st...
-
[19]
Setup We now validate the parameter-averaged harmonic QNTK reconstruction derived in Section V. Let Xωa(θ) :=∂θ�aω(θ), H(θ) :=X(θ)X(θ)†,(95) so that the parameter-averaged harmonic QNTK is ¯H:=� θ∼Unif(�� )[H(θ)] =�θ �X(θ)X(θ)†�.(96) Section V shows that due to independent tra...
-
[20]
As in the correlation-matrix comparisons of Figures 4–6, we report both normalised Frobenius errors and cosine similarities
Results Figures 7–9 compare the absolute values of the variance-normalised, parameter-averaged harmonic QNTKs derived from theC-matrix construction against direct Monte Carlo estimation of the Gram matrix of coefficient-space Jacobians. As in the correlation-matrix comparisons...
-
[21]
Architectures employing non-commuting feature maps or qualitatively different data-loading schemes may not admit the same finite harmonic description without modification
Restricted to commuting phase encoders The framework is developed for the Quantum Fourier Model setting, where commuting phase-style encoders yield a finite encoder-accessible input spectrum and an explicit harmonic factorisation of the learned function. Architectures employin...
-
[22]
Single-use parameters A key simplification in the present framework is the single-use regime, where each trainable parameter θa appears uniquely in a one-parameter rotation. Although typical, in this case, the circuit output is at most degree-one trigonometric in each coordina...
-
[23]
Gate-set and propagation assumptions Our analytic derivations exploit a setting where trainable blocks are Pauli rotations and the circuit structure supports a clean Pauli-propagation picture. For more general non-Pauli trainables and/or less Pauli-trackable interleavings, the...
-
[24]
Uniform parameter averaging versus trained parameter distributions Many closed-form identities in this paper are population statements under uniform parameter sampling,θ�Unif(� m), where character orthogonality collapses mixed terms. Objects of this type, such as coefficient v...
-
[25]
Concretely, second-order quantities control the linearised training dynamics around a parameter point (e.g
Second-order statistics In this paper, we emphasisesecond-orderobjects, covariances, correlation matrices, and Gram/kernel constructions, because they constitute the next natural level of structure beyond first-moment (mean) behaviour and already capture a substantial portion ...
-
[26]
sampling, conditioning, and lattice effects)
Finite input lattice effects Data-space kernels inherit additional structure from the finite input design (e.g. sampling, conditioning, and lattice effects). Consequently, some phenomena observed in data space reflect the interaction between architecture and the chosen input d...
-
[27]
Noise and hardware effects All results are derived for idealised circuits and exact expectations (or noiseless parameter-shift estimates). Device noise, finite sampling in expectation value estimates, and compilation constraints can suppress 19 gradients and alter effective ha...
-
[28]
Numerical limitations The numerical results in this paper should be interpreted as evidence that the predicted structures are visible in compute-limited regimes, rather than as demonstrations of convergence to the full, untruncated objects. The dominant limitation is the sever...
-
[29]
Architectural bias during training A primary motivation for introducing a joint input–parameter harmonic representation is to provide a description in which architectural contributions to learning dynamics can be stated explicitly. Within this formalism, the joint-coefficient ...
-
[30]
The dependence of data-space kernels on the input design viaVinK(θ) =VH(θ)V †, is therefore a relevant consideration
Finite input lattice effects The harmonic-space description isolates architecture-dependent coupling structure, whereas supervised learning is determined by data-space objects evaluated on a finite set of inputs. The dependence of data-space kernels on the input design viaVinK...
-
[31]
Extending to multi-dimensional inputs with multi-dimensional frequencies Many applications involve multivariate inputsx�� d and therefore admit Fourier representations indexed by multi-indicesω�� d [19, 21]. An extension of the present analysis to this setting requires (i) for...
-
[32]
Connections to generative quantum machine learning A natural domain of application for the present framework is quantum generative modelling, and in particular quantum circuit Born machines (QCBMs) [31], where a parametrised circuitU(θ) acting on�0�produces a probability distr...
-
[33]
As noted in the limitations, this focus leaves open the systematic role of higher-order statistics
Higher-order statistics and training dynamics beyond the kernel regime This paper develops a joint input–parameter harmonic representation and uses it to obtain closed expressions for second-order objects, including covariances and Gram/kernel constructions, which govern linea...
-
[34]
conservation laws, permutation invariances, and problem-specific equivariances)
Circuit symmetries and their effect on joint coefficients Symmetries in parametrised quantum circuits constrain both expressibility and training dynamics and are widely used to incorporate prior structural information (e.g. conservation laws, permutation invariances, and probl...
-
[35]
Endowing parameter space with curvature In this paper, the parameter domain is modelled as an m-torus Θm �� m with independent angular coordinates and the uniform reference measure. This choice underlies the character orthogonality identities used throughout (Appendix D) and i...
-
[36]
Quantum machine learning,
J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, and S. Lloyd, “Quantum machine learning,” Nature, vol. 549, pp. 195–202, Sept. 2017
2017
-
[37]
Schuld and F
M. Schuld and F. Petruccione,Supervised Learning with Quantum Computers. Quantum Science and Technology, Cham: Springer International Publishing, 2018
2018
-
[38]
Variational quantum algorithms,
M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, and P. J. Coles, “Variational quantum algorithms,”Nature Reviews Physics, vol. 3, pp. 625–644, Sept. 2021
2021
-
[39]
Supervised learning with quantum-enhanced feature spaces,
V. Havl´ ıˇ cek, A. D. C´ orcoles, K. Temme, A. W. Harrow, A. Kandala, J. M. Chow, and J. M. Gambetta, “Supervised learning with quantum-enhanced feature spaces,”Nature, vol. 567, pp. 209–212, Mar. 2019
2019
-
[40]
Evaluating analytic gradients on quantum hardware,
M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, and N. Killoran, “Evaluating analytic gradients on quantum hardware,”Physical Review A, vol. 99, p. 032331, Mar. 2019
2019
-
[41]
Barren plateaus in quantum neural network training landscapes,
J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, “Barren plateaus in quantum neural network training landscapes,”Nature Communications, vol. 9, p. 4812, Nov. 2018
2018
-
[42]
A Review of Barren Plateaus in Variational Quantum Computing,
M. Larocca, S. Thanasilp, S. Wang, K. Sharma, J. Biamonte, P. J. Coles, L. Cincio, J. R. McClean, Z. Holmes, and M. Cerezo, “A Review of Barren Plateaus in Variational Quantum Computing,” May
-
[43]
arXiv:2405.00781 [quant-ph, stat]
-
[44]
A Lie Algebraic Theory of Barren Plateaus for Deep Parameterized Quantum Circuits,
M. Ragone, B. N. Bakalov, F. Sauvage, A. F. Kemper, C. O. Marrero, M. Larocca, and M. Cerezo, “A Lie Algebraic Theory of Barren Plateaus for Deep Parameterized Quantum Circuits,”Nature Communications, vol. 15, p. 7172, Aug. 2024. arXiv:2309.09342 [quant-ph]
2024 arXiv
-
[45]
Noise-induced barren plateaus in variational quantum algorithms,
S. Wang, E. Fontana, M. Cerezo, K. Sharma, A. Sone, L. Cincio, and P. J. Coles, “Noise-induced barren plateaus in variational quantum algorithms,”Nature communications, vol. 12, no. 1, p. 6961, 2021. ISBN: 2041-1723
2021
-
[46]
Expressibility and Entangling Capability of Parameterized Quantum Circuits for Hybrid Quantum-Classical Algorithms,
S. Sim, P. D. Johnson, and A. Aspuru-Guzik, “Expressibility and Entangling Capability of Parameterized Quantum Circuits for Hybrid Quantum-Classical Algorithms,”Advanced Quantum Technologies, vol. 2, no. 12, p. 1900070, 2019
2019
-
[47]
Theory of overparametrization in quantum neural networks,
M. Larocca, N. Ju, D. Garc´ ıa-Mart´ ın, P. J. Coles, and M. Cerezo, “Theory of overparametrization in quantum neural networks,”Nature Computational Science, vol. 3, pp. 542–551, June 2023. arXiv:2109.11676 [quant-ph]
2023 arXiv
-
[48]
Representation Learning via Quantum Neural Tangent Kernels,
J. Liu, F. Tacchino, J. R. Glick, L. Jiang, and A. Mezzacapo, “Representation Learning via Quantum Neural Tangent Kernels,”PRX Quantum, vol. 3, p. 030323, Aug. 2022
2022
-
[49]
Quantum Lazy Training,
E. Abedi, S. Beigi, and L. Taghavi, “Quantum Lazy Training,”Quantum, vol. 7, p. 989, Apr. 2023. arXiv:2202.08232 [quant-ph]
2023 arXiv
-
[50]
Data re-uploading for a universal quantum classifier,
A. P´ erez-Salinas, A. Cervera-Lierta, E. Gil-Fuster, and J. I. Latorre, “Data re-uploading for a universal quantum classifier,”Quantum, vol. 4, p. 226, Feb. 2020. arXiv:1907.02085 [quant-ph]
2020 arXiv
-
[51]
Effect of data encoding on the expressive power of variational quantum-machine-learning models,
M. Schuld, R. Sweke, and J. J. Meyer, “Effect of data encoding on the expressive power of variational quantum-machine-learning models,”Physical Review A, vol. 103, p. 032430, Mar. 2021
2021
-
[52]
The Heisenberg Representation of Quantum Computers,
D. Gottesman, “The Heisenberg Representation of Quantum Computers,” July 1998. arXiv:quant-ph/9807006
1998 arXiv
-
[53]
Improved Simulation of Stabilizer Circuits,
S. Aaronson and D. Gottesman, “Improved Simulation of Stabilizer Circuits,”Physical Review A, vol. 70, p. 052328, Nov. 2004. arXiv:quant-ph/0406196
2004 arXiv
-
[54]
Simulation of qubit quantum circuits via Pauli propagation,
P. Rall, “Simulation of qubit quantum circuits via Pauli propagation,”Physical Review A, vol. 99, no. 6, 2019
2019
-
[55]
Multidimensional Fourier series with quantum circuits,
B. Casas and A. Cervera-Lierta, “Multidimensional Fourier series with quantum circuits,”Physical Review A, vol. 107, p. 062612, June 2023
2023
-
[56]
Constrained and Vanishing Expressivity of Quantum Fourier Models,
H. Mhiri, L. Monbroussou, M. Herrero-Gonzalez, S. Thabet, E. Kashefi, and J. Landman, “Constrained and Vanishing Expressivity of Quantum Fourier Models,” Quantum, vol. 9, p. 1847, Sept. 2025. arXiv:2403.09417 [quant-ph]
2025 arXiv
-
[57]
Fourier Fingerprints of Ansatzes in Quantum Machine Learning,
M. Strobl, M. E. Sahin, L. v. d. Horst, E. Kuehn, A. Streit, and B. Jaderberg, “Fourier Fingerprints of Ansatzes in Quantum Machine Learning,” Aug. 2025. arXiv:2508.20868 [quant-ph]
2025 arXiv
-
[58]
Fourier Analysis of Variational Quantum Circuits for Supervised Learning,
M. Wiedmann, M. Periyasamy, and D. D. Scherer, “Fourier Analysis of Variational Quantum Circuits for Supervised Learning,” Nov. 2024. arXiv:2411.03450 [cs]
2024 arXiv
-
[59]
Spectral Bias in Variational Quantum Machine Learning,
C. Duffy and M. Jastrzebski, “Spectral Bias in Variational Quantum Machine Learning,” Jan. 2026. arXiv:2506.22555 [quant-ph]
2026
-
[60]
Quantum tangent kernel,
N. Shirai, K. Kubo, K. Mitarai, and K. Fujii, “Quantum tangent kernel,”Physical Review Research, vol. 6, Aug. 2024
2024
-
[61]
Katznelson,An Introduction to Harmonic Analysis
Y. Katznelson,An Introduction to Harmonic Analysis. Cambridge Mathematical Library, Cambridge: Cambridge University Press, 3 ed., 2004. 23
2004
-
[62]
E. M. Stein,Fourier Analysis: An Introduction. New Jersey: Princeton University Press, 1st ed ed., 2003
2003
-
[63]
Neural tangent kernel: Convergence and generalization in neural networks,
A. Jacot, F. Gabriel, and C. Hongler, “Neural tangent kernel: Convergence and generalization in neural networks,”Advances in neural information processing systems, vol. 31, 2018
2018
-
[64]
C. E. Rasmussen and C. K. I. Williams,Gaussian Processes for Machine Learning. The MIT Press, Nov. 2005
2005
-
[65]
Random Features for Large-Scale Kernel Machines,
A. Rahimi and B. Recht, “Random Features for Large-Scale Kernel Machines,” inAdvances in Neural Information Processing Systems, vol. 20, Curran Associates, Inc., 2007
2007
-
[66]
On the Spectral Bias of Neural Networks,
N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. A. Hamprecht, Y. Bengio, and A. Courville, “On the Spectral Bias of Neural Networks,” June 2018
2018
-
[67]
The Born Ultimatum: Conditions for Classical Surrogation of Quantum Generative Models with Correlators,
M. Herrero-Gonzalez, B. Coyle, K. McDowall, R. Grassie, S. Beentjes, A. Khamseh, and E. Kashefi, “The Born Ultimatum: Conditions for Classical Surrogation of Quantum Generative Models with Correlators,” Nov
-
[68]
arXiv:2511.01845 [quant-ph]
-
[69]
A Unified Theory of Quantum Neural Network Loss Landscapes,
E. R. Anschuetz, “A Unified Theory of Quantum Neural Network Loss Landscapes,” Oct. 2024. arXiv:2408.11901 [quant-ph] version: 2
2024 arXiv
-
[70]
R. A. Horn and C. R. Johnson,Matrix analysis. New York, NY: Cambridge University Press, second edition, corrected reprint ed., 2017
2017
-
[71]
Stabilizer Codes and Quantum Error Correction,
D. Gottesman, “Stabilizer Codes and Quantum Error Correction,” May 1997. arXiv:quant-ph/9705052
1997 arXiv
-
[72]
The Clifford group, stabilizer states, and linear and quadratic operations over GF(2),
J. Dehaene and B. D. Moor, “The Clifford group, stabilizer states, and linear and quadratic operations over GF(2),”Physical Review A, vol. 68, p. 042318, Oct. 2003. arXiv:quant-ph/0304125
2003 arXiv
-
[73]
M. A. Nielsen and I. L. Chuang,Quantum computation and quantum information. Cambridge: Cambridge university press, 10th anniversary edition ed., 2010. 24 Appendix CONTENTS A. Table of notation used 25 B. Additional Results 26
2010
-
[74]
Quantum Fourier models and encoder-accessible harmonics 31
Parameter Averaged Harmonic QNTKs 29 C. Quantum Fourier models and encoder-accessible harmonics 31
-
[75]
Encoder assumption and difference sets 31
-
[76]
A single encoder insertion yields only difference frequencies 31
-
[77]
Re-uploading: Minkowski-sum construction of Ω 32
-
[78]
Path sets and redundancy 32
-
[79]
Proofs for coefficient statistics 33
Examples 32 D. Proofs for coefficient statistics 33
-
[80]
Character orthogonality on the parameter torus 33
-
[81]
Pauli propagation and node expansion 34
Mean, second moment, and covariance as row Gram matrices 33 E. Pauli propagation and node expansion 34
-
[82]
Setup: single-use Pauli rotations with Clifford interleavings 35
-
[83]
Pauli conjugation branching 35
-
[84]
Heisenberg back-propagation and node branching 35
-
[85]
Trig-to-character conversion andK� ��1,0,1� m 36
-
[86]
Joint coefficients from node expansion 37 F.k-mode support growth and scaling lower bounds 37
-
[87]
Active set and per-node character count 37
-
[88]
��� � Trainable parameter vector on the�-torus;�is the number of trainable parameters
Globalk-support and node-generated lower bounds 38 25 Appendix A: Table of notation used Symbol Meaning / definition ��� � Input variable;�is the input dimension. ��� � Trainable parameter vector on the�-torus;�is the number of trainable parameters. �� �Input state and measure...
-
[89]
Note the similarity in structure to the circuit’s correlation matrix in Figure 4
Parameter Averaged Harmonic QNTKs Figure 7: Averaged harmonic QNTK for the YZY circuit with no entanglers and� � encoders for depths 1 to 5. Note the similarity in structure to the circuit’s correlation matrix in Figure 4. Also note that for the QNTK, the�matrix approximation ...
-
[90]
Encoder assumption and difference sets We assume each encoding block takes the commuting phase-generator form Sℓ(x) = exp��ix�G (ℓ)�, G (ℓ)= (G(ℓ) 1 ,...,G (ℓ) d ),(C1) where [G(ℓ) α,G (ℓ) β] = 0 for allα,β. Since theG(ℓ) α are commuting Hermitian operators, they admit a commo...
-
[91]
����� ���(Difference-frequency expansion)�LetS(x) = exp(�ix�G)with commuting Hermitian generators G= (G 1,...,Gd)and joint eigenbasis��λj��j
A single encoder insertion yields only difference frequencies The basic mechanism is that conjugation byS ℓ(x) produces phases given by eigenvalue differences (equivalently, transition frequencies between eigenspaces). ����� ���(Difference-frequency expansion)�LetS(x) = exp(�i...
-
[92]
Re-uploading: Minkowski-sum construction ofΩ Consider a depth-Lre-uploading circuit of the form U(x;θ) =WL(θ(L))SL(x)� � �W1(θ(1))S 1(x),(C5) and an observable modelf(x;θ) = Tr�OU(x;θ)ρU(x;θ)†�. Fixθ. By repeatedly expanding each encoder insertion using Lemma C.1, one finds th...
-
[93]
��������� ���(Path set and redundancy)�Let Ω (ℓ)be the difference set of layerℓ
Path sets and redundancy The Minkowski-sum picture suggests a natural decomposition of input harmonics into layer-wise spectral choices (or paths), which is particularly useful for discussing redundancy and frequency-dependent multiplicities [20, 21]. ��������� ���(Path set an...
-
[94]
Single-qubit Pauli encoder.LetS(x) =e −ixZ/2
Examples a. Single-qubit Pauli encoder.LetS(x) =e −ixZ/2. The generator has eigenvalues� 1 2, so the difference set is Ω (1) =�0,�1�. WithLre-uploadings, Ω� ��L,�L+ 1,...,L�1,L�.This recovers the standard band-limit scaling for trigonometric polynomials produced by repeated Pa...
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.