Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

A quantum model can be valuable for how its design shapes behavior, not just for how well it predicts.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 03:35 UTC pith:5BRHACO3

load-bearing objection A clear, honest perspective on why QML value might lie in interpretability; the motivating QFM/GP example is hypothetical without control of Fourier coefficients, but the inductive-bias review is genuinely useful. the 3 major comments →

arxiv 2607.13827 v2 pith:5BRHACO3 submitted 2026-07-15 quant-ph

Inherent interpretability provides inherent value in quantum machine learning

classification quant-ph PACS 03.67.Lx
keywords quantum machine learninginherent interpretabilityquantum Fourier modelsrandom Fourier featuresGaussian processeskernel designuncertainty quantificationinductive bias
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that quantum machine learning should be evaluated not only by predictive performance or classical simulability but by inherent interpretability: the mathematical structure that meaningfully shapes model behavior for a given task. As a motivating example, it shows that quantum Fourier models and random Fourier features induce the same kind of Gaussian process kernel for uncertainty quantification, but through opposite design routes. Random Fourier features require choosing a spectral distribution from the top down; quantum Fourier models generate an empirical frequency distribution from the bottom up, from the eigenvalue gaps of data-encoding Hamiltonians. If this framing is right, a quantum model can provide value even when it is classically simulatable or matched in accuracy, because its design space enables principled kernel design and interpretable discovery. The paper also surveys symmetry, metric, and topological inductive biases that give quantum models parameter-independent behavioral guarantees.

Core claim

Central claim: a quantum ML model's value can reside in its inherent interpretability—the mathematical structure that contributes meaningfully to desired model behavior—not just its accuracy. The motivating example rewrites a data-re-uploading circuit as a Fourier series, showing frequencies are sums of eigenvalue gaps of the data-encoding Hamiltonians. Under Gaussian assumptions on Fourier coefficients, the model induces a shift-invariant GP kernel structurally identical to random Fourier features; the difference is the design route. RFF chooses a spectral distribution top-down; quantum Fourier models generate an empirical frequency distribution bottom-up from Hamiltonian spectra. The paper

What carries the argument

The quantum Fourier model representation: a data-re-uploading circuit expanded as a Fourier series whose frequency vectors are sums over eigenvalue gaps λ(l)k(m*)−λ(l)k(m) of the data-encoding Hamiltonians across L blocks. With a Gaussian prior on the Fourier coefficients, the feature-map inner product induces the GP kernel (1/|Ω|)Σ_z 2cos(ω_z^T(x−x')), determined by the empirical distribution over frequencies created by Hamiltonian spectra. This machinery converts kernel design from choosing a spectral density (top-down, as in random Fourier features) into choosing data-encoding Hamiltonians (bottom-up), making Hamiltonian engineering the design lever for GP uncertainty quantification.

Load-bearing premise

The motivating GP-kernel construction assumes a quantum Fourier model's Fourier coefficients are independent standard Gaussians—an assumption the authors state does not hold for many practical circuits and that they do not yet know how to enforce on hardware.

What would settle it

Estimate the Fourier coefficients of a trained shallow quantum Fourier model by quantum tomography or classical simulation; if their distribution deviates from N(0,I) for the Hamiltonian choices used, the induced kernel formula will not match the kernel the model actually realizes, and the bottom-up kernel-design claim fails for real circuits.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Dequantization results showing that random Fourier features can match quantum Fourier model performance become an expected, even positive, outcome: numerical equivalence leaves the interpretability complementarity intact.
  • GP kernel design for uncertainty quantification can be approached bottom-up by choosing Hamiltonian families whose eigenvalue spectra generate frequency distributions with desired properties such as smoothness, periodicity, or sharpness.
  • Resource tradeoffs become explicit: small-dimensional dense feature matrices are classically simulatable; larger dimensions motivate computing pairwise kernel overlaps on a quantum device, limited by control of Fourier coefficients and logical qubit counts.
  • Inherently interpretable quantum models can carry guarantees that hold before training—symmetry invariance, certified robustness radii from bi-Lipschitz encodings, and homology preservation—so interpretability can be a design target rather than a post-hoc explanation.
  • QML evaluation should include design-space tools as a valued quantity, giving a reason to use a quantum model even when a classical model reaches comparable accuracy.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the coefficient-control problem is eventually solved, the quantum-Fourier-to-GP construction becomes a concrete kernel discovery tool: one could search Hamiltonian families for induced frequency distributions that beat hand-chosen spectral densities on calibration benchmarks—a testable extension not run in the paper.
  • The paper's motivating example depends on a Gaussian prior on Fourier coefficients that it admits is not enforceable; the symmetry, metric, and topological inductive-bias arguments do not depend on that assumption and may carry the interpretability thesis even if the kernel example remains hypothetical.
  • The framing suggests a new comparative metric for QML: design-space complementarity rather than performance delta. Two models with identical predictive behavior could still differ in the questions their internal structure lets a user answer.
  • A natural experimental check of the narrower example: estimate a trained shallow quantum Fourier model's actual kernel from samples and compare it to the formula (1/|Ω|)Σ 2 cos(ω_z^T(x−x')) derived from the accessible frequency set; deviations would indicate the Gaussian coefficient assumption is doing the work.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper argues that quantum machine learning (QML) model value should be assessed not primarily through predictive performance or classical simulability, but through a model's "inherent interpretability offerings"—the mathematical structure that contributes meaningfully to desired behavior in a specific ML task. The central motivating example compares random Fourier features (RFF) and quantum Fourier models (QFM) as two routes to constructing Gaussian-process kernels for uncertainty quantification: RFFs specify a spectral distribution top-down, while QFMs generate an empirical frequency distribution bottom-up from Hamiltonian choices. The paper then reviews symmetry, metric-geometric, and topological inductive biases from the QML literature as further instances of inherently interpretable design. The overall thesis is normative and programmatic rather than an empirical claim, and the mathematical examples are worked out in detail.

Significance. If the thesis were established, it would broaden how the QML community evaluates models, giving a principled counterweight to dequantization arguments and to benchmarks based solely on accuracy. The paper has real strengths: the RFF/QFM comparison is clean, the worked symmetry, metric, and topology examples are internally consistent and expose genuinely legible structure, and the authors are unusually candid about the assumptions behind their main example. However, the motivating example is not currently realized: Eq. (23) requires a Gaussian prior over Fourier coefficients with identity covariance, and the authors explicitly state that controlling these coefficients is an unsolved problem. In addition, the central value claim is largely definitional. The paper therefore reads as a promising perspective whose primary concrete support remains a hypothetical construction, with the review sections providing independent but softer evidence for the general thesis.

major comments (3)
  1. [§II.C, Eq. (23) and §II.D.1] The induced GP kernel for a QFM is derived by imposing the same Gaussian prior on Fourier coefficients as in the RFF setting, so that Eq. (23) depends only on the generated frequency set. The text itself concedes that "these assumptions do not hold for many practical quantum Fourier models" and that "we currently do not know how to control the Fourier coefficients in the quantum circuit." This is a load-bearing issue: without control of the coefficients, the actual kernel of a QFM is a weighted sum over frequencies whose weights are set by the circuit parameters and observable, not by the Hamiltonian alone. Thus the "bottom-up route to kernel design" is a hypothetical construction rather than a property of actual QFMs. The authors must either provide a concrete circuit/observable family for which the coefficients are (exactly or approximately) uniform, or explicitly reframe Eq. (23) as a
  2. [§I and §IV] The central claim—that model value can be found through characterization of inherent interpretability offerings—is stated in a way that is largely definitional. Since "inherent interpretability" is defined as "mathematical structure that contributes meaningfully to desired model behavior," the conclusion that such structure is valuable partially restates the premise. To make the thesis testable, the paper should specify what evidence could count against it, or operationalize value through observable outcomes (e.g., successful co-design with domain experts, maintainability, calibration of uncertainty estimates, or adoption decisions). Without such an operationalization, the review sections illustrate mathematical structure but do not empirically validate the value claim.
  3. [§III.A–III.B] The term "interpretability" is used expansively. A Lipschitz bound or a lower bound on trace distance is a robustness/stability guarantee; it is not, by itself, an explanation in the Rudin sense referenced in the introduction. Similarly, an equivariance constraint fixes model behavior by construction but only becomes interpretable when paired with a vocabulary (e.g., isotypic decomposition) that lets a user ask and answer questions about the model. The paper should distinguish between "mathematical structure that contributes to desired behavior" and "mechanisms that support human understanding or co-design." Without this distinction, the thesis risks reducing to the trivial claim that all useful inductive biases are valuable.
minor comments (5)
  1. [Abstract and §II.D] There is a typo in the abstract: "offerdifferenttools" should be "offer different tools."
  2. [References] Reference [43] lists "G. M. Canelles" as the author of "Gaussian Processes for Machine Learning"; this should be Rasmussen and Williams (2006). Reference [52] appears garbled as "L. Strauss, Pattern Recognition and Machine Learning"; the actual author is Bishop. Please verify these and all other references for correctness.
  3. [Fig. 1] Figure 1 is illustrative but hard to parse in the text version: the arrow from "Choose Hamiltonian families per block" to the frequency vector is broken, and the relationship between the two rows would benefit from a cleaner layout or an explicit caption explaining the top-down vs bottom-up contrast.
  4. [§II.C] The statement "each entry can be written as e^{-ix_i λ^T_{k^{(s)}}}" uses k^{(s)} for a binary string, but the inner-product notation is introduced informally; a short note that λ is the vector of eigenvalues associated with that binary string would improve readability for non-specialists.
  5. [Appendix A.2] "Kullback Libeler divergence" is a typo; it should be "Kullback–Leibler divergence."

Circularity Check

0 steps flagged

No significant circularity: the motivating example is a disclosed conditional derivation, and the review sections rest on textbook results and independent citations.

full rationale

Walking the paper's derivation chain: (1) the RFF kernel in Eq. 8 follows from a standard trigonometric identity and Bochner's theorem, cited to Rahimi & Recht; (2) the quantum Fourier model's Fourier-series form in Eqs. 19-21 is a direct linear-algebra expansion of the layered encoding, reproducing Schuld et al. [41] from independent published work; (3) the QFM kernel in Eq. 23 is obtained only after explicitly imposing the same Gaussian prior on Fourier coefficients as in the RFF setting: 'If we impose the same Gaussian prior distribution assumptions on the quantum model, we induce a GP as in the RFF setting.' That assumption is disclosed and then explicitly flagged as unrealized: 'these assumptions do not hold for many practical quantum Fourier models and may be difficult to enforce on quantum hardware' (Sec. II C), and 'we currently do not know how to control the Fourier coefficients in the quantum circuit' (Sec. II D.1). This is a feasibility gap, not circularity: Eq. 23 is a conditional derivation from a stated assumption, not a fitted parameter relabeled as a prediction nor a target defined in terms of the conclusion. The symmetry, metric-geometry, and topology sections are self-contained mathematical arguments relying on Schur's lemma, bi-Lipschitz bounds, Peter-Weyl/SNAG, and homology preservation, with standard textbook citations. The only self-citations ([83] and [84]) appear as illustrative examples in the review sections, not as load-bearing justifications of the central thesis. The central valuation claim is a normative framing--'value can be found through the characterization of its inherent interpretability offerings'--which is definitional in character, but the paper presents it as a perspective rather than as a derived prediction. A definitional normative premise does not, by itself, make the technical comparisons circular. No fitted-input-called-prediction, no imported uniqueness theorem, and no ansatz-smuggled-via-self-citation are present.

Axiom & Free-Parameter Ledger

0 free parameters · 7 axioms · 0 invented entities

The paper's core claim rests on a definitional value premise plus standard mathematics. The most fragile assumption is the hypothetical treatment of QFM Fourier coefficients as a Gaussian prior for GP kernel design; the authors flag this limitation themselves. No new physical or mathematical entities are proposed.

axioms (7)
  • ad hoc to paper A model's value can (and should) be assessed through its inherent interpretability offerings.
    Central normative premise, stated in the abstract and Sec. I; the paper does not derive it from external benchmarks.
  • standard math A Gaussian prior over model coefficients with identity covariance induces a Gaussian process whose kernel is the feature inner product.
    Sec. II A, Eqs. (2)-(4); standard weight-space/function-space duality in GP regression.
  • domain assumption Data-encoding Hamiltonian eigenvalue differences determine the accessible frequency spectrum of a quantum Fourier model.
    Sec. II C, Eqs. (21)-(23), following Schuld et al. [41]; assumes Hamiltonian selection is a practical design lever.
  • standard math The empirical frequency distribution generated by quantum Fourier models is symmetric and yields a shift-invariant kernel.
    Sec. II C; frequencies appear in +/- pairs from Eq. (21), so the kernel in Eq. (23) matches the RFF form.
  • ad hoc to paper The Gaussian-prior assumption on quantum Fourier coefficients can be ignored for GP kernel design.
    Sec. II C, paragraph beginning 'Although these assumptions do not hold...'; this is the paper's acknowledged hypothetical step.
  • domain assumption Interpretability is necessary for domain-adapted co-design and human adoption in high-risk ML.
    Sec. I, citing [39,40]; external empirical premise adopted from classical ML, not established here.
  • standard math Unitary transformations preserve trace distance and do not change the topological features of the embedded data manifold.
    Sec. III B-C, Eq. (43) and homology invariance; standard unitarity/homeomorphism facts.

pith-pipeline@v1.3.0-alltime-deepseek · 32480 in / 15357 out tokens · 156501 ms · 2026-08-02T03:35:22.876851+00:00 · methodology

0 comments
read the original abstract

The field of quantum machine learning (QML) evolved to value models believed to most directly rival those providing utility in classical ML, namely large-scale neural networks. Although more recently, classical ML has been learning a hard lesson with respect to deploying un-interpretable neural networks in the wild: model interpretability matters for domain-adapted co-design and human adoption. We adopt this larger ML perspective to argue that quantum ML model value can be found through the characterization of its inherent interpretability offerings -- i.e. its mathematical structure that contributes meaningfully to desired model behavior for the specific ML task. To support our perspective, we provide a motivating example of a characterization with quantum Fourier models and random Fourier features (RFF) as approaches to approximate Gaussian process (GP) kernels for uncertainty quantification tasks in ML. The top-down and bottom-up complementarity of the two mathematical constructions reveals that quantum Fourier models offer different tools than RFFs for principled GP kernel design and interpretable discovery for uncertainty quantification with real-world data. To showcase the rich variety of inductive biases enabled by quantum information tools, we review examples from the QML literature -- including symmetry, metric geometry, and topology -- that can be used to design inherently interpretable ML models for specific tasks. We hope this framing encourages the QML community to value the inherent components and mechanisms of quantum models separately from task performance, as inherent interpretability might be the reason that a quantum model, and potentially a quantum computer, gets used in practice for ML.

Figures

Figures reproduced from arXiv: 2607.13827 by Kaitlin Gili, Zachary P. Bradshaw.

Figure 1
Figure 1. Figure 1: FIG. 1: Visual representation of the inherent interpretability [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From hyperplanes to hyperellipsoids: characterizing the inherent interpretability of linear and single-qubit mixed-state binary classification models

    cs.LG 2026-07 conditional novelty 4.0

    A single-qubit mixed-state binary classifier is a centered hyperellipsoid classifier — equivalent to a linear model on squared features with normalized, non-negative weights.

Reference graph

Works this paper leans on

110 extracted references · 3 canonical work pages · cited by 1 Pith paper · 1 internal anchor

  1. [1]

    frequencies

    Computational resource considerations RFF models were originally introduced as a more scalable construction for GPs, allowing inference to scale with the number of features rather than directly with the number of datapoints [42]. When viewed through this lens, the quantum Fourier model does not appear to offer the same computational benefit. The number of...

  2. [2]

    Consider a training datasetD={x i, yi}N i=1, where xi ∈R d andy i ∈ {l1,

    Probabilistic machine learning Bayesian methods offer pathways to bake structure into ML settings, but they require one to view a standard regression or classification problem as one that learns parameters of a probability distribution or density func- tion. Consider a training datasetD={x i, yi}N i=1, where xi ∈R d andy i ∈ {l1, . . . , lK}. Here,y i is ...

  3. [3]

    updated prior

    Posterior estimation for prediction As stated previously, the posterior in Eq. A4 reveals the uncertainty over the base model parameters, given all of the training observations. Our aim is to estimate this posterior, so that it can be used as the “updated prior” when making predictions on unseen data. For cer- tain likelihood and prior distribution famili...

  4. [4]

    A4 is most useful when one is in search of agoodprobabilistic model, where good is defined by Occam’s razor – a simple model that fits the data well

    Bayesian evidence for hyper-parameter optimization The evidence term in Eq. A4 is most useful when one is in search of agoodprobabilistic model, where good is defined by Occam’s razor – a simple model that fits the data well. The evidence is typically written as the integral over the model parameters with respect to the data: p(y) = Z p(y|θ)p(θ)dθ(A10) Th...

  5. [5]

    In exact GP inference, one typically starts by defining a GP prior over functions f(x)∼ GP(0 N , kxx′), and a corresponding likelihood p(y|f(x))

    Exact GP inference A GP is defined as a Gaussian distribution over func- tions,f(x)∼ GP(µ, k(x i, x′ i)), whereµis a vector of mean function evaluations and the covariance matrix is determined by pairwise input similarities encoded by the kernel functionk(x i, x′ i). In exact GP inference, one typically starts by defining a GP prior over functions f(x)∼ G...

  6. [6]

    Preskill, Quantum computing in the NISQ era and beyond, Quantum2, 79 (2018)

    J. Preskill, Quantum computing in the NISQ era and beyond, Quantum2, 79 (2018)

  7. [7]

    Bharti, A

    K. Bharti, A. Cervera-Lierta, T. H. Kyaw, T. Haug, S. Alperin-Lea, A. Anand, M. Degroote, H. Heimonen, J. S. Kottmann, T. Menke,et al., Noisy intermediate- scale quantum algorithms, Reviews of Modern Physics 94, 015004 (2022)

  8. [8]

    Cerezo, A

    M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio,et al., Variational quantum algorithms, Na- ture Reviews Physics3, 625 (2020)

  9. [9]

    N. Moll, P. Barkoutsos, L. S. Bishop, J. M. Chow, A. Cross, D. J. Egger, S. Filipp, A. Fuhrer, J. M. Gam- betta, M. Ganzhorn,et al., Quantum optimization using variational algorithms on near-term quantum devices, Quantum Science and Technology3, 030503 (2017)

  10. [10]

    Wecker, M

    D. Wecker, M. B. Hastings, and M. Troyer, Progress to- wards practical quantum variational algorithms, Physi- cal Review A92, 042303 (2015)

  11. [11]

    Lubasch, J

    M. Lubasch, J. Joo, P. Moinier, M. Kiffner, and D. Jaksch, Variational quantum algorithms for nonlin- ear problems, Physical Review A101, 010301 (2019)

  12. [12]

    Farhi and H

    E. Farhi and H. Neven, Classification with quan- tum neural networks on near term processors, arXiv:1802.06002 10.37686/qrl.v1i2.80 (2018)

  13. [13]

    I. Cong, S. Choi, and M. D. Lukin, Quantum convolu- tional neural networks, Nature Physics15, 1273 (2018)

  14. [14]

    Bausch, Recurrent quantum neural networks, inNeu- 20 ral Information Processing Systems, Vol

    J. Bausch, Recurrent quantum neural networks, inNeu- 20 ral Information Processing Systems, Vol. 33 (Curran As- sociates, Inc., 2020)

  15. [15]

    Benedetti, B

    M. Benedetti, B. Coyle, M. Fiorentini, M. Lubasch, and M. Rosenkranz, Variational inference with a quantum computer, Physical Review Applied16, 044057 (2021)

  16. [16]

    K. Gili, M. Sveistrys, and C. Ballance, Introducing nonlinear activations into quantum generative models, Physical Review A107, 012406 (2022)

  17. [17]

    S. Y.-C. Chen, S. Yoo, and Y.-L. L. Fang, Quantum long short-term memory, inIEEE International Conference on Acoustics, Speech, and Signal Processing(IEEE,

  18. [18]

    E. A. Cherrat, I. Kerenidis, N. Mathur, J. Landman, M. Strahm, and Y. Y. Li, Quantum Vision Transform- ers, Quantum8, 1265 (2022)

  19. [19]

    Abbas, D

    A. Abbas, D. Sutter, C. Zoufal, A. Lucchi, A. Figalli, and S. Woerner, The power of quantum neural networks, Nature Computational Science1, 403 (2020)

  20. [20]

    K. Gili, M. Hibat-Allah, M. Mauri, C. Ballance, and A. Perdomo-Ortiz, Do quantum circuit born machines generalize?, Quantum Science and Technology8, 035021 (2022)

  21. [21]

    T. Hur, L. Kim, and D. K. Park, Quantum convolu- tional neural network for classical data classification, Quantum Machine Intelligence4, 3 (2021)

  22. [22]

    J. Y. Khoo, C. K. Gan, W. Ding, S. Carrazza, J. Ye, and J. F. Kong, Benchmarking quantum convolutional neural networks for classification and data compression tasks, arXiv preprint arXiv:2411.13468 (2024)

  23. [23]

    Bowles, S

    J. Bowles, S. Ahmed, and M. Schuld, Better than classi- cal? the subtle art of benchmarking quantum machine learning models, arXiv.org 10.48550/arXiv.2403.07059 (2024)

  24. [24]

    Basilewitsch, J

    D. Basilewitsch, J. F. Bravo, C. Tutschku, and F. Struckmeier, Quantum neural networks in practice: A comparative study with classical models from stan- dard data sets to industrial images, Quantum Machine Intelligence7, 110 (2024)

  25. [25]

    Ill´ esov´ a, T

    S. Ill´ esov´ a, T. Rybotycki, and M. Beseda, Qmetric: Benchmarking quantum neural networks across circuits, features, and training dimensions, QualITA (2025)

  26. [26]

    Schuld, V

    M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, and N. Killoran, Evaluating analytic gradients on quantum hardware, Physical Review A99, 032331 (2018)

  27. [27]

    Mcclean, S

    J. Mcclean, S. Boixo, V. Smelyanskiy, R. Babbush, and H. Neven, Barren plateaus in quantum neural net- work training landscapes, Nature Communications9, 10.1038/s41467-018-07090-4 (2018)

  28. [28]

    Holmes, A

    Z. Holmes, A. Arrasmith, B. Yan, P. J. Coles, A. Al- brecht, and A. T. Sornborger, Barren plateaus pre- clude learning scramblers, Physical Review Letters126, 190501 (2020)

  29. [29]

    S. Wang, E. Fontana, M. Cerezo, K. Sharma, A. Sone, L. Cincio, and P. J. Coles, Noise-induced barren plateaus in variational quantum algorithms, Nature communications12, 1 (2020)

  30. [30]

    Holmes, K

    Z. Holmes, K. Sharma, M. Cerezo, and P. J. Coles, Con- necting ansatz expressibility to gradient magnitudes and barren plateaus, PRX Quantum3, 010313 (2021)

  31. [31]

    Cerezo, A

    M. Cerezo, A. Sone, T. Volkoff, L. Cincio, and P. J. Coles, Cost function dependent barren plateaus in shal- low parametrized quantum circuits, Nature communi- cations12, 1 (2021)

  32. [32]

    Arrasmith, Z

    A. Arrasmith, Z. Holmes, M. Cerezo, and P. J. Coles, Equivalence of quantum barren plateaus to cost concen- tration and narrow gorges, Quantum Science and Tech- nology7, 045015 (2021)

  33. [33]

    Larocca, S

    M. Larocca, S. Thanasilp, S. Wang, K. Sharma, J. Bia- monte, P. J. Coles, L. Cincio, J. R. McClean, Z. Holmes, and M. Cerezo, A review of barren plateaus in varia- tional quantum computing, Nature Reviews Physics7, 174 (2025)

  34. [34]

    Shi and Y

    X. Shi and Y. Shang, Avoiding barren plateaus via gaus- sian mixture model, New Journal of Physics27, 104501 (2024)

  35. [35]

    X. Hao, G. Zhang, and S. Ma,Deep Learning(MIT Press, 2016)

  36. [36]

    Bengio, Y

    Y. Bengio, Y. LeCun, and G. Hinton, Deep learning for ai, Communications of the ACM64, 58 (2021)

  37. [37]

    Paleyes, R.-G

    A. Paleyes, R.-G. Urma, and N. D. Lawrence, Chal- lenges in deploying machine learning: A survey of case studies, ACM Computing Surveys55, 1 (2020)

  38. [38]

    Hendrickx, L

    K. Hendrickx, L. Perini, D. V. d. Plas, W. Meert, and J. Davis, Machine learning with a reject option: a survey, Machine-mediated learning 10.1007/s10994-024- 06534-x (2021)

  39. [39]

    Hasan, M

    M. Hasan, M. Abdar, A. Khosravi, U. Aickelin, P. Lio’, I. Hossain, A. Rahman, and S. Nahavandi, Survey on leveraging uncertainty estimation towards trustworthy deep neural networks: The case of reject option and post-training processing, ACM Computing Surveys57, 236 (2023)

  40. [40]

    Norori, Q

    N. Norori, Q. Hu, F. M. Aellen, F. D. Faraci, and A. Tzovara, Addressing bias in big data and ai for health care: A call for open science, Patterns2, 100347 (2021)

  41. [41]

    B. E. Turner, J. R. Steinberg, B. T. Weeks, F. Ro- driguez, and M. R. Cullen, Race/ethnicity reporting and representation in us clinical trials: A cohort study, The Lancet Regional Health – Americas11, 100252 (2022)

  42. [42]

    Bereska and E

    L. Bereska and E. Gavves, Mechanistic interpretabil- ity for AI safety – a review, Trans. Mach. Learn. Res. 10.48550/arXiv.2404.14082 (2024)

  43. [43]

    Somvanshi, M

    S. Somvanshi, M. M. Islam, A. Rafe, A. G. Tusti, A. Chakraborty, A. Baitullah, T. I. Chowdhury, N. Al- nawmasi, A. Dutta, and S. Das, Bridging the black box: A survey on mechanistic interpretability in ai, ACM Computing Surveys58, 1 (2026)

  44. [44]

    Zschech, S

    P. Zschech, S. Weinzierl, and M. Kraus, Inherently in- terpretable machine learning: A contrasting paradigm to post-hoc explainable ai, Business & Information Sys- tems Engineering , 1 (2025), received 17 Dec 2024; ac- cepted 24 Jul 2025; published 15 Sep 2025

  45. [45]

    C. Rudin, Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead, Nature Machine Intelligence1, 206 (2018), received 30 Dec 2018; accepted 26 Mar 2019; published 13 May 2019

  46. [46]

    Schuld, R

    M. Schuld, R. Sweke, and J. J. Meyer, The effect of data encoding on the expressive power of variational quantum-machine-learning models, Physical Review A 103, 032430 (2020)

  47. [47]

    Rahimi and B

    A. Rahimi and B. Recht, Random features for large- scale kernel machines, inNeural Information Processing Systems, Vol. 20 (Curran Associates, Inc., 2007)

  48. [48]

    G. M. Canelles,Gaussian Processes for Machine Learn- ing(MIT Press, Cambridge, MA, 2017)

  49. [49]

    Li and H

    J. Li and H. Wang, Gaussian processes regres- sion for uncertainty quantification: An intro- 21 ductory tutorial, arXiv preprint arXiv:2502.03090 10.48550/arXiv.2502.03090 (2025)

  50. [50]

    A. J. Parzygnat, T.-D. Bradley, A. Vlasic, and A. Pham, Toward structure-preserving quantum encodings, Phys- ical Review Research7, 041001 (2025)

  51. [51]

    Tang, Dequantizing algorithms to understand quan- tum advantage in machine learning, Nature Reviews Physics4, 692 (2022)

    E. Tang, Dequantizing algorithms to understand quan- tum advantage in machine learning, Nature Reviews Physics4, 692 (2022)

  52. [52]

    Belis, J

    V. Belis, J. Bowles, R. Gupta, E. Peters, and M. Schuld, Spectral methods: Crucial for machine learning, natural for quantum computers? (2026)

  53. [53]

    D. Sam, R. Pukdee, D. P. Jeong, Y. Byun, and J. Z. Kolter, Bayesian neural networks with domain knowledge priors, arXiv.org 10.48550/arXiv.2402.13410 (2024)

  54. [54]

    Harvey and M

    E. Harvey and M. C. Hughes, Occam’s razor is only as sharp as your elbo (2026)

  55. [55]

    A. G. Wilson and R. P. Adams, Gaussian process ker- nels for pattern discovery and extrapolation, inInter- national Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 28 (PMLR, 2013) pp. 1067–1075

  56. [56]

    Tompkins and F

    A. Tompkins and F. Ramos, Fourier feature approxi- mations for periodic kernels in time-series modelling, in AAAI Conference on Artificial Intelligence, Vol. 32 (As- sociation for the Advancement of Artificial Intelligence (AAAI), 2018)

  57. [57]

    Strauss,Pattern Recognition and Machine Learning, Information Science and Statistics (Springer, New York, NY, 2016)

    L. Strauss,Pattern Recognition and Machine Learning, Information Science and Statistics (Springer, New York, NY, 2016)

  58. [58]

    L´ azaro-Gredilla, J

    M. L´ azaro-Gredilla, J. Q. Candela, C. Rasmussen, and A. Figueiras-Vidal, Sparse spectrum gaussian process regression, Journal of Machine Learning Research11, 1865 (2010)

  59. [59]

    Gal and R

    Y. Gal and R. Turner, Improving the gaussian process sparse spectrum approximation by representing uncer- tainty in frequency inputs, inInternational Conference on Machine Learning, Proceedings of Machine Learn- ing Research, Vol. 37, edited by F. Bach and D. Blei (PMLR, Lille, France, 2015) pp. 655–664

  60. [60]

    A. G. Wilson, Z. Hu, R. Salakhutdinov, and E. P. Xing, Deep kernel learning, inInternational Conference on Artificial Intelligence and Statistics, Proceedings of Ma- chine Learning Research, Vol. 51, edited by A. Gretton and C. C. Robert (PMLR, Cadiz, Spain, 2015) pp. 370– 378

  61. [61]

    D. J. C. MacKay,Information Theory, Inference, and Learning Algorithms, Vol. 50 (Institute of Electrical and Electronics Engineers (IEEE), Cambridge, UK, 2004) pp. 2544–2545

  62. [62]

    J. Z. Liu, S. Padhy, J. Ren, Z. Lin, Y. Wen, G. Jerfel, Z. Nado, J. Snoek, D. Tran, and B. Lakshminarayanan, A simple approach to improve single-model deep un- certainty via distance-awareness, Journal of Machine Learning Research24, 1 (2022)

  63. [63]

    Harvey, M

    E. Harvey, M. Petrov, and M. C. Hughes, Learning hy- perparameters via a data-emphasized variational objec- tive, inarXiv.org, Proceedings of Machine Learning Re- search (2025)

  64. [64]

    Jaderberg, A

    B. Jaderberg, A. A. Gentile, Y. Berrada, E. Shishenina, and V. Elfving, Let quantum neural networks choose their own frequencies, Physical Review A109, 042421 (2023)

  65. [65]

    Landman, S

    J. Landman, S. Thabet, C. Dalyac, H. Mhiri, and E. Kashefi, Classically approximating variational quan- tum machine learning with random fourier features, in International Conference on Learning Representations (2022)

  66. [66]

    Shin, Y.-S

    S. Shin, Y.-S. Teo, and H. Jeong, Exponential data en- coding for quantum supervised learning, Physical Re- view A107, 012422 (2022)

  67. [67]

    Peters and M

    E. Peters and M. Schuld, Generalization despite overfit- ting in quantum machine learning models, Quantum7, 1210 (2022)

  68. [68]

    Mhiri, L

    H. Mhiri, L. Monbroussou, M. Herrero-Gonzalez, S. Thabet, E. Kashefi, and J. Landman, Constrained and vanishing expressivity of quantum fourier models, Quantum9, 1847 (2024)

  69. [69]

    Sweke, E

    R. Sweke, E. Recio, S. Jerbi, E. Gil-Fuster, B. Fuller, J. Eisert, and J. J. Meyer, Potential and limitations of random fourier features for dequantizing quantum ma- chine learning, Quantum9, 1640 (2023)

  70. [70]

    Strobl, M

    M. Strobl, M. E. Sahin, L. Horst, E. Kuehn, A. Streit, B. Jaderberg, and C.-C. Corr, Fourier fingerprints of ansatzes in quantum machine learning, arXiv preprint arXiv:2508.20868 10.48550/arXiv.2508.20868 (2025)

  71. [71]

    Sahebi, A

    M. Sahebi, A. Barthe, Y. Suzuki, Z. Holmes, and M. Grossi, On dequantization of supervised quantum machine learning via random fourier features (2025)

  72. [72]

    Tuysuz, O

    C. Tuysuz, O. Kyriienko, and M. Grossi, Quantum fourier generative models trainable at large scale, arXiv preprint arXiv:2606.28483 10.48550/arXiv.2606.28483 (2026)

  73. [73]

    S. Oh, E. J. Roh, A. V. Vasilakos, S. Park, and J. Kim, Fourier analysis perspective on quantum neural net- works, Communications Physics9, 176 (2026)

  74. [74]

    McArdle, S

    S. McArdle, S. Endo, A. Aspuru-Guzik, S. C. Benjamin, and X. Yuan, Quantum computational chemistry, Re- views of Modern Physics92, 015003 (2018)

  75. [75]

    Shawe-Taylor and N

    J. Shawe-Taylor and N. Cristianini,Kernel Methods for Pattern Analysis(Cambridge University Press, Cam- bridge, 2004)

  76. [76]

    Hofmann, B

    T. Hofmann, B. Scholkopf, and A. Smola, Kernel meth- ods in machine learning, The Annals of Statistics36, 1171 (2007)

  77. [77]

    Schuld, Supervised quantum machine learning mod- els are kernel methods (2021)

    M. Schuld, Supervised quantum machine learning mod- els are kernel methods (2021)

  78. [78]

    Fomichev, K

    S. Fomichev, K. Hejazi, M. S. Zini, M. Kiser, J. Morales, P. A. M. Casares, A. Delgado, J. Huh, A.-C. Voigt, J. E. Mueller,et al., Initial state preparation for quantum chemistry on quantum computers, PRX Quantum5, 040339 (2023)

  79. [79]

    J. J. Meyer, M. Mularski, E. Gil-Fuster, A. A. Mele, F. Arzani, A. Wilms, and J. Eisert, Exploiting sym- metry in variational quantum machine learning, PRX Quantum4, 010328 (2022)

  80. [80]

    Larocca, F

    M. Larocca, F. Sauvage, F. M. Sbahi, G. Verdon, P. J. Coles, and M. Cerezo, Group-invariant quantum machine learning, PRX Quantum3, 10.1103/prxquan- tum.3.030341 (2022)

Showing first 80 references.