Pith. sign in

REVIEW 4 major objections 6 minor 70 references

Swap-test attention matrices built from a shallow quantum circuit classify quantum phases from tens of training pairs and reveal phase-dependent correlation lengths.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Swap-test measurements of qubit-pair correlations plus a classical network classify cluster-Ising ground states into AFM, SPT, and paramagnetic phases with high accuracy from small training sets.

T0 review reviewed 2026-08-03 challenge →

load-bearing objection A clean swap-test-based attention construction for phase recognition, undermined by an incomplete evaluation protocol that needs fixing before the data-efficiency claims land. the 4 major comments →

arxiv 2602.00473 v1 pith:IO65DHQS submitted 2026-01-31 quant-ph cs.AIcs.LG

Quantum Phase Recognition via Quantum Attention Mechanism

classification quant-ph cs.AIcs.LG
keywords quantum phase recognitionattention mechanismswap testcluster-Ising modelsymmetry-protected topological phasehybrid quantum-classical modeldata-efficient learningcorrelation length
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes a hybrid quantum-classical model that classifies quantum phases by building an attention matrix from pairwise swap tests on the input quantum state. The central claim is that after a shallow parameterized circuit, the swap-test attention matrix contains enough correlation structure for a classical neural network to label ground states of the cluster-Ising model as antiferromagnetic, symmetry-protected topological, or paramagnetic. The model reportedly reaches about 98% accuracy on 9-qubit systems using only 20 training pairs, with similar behavior on 15 qubits. The paper further claims the learned attention matrices are interpretable: their patterns differ by phase, and derived quantities (a contrast value and an effective correlation length) track phase boundaries. If correct, this offers a data-efficient and inspectable route to quantum phase recognition that does not presuppose an order parameter.

Core claim

The core claim is that the correlation structure needed to distinguish quantum phases is directly accessible through a swap-test-based attention matrix. For a ground state loaded into n qubits, the model applies a shallow parameterized circuit, then measures the probability that an ancilla stays in |0> for each pair of qubits under a controlled swap, defining q_ij = 2P(0)-1. These q_ij form a symmetric attention matrix whose upper triangle is fed to a classical feedforward network. The paper shows that on the cluster-Ising model this pipeline reconstructs the phase diagram with high accuracy from very few labeled samples, and that the learned matrices exhibit distinct structural signatures i

What carries the argument

The central object is the swap-test attention matrix q with entries q_ij = 2P_ij(0) - 1, measuring how much the global state changes when qubits i and j are exchanged. This matrix is symmetric, learned implicitly through the parameterized circuit that shapes the state before the swap tests, and then processed by a classical classifier. Its role is to encode pairwise correlation structure in a physically interpretable form, replacing the query-key-value attention of classical transformers with a directly measurable quantity.

Load-bearing premise

The reported accuracy and phase-dependent attention patterns rest on the way ground states are labeled as AFM, SPT, or paramagnetic from the string order parameter ⟨S⟩, whose threshold for assigning labels is not stated, and on the test set being free of training states.

What would settle it

Re-running the classification with a training set explicitly disjoint from the test set, and with the string-order-labeling threshold varied across the crossover region, would settle whether the claimed 98% accuracy at 20 training pairs reflects generalization or label leakage.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Phase classification for the cluster-Ising model saturates near 98% accuracy with about 20 training pairs for 9 qubits, and similar behavior for 15 qubits, implying strong data efficiency for one-dimensional symmetry-protected topological systems.
  • The learned attention matrices show phase-specific patterns that are reproducible across independent training runs, so the model can serve as a diagnostic that highlights which spatial correlations carry phase information.
  • The contrast value C defined from near and long-range attention elements changes sign sharply at phase boundaries, providing a learned order-parameter-like signal that does not require a known order parameter.
  • The effective correlation length ξ extracted from the attention matrix is phase-dependent, with the largest values in the SPT phase, suggesting the model captures nonlocal correlation scales as well as the phase label.
  • Because the attention matrix encodes intrinsic correlations rather than a specified order parameter, the same architecture is claimed to extend to other models with topological or symmetry-protected phases.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to apply the same swap-test attention construction to a two-dimensional model or a model with a gapless phase, where the distance-dependent correlation function and contrast measure may behave differently.
  • If the attention matrix truly reflects entanglement structure, the fitted correlation length ξ might track known correlation lengths or entanglement length scales; comparing ξ to the conventional correlation length of the cluster-Ising model would validate the physics.
  • The paper leaves open whether the 15-qubit results use the same training-set sizes and labeling thresholds; a careful ablation varying the string-order labeling threshold would clarify the boundary of the data-efficiency claim.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a hybrid quantum-classical model for quantum phase recognition. Ground states of the cluster-Ising model are passed through a shallow parameterized quantum circuit, after which pairwise swap tests are used to build a symmetric attention matrix; the upper-triangular elements feed a classical feed-forward network that outputs a phase label. For N=9 and N=15, the authors report high accuracy with fewer than 100 training pairs, compare against a QCNN, and analyze the learned attention matrices to define a contrast measure and an effective correlation length that behave differently across the AFM, SPT, and paramagnetic phases. Appendix A derives the swap-test measurement, and Appendix B shows additional attention matrices.

Significance. If the evaluation is made rigorous, the work would be a useful contribution to quantum machine learning for many-body physics. The swap-test feature construction is simple, physically motivated, and analytically clean; the O(n^2) scaling is reasonable for near-term demonstrations. The claim that the model learns interpretable, phase-sensitive correlation structures is appealing. However, the manuscript does not currently establish the central data-efficiency claim because the evaluation protocol is underspecified, and the interpretability analysis lacks quantitative rigor. The strengths are the clear swap-test derivation, the use of exact diagonalization ground states, and the reporting of multiple independent training runs with confidence intervals. The missing elements are standard and fixable: a disclosed train/test split, a classical baseline, and a precise definition of the label and metric extraction procedures.

major comments (4)
  1. [Sec. III / Fig. 2] The evaluation does not establish held-out generalization. The text states that 2500 ground states are generated and that "the training set is then constructed by randomly sampling such data pairs from the full dataset," while Fig. 2 is described as "classification accuracy over the entire phase space." No sentence states that the test set is disjoint from the training set. For N=15 the model has roughly 60 PQC parameters plus about 315 FFN weights (105 upper-triangular attention entries times three output classes), so 20 training pairs without a disjoint test set cannot support a generalization claim. Please specify the exact train/test split, report accuracy on held-out states only, and provide class-wise results or confusion matrices.
  2. [Sec. III / Eq. (4)] The phase-label assignment rule is not specified. The labels are determined from the expectation value of the string order parameter <S>, but no threshold or decision rule is given, even though finite-size crossover makes the labeling ambiguous near phase boundaries. The crossover lines in Fig. 1(c) are located from maxima of the second derivative of <H>, but the mapping from continuous <S> values to the three discrete labels is never stated. Please state the labeling threshold (or the procedure used for each sampled point), and quantify how sensitive the reported accuracies are to states near the crossover boundaries.
  3. [Sec. III / Fig. 2] The comparison with QCNN is inadequately documented, and no classical baseline is provided. Fig. 2 shows QCNN curves, but the QCNN architecture, hyperparameters, and training procedure are not described. More importantly, there is no classical baseline trained on the same swap-test features (e.g., logistic regression or a classical FFN on q_ij), nor a baseline using conventional two-point correlators. Without such a baseline, the reader cannot tell whether the PQC and attention mechanism add anything beyond simple pairwise correlations. Please add these baselines and describe the QCNN implementation in enough detail to be reproducible.
  4. [Sec. III.A / Eq. (5) and Sec. III.B / Eq. (6)] The quantitative interpretability claims are not backed by a precise protocol. The contrast C in Eq. (5) requires selecting a "qubit with significantly reduced correlations" and a "distant qubit," but the selection criterion is not algorithmic; Appendix B shows that the weakly correlated qubit can change between runs (e.g., qubit 7 vs. qubit 1 for N=9, and qubits 5, 8, or 11 for N=15). Likewise, the effective correlation length xi is obtained by fitting f(r) to an exponential, but the fitting range, normalization, and uncertainty are not stated, and Fig. 5 has no error bars. The claims that xi is "one order smaller" in AFM than in SPT, or that sharp sign changes in C delineate boundaries, need a well-defined extraction procedure and statistics over independent runs.
minor comments (6)
  1. [Section II.A heading] Typo: "Quautum" should be "Quantum."
  2. [Sec. III / text before Fig. 3] Grammatical error: "with a the rest" should be "with the rest."
  3. [Appendix A, Eq. (A4)] The displayed equation omits normalization: the first equality should include 1/4 and the norm squared; the final result is correct for normalized states but the intermediate expression as written is not rigorous.
  4. [Sec. III.B] The abstract and conclusion call xi a "physical length scale," but Sec. III.B correctly cautions that f(r) is a learned attention proxy, not a conventional two-point correlator. Please align the language in the abstract and conclusion with that caveat.
  5. [General] No data or code availability statement is provided. Given the reproducibility claims, the authors should state whether simulation code and datasets will be made available.
  6. [Sec. III / Fig. 5] The numerical values quoted for xi (10^8, 10^9, ~10) appear without units and without description of the fitting procedure; please clarify whether these are dimensionless fit coefficients and specify the fit form and the r-range used.

Circularity Check

0 steps flagged

No formal circularity: the classifier is a supervised model operating on swap-test features, not on the order-parameter labels; the interpretability analysis is post-hoc but does not reduce to a tautology.

full rationale

The paper's central pipeline is a standard supervised learning task: ground states are labeled by the string order parameter ⟨S⟩ (Sec. III), and the model is trained to map swap-test-derived attention matrices to those labels. The attention matrix q_ij is obtained from a physical swap-test overlap, not from ⟨S⟩, so classification accuracy is not guaranteed by construction; it is an empirical result about the informativeness of these features. The subsequent analyses of attention matrices, contrast C, and correlation length ξ are interpretability tools applied to the trained model. They are influenced by the training labels, so the phrase 'without prior knowledge of the optimal order parameter' is somewhat overstated, but this is a framing caveat rather than a circular derivation: no equation in the paper makes ξ or q equal to the label b by construction. The only self-citation [42] is background motivation and is not load-bearing. Concerns about the missing explicit train/test split and the absence of a classical baseline are methodological and affect generalization claims, but they do not constitute circularity. The paper also self-acknowledges its limitation to small sizes and ideal simulators. Overall, no step in the derivation chain reduces to its own input by definition or by fitted-parameter renaming.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 0 invented entities

The model's central claims rest on supervised labels from the string order parameter, a hand-selected PQC ansatz, and several fitted diagnostics (ξ, the reference qubit for C). No new physical entities are introduced. The main uncharged assumptions are the phase-labeling rule, the expressivity of the shallow circuit, and the stability of the trained attention patterns.

free parameters (6)
  • PQC parameters θ (4nl) = not reported; 36 (N=9) / 60 (N=15) parameters for l=1
    Trained by gradient descent on labeled phase data; central to feature extraction.
  • FFN weights w and bias b = not reported
    Trained classifier weights mapping the attention matrix to phase labels.
  • Model hyperparameters (l=1, optimizer, learning rate, epochs, initialization) = not reported
    Chosen by hand; not specified, and they directly affect all reported results.
  • Phase-label threshold for ⟨S⟩ = not stated
    No explicit rule is given for converting ⟨S⟩ expectation into one of three phase labels; labels near crossover are ambiguous.
  • Effective correlation length ξ = ≈10^8–10^9 (AFM/SPT), ≈10 (paramagnetic)
    Obtained by exponential fit f(r)∼e^{-r/ξ}; no error bars; used to support the length-scale claims.
  • Reference qubit choice for contrast C = e.g., q12 and q19 for N=9 at h1≈0.39
    Hand-selected 'qubit with significantly reduced correlations'; not algorithmically defined.
axioms (6)
  • domain assumption Exact diagonalization of H in Eq. (3) yields exact ground states for N=9 and N=15.
    Used to generate the 2500-state datasets; exact for these sizes but only in noiseless simulation.
  • domain assumption The string order parameter ⟨S⟩ (Eq. 4) reliably distinguishes SPT, AFM, and paramagnetic phases in finite systems.
    The phase labels are derived from ⟨S⟩, but no threshold is specified.
  • domain assumption Crossover boundaries are identified by maxima of the second derivative of ⟨H⟩ with respect to h2.
    Used to construct the phase diagram and to define regions for labels.
  • standard math Swap-test outcome q_ij = 2P(0)-1 equals ⟨ψ|SWAP_ij|ψ⟩ and is a meaningful pairwise correlation measure.
    Derived in Appendix A for pure states; assumes ideal measurements with no noise.
  • domain assumption A single-layer RY+CRX PQC (l=1) is expressive enough for phase-discriminative feature extraction.
    The model's success depends on this heuristic; no expressibility analysis is provided.
  • domain assumption Learned attention patterns are stable across random initializations.
    Supported only qualitatively by Appendix B; no quantitative stability measure is given.

reviewed 2026-08-03 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Quantum Phase Recognition via Quantum Attention Mechanism." pith.science (2026). https://pith.science/paper/IO65DHQS

@misc{pith2026260200473,
  author       = {Pith},
  title        = {Pith review of: Quantum Phase Recognition via Quantum Attention Mechanism},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IO65DHQS}},
  note         = {Machine review of arXiv:2602.00473}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Quantum phase transitions in many-body systems are fundamentally characterized by complex correlation structures, which pose computational challenges for conventional methods in large systems. To address this, we propose a hybrid quantum-classical attention model. This model uses an attention mechanism, realized through swap tests and a parameterized quantum circuit, to extract correlations within quantum states and perform ground-state classification. Benchmarked on the cluster-Ising model with system sizes of 9 and 15 qubits, the model achieves high classification accuracy with less than 100 training data and demonstrates robustness against variations in the training set. Further analysis reveals that the model successfully captures phase-sensitive features and characteristic physical length scales, offering a scalable and data-efficient approach for quantum phase recognition in complex many-body systems.

Figures

Figures reproduced from arXiv: 2602.00473 by Jin-Long Chen, Xin Li, Zhang-qi Yin.

Figure 1
Figure 1. Figure 1: FIG. 1. (a) Over framework of attention model. The input state [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. Classification accuracy of QCNN and attention-based [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Heatmaps of attention matrices revealing distinct correlation patterns in different quantum phases. The axes correspond [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4. Contrast value [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5. Effective correlation length [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6. The swap test circuit [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: FIG. 7. Measured attention matrix for the qubit register for 9-qubit and (b) 15-qubit system. Each element represents the [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

70 extracted references · 8 linked inside Pith

  1. [1]

    Sachdev, Quantum phase transitions, Physics World 12, 33 (1999)

    S. Sachdev, Quantum phase transitions, Physics World 12, 33 (1999)

  2. [2]

    Vojta, Quantum phase transitions, Reports on Progress in Physics66, 2069 (2003)

    M. Vojta, Quantum phase transitions, Reports on Progress in Physics66, 2069 (2003)

  3. [3]

    S. L. Sondhi, S. Girvin, J. Carini, and D. Shahar, Con- tinuous quantum phase transitions, Reviews of modern physics69, 315 (1997)

  4. [4]

    Amico, R

    L. Amico, R. Fazio, A. Osterloh, and V. Vedral, En- tanglement in many-body systems, Reviews of modern physics80, 517 (2008)

  5. [5]

    T. J. Osborne and M. A. Nielsen, Entanglement in a simple quantum phase transition, Physical Review A66, 032110 (2002)

  6. [6]

    Osterloh, L

    A. Osterloh, L. Amico, G. Falci, and R. Fazio, Scaling of entanglement close to a quantum phase transition, Na- ture416, 608 (2002)

  7. [7]

    Goldenfeld,Lectures on phase transitions and the renormalization group(Addison-Wesley, Boston, 1992)

    N. Goldenfeld,Lectures on phase transitions and the renormalization group(Addison-Wesley, Boston, 1992)

  8. [8]

    Malpetti and T

    D. Malpetti and T. Roscilde, Quantum mean-field ap- proximation for lattice quantum models: Truncating quantum correlations and retaining classical ones, Phys- ical Review B95, 075112 (2017)

  9. [9]

    Kohn, Nobel lecture: Electronic structure of mat- ter—wave functions and density functionals, Reviews of Modern Physics71, 1253 (1999)

    W. Kohn, Nobel lecture: Electronic structure of mat- ter—wave functions and density functionals, Reviews of Modern Physics71, 1253 (1999)

  10. [10]

    Schuch and F

    N. Schuch and F. Verstraete, Computational complexity of interacting electrons and fundamental limitations of density functional theory, Nature Physics5, 732 (2009)

  11. [11]

    S. R. White, Density matrix formulation for quantum renormalization groups, Physical Review Letters69, 2863 (1992)

  12. [12]

    S. R. White, Density-matrix algorithms for quantum renormalization groups, Physical Review b48, 10345 (1993)

  13. [13]

    Becca and S

    F. Becca and S. Sorella,Quantum Monte Carlo ap- proaches for correlated systems(Cambridge University Press, 2017)

  14. [14]

    Carlson, S

    J. Carlson, S. Gandolfi, F. Pederiva, S. C. Pieper, R. Schi- avilla, K. E. Schmidt, and R. B. Wiringa, Quantum monte carlo methods for nuclear physics, Reviews of Modern Physics87, 1067 (2015)

  15. [15]

    Eisert, M

    J. Eisert, M. Cramer, and M. B. Plenio, Colloquium: Area laws for the entanglement entropy, Reviews of Mod- ern Physics82, 277 (2010)

  16. [16]

    Troyer and U.-J

    M. Troyer and U.-J. Wiese, Computational complexity and fundamental limitations to fermionic quantum monte carlo simulations, Physical Review Letters94, 170201 (2005)

  17. [17]

    R. P. Feynman, Simulating physics with computers, in Feynman and computation(cRc Press, 2018) pp. 133– 153

  18. [18]

    Fauseweh, Quantum many-body simulations on digi- tal quantum computers: State-of-the-art and future chal- lenges, Nature Communications15, 2123 (2024)

    B. Fauseweh, Quantum many-body simulations on digi- tal quantum computers: State-of-the-art and future chal- lenges, Nature Communications15, 2123 (2024)

  19. [19]

    Peruzzo, J

    A. Peruzzo, J. McClean, P. Shadbolt, M.-H. Yung, X.-Q. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’brien, A variational eigenvalue solver on a photonic quantum processor, Nature communications5, 4213 (2014)

  20. [20]

    Tilly, H

    J. Tilly, H. Chen, S. Cao, D. Picozzi, K. Setia, Y. Li, E. Grant, L. Wossnig, I. Rungger, G. H. Booth,et al., The variational quantum eigensolver: a review of meth- ods and best practices, Physics Reports986, 1 (2022)

  21. [21]

    Huang, R

    H.-Y. Huang, R. Kueng, and J. Preskill, Predicting many properties of a quantum system from very few measure- ments, Nature Physics16, 1050 (2020)

  22. [22]

    Huang, R

    H.-Y. Huang, R. Kueng, G. Torlai, V. V. Albert, and J. Preskill, Provably efficient machine learning for quan- tum many-body problems, Science377, eabk3333 (2022)

  23. [23]

    Carleo, I

    G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborov´ a, Machine learning and the physical sciences, Reviews of Modern Physics91, 045002 (2019)

  24. [24]

    E. P. Van Nieuwenburg, Y.-H. Liu, and S. D. Huber, Learning phase transitions by confusion, Nature Physics 13, 435 (2017)

  25. [25]

    Carrasquilla and R

    J. Carrasquilla and R. G. Melko, Machine learning phases of matter, Nature Physics13, 431 (2017)

  26. [26]

    Schindler, N

    F. Schindler, N. Regnault, and T. Neupert, Probing many-body localization with neural networks, Physical Review B95, 245134 (2017)

  27. [27]

    Greplova, A

    E. Greplova, A. Valenti, G. Boschung, F. Sch¨ afer, N. L¨ orch, and S. D. Huber, Unsupervised identification of topological phase transitions using predictive models, New Journal of Physics22, 045003 (2020)

  28. [28]

    D.-L. Deng, X. Li, and S. Das Sarma, Machine learning topological states, Physical Review B96, 195145 (2017)

  29. [29]

    Biamonte, P

    J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, and S. Lloyd, Quantum machine learning, Na- ture549, 195 (2017)

  30. [30]

    Schuld, I

    M. Schuld, I. Sinayskiy, and F. Petruccione, An intro- duction to quantum machine learning, Contemporary Physics56, 172 (2015)

  31. [31]

    Schuld and F

    M. Schuld and F. Petruccione, Supervised learning with quantum computers, Quantum Science and Technology 17(2018)

  32. [32]

    Dunjko and H

    V. Dunjko and H. J. Briegel, Machine learning & artifi- cial intelligence in the quantum domain: a review of re- cent progress, Reports on Progress in Physics81, 074001 (2018)

  33. [33]

    Y. Du, Z. Tu, X. Yuan, and D. Tao, Efficient measure for the expressivity of variational quantum algorithms, Physical Review Letters128, 080506 (2022)

  34. [34]

    Abbas, D

    A. Abbas, D. Sutter, C. Zoufal, A. Lucchi, A. Figalli, and S. Woerner, The power of quantum neural networks, Nature Computational Science1, 403 (2021)

  35. [35]

    Holmes, K

    Z. Holmes, K. Sharma, M. Cerezo, and P. J. Coles, Con- necting ansatz expressibility to gradient magnitudes and barren plateaus, PRX quantum3, 010313 (2022)

  36. [36]

    Huang, M

    H.-Y. Huang, M. Broughton, J. Cotler, S. Chen, J. Li, M. Mohseni, H. Neven, R. Babbush, R. Kueng, J. Preskill,et al., Quantum advantage in learning from experiments, Science376, 1182 (2022). 10

  37. [37]

    Y. Liu, S. Arunachalam, and K. Temme, A rigorous and robust quantum speed-up in supervised machine learn- ing, Nature Physics17, 1013 (2021)

  38. [38]

    Farhi and H

    E. Farhi and H. Neven, Classification with quantum neu- ral networks on near term processors, arXiv preprint arXiv:1802.06002 (2018)

  39. [39]

    I. Cong, S. Choi, and M. D. Lukin, Quantum convolu- tional neural networks, Nature Physics15, 1273 (2019)

  40. [40]

    J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Bab- bush, and H. Neven, Barren plateaus in quantum neural network training landscapes, Nature Communications9, 4812 (2018)

  41. [41]

    Pesah, M

    A. Pesah, M. Cerezo, S. Wang, T. Volkoff, A. T. Sorn- borger, and P. J. Coles, Absence of barren plateaus in quantum convolutional neural networks, Physical Review X11, 041011 (2021)

  42. [42]

    X. Li, D. Zhang, and Z.-Q. Yin, Unsupervised detection of topological phase transitions with a quantum reservoir, Physical Review A , 012422 (2025)

  43. [43]

    Bahdanau, K

    D. Bahdanau, K. Cho, and Y. Bengio, Neural machine translation by jointly learning to align and translate, arXiv preprint arXiv:1409.0473 (2014)

  44. [44]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, At- tention is all you need, Advances in neural information processing systems30(2017)

  45. [45]

    Zhang and Q

    H. Zhang and Q. Zhao, A survey of quantum trans- formers: Architectures, challenges and outlooks, arXiv preprint arXiv:2504.03192 (2025)

  46. [46]

    G. Li, X. Zhao, and X. Wang, Quantum self-attention neural networks for text classification, Science China In- formation Sciences67, 142501 (2024)

  47. [47]

    Xue, Z.-Y

    C. Xue, Z.-Y. Chen, X.-N. Zhuang, Y.-J. Wang, T.-P. Sun, J.-C. Wang, H.-Y. Liu, Y.-C. Wu, Z.-L. Wang, and G.-P. Guo, End-to-end quantum vision transformer: To- wards practical quantum speedup in large-scale models, arXiv preprint arXiv:2402.18940 (2024)

  48. [48]

    Zhang, Q

    H. Zhang, Q. Zhao, M. Zhou, and L. Feng, Hqvit: Hy- brid quantum vision transformer for image classification, arXiv preprint arXiv:2504.02730 (2025)

  49. [49]

    F. Chen, Q. Zhao, L. Feng, C. Chen, Y. Lin, and J. Lin, Quantum mixed-state self-attention network, Neural Networks185, 107123 (2025)

  50. [50]

    J. He, Y. Kan, and C. Xue, Training quantum self- attention model in near-term quantum computer, in2024 16th International Conference on Wireless Communica- tions and Signal Processing (WCSP)(IEEE, 2024) pp. 139–144

  51. [51]

    Heredge, M

    J. Heredge, M. West, L. Hollenberg, and M. Sevior, Nonunitary quantum machine learning, Physical Review Applied23, 044046 (2025)

  52. [52]

    A. M. Childs and N. Wiebe, Hamiltonian simulation using linear combinations of unitary operations, arXiv preprint arXiv:1202.5822 (2012)

  53. [53]

    Gily´ en, Y

    A. Gily´ en, Y. Su, G. H. Low, and N. Wiebe, Quantum singular value transformation and beyond: exponential improvements for quantum matrix arithmetics, inPro- ceedings of the 51st annual ACM SIGACT symposium on theory of computing(2019) pp. 193–204

  54. [54]

    E. A. Cherrat, I. Kerenidis, N. Mathur, J. Landman, M. Strahm, and Y. Y. Li, Quantum vision transformers, Quantum8, 1265 (2024)

  55. [55]

    Khatri, G

    N. Khatri, G. Matos, L. Coopmans, and S. Clark, Quixer: A quantum transformer model, arXiv preprint arXiv:2406.04305 (2024)

  56. [56]

    Comajoan Cara, G

    M. Comajoan Cara, G. R. Dahale, Z. Dong, R. T. Forestano, S. Gleyzer, D. Justice, K. Kong, T. Magorsch, K. T. Matchev, K. Matcheva,et al., Quantum vision transformers for quark–gluon classification, Axioms13, 323 (2024)

  57. [57]

    E. B. Unlu, M. Comajoan Cara, G. R. Dahale, Z. Dong, R. T. Forestano, S. Gleyzer, D. Justice, K. Kong, T. Magorsch, K. T. Matchev,et al., Hybrid quantum vision transformers for event classification in high energy physics, Axioms13, 187 (2024)

  58. [58]

    S. Sim, P. D. Johnson, and A. Aspuru-Guzik, Express- ibility and entangling capability of parameterized quan- tum circuits for hybrid quantum-classical algorithms, Ad- vanced Quantum Technologies2, 1900070 (2019)

  59. [59]

    Smacchia, L

    P. Smacchia, L. Amico, P. Facchi, R. Fazio, G. Flo- rio, S. Pascazio, and V. Vedral, Statistical mechanics of the cluster ising model, Physical Review A84, 022304 (2011)

  60. [60]

    Verresen, R

    R. Verresen, R. Moessner, and F. Pollmann, One- dimensional symmetry protected topological phases and their transitions, Physical Review B96, 165124 (2017)

  61. [61]

    F. D. M. Haldane, Nonlinear field theory of large-spin heisenberg antiferromagnets: semiclassically quantized solitons of the one-dimensional easy-axis n´ eel state, Phys- ical Review Letters50, 1153 (1983)

  62. [62]

    Pollmann and A

    F. Pollmann and A. M. Turner, Detection of symmetry- protected topological phases in one dimension, Physical Review B86, 125441 (2012)

  63. [63]

    Javadi-Abhari, M

    A. Javadi-Abhari, M. Treinish, K. Krsulich, C. J. Wood, J. Lishman, J. Gacon, S. Martiel, P. D. Nation, L. S. Bishop, A. W. Cross,et al., Quantum computing with qiskit, arXiv preprint arXiv:2405.08810 (2024)

  64. [64]

    Bergholm, J

    V. Bergholm, J. Izaac, M. Schuld, C. Gogolin, S. Ahmed, V. Ajith, M. S. Alam, G. Alonso-Linaje, B. Akash- Narayanan, A. Asadi,et al., Pennylane: Automatic dif- ferentiation of hybrid quantum-classical computations, arXiv preprint arXiv:1811.04968 (2018)

  65. [65]

    M. C. Caro, H.-Y. Huang, M. Cerezo, K. Sharma, A. Sornborger, L. Cincio, and P. J. Coles, Generalization in quantum machine learning from few training data, Na- ture Communications13, 4919 (2022)

  66. [66]

    Malpetti and T

    D. Malpetti and T. Roscilde, Quantum correlations, separability, and quantum coherence length in equilib- rium many-body systems, Physical Review Letters117, 130401 (2016)

  67. [67]

    Y.-J. Liu, A. Smith, M. Knap, and F. Pollmann, Model- independent learning of quantum phases of matter with quantum convolutional neural networks, Physical Review Letters130, 220603 (2023)

  68. [68]

    Caron, H

    M. Caron, H. Touvron, I. Misra, H. J´ egou, J. Mairal, P. Bojanowski, and A. Joulin, Emerging properties in self-supervised vision transformers, inProceedings of the IEEE/CVF international conference on computer vision (2021) pp. 9650–9660

  69. [69]

    Y. Shao, F. Wei, S. Cheng, and Z. Liu, Simulating noisy variational quantum algorithms: A polynomial approach, Physical Review Letters133, 120603 (2024)

  70. [70]

    Preskill, Quantum computing in the nisq era and be- yond, Quantum2, 79 (2018)

    J. Preskill, Quantum computing in the nisq era and be- yond, Quantum2, 79 (2018)

This paper was first reviewed by deepseek-v4-flash on August 3, 2026.