REVIEW 4 major objections 4 minor 79 references
Attention in Krylov Space: Transformer-Based Extrapolation of Lanczos Coefficients
T0 review · 4 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read A transformer trained on short Lanczos prefixes extrapolates operator-growth coefficients with roughly ten times lower error than asymptotic fits, and transfers without retraining from small to larger systems.
desk verdict Solid in-horizon forecasting of Lanczos coefficients, but the 'deep Krylov space' claim is untested: the model never predicts beyond the largest index it saw in training. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a decoder-only transformer with masked multi-head self-attention—a sequence model that assigns learned weights to all past coefficients when predicting the next one—applied to the causal sequence of coefficient differences Δb_n = b_n − b_{n−1}. Predicting differences rather than raw b_n stabilizes the target scale, and the causal mask enforces that predictions depend only on earlier coefficients. Attention then captures long-range correlations that the iterative Lanczos recursion implicitly propagates.
What would settle it
Recompute the long-tail coefficients for the same L=8 Hamiltonians with arbitrary-precision arithmetic or via an independent moment-based recursion, then rerun the RMSE comparison; if the transformer's order-of-magnitude advantage over the asymptotic fit disappears or shrinks substantially, the central claim fails. Also check whether the L=12 'true' sequences used for zero-shot evaluation saturate before n=30—if they do, the benchmark is against truncated labels.
Extended reading notes
Core claim
The paper claims that Lanczos coefficients are not just an asymptotic ramp; they carry subleading, history-dependent structure—even-odd staggering, slow deviations, and early-coefficient anchors—that standard fits miss. A small decoder-only transformer, trained on the increment sequence Δb_n = b_n − b_{n−1} from thousands of Hamiltonian instances, predicts coefficients for n>10 autoregressively. On unseen instances of the classical XYZ spin top and a chaotic transverse-field Ising chain with L=8, the transformer's coefficient RMSE is consistently an order of magnitude smaller than the universal-operator-growth-hypothesis asymptotic fit with staggering; reconstructed Krylov complexity and aut
Load-bearing premise
The full-reorthogonalization Lanczos output used as ground truth is exact over the entire horizon (T=100 for XYZ, T=30 for TFIM); if loss of orthogonality or finite-size saturation corrupts the long-tail labels, the reported error reductions are measured against corrupted targets.
Editorial extensions
If this is right
- Coefficient extrapolation can be treated as sequence prediction: a model trained on short prefixes reproduces long Lanczos tails with order-of-magnitude lower RMSE than linear asymptotic fits on chaotic systems.
- Zero-shot transfer from L=8 to L=12 means training data can be collected at accessible sizes and deployed to larger systems where direct Lanczos iteration is prohibitively expensive.
- Observable-level reconstruction—Krylov complexity and autocorrelation—inherits the extrapolation accuracy, with errors orders of magnitude smaller at late times.
- Attention ablations show that removing long-range context, parity information, or the first few coefficients degrades accuracy by an order of magnitude, so Lanczos sequences should be modeled as long-range causal sequences rather than fitted to a local asymptotic form.
Reading between the lines
- One passage to weigh: the abstract credits the model with extrapolating integrable regimes, but the body's Sec. VI lists that as an open direction; a reader should treat the integrable-regime claim as stated but unsupported in the current text.
- The success of the increment-sequence representation suggests that other spectral quantities derived from autocorrelation moments—not just Lanczos coefficients—may be forecastable by the same autoregressive scheme.
- If attention weights on early coefficients act as global anchors, then an even shorter input prefix plus system-size metadata might suffice for transfer; this is a testable ablation.
- A practical deployment would train on small chains and then predict coefficients for a Hamiltonian one cannot iterate; verifying against exactly solvable or sparse Hamiltonians at intermediate sizes would sharpen the zero-shot claim.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper frames the Lanczos coefficients {b_n} of operator growth as a causal time series and trains a masked decoder-only transformer to autoregressively predict future coefficient differences Δb_n from a short prefix. The model is benchmarked on the classical XYZ top (T=100) and on a chaotic transverse-field Ising chain (L=8, T=30, with an L=12 transfer), and it is compared against the standard asymptotic UOGH fit with an even-odd staggering term. Reported results include RMSE reductions in coefficient prediction, much smaller reconstruction errors in Krylov complexity K(t) and autocorrelation C(t), zero-shot transfer from L=8 to L=12, and attention-map/ablation analyses. The paper's central claim is that a transformer can serve as a practical surrogate for probing operator dynamics deep in Krylov space where brute-force Lanczos iteration is prohibitive.
Significance. If fully substantiated, the work would provide a useful ML-based tool for extending Lanczos sequences, and the zero-shot system-size transfer is a genuinely valuable feature. The paper is clearly written, the architecture is standard and reproducible from the hyperparameter table, and the attention ablation is a nice mechanistic addition. The main strengths are the direct comparison to the standard asymptotic baseline, the validation on two different chaotic models, and the transfer experiment. However, the support for the 'deep Krylov space' claim is incomplete: all evaluations stay within the sequence length T used in training, and the abstract claims an integrable-regime result that the body of the paper does not report. These issues are fixable with additional experiments or appropriately reworded claims, but they are load-bearing for the paper's central contribution as stated.
major comments (4)
- [Sec. III, Eq. (24); Sec. IV, Figs. 3-6] The evaluation horizon equals the training horizon. The loss in Eq. (24) trains next-token predictions for all m from n_in to T-1, so every target index m+1 <= T is present in the training sequences. All reported coefficient RMSEs (Figs. 3(b), 4(b), 6) and observable reconstructions (Fig. 5) are for n <= T. Thus the model is tested on continuation within the length distribution seen in training, not on forecasting coefficients for n > T. The L=12 transfer in Fig. 6 is a system-size extrapolation, not a depth extrapolation: the n range is still 1..30. Because the abstract and Secs. I and VI claim the model can probe 'deep in Krylov space' and act 'as a practical surrogate' where direct Lanczos is prohibitive, this is a central unsupported claim. I request an explicit test in which the training sequences are truncated at T_train < T_max and the model is evaluated for T_train < n <= T_max,
- [Abstract; Sec. VI] The abstract states: 'The model also accurately extrapolates coefficients in integrable regimes, where no universal asymptotic fit exists.' The full text contains no integrable-regime experiment; Sec. VI lists testing the model on sub-linear asymptotics as a future direction ('Another interesting direction would be to test the model's ability to forecast Lanczos sequences that are not asymptotically linear'). This is an unsupported claim in the abstract and should be removed or backed by the missing experiments.
- [Sec. III, Eq. (4); Sec. IV] All ground-truth labels are generated by finite-precision Lanczos with full reorthogonalization, yet the introduction motivates the work by the numerical instability of the same procedure. No independent validation of the long-tail labels is reported (e.g., high-precision arithmetic, orthogonality residuals, or comparison with direct moment computations). If loss of orthogonality or finite-size saturation silently corrupts the labels for n near T (T=100 for XYZ, T=30 for TFIM), the RMSE improvements are measured against corrupted labels and the 'order-of-magnitude' claim would not reflect true coefficient extrapolation. Please add a validation of at least a subset of long sequences, reporting, for example, the norm of the reorthogonalization residual as a function of n.
- [Sec. IV, Eq. (27); Figs. 4-6] The RMSE curves are point estimates averaged over N_T=100 test sequences without error bars, bootstrap confidence intervals, or per-sequence quantiles. The central quantitative claims ('order-of-magnitude reduction,' 'orders of magnitude smaller errors') are therefore not accompanied by any uncertainty estimate. Given the large reported separations this is not necessarily fatal, but for a quantitative ML benchmark in a physics journal the authors should report at least bootstrap CIs or the spread of per-sequence errors.
minor comments (4)
- [Sec. III, Eq. (13)] The text says 'Under UOGH, the raw coefficients grow without bound, whereas Δb_n approaches an n-independent constant in the asymptotic limit.' This is true for the d != 1 linear asymptote, but for the d=1 case in Eq. (9), b_n ~ α n / log n, so Δb_n ~ α / log n, which tends to 0 rather than a constant. Since the TFIM benchmark uses d=1, please qualify this statement.
- [Sec. V, Eq. (20), Fig. 7] The triangular support in the attention maps is a direct consequence of the hard causal mask M=-∞ in Eq. (20), not a learned property of the trained model. The meaningful learned patterns are the diagonal concentration, the checkerboard pattern, and the early-token attention. Please rephrase the sentence 'The observed triangular structure... provides a direct check that the causal constraint is indeed being enforced in the trained model' to avoid implying this is learned.
- [Abstract] The abstract says the method achieves 'an order-of-magnitude reduction in error' for 'both classical and quantum chaotic systems.' In the classical XYZ case (Fig. 3(b)) the reported reduction is 'a few times,' not an order of magnitude. Please make the abstract's quantitative claim match the figures.
- [General] No code or data availability statement is included. For reproducibility of a numerical ML benchmark, please indicate whether training/test data and code will be released.
Circularity Check
No significant circularity: supervised benchmark uses held-out test sequences and no central claim reduces to its inputs.
full rationale
The derivation chain is: generate Lanczos sequences by the recursion (Eq. 4) with full reorthogonalization; train a transformer with the next-step loss (Eq. 24) on 10,000 training sequences; evaluate RMSE (Eq. 27) on 100 unseen sequences; reconstruct observables via Eq. (30). Test sequences are not used in training, so the reported coefficient and observable errors are genuine out-of-sample measurements, not fitted inputs renamed as predictions. The UOGH asymptotic-fit baseline is fit per test sequence, but it is the comparator, not the paper's central claim. The attention ablation fixes the trained weights and alters inference-time masks; this is interpretability, not circular derivation. Self-citations ([25], [54]) appear only in motivating/outlook contexts and are not load-bearing for the transformer extrapolation result. No uniqueness theorem is imported from the authors' prior work. Two non-circular caveats should be weighed: (i) the test horizon is never longer than the training sequence length (T=100 for XYZ, T=30 for TFIM) — Eq. (24) supervises positions n_in..T-1, so 'deep Krylov space' beyond T is not actually tested; (ii) the abstract's claim of accurate extrapolation 'in integrable regimes' has no corresponding experiment in the body. These are support/scope gaps, not circular steps.
Assumptions & free parameters
free parameters (3)
- Transformer parameters Θ =
trained on 10,000 sequences; weights not released
- Model hyperparameters (L=3, d_model=64, H=4, n_in=10, dropout=0.1, lr=1e-3, epochs=300) =
listed in Table I; d_model=64 chosen empirically via Fig. 9
- Asymptotic fit parameters α, γ, γ* (baseline) =
least-squares fit per test sequence to first n_in=10 coefficients
assumptions (4)
- domain assumption The Lanczos algorithm with full reorthogonalization produces numerically exact coefficients up to T=30/100
- domain assumption The UOGH asymptotic forms (Eqs. 9, 25-26) are the correct standard baseline for these systems
- domain assumption Sampled TFIM parameters (g/J ∈ [1,2], h/J ∈ [0.1,1]) are all in the chaotic regime with asymptotically linear growth
- domain assumption Lanczos sequences across system sizes share learnable statistical structure sufficient for zero-shot transfer
Cite this review
Pith. "Pith review of Attention in Krylov Space: Transformer-Based Extrapolation of Lanczos Coefficients." pith.science (2026). https://pith.science/paper/NLXOMAAM
@misc{pith2026260107937,
author = {Pith},
title = {Pith review of: Attention in Krylov Space: Transformer-Based Extrapolation of Lanczos Coefficients},
year = {2026},
howpublished = {\url{https://pith.science/paper/NLXOMAAM}},
note = {Machine review of arXiv:2601.07937}
}
read the original abstract
The Universal Operator Growth Hypothesis formulates time evolution of operators through Lanczos coefficients. In practice, however, numerical instability and memory cost limit the number of coefficients that can be exactly computed. In response to these challenges, the standard approach relies on fitting early coefficients to asymptotic forms, but such procedures can miss subleading, history-dependent structures in the coefficients that subsequently affect reconstructed observables. In this work, we treat the Lanczos coefficients as a causal time sequence and introduce a transformer-based model to autoregressively predict future Lanczos coefficients from short prefixes. For classical and quantum chaotic systems, our model outperforms asymptotic fits in both coefficient extrapolation and physical observable reconstruction, and achieves an order-of-magnitude reduction in error. The model also accurately extrapolates coefficients in integrable regimes, where no universal asymptotic fit exists. Remarkably, our model transfers across system sizes: it can be trained on smaller systems and then be used to extrapolate coefficients on a larger system \emph{without retraining}. By probing the learned attention patterns and performing targeted attention ablations, we identify portions of the coefficient history that are most influential for accurate forecasts. Our results demonstrate that modern sequence models can serve as practical surrogates for probing operator dynamics deep in Krylov space, where brute-force Lanczos iteration can be computationally prohibitive.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
zero-shot transferability
We use open boundary conditions and set J= 1. We sample the dimensionless parametersg/Jand h/Juniformly from [1.0,2.0] and [0.1,1.0], respectively. Becauseh̸= 0 breaks integrability, all sampled Hamil- tonians are in the chaotic regime. Throughout training, the system size is fixed to beL= 8 and the sequence length to beT= 30. We first evaluate performanc...
-
[2]
Srednicki, Chaos and quantum thermalization, Phys
M. Srednicki, Chaos and quantum thermalization, Phys. Rev. E50, 888 (1994)
1994
-
[3]
J. M. Deutsch, Quantum statistical mechanics in a closed system, Phys. Rev. A43, 2046 (1991)
-
[4]
Rigol, V
M. Rigol, V. Dunjko, and M. Olshanii, Thermalization and its mechanism for generic isolated quantum systems, Nature452, 854 (2008)
2008
-
[5]
L. D’Alessio, Y. Kafri, A. Polkovnikov, and M. Rigol, From quantum chaos and eigenstate thermalization to statistical mechanics and thermodynamics, Advances in Physics65, 239 (2016), arXiv:1509.06411 [cond-mat.stat- mech]. 10
arXiv 2016
-
[6]
C. Gogolin and J. Eisert, Equilibration, thermalisation, and the emergence of statistical mechanics in closed quantum systems, Reports on Progress in Physics79, 056001 (2016), arXiv:1503.07538 [quant-ph]
arXiv 2016
- [7]
-
[8]
A. Polkovnikov, K. Sengupta, A. Silva, and M. Vengalat- tore, Colloquium: Nonequilibrium dynamics of closed in- teracting quantum systems, Reviews of Modern Physics 83, 863 (2011), arXiv:1007.5331 [cond-mat.stat-mech]
arXiv 2011
Show all 79 references
-
[9]
J. M. Deutsch, Eigenstate thermalization hypothesis, Re- ports on Progress in Physics81, 082001 (2018)
2018
-
[10]
Maldacena, S
J. Maldacena, S. H. Shenker, and D. Stanford, A bound on chaos, Journal of High Energy Physics2016, 1 (2016)
2016
-
[11]
A. I. Larkin and Y. N. Ovchinnikov, Quasiclassical method in the theory of superconductivity, Soviet Physics JETP28, 1200 (1969)
1969
-
[12]
Hayden and J
P. Hayden and J. Preskill, Black holes as mirrors: quan- tum information in random subsystems, Journal of High Energy Physics2007, 120 (2007), arXiv:0708.4025 [hep- th]
2007 arXiv
-
[13]
Sekino and L
Y. Sekino and L. Susskind, Fast scramblers, Journal of High Energy Physics2008, 065 (2008), arXiv:0808.2096 [hep-th]
2008 arXiv
-
[14]
Hosur, X.-L
P. Hosur, X.-L. Qi, D. A. Roberts, and B. Yoshida, Chaos in quantum channels, Journal of High Energy Physics 2016, 004 (2016), arXiv:1511.04021 [hep-th]
2016 arXiv
-
[15]
Swingle, Unscrambling the physics of out-of-time- order correlators, Nature Physics14, 988 (2018)
B. Swingle, Unscrambling the physics of out-of-time- order correlators, Nature Physics14, 988 (2018)
2018
-
[16]
D. E. Parker, X. Cao, A. Avdoshkin, T. Scaffidi, and E. Altman, A universal operator growth hypothesis, Phys. Rev. X9, 041017 (2019)
2019
-
[17]
Caputa, J
P. Caputa, J. M. Magan, and D. Patramanis, Geome- try of Krylov complexity, Physical Review Research4, 013041 (2022), arXiv:2109.03824 [hep-th]
2022 arXiv
-
[18]
Nandy, A
P. Nandy, A. S. Matsoukas-Roubeas, P. Mart ´ ınez- Azcona, A. Dymarsky, and A. del Campo, Quan- tum dynamics in Krylov space: Methods and applica- tions, Physics Reports1125, 1 (2025), arXiv:2405.09628 [quant-ph]
2025 arXiv
-
[19]
J. L. F. Barb´ on, E. Rabinovici, R. Shir, and S. S. K. Ali, On the evolution of operator complexity in quantum field theory, Journal of High Energy Physics2019, 1 (2019)
2019
-
[20]
Bhattacharya, A
A. Bhattacharya, A. Bhattacharyya, P. Hunt, and N. Kundu, Krylov complexity in open quantum systems, Journal of High Energy Physics2022, 1 (2022)
2022
-
[21]
Bhattacharjee, X
B. Bhattacharjee, X. Cao, P. Nandy, and T. Pathak, Op- erator growth in open quantum systems: lessons from the dissipative syk, Journal of High Energy Physics2023, 10.1007/jhep03(2023)054 (2023)
2023 doi
-
[22]
Bhattacharya, P
A. Bhattacharya, P. Nandy, P. P. Nath, and H. Sahu, On krylov complexity in open systems: an approach via bi- lanczos algorithm, Journal of High Energy Physics2023, 10.1007/jhep12(2023)066 (2023)
2023 doi
-
[23]
Bhattacharya, P
A. Bhattacharya, P. P. Nath, and H. Sahu, Speed lim- its to the growth of krylov complexity in open quantum systems, Phys. Rev. D109, L121902 (2024)
2024
-
[24]
N. S. Srivatsa and C. von Keyserlingk, Operator growth hypothesis in open quantum systems, Phys. Rev. B109, 125149 (2024)
2024
-
[25]
D. J. Yates, A. G. Abanov, and A. Mitra, Lifetime of al- most strong edge-mode operators in one-dimensional, in- teracting, symmetry protected topological phases, Phys. Rev. Lett.124, 206803 (2020)
2020
-
[26]
Z. Qi, T. Scaffidi, and X. Cao, Surprises in the deep hilbert space of all-to-all systems: From superexponen- tial scrambling to slow entanglement growth, Phys. Rev. B108, 054301 (2023)
2023
-
[27]
Cao, A statistical mechanism for operator growth, Journal of Physics A: Mathematical and Theoretical54, 144001 (2021)
X. Cao, A statistical mechanism for operator growth, Journal of Physics A: Mathematical and Theoretical54, 144001 (2021)
2021
-
[28]
Yi-Thomas, B
S. Yi-Thomas, B. Ware, J. D. Sau, and C. D. White, Comparing numerical methods for hydrodynamics in a one-dimensional lattice spin model, Phys. Rev. B110, 134308 (2024)
2024
-
[29]
Y. Zhou, W. Xia, L. Li, and W. Li, Diagnosing quantum many-body chaos in non-hermitian quantum spin chain via krylov complexity, Phys. Rev. Res.7, 033281 (2025)
2025
-
[30]
Bartsch, A
C. Bartsch, A. Dymarsky, M. H. Lamann, J. Wang, R. Steinigeweg, and J. Gemmer, Estimation of equili- bration time scales from nested fraction approximations, Phys. Rev. E110, 024126 (2024)
2024
-
[31]
Teretenkov, F
A. Teretenkov, F. Uskov, and O. Lychkovskiy, Pseu- domode expansion of many-body correlation functions, Phys. Rev. B111, 174308 (2025)
2025
-
[32]
Pinna, O
G. Pinna, O. Lunt, and C. von Keyserlingk, Approxi- mation theory for green’s functions via the lanczos algo- rithm, Phys. Rev. B112, 054435 (2025)
2025
-
[33]
Ballar Trigueros and C.-J
F. Ballar Trigueros and C.-J. Lin, Krylov complexity of many-body localization: Operator localization in krylov basis, SciPost Physics13, 10.21468/scipostphys.13.2.037 (2022)
2022 doi
-
[34]
D. J. Yates, F. H. L. Essler, and A. Mitra, Almost strong (0, π) edge modes in clean interacting one-dimensional floquet systems, Phys. Rev. B99, 205419 (2019)
2019
-
[35]
Tausendpfund, A
N. Tausendpfund, A. Mitra, and M. Rizzi, Almost strong zero modes at finite temperature, Phys. Rev. Res.7, 023245 (2025)
2025
-
[36]
D. J. Yates and A. Mitra, Strong and almost strong modes of floquet spin chains in krylov subspaces, Phys. Rev. B104, 195121 (2021)
2021
-
[37]
G. F. Scialchi, A. J. Roncaglia, and D. A. Wisniacki, Integrability-to-chaos transition through the krylov ap- proach for state evolution, Phys. Rev. E109, 054209 (2024)
2024
-
[38]
Heveling, J
R. Heveling, J. Wang, and J. Gemmer, Numerically prob- ing the universal operator growth hypothesis, Physical Review E106, 014152 (2022), arXiv:2203.00533 [cond- mat.stat-mech]
2022 arXiv
-
[39]
Eckseler, M
J. Eckseler, M. Pieper, and J. Schnack, Escaping the krylov space during the finite-precision lanczos algo- rithm, Physical Review E112, 10.1103/k94p-vls8 (2025)
2025 doi
-
[40]
Rabinovici, A
E. Rabinovici, A. S´ anchez-Garrido, R. Shir, and J. Son- ner, Operator complexity: a journey to the edge of krylov space, Journal of High Energy Physics2021, 10.1007/jhep06(2021)062 (2021)
2021 doi
-
[41]
H. D. Simon, The lanczos algorithm with partial re- orthogonalization, Mathematics of computation42, 115 (1984)
1984
-
[42]
Van der Veen and K
H. Van der Veen and K. Vuik, Bi-lanczos with partial or- thogonalization, Computers & structures56, 605 (1995)
1995
-
[43]
B. N. Parlett and D. S. Scott, The Lanczos algorithm with selective orthogonalization, Mathematics of Com- putation33, 217 (1979)
1979
-
[44]
Capizzi, L
L. Capizzi, L. Mazza, and S. Murciano, Universal prop- erties of the many-body lanczos algorithm at finite size (2025), arXiv:2507.17424 [quant-ph]. 11
2025
-
[45]
O. Lunt, T. Kriecherbauer, K. T.-R. McLaughlin, and C. von Keyserlingk, Emergent random matrix universality in quantum operator dynamics (2025), arXiv:2504.18311 [quant-ph]
2025
-
[46]
Uskov and O
F. Uskov and O. Lychkovskiy, Quantum dynamics in one and two dimensions via the recursion method, Phys. Rev. B109, L140301 (2024)
2024
-
[47]
Dymarsky and M
A. Dymarsky and M. Smolkin, Krylov complexity in con- formal field theory, Phys. Rev. D104, L081702 (2021)
2021
-
[48]
J. Wang, M. H. Lamann, R. Steinigeweg, and J. Gemmer, Diffusion constants from the recursion method, Physical Review B110, 10.1103/physrevb.110.104413 (2024)
2024 doi
-
[49]
L. E. Herrera Rodr ´ ıguez and A. A. Kananenka, Convolu- tional neural networks for long time dissipative quantum dynamics, The Journal of Physical Chemistry Letters12, 2476 (2021)
2021
-
[50]
L. E. Herrera Rodr ´ ıguez and A. A. Kananenka, A short trajectory is all you need: A transformer-based model for long-time dissipative quantum dynamics, The Journal of Chemical Physics161, 171101 (2024)
2024
-
[51]
Kaneko, M
R. Kaneko, M. Imada, Y. Kabashima, and T. Ohtsuki, Forecasting long-time dynamics in quantum many-body systems by dynamic mode decomposition, Physical Re- view Research7, 013085 (2025)
2025
-
[52]
Pescia, J
G. Pescia, J. Nys, J. Kim, A. Lovato, and G. Carleo, Message-passing neural quantum states for the homoge- neous electron gas, Phys. Rev. B110, 035108 (2024)
2024
-
[53]
Denis and G
Z. Denis and G. Carleo, Accurate neural quantum states for interacting lattice bosons, Quantum9, 1772 (2025)
2025
-
[54]
Rende, L
R. Rende, L. L. Viteritti, F. Becca, A. Scardicchio, A. Laio, and G. Carleo, Foundation neural-networks quantum states as a unified ansatz for multiple hamil- tonians, Nature Communications16, 7213 (2025)
2025
-
[55]
Z. Qi, Y. Peng, and C. Earls, Fourier neural operators for time-periodic quantum systems: Learning floquet hamiltonians, observable dynamics, and operator growth (2025), arXiv:2509.07084 [quant-ph]
2025 arXiv
-
[56]
Carleo and M
G. Carleo and M. Troyer, Solving the quantum many- body problem with artificial neural networks, Science 355, 602 (2017)
2017
-
[57]
Carrasquilla and R
J. Carrasquilla and R. G. Melko, Machine learning phases of matter, Nature Physics13, 431 (2017)
2017
-
[58]
D. Luo, Z. Chen, K. Hu, Z. Zhao, V. M. Hur, and B. K. Clark, Gauge-invariant and anyonic-symmetric au- toregressive neural network for quantum lattice models, Phys. Rev. Res.5, 013216 (2023)
2023
-
[59]
Viteritti, R
L. Viteritti, R. Rende, and S. Goldt, Transformer quan- tum states: A new class of expressive ansatzes for quan- tum many-body systems, Physical Review Letters130, 096401 (2023)
2023
-
[60]
Zhang and M
Y.-H. Zhang and M. Di Ventra, Transformer quantum state: A multipurpose model for quantum many-body problems, Phys. Rev. B107, 075147 (2023)
2023
-
[61]
P. Cha, P. Ginsparg, F. Wu, J. Carrasquilla, P. L. McMa- hon, and E.-A. Kim, Attention-based quantum tomog- raphy, Machine Learning: Science and Technology3, 01LT01 (2021)
2021
-
[62]
Carleo, I
G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborov´ a, Machine learning and the physical sciences, Reviews of Modern Physics91, 045002 (2019)
2019
-
[63]
Carrasquilla, Machine learning for quantum matter, Advances in Physics: X5, 1797528 (2020)
J. Carrasquilla, Machine learning for quantum matter, Advances in Physics: X5, 1797528 (2020)
2020
-
[64]
H. Kim, A. Kumar, Y. Zhou, Y. Xu, R. Vasseur, and E.-A. Kim, Learning measurement-induced phase tran- sitions using attention (2025), arXiv:2508.15895 [quant- ph]
2025 arXiv
-
[65]
Y. Zhou, C. Wan, Y. Xu, J. P. Zhou, K. Q. Weinberger, and E.-A. Kim, Learning to decode logical circuits, Na- ture Computational Science5, 1158 (2025)
2025
-
[66]
H. Kim, Y. Zhou, Y. Xu, K. Varma, A. H. Karamlou, I. T. Rosen, J. C. Hoke, C. Wan, J. P. Zhou, W. D. Oliver, Y. D. Lensky, K. Q. Weinberger, and E.-A. Kim, Attention to quantum complexity, Science Advances11, eadu0059 (2025)
2025
-
[67]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, At- tention is all you need, Advances in Neural Information Processing Systems30(2017)
2017
-
[68]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, BERT: Pre-training of deep bidirectional trans- formers for language understanding, arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[69]
H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and W. Zhang, Informer: Beyond efficient transformer for long sequence time-series forecasting, inProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35 (2021) pp. 11106–11115
2021
-
[70]
Liu and G
J.-M. Liu and G. M¨ uller, Infinite-temperature dynamics of the equivalent-neighbor xyz model, Phys. Rev. A42, 5854 (1990)
1990
-
[71]
Pfeuty, The transverse-field ising model, Annals of Physics57, 79 (1970)
P. Pfeuty, The transverse-field ising model, Annals of Physics57, 79 (1970)
1970
-
[72]
Roca-Jerat, M
S. Roca-Jerat, M. Gallego, F. Luis, J. Carrete, and D. Zueco, Transformer wave function for quantum long- range models, Physical Review B110, 10.1103/phys- revb.110.205147 (2024)
2024 doi
-
[73]
Hashimoto, K
K. Hashimoto, K. Murata, N. Tanahashi, and R. Watan- abe, Krylov complexity and chaos in quantum me- chanics, Journal of High Energy Physics2023, 10.1007/jhep11(2023)040 (2023)
2023 doi
-
[74]
B. L. Espa˜ nol and D. A. Wisniacki, Assessing the satu- ration of krylov complexity as a measure of chaos, Phys. Rev. E107, 024217 (2023)
2023
-
[75]
J. D. Noh, Operator growth in the transverse-field ising spin chain with integrability-breaking longitudinal field, Phys. Rev. E104, 034112 (2021)
2021
-
[76]
Y. Xian, B. Schiele, and Z. Akata, Zero-shot learning - the good, the bad and the ugly, inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)(2017)
2017
-
[77]
D. P. Jatkar, S. Mahato, S. Mondkar, and P. Thalore, Violation of universal operator growth hypothesis inw 3 conformal field theories (2025), arXiv:2506.01957 [hep- th]
2025 arXiv
-
[78]
Sprague and S
K. Sprague and S. Czischek, Variational monte carlo with large patched transformers, Communications Physics7, 10.1038/s42005-024-01584-y (2024)
2024 doi
-
[79]
Loshchilov and F
I. Loshchilov and F. Hutter, Decoupled weight decay reg- ularization (2019), arXiv:1711.05101 [cs.LG]. 12 Appendix A: Lanczos Algorithm for Classical Chaotic Systems In Sec. II, we review the Lanczos algorithm for quantum operators acting on finite-dimensional Hilbert spaces. ...
2019 arXiv
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.