REVIEW 4 major objections 4 minor 47 references
Padé rational fits to qubit trajectory data learn effective Volterra memory kernels for non-Markovian dynamics, and the regularized learning problem is provably well-posed.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 10:44 UTC pith:JPN64FGK
load-bearing objection Trajectory prediction works, but the paper's central claim of identifying memory kernels is not supported: the model assumes a convolution form that the exact kernels of two test problems violate, and the abstract promises more than the body delivers. the 4 major comments →
Learning Volterra Memory Kernels for Non-Markovian Qubit Dynamics
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that a low-parameter rational representation—fixed-order Padé approximants for each entry of the matrix-valued memory kernel—suffices to capture nontrivial non-Markovian features such as oscillatory memory, algebraic tails, and phase-sensitive coherence transfer. Learning is posed as a constrained optimization over an admissible operator space, regularized by an H^1 penalty, and the Volterra state equation is solved by a nonlocal Crank-Nicolson scheme. The paper proves existence of minimizers and sequential continuity of the loss along bounded operator sequences. Across three increasingly complex synthetic testbeds, the learned models accurately reproduce and generalize
What carries the argument
The central object is the vectorized Volterra equation dx/dt = A x(t) + ∫_0^t B(t−τ)x(τ)dτ, where x ∈ C^4 is the vectorized qubit density matrix, A is the instantaneous generator, and B is the operator-valued memory kernel. Each scalar entry B_ij is modeled as a [q/r] Padé approximant in the lag variable—a fixed-order rational function with coefficients ξ. A nonlocal Crank-Nicolson update solves the forward dynamics, and the learning objective is a Tikhonov-regularized L^2 trajectory misfit. The paper proves existence of minimizers via coercivity of the H^1 penalty, Sobolev compactness, and weak lower semicontinuity.
Load-bearing premise
The load-bearing premise is that the true reduced dynamics are exactly representable as a time-translation-invariant Volterra convolution, dx/dt = A x(t) + ∫_0^t B(t−τ)x(τ)dτ with B depending only on the lag; but the exact kernels of the paper's own Test Problems 2 and 3 contain explicit absolute-time factors (e.g., e^{-2iε0t} in Eq. (15b) and the t-dependent matrices B_αβ(t) in Eq. (20)), so the convolution ansatz is already violated in the testbeds.
What would settle it
A decisive test: use the trained kernel from Test Problem 2 to predict the state at a time well beyond the training horizon (say t=10) and compare to exact integration of Eqs. (15); a divergence would confirm that the learned B is an effective finite-window kernel rather than the true memory kernel. A more local falsifier is to compare the learned kernel's prediction for d²ρ/dt² at t=0 (which is governed by B(0)) against the exact short-time expansion.
If this is right
- For pure dephasing, the Padé model reproduces the correlation function accurately in sub-Ohmic, Ohmic, and super-Ohmic regimes, so the method can be used as a fast emulator (about 0.01 s per evaluation) for spectral-density sweeps.
- Learned Padé-kernel models can predict non-Markovian qubit trajectories on moderate time scales from a modest number of training trajectories (30 for the multi-channel problems).
- Because state trajectories are insensitive to the unrecoverable parts of the kernel, trajectory-level prediction is robust to the severe ill-conditioning of kernel recovery.
- The approach avoids symbolic libraries of special functions and deep-network architectures, needing only fixed-order rational functions plus a standard quasi-Newton optimizer.
- The well-posedness theorem guarantees that the regularized learning objective always has a minimizer, so numerical fitting is on solid ground even when the kernel is not unique.
Where Pith is reading between the lines
- A consequence the authors leave implicit: because the learned kernel is effective rather than pointwise-correct, any downstream quantity that depends on the kernel's values—multi-time correlations, fidelities under pulse sequences, or entanglement with the environment—must be validated separately; trajectory fidelity alone does not certify them.
- The observed non-identifiability suggests a general limit for Nakajima-Zwanzig kernel fitting from state trajectories alone: without additional structural assumptions (complete positivity, known commutator structure, longer observation windows), the true memory kernel cannot be disentangled from the effective one.
- A natural experimental extension would be to feed noisy process-tomography data from a real qubit into this pipeline; the noise-sensitivity analysis in Test Problem 1 indicates that a small H^1 regularization can suppress spurious kernel oscillations, which would be essential for experimental data.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a data-driven framework for extracting the linear operator A and the memory kernel B in the Volterra integro-differential equation dx/dt = Ax + ∫_0^t B(t−τ)x(τ)dτ from qubit trajectory data. Each component of B is parameterized as a Padé rational function; the parameters are optimized with a BFGS quasi-Newton method against a sum of trajectory misfits plus an H^1 regularization. The method is applied to three synthetic models: pure dephasing, an energy-exchange model, and a noncommuting σx/σz coupling model. The paper includes a well-posedness proof for the regularized optimization problem and a Crank-Nicolson discretization of the Volterra equation.
Significance. If the claims were fully supported, the paper would provide a simple, interpretable baseline for memory-kernel identification in open quantum systems. The pure-dephasing test (Sec. III) is a clean proof-of-principle: the convolution assumption is exact there, and the learned Padé kernel appears to reproduce both the trajectory and the correlation function. The existence theorem (App. B), while not novel, is a useful guarantee for the regularized finite-dimensional problem. The paper also deserves credit for explicitly acknowledging the ill-posedness of kernel recovery. However, the main claim of 'identifying A and B directly from data' is established only for the first test problem. For the other two problems the data-generating dynamics are not in the assumed convolution class, so the learned B is an effective object rather than the physical memory kernel. This confines the actual contribution to trajectory forecasting with effective convolution models.
major comments (4)
- [Sec. II Eq. (3); Sec. IV Eq. (15b); Sec. V Eq. (20)] The central hypothesis (3) requires B(t−τ) to depend only on the lag. The exact dynamics used for Test Problems 2 and 3 violate this. In Eq. (15b), the second term contains e^{−2iε0t} e^{i(ε0−ω)(t−s)}; for ε0=1 this equals e^{−i(ε0+ω)t} e^{−i(ε0−ω)s}, which depends on t and s separately, not on t−s alone. Thus no matrix kernel B(t−τ) can reproduce Eq. (15). The statement in Sec. IV that 'the dynamics can be written in the form of eq. (3)' is therefore incorrect. For Test Problem 3, Eq. (20) explicitly defines B(t,τ) with t-dependent matrices B_αβ(t). The optimization over O in Eq. (5) is searching the wrong class for these problems.
- [Sec. IV Fig. 3; Sec. V; Abstract; Sec. II] The paper's own results show that the learned B is not the physical kernel for Problems 2 and 3. The caption of Fig. 3 states 'the learned correlation functions are not close by any metric to the numerically evaluated correlation functions embedded in eqs (15).' Sec. V states 'Pointwise kernel recovery is, as in problem 2, significantly more challenging.' Since the true kernel is outside the model class, a good trajectory fit is not evidence of kernel identification; it is an effective-fit phenomenon. The abstract's 'identifying non-Markovian kernels' and Sec. II's 'central objective ... identify A and B directly from data' are therefore overstated and should be rephrased, or the experiments must be restricted to convolution-compatible models.
- [Appendix A, Eqs. (A1)–(A5)] The numerical scheme is derived for Eq. (A1) with B(t−τ), and the update (A4)–(A5) evaluates B at entries such as B(t_n−t_k). For the two-argument kernel B(t,τ) in Eq. (19), the scheme as written is not defined. The manuscript does not describe how the nonlocal Crank–Nicolson update is modified for Test Problems 2 and 3. This is a reproducibility gap in the numerical sections and also indicates that the implemented solver does not solve the stated state equation for the nonconvolution part.
- [Appendix B, Eq. (B1)] The existence and regularity theorem is proved only for the convolution state equation (B1). It does not cover the nonconvolution kernels B(t,τ) of Test Problem 3 or the second term of Eq. (15b) in Test Problem 2. Thus the claimed 'well-posedness of the learning problem' does not extend to two of the three numerical test problems. The theorem should either be extended or the theoretical claims qualified.
minor comments (4)
- [Throughout] Spelling is inconsistent: both 'Crank-Nicholson' and 'Crank-Nicolson' appear, and 'Padé' accents are inconsistent.
- [Sec. III, Fig. 1] The text says the sub-Ohmic, Ohmic, and super-Ohmic regimes are all studied, but the figure caption lists only p=1/2 and p=2. The Ohmic case is missing from the displayed results.
- [Eq. (4)] The [q/r] notation is nonstandard because the denominator degree appears to be q+r+1; the convention should be defined explicitly.
- [Sec. III] The sentence 'therefore, we do not investigate how our trained model generalizes given the uniqueness of integral curves from the dynamics' is unclear and should be rewritten.
Circularity Check
No circularity: the learning pipeline is an independent data fit with honest out-of-sample trajectory evaluation; kernel-recovery failures are admitted limitations, not circular steps.
full rationale
The paper does not reduce its output to its inputs by construction. Equation (3) is an explicit Volterra ansatz, and the Padé parameterization in Eq. (4) is introduced as a stated hypothesis class ('the use of a Pade approximant at this stage is a computational choice whose viability will be demonstrated throughout this paper'). The loss functional (6) and regularized objective (8) define a standard least-squares fit of (A,B) to trajectory data; there is no fitted parameter that is then renamed as a prediction of the same quantity. The out-of-training initial-state evaluations in Secs. IV and V (empirical risk in Eq. (17), held-out trajectories in Figs. 3 and 4) are genuine out-of-sample tests: the predicted trajectories are generated from initial states not used in training. The well-posedness theorem in Appendix B is a classical existence and sequential-continuity proof, explicitly assembled from standard ingredients (coercivity, Sobolev compactness, Grönwall, convex lower semicontinuity); it is not a uniqueness theorem and is not imported from the authors' prior work. The self-citations [31,32,41] concern pure-dephasing and entanglement background and are not load-bearing for the learning claim. The paper itself asserts the key limitation that 'the learned correlation functions are not close by any metric to the numerically evaluated correlation functions' (Sec. IV, Fig. 3 caption) and that 'accurate state fits do not guarantee pointwise kernel recovery' (Sec. VI). That admission undermines the abstract's 'identifying kernels' language, but it is a correctness/identifiability concern, not circularity. The separate observation that the exact kernels of Test Problems 2 and 3 contain explicit t-dependence (e.g., the e^{-2iε0 t} factor in Eq. (15b) and the t-dependent matrices B_αβ(t) in Eq. (20)) and therefore are not of the pure convolution form B(t−τ) is a model-misspecification issue; the learned B is then an effective kernel, not the true kernel, but the paper does not present that effective kernel as being identical to the input by construction. No circular step can be exhibited from the quoted equations.
Axiom & Free-Parameter Ledger
free parameters (6)
- α (regularization weight) =
0 (Problem 1), 1e-4 (Problem 2), 'extremely small' (Problem 3)
- β (H1 split weight) =
1 (noisy Problem 1), 0.95 (Problem 2)
- Padé orders [q/r] =
[4/4] (Problem 1), [3/3] (Problems 2, 3)
- Padé coefficient vector ξ =
10 real parameters for [4/4], 8 for [3/3] per scalar kernel
- Number of training trajectories =
1 (Problem 1), 30 (Problem 2), not stated (Problem 3)
- Time grid and domain =
T=3, M=32/64 grid points, t0=1e-6
axioms (6)
- domain assumption Product initial system-environment state and vanishing odd bath moments
- domain assumption Born (second-order) approximation for Test Problems 2 and 3
- domain assumption Time-translation-invariant convolution kernel B(t−τ)
- ad hoc to paper Padé rational functions can represent the relevant correlation functions on [0,T]
- ad hoc to paper Rational functions in the admissible set O have no poles on [0,T]
- ad hoc to paper Rank-1 factorization C_αβ = f_α f_β* enforces complete positivity
read the original abstract
We develop a data-driven framework for identifying non-Markovian equations of motion for open quantum systems, demonstrated here for qubit-environment dynamics. Starting from the Nakajima-Zwanzig formalism, we vectorize the reduced density matrix into a four-dimensional state vector and cast the dynamics as a Volterra integro-differential equation with an operator-valued memory kernel. The learning task is then formulated as a constrained optimization problem over the admissible operator space, where correlation functions are approximated by rational functions using Pade approximants. We establish well-posedness of the learning problem, ensuring existence of minimizers. To assess performance, we construct synthetic data sets from representative test problems of increasing complexity: (i) exactly solvable pure dephasing, with correlation functions expressed in terms of special functions, (ii) a damped Jaynes-Cummings model with an analytic coherence kernel, (iii) a transverse Born model with frequency-resolved bath integrals and population-coherence coupling, and (iv) a non-rotating-wave quantum Rabi model whose memory kernel has no closed form. Numerical experiments demonstrate that Pade captures nontrivial temporal structures such as oscillatory memory, algebraic tails, and phase-sensitive coherence transfer, and that the learned models generalize across ensembles of physically admissible initial states. We perform a parametrization-invariant sensitivity analysis and show that the trajectories are insensitive to the unrecoverable parts of the kernel, so the learned models stay predictive despite severe ill-conditioning in kernel recovery. These results together illustrate that data-driven rational approximation provides an effective route to identifying non-Markovian kernels of practical relevance in quantum technologies.
Figures
Reference graph
Works this paper leans on
-
[1]
Breuer and F
H.-P. Breuer and F. Petruccione,The Theory of Open Quantum Systems(Oxford University Press, 2002)
2002
-
[2]
de Vega and D
I. de Vega and D. Alonso, Dynamics of non-markovian open quantum systems, Rev. Mod. Phys.89, 015001 (2017)
2017
-
[3]
Rivas and S
A. Rivas and S. F. Huelga,Open Quantum Systems: An Introduction(Springer, 2012)
2012
-
[4]
Nakajima, On quantum theory of transport phenom- ena, Prog
S. Nakajima, On quantum theory of transport phenom- ena, Prog. Theor. Phys.20, 948 (1958)
1958
-
[5]
Zwanzig, Ensemble method in the theory of irre- versibility, J
R. Zwanzig, Ensemble method in the theory of irre- versibility, J. Chem. Phys.33, 1338 (1960)
1960
-
[6]
Mori, Transport, collective motion, and brownian mo- tion, Prog
H. Mori, Transport, collective motion, and brownian mo- tion, Prog. Theor. Phys.33, 423 (1965)
1965
-
[7]
Ciccotti, M
G. Ciccotti, M. H. Kalos, and J. C. Tully, Projection methods in molecular dynamics, J. Stat. Phys.118, 373 (2005)
2005
-
[8]
Zwanzig,Nonequilibrium Statistical Mechanics(Ox- ford University Press, 2001)
R. Zwanzig,Nonequilibrium Statistical Mechanics(Ox- ford University Press, 2001)
2001
-
[9]
A. J. Chorin and O. H. Hald,Stochastic Tools in Math- ematics and Science, 3rd ed. (Springer, 2013)
2013
-
[10]
Z. Li, X. Bian, T. A. Caswell, and G. E. Karniadakis, Data-driven parametrization of memory kernels in gen- eralized langevin dynamics, J. Chem. Phys.146, 014104 (2017)
2017
-
[11]
G. Jung, M. Hanke, and F. Schmid, Coarse-grained molecular dynamics with memory kernels, J. Chem. The- ory Comput.13, 2481 (2017)
2017
-
[12]
J. Wang, S. Olsson, C. Wehmeyer, and F. No´ e, Ma- chine learning of coarse-grained molecular dynamics force fields, J. Chem. Phys.152, 194106 (2020)
2020
-
[13]
No´ e and C
F. No´ e and C. Clementi, Collective variables for machine learning in molecular dynamics, Nature Reviews Physics 2, 42 (2020)
2020
-
[14]
F. A. Pollock, C. Rodr ´ ıguez-Rosario, T. Frauenheim, M. Paternostro, and K. Modi, Operational markov condi- tion for quantum processes, Phys. Rev. Lett.120, 040405 (2018)
2018
-
[15]
Cerrillo and J
J. Cerrillo and J. Cao, Non-markovian dynamical maps: Numerical processing of open quantum trajectories, Phys. Rev. Lett.112, 110401 (2014)
2014
-
[16]
Krastanov and L
S. Krastanov and L. Jiang, Stochastic modeling of quan- tum dynamics with process tensors, Phys. Rev. A103, 032605 (2021)
2021
-
[17]
P. Sentz, S. Nicholson, Y. Cho, S. Reddy, B. Keith, and S. G¨ unther, Learning thermodynamic master equations for open quantum systems, arXiv preprint arXiv:2506.01882 (2025), submitted; includes thermody- namically consistent machine learning of open quantum dynamics
Pith/arXiv arXiv 2025
-
[19]
Weiss,Quantum Dissipative Systems, 2nd ed
U. Weiss,Quantum Dissipative Systems, 2nd ed. (World Scientific, 1999)
1999
-
[20]
Krummheuer, V
B. Krummheuer, V. M. Axt, and T. Kuhn, Theory of pure dephasing and the resulting absorption line shape in semiconductor quantum dots, Phys. Rev. B65, 195313 (2002)
2002
-
[21]
A. J. Ramsay, A. V. Gopal, E. M. Gauger, A. Nazir, B. W. Lovett, A. M. Fox, and M. S. Skolnick, Phonon-induced rabi-frequency renormalization of opti- cally driven single ingaas/gaas quantum dots, Phys. Rev. Lett.104, 017402 (2010)
2010
-
[22]
Clos and H.-P
G. Clos and H.-P. Breuer, Quantification of memory ef- fects in the spin-boson model, Phys. Rev. A86, 012115 (2012)
2012
-
[23]
Flindt, T
C. Flindt, T. Novotn´ y, and A.-P. Jauho, Full counting statistics of nano-electromechanical systems, Europhys. Lett.69, 475 (2005)
2005
-
[24]
Gorini, A
V. Gorini, A. Kossakowski, and E. C. G. Sudarshan, Completely positive dynamical semigroups of n-level sys- tems, Journal of Mathematical Physics17, 821 (1976)
1976
-
[25]
Lindblad, On the generators of quantum dynamical semigroups, Communications in Mathematical Physics 48, 119 (1976)
G. Lindblad, On the generators of quantum dynamical semigroups, Communications in Mathematical Physics 48, 119 (1976)
1976
-
[26]
Shibata, Y
F. Shibata, Y. Takahashi, and N. Hashitsume, General- ized projection operator formalism for non-equilibrium statistical mechanics, Journal of Statistical Physics17, 171 (1977)
1977
-
[27]
Breuer, B
H.-P. Breuer, B. Kappler, and F. Petruccione, Stochastic wave-function method for non-markovian quantum mas- ter equations, Phys. Rev. A59, 1633 (1999)
1999
-
[28]
Breuer, B
H.-P. Breuer, B. Kappler, and F. Petruccione, The time-convolutionless projection operator technique in the quantum theory of dissipation and decoherence, Annals of Physics291, 36 (2001)
2001
-
[29]
Breuer, J
H.-P. Breuer, J. Gemmer, and M. Michel, Non-markovian quantum dynamics: Correlated projection superopera- tors and hilbert space averaging, Phys. Rev. E73, 016139 (2006)
2006
-
[30]
W. H. Zurek, Decoherence, einselection, and the quan- tum origins of the classical, Rev. Mod. Phys.75, 715 (2003)
2003
-
[31]
Roszak and L
K. Roszak and L. Cywi´ nski, Characterization and mea- surement of qubit-environment-entanglement generation during pure dephasing, Phys. Rev. A92, 032310 (2015). 12
2015
-
[32]
Roszak, Criteria for system-environment entangle- ment generation for systems of any size in pure-dephasing evolutions, Phys
K. Roszak, Criteria for system-environment entangle- ment generation for systems of any size in pure-dephasing evolutions, Phys. Rev. A98, 052344 (2018)
2018
-
[33]
A. J. Leggett, S. Chakravarty, A. T. Dorsey, M. P. A. Fisher, A. Garg, and W. Zwerger, Dynamics of the dissi- pative two-state system, Rev. Mod. Phys.59, 1 (1987)
1987
-
[34]
C. Guo, A. Weichselbaum, J. von Delft, and M. Vojta, Critical and strong-coupling phases in one- and two-bath spin-boson models, Phys. Rev. Lett.108, 160401 (2012)
2012
-
[35]
Z. Cai, U. Schollw¨ ock, and L. Pollet, Identifying a bath- induced bose liquid in interacting spin-boson models, Phys. Rev. Lett.113, 260403 (2014)
2014
-
[36]
Ferialdi, Exact non-markovian master equation for the spin-boson and jaynes-cummings models, Phys
L. Ferialdi, Exact non-markovian master equation for the spin-boson and jaynes-cummings models, Phys. Rev. A 95, 020101 (2017)
2017
-
[37]
Lampo, J
A. Lampo, J. Tuziemski, M. Lewenstein, and J. K. Kor- bicz, Objectivity in the non-markovian spin-boson model, Phys. Rev. A96, 012120 (2017)
2017
-
[38]
D. P. DiVincenzo and D. Loss, Rigorous born approxi- mation and beyond for the spin-boson model, Phys. Rev. B71, 035318 (2005)
2005
-
[39]
Wu and M
W. Wu and M. Liu, Effects of counter-rotating-wave terms on the non-markovianity in quantum open systems, Phys. Rev. A96, 032125 (2017)
2017
-
[40]
Gul´ acsi and G
B. Gul´ acsi and G. Burkard, Signatures of non- markovianity of a superconducting qubit, Phys. Rev. B 107, 174511 (2023)
2023
-
[41]
Strza lka, R
M. Strza lka, R. Filip, and K. Roszak, Qubit-environment entanglement in time-dependent pure dephasing, Phys. Rev. A109, 032412 (2024)
2024
-
[42]
Vacchini, A
B. Vacchini, A. Smirne, E.-M. Laine, J. Piilo, and H.-P. Breuer, Generalized master equations for non-markovian dynamics, New Journal of Physics13, 093004 (2011)
2011
-
[43]
Berrut and L
J.-P. Berrut and L. N. Trefethen, Barycentric lagrange interpolation, SIAM Rev.46, 501 (2004)
2004
-
[44]
Nakatsukasa, O
Y. Nakatsukasa, O. S` ete, and L. N. Trefethen, The aaa algorithm for rational approximation, SIAM J. Sci. Com- put.40, A1494 (2018)
2018
-
[45]
Lubich, Convolution quadrature and discretized oper- ational calculus
C. Lubich, Convolution quadrature and discretized oper- ational calculus. i, Numer. Math.52, 129 (1988)
1988
-
[46]
Brezis,Functional Analysis, Sobolev Spaces and Par- tial Differential Equations(Springer, 2010)
H. Brezis,Functional Analysis, Sobolev Spaces and Par- tial Differential Equations(Springer, 2010)
2010
-
[47]
Zeidler,Nonlinear Functional Analysis and its Ap- plications II: Variational Methods and Optimization (Springer, 1990)
E. Zeidler,Nonlinear Functional Analysis and its Ap- plications II: Variational Methods and Optimization (Springer, 1990)
1990
-
[48]
L. C. Evans,Partial Differential Equations, 2nd ed. (American Mathematical Society, 2010)
2010
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.