REVIEW 3 major objections 5 minor 72 references
Quantum Error Mitigation with Diffusion-Like Models
T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Neural networks trained on synthetic weak-measurement trajectories reconstruct pre-noise quantum states with fidelity above 0.99.
desk verdict Novel weak-measurement diffusion framing for error mitigation, but the headline fidelities may be inflated by channel sharing in the split and there are no baselines. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing construction is the monitoring quantum channel $M(\rho) = (1-\epsilon)\rho + \epsilon \Phi_A(\rho)$, where $\Phi_A$ is non-selective dephasing in the eigenbasis of a Pauli observable. Composing randomly sampled local copies of this channel along a trajectory produces the forward diffusion-like process, and uniform sampling over $x,y,z$ makes the ensemble-averaged single-qubit step the locally depolarizing channel $D_{1-2\epsilon/3}$. The learned object is the effective inverse map $M^{-1}_{\mathrm{eff}}$, trained with a geometric-mean-fidelity loss and a small L1 regularization term, with outputs projected onto the set of valid density matrices. Architectures are chosen by representation: LSTM for separable registers whose feature count scales linearly with qubit number, and Vision Transformer or enhanced U-Net for full density-matrix representations of entangled registers.
What would settle it
Train the same models on weak-measurement trajectories generated with one interaction strength and state ensemble, then evaluate them on states produced by a different noise model (such as amplitude damping or gate-dependent dephasing reaching the same purity) or on global states whose single-qubit marginals are identical but whose correlations differ; a significant drop in geometric-mean fidelity would show the learned inverse map is tied to the synthetic training channel.
Extended reading notes
Core claim
The central claim is that a neural network, trained purely on synthetic density matrices produced by a fixed sequential weak-measurement channel, can approximate the inverse of that channel and reconstruct pre-noise states with high fidelity. The forward process is built from random non-selective partial dephasing maps in the $x$, $y$, and $z$ Pauli bases, which on average act as a locally depolarizing channel, while individual trajectories retain basis-dependent structure. The trained models learn an effective data-driven inverse map rather than a physically implementable CPTP inverse channel, and the paper explicitly notes that the local-to-global reconstruction rule is distribution-dependent: one-qubit marginals do not determine a general global density matrix, so the learned map reflects the specific state-preparation and weak-measurement ensemble.
Load-bearing premise
The load-bearing premise is that the noise a real device produces is well approximated by the synthetic sequential weak-measurement channel with a fixed interaction strength, and that the states seen at test time come from the same ensemble used in training; the authors state that the local-to-global reconstruction rule is therefore not a universal map.
Editorial extensions
If this is right
- For separable registers, local noise channels can be calibrated with a model whose input features grow linearly with qubit count, since each qubit's density matrix is represented by four real parameters and the learned inverse map reaches $F_{\mathrm{GM}} > 0.99$ on a six-qubit example.
- For entangled registers, the full density-matrix representation is necessary; an LSTM fails with fidelity near 0.5 regardless of dataset size, while Vision Transformer and enhanced U-Net models learn the inverse map once enough training data are supplied.
- The local-to-global setting, which reconstructs the global state from single-qubit reduced density matrices, is experimentally cheaper than full tomography but is not a universal map from marginals to global states.
- Because the formal inverse of the depolarizing channel is not completely positive, the learned reconstruction is an effective data-driven inverse applied as classical post-processing rather than a physically implementable inverse channel.
- The authors view the sequential weak-measurement process as measurement-induced and contextual, making the learned inverse a step toward data-driven characterization of non-unitary dynamics in distributed quantum architectures.
Reading between the lines
- A natural next step outside the paper is to test whether the same training procedure transfers to hardware noise: if the real device's effective channel is not close to sequential Pauli-basis dephasing, the learned inverse map should degrade, and retraining on device-calibrated trajectories could restore performance.
- Because local-to-global reconstruction is distribution-dependent, the method is best understood as a prior-aware estimator rather than a tomographic protocol; it answers what global state is most plausible under the training ensemble, which could be useful for state-preparation verification where the ensemble is known.
- The ensemble-averaged equivalence to a locally depolarizing channel suggests that deliberately randomizing or twirling measurement bases on hardware could bring real noise closer to the training distribution and improve transfer without changing the architecture.
- The components, including the geometric-mean-fidelity loss, the density-matrix projection filter, and trajectory-based training with teacher forcing, are reusable as generic post-processing in other learning-based reconstruction pipelines beyond the weak-measurement diffusion picture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a diffusion-like quantum error mitigation framework in which the forward process is generated by sequential weak measurements in randomly selected Pauli bases, driving the state toward decoherence. Neural networks (LSTM, Vision Transformer, enhanced U-Net) are trained on exact synthetic density-matrix trajectories to learn an effective inverse map that reconstructs the pre-noise state. The method is benchmarked on single-qubit Bloch-vector states, separable multi-qubit registers, entangled registers, and a local-to-global reconstruction task where only reduced density matrices are inputs. The reported results include F_GM > 0.99 for a six-qubit separable register and a mean test fidelity of 0.9694 for a five-qubit local-to-global task.
Significance. If the experimental claims are robust, the work would provide a novel hybrid classical-quantum tool for approximating non-unitary dynamics and mitigating measurement-induced decoherence, with potential applications in distributed and noisy quantum systems. The theoretical framing connecting sequential weak measurements to diffusion processes is conceptually appealing and the paper includes a clear derivation of the averaged depolarizing behavior. The main value rests on the generalization claim that the trained networks learn an effective inverse map for the forward process. However, the current evidence for this generalization is weakened by the absence of channel-disjoint evaluation and the lack of any baseline comparisons, so the significance cannot be fully assessed from the manuscript as written.
major comments (3)
- [Sec. IV, Figs. 4 and 5] The data-splitting protocol is not described in a way that excludes a serious confound. Algorithm 1 generates one weak-measurement channel by evolving a single Haar-random state to the purity threshold, and then applies that same recorded channel to many additional random initial states. The paper states an 80/10/10 split over "base states" but never states that all examples generated from one recorded channel are kept in the same partition. If channels are shared between training and test, the model can memorize the inverse of each specific basis sequence (which is identifiable from the trajectory) and report artificially high test fidelity. The central claim of Eq. (10) that the model learns an effective inverse map requires evaluation on channels not seen during training. The authors should report the number of channels, the number of states per channel, and perform a channel-disjoint split, or justify why channel overlap does not inflate the reported mean test fidelity of 0.9694.
- [Sec. IV] No baseline comparisons are reported. The paper presents absolute fidelities such as F_GM > 0.99 for a six-qubit separable register, but without comparing to (i) the trivial no-mitigation baseline (fidelity of the final noisy state to the original state), (ii) a standard analytic inversion of the averaged local depolarizing channel, or (iii) established methods such as classical shadows, the reader cannot judge whether the proposed approach provides any practical advantage. For separable states under local depolarizing noise, a simple per-qubit inverse of the averaged channel may already achieve high fidelity, and the LSTM result would then be unsurprising. Adding such baselines is essential to support the claim that the AI models are useful for error mitigation.
- [Sec. IV] The paper acknowledges that the local-to-global reconstruction is distribution dependent and that the network learns a rule induced by the specific state-preparation and weak-measurement process. This is an honest and important qualification, but it also means the method is not a general error-mitigation technique; it is a tailored fit to a specific synthetic noise model. The practical relevance depends on whether the trained model can transfer to a real device whose noise is only approximately described by the sequential weak-measurement channel with fixed epsilon. The paper should either provide experiments with a mismatched training/deployment distribution (e.g., different epsilon or different state ensemble) or clearly articulate the conditions under which the method is expected to work. Without this, the claim of "quantum error mitigation" is overly broad.
minor comments (5)
- [Sec. IIA] The sentence "Equation 3 assumes this projective-dephasing channel" is confusing because Eq. (3) is the iterative map; the projective-dephasing channel is defined in Eq. (2). Please clarify the cross-reference.
- [Sec. IIIB] The line "P_stop = eta(Tr(rho^2) = eta)" is ambiguous; it should read "P_stop is the purity threshold such that the loop stops when Tr(rho^2) <= eta".
- [Sec. IV] The caption states "test set (n=18,000 states)" but the text says 1,500 test base states corresponding to 18,000 noisy-clean pairs. Please make the distinction between base states and trajectory pairs consistent.
- [Additional Information] The data availability statement says the data are provided within the article, but the actual datasets are not included; only descriptions and aggregated statistics are given. The code is not publicly available due to IP restrictions. For a methodological paper, a public release of code and synthetic-data generation scripts would greatly aid reproducibility; at minimum, the authors should provide the exact hyperparameters and data-generation seeds for all reported experiments.
- [Sec. IV] Several reported results (e.g., the six-qubit separable register, the entangled-register scaling curves) are given without error bars or confidence intervals. Only Fig. 5b reports a mean over 10 runs. Error bars or multiple-seed statistics should be provided for all headline fidelities.
Circularity Check
No significant circularity: the reported reconstruction is a supervised benchmark on synthetic trajectories generated from the paper's own stated weak-measurement model, and the only self-citation is non-load-bearing.
full rationale
The paper's central claim is that neural networks trained on weak-measurement trajectories learn an effective inverse map for the specific synthetic channel ensemble. This is a standard supervised-learning evaluation: the forward process in Eqs. (2) and (11) generates training and test pairs, the loss in Eq. (16) compares the model output to the independently known initial state, and the reported test fidelities are computed on held-out pairs. The target state is not a fitted parameter of the loss, and the test pairs are not identical to the training pairs, so the prediction does not reduce to its inputs by construction. The ideal reconstruction relation in Eq. (10), rho approx M_eff^{-1}(M(rho)), is a definition of the goal, not a proof that the learned map succeeds. The only self-citation is ref. [3] (Tamir and Cohen, with E. Cohen as a co-author), used for the standard von Neumann measurement interaction in Eq. (1); the actual monitoring channel in Eq. (2) is cited to the independent ref. [59], and the numerical results do not depend on ref. [3]. The paper explicitly acknowledges the distribution-dependent nature of the local-to-global task: 'for arbitrary quantum states, one-qubit marginals do not uniquely determine the global density matrix. The network therefore learns a reconstruction rule induced by the specific state-preparation and weak-measurement process, not a universal map.' This honest scoping further supports that the work is a self-contained simulation study rather than a circular derivation. The data-split ambiguity noted in the skeptic materials (whether channels are shared between train and test base states) is a potential evaluation-leakage concern about generalization, but it cannot be exhibited from the paper's equations as a reduction of the claimed prediction to its inputs, so it does not constitute circularity under the stated rules.
Assumptions & free parameters
free parameters (3)
- measurement strength epsilon =
0.1 to 0.15 (per experiment)
- purity threshold eta (P_stop) =
0.55 or 0.6
- loss weights lambda_F and lambda_1 =
10 and 1e-2
assumptions (3)
- domain assumption The monitoring channel M(ρ) = (1-ε)ρ + ε Φ_A(ρ), with non-selective dephasing Φ_A, describes the effect of a weak measurement with discarded outcomes (Eq. 2).
- domain assumption The training state distribution consists of Haar-random single-qubit states and, for multi-qubit registers, random product states followed by n-1 random CNOT gates.
- domain assumption The synthetic weak-measurement process is a faithful surrogate for the decoherence that error mitigation is intended to correct.
Cite this review
Pith. "Pith review of Quantum Error Mitigation with Diffusion-Like Models." pith.science (2026). https://pith.science/paper/5E55FQ3I
@misc{pith2026260805382,
author = {Pith},
title = {Pith review of: Quantum Error Mitigation with Diffusion-Like Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/5E55FQ3I}},
note = {Machine review of arXiv:2608.05382}
}
read the original abstract
Coupling between a quantum system and its environment causes decoherence by transferring information from the system to environmental degrees of freedom. When discretized in time, such interactions can be interpreted as sequences of weak measurements that provide an effective model of noisy quantum dynamics. Motivated by this picture, we propose an AI-assisted error-mitigation framework for quantum diffusion processes generated by sequential local weak measurements. The forward process progressively erases information from the input state through weak measurements performed in randomly selected Pauli bases, producing basis-dependent local dephasing and locally depolarizing dynamics on average. Machine-learning models are trained on exact synthetic density matrices to learn a channel- and distribution-specific denoising map and estimate the corresponding pre-noise state. We benchmark the approach on single-qubit states and separable and entangled multi-qubit registers. We also study distribution-dependent local-to-global reconstruction, in which local reduced density matrices are used to reconstruct the global state. This experimentally motivated setting relies on locally accessible information and is therefore compatible with noisy and distributed quantum systems. More broadly, the framework provides a hybrid classical-quantum approach for approximating non-unitary dynamics and mitigating coherence loss.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Single-qubit Bloch-vector dataset We begin with the simplest reconstruction task: the inverse evolution of a single qubit represented by its Bloch vector, ρ= 1 2 (I+⃗ r·⃗ σ),(18) where⃗ ri = (rx i , ry i , rz i )contains three real features per time step. This representation is useful for building intuition, since the noisy evolution can be visualized dir...
-
[2]
QERNEL” and the “Interdisciplinary Center for the Theory of Quantum Computing
Separable multi-qubit dataset We next consider the extension from a single qubit to a register ofnseparable, non-interacting qubits. Each local qubit state is represented by a2×2density matrix, ρ(t,k) i = α(k,t) 1,i β(k,t) 1,i +iβ (k,t) 2,i β(k,t) 1,i −iβ (k,t) 2,i δ(k,t) 1,i ! ,(24) 10 where Hermiticity implies that only four real parameters are required...
-
[3]
B. Coecke and A. Kissinger,Picturing Quantum Processes: A First Course in Quantum Theory and Diagrammatic Rea- soning. Cambridge: Cambridge University Press, 2017
work page 2017
-
[4]
Dynamics of quantum causal structures,
E. Castro-Ruiz, F. Giacomini, and Č. Brukner, “Dynamics of quantum causal structures,”Physical Review X, vol. 8, no. 1, p. 011047, 2018
work page 2018
-
[5]
Introduction to weak measurements and weak values,
B. Tamir and E. Cohen, “Introduction to weak measurements and weak values,”Quanta, vol. 2, no. 1, pp. 7–17, 2013
work page 2013
-
[6]
Decoherence, einselection, and the quantum origins of the classical,
W. H. Zurek, “Decoherence, einselection, and the quantum origins of the classical,”Reviews of modern physics, vol. 75, no. 3, p. 715, 2003
work page 2003
-
[7]
A. Paetznick, M. P. da Silva, C. Ryan-Anderson, J. M. Bello-Rivas, J. P. Campora III, A. Chernoguzov, J. M. Dreiling, C. Foltz, F. Frachon, J. P. Gaebler,et al., “Demonstration of logical qubits and repeated error correction with better-than- physical error rates,” 2024
work page 2024
-
[8]
Quantum error correction below the surface code threshold,
Google Quantum AI and Collaborators, “Quantum error correction below the surface code threshold,”Nature, vol. 638, pp. 920–926, 2025
2025
Show all 72 references
-
[9]
A fault-tolerant neutral-atom architecture for universal quantum computation,
D. Bluvstein, A. A. Geim, S. H. Li, S. J. Evered, J. P. Bonilla Ataides, G. Baranes, A. Gu, T. Manovitz, M. Xu, M. Kalinowski,et al., “A fault-tolerant neutral-atom architecture for universal quantum computation,”Nature, vol. 649, no. 8095, pp. 39–46, 2026
2026
-
[10]
Lattice surgery realized on two distance-three repetition codes with superconducting qubits,
I. Besedin, M. Kerschbaum, J. Knoll, I. Hesner, L. Bödeker, L. Colmenarez, L. Hofele, N. Lacroix, C. Hellings, F. Swiadek, et al., “Lattice surgery realized on two distance-three repetition codes with superconducting qubits,”Nature physics, vol. 22, no. 2, pp. 189–194, 2026
2026
-
[11]
Hardware-efficient quantum error correction via concatenated bosonic qubits,
H. Putterman, K. Noh, C. T. Hann, G. S. MacCabe, S. Aghaeimeibodi, R. N. Patel, M. Lee, W. M. Jones, H. Moradinejad, R. Rodriguez,et al., “Hardware-efficient quantum error correction via concatenated bosonic qubits,”Nature, vol. 638, no. 8052, pp. 927–934, 2025
2025
-
[12]
On the importance of error mitigation for quantum computation,
D. Aharonov, O. Alberton, I. Arad, Y. Atia, E. Bairey, Z. Brakerski, I. Cohen, O. Golan, I. Gurwich, O. Kenneth, E. Leviatan, N. H. Lindner, R. A. Melcer, A. Meyer, G. Schul, and M. Shutman, “On the importance of error mitigation for quantum computation,” 2025
2025
-
[13]
Noisy intermediate-scale quantum algorithms,
K. Bharti, A. Cervera-Lierta, T. H. Kyaw, T. Haug, S. Alperin-Lea, A. Anand, M. Degroote, H. Heimonen, J. S. Kottmann, T. Menke, W.-K. Mok, S. Sim, L.-C. Kwek, and A. Aspuru-Guzik, “Noisy intermediate-scale quantum algorithms,”Reviews of Modern Physics, vol. 94, p. 015004, 2022
2022
-
[14]
M. A. Nielsen and I. L. Chuang,Quantum computation and quantum information. Cambridge university press, 2010
2010
-
[15]
On the complementary quantum capacity of the depolarizing channel,
D. Leung and J. Watrous, “On the complementary quantum capacity of the depolarizing channel,”Quantum, vol. 1, p. 28, 2017
2017
-
[16]
The capacity of the quantum depolarizing channel,
C. King, “The capacity of the quantum depolarizing channel,”IEEE Transactions on Information Theory, vol. 49, no. 1, pp. 221–229, 2003
2003
-
[17]
Denoising diffusion probabilistic models,
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”Advances in neural information processing systems, vol. 33, pp. 6840–6851, 2020
2020
-
[18]
Improved denoising diffusion probabilistic models,
A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” inInternational Conference on Machine Learning, pp. 8162–8171, PMLR, 2021
2021
-
[19]
Deep unsupervised learning using nonequilibrium thermodynamics,
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” inInternational conference on machine learning, pp. 2256–2265, PMLR, 2015
2015
-
[20]
Shadow tomography of quantum states,
S. Aaronson, “Shadow tomography of quantum states,” inProceedings of the 50th annual ACM SIGACT symposium on theory of computing, pp. 325–338, 2018
2018
-
[21]
Predicting many properties of a quantum system from very few measurements,
H.-Y. Huang, R. Kueng, and J. Preskill, “Predicting many properties of a quantum system from very few measurements,” Nature Physics, vol. 16, no. 10, pp. 1050–1057, 2020
2020
-
[22]
Classical shadows with noise,
D. E. Koh and S. Grewal, “Classical shadows with noise,”Quantum, vol. 6, p. 776, 2022
2022
-
[23]
Quantum error mitigation,
Z. Cai, R. Babbush, S. C. Benjamin, S. Endo, W. J. Huggins, Y. Li, J. R. McClean, and T. E. O’Brien, “Quantum error mitigation,”Reviews of Modern Physics, vol. 95, no. 4, p. 045005, 2023
2023
-
[24]
Error mitigation extends the computational reach of a noisy quantum processor,
A. Kandala, K. Temme, A. D. Córcoles, A. Mezzacapo, J. M. Chow, and J. M. Gambetta, “Error mitigation extends the computational reach of a noisy quantum processor,”Nature, vol. 567, no. 7749, pp. 491–495, 2019
2019
-
[25]
Learning-based quantum error mitigation,
A. Strikis, D. Qin, Y. Chen, S. C. Benjamin, and Y. Li, “Learning-based quantum error mitigation,”PRX Quantum, vol. 2, no. 4, p. 040330, 2021
2021
-
[26]
Error mitigation with clifford quantum-circuit data,
P. Czarnik, A. Arrasmith, P. J. Coles, and L. Cincio, “Error mitigation with clifford quantum-circuit data,”Quantum, vol. 5, p. 592, 2021
2021
-
[27]
Differentiable quantum architecture search,
S.-X. Zhang, C.-Y. Hsieh, S. Zhang, and H. Yao, “Differentiable quantum architecture search,”Quantum Sci. Technol., vol. 7, no. 4, p. 045023, 2022. 14
2022
-
[28]
Stability of classical shadows under gate-dependent noise,
R. Brieger, M. Heinrich, I. Roth, and M. Kliesch, “Stability of classical shadows under gate-dependent noise,”Phys. Rev. Lett., vol. 134, p. 090801, Mar 2025
2025
-
[29]
Vaqem: A variational ap- proach to quantum error mitigation,
G. S. Ravi, K. N. Smith, P. Gokhale, A. Mari, N. Earnest, A. Javadi-Abhari, and F. T. Chong, “Vaqem: A variational ap- proach to quantum error mitigation,” in2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pp. 288–303, IEEE, 2022
2022
-
[30]
Efficient quantum tomography,
R. O’Donnell and J. Wright, “Efficient quantum tomography,” inProceedings of the forty-eighth annual ACM symposium on Theory of Computing, pp. 899–912, 2016
2016
-
[31]
Quantum resource theories,
E. Chitambar and G. Gour, “Quantum resource theories,”Rev. Mod. Phys., vol. 91, p. 025001, Apr 2019
2019
-
[32]
How the result of a measurement of a component of the spin of a spin-1/2 particle can turn out to be 100,
Y. Aharonov, D. Z. Albert, and L. Vaidman, “How the result of a measurement of a component of the spin of a spin-1/2 particle can turn out to be 100,”Physical review letters, vol. 60, no. 14, p. 1351, 1988
1988
-
[33]
Weak-value amplification of the nonlinear effect of a single photon,
M. Hallaji, A. Feizpour, G. Dmochowski, J. Sinclair, and A. M. Steinberg, “Weak-value amplification of the nonlinear effect of a single photon,”Nature Physics, vol. 13, no. 6, pp. 540–544, 2017
2017
-
[34]
Quantum cryptography with weak measurements,
J. E. Troupe and J. M. Farinholt, “Quantum cryptography with weak measurements,”arXiv preprint arXiv:1702.04836, 2017
2017 arXiv
-
[35]
Error mitigation for short-depth quantum circuits,
K. Temme, S. Bravyi, and J. M. Gambetta, “Error mitigation for short-depth quantum circuits,”Physical review letters, vol. 119, no. 18, p. 180509, 2017
2017
-
[36]
Quantum error mitigation with artificial neural network,
C. Kim, K. D. Park, and J.-K. Rhee, “Quantum error mitigation with artificial neural network,”IEEE Access, vol. 8, pp. 188853–188860, 2020
2020
-
[37]
Quantum readout error mitigation via deep learning,
J. Kim, B. Oh, Y. Chong, E. Hwang, and D. K. Park, “Quantum readout error mitigation via deep learning,”New Journal of Physics, vol. 24, no. 7, p. 073009, 2022
2022
-
[38]
Reservoir computing approach to quantum state measurement,
G. Angelatos, S. A. Khan, and H. E. Türeci, “Reservoir computing approach to quantum state measurement,”Physical Review X, vol. 11, no. 4, p. 041062, 2021
2021
-
[39]
The variational quantum eigensolver: a review of methods and best practices,
J. Tilly, H. Chen, S. Cao, D. Picozzi, K. Setia, Y. Li, E. Grant, L. Wossnig, I. Rungger, G. H. Booth,et al., “The variational quantum eigensolver: a review of methods and best practices,”Physics Reports, vol. 986, pp. 1–128, 2022
2022
-
[40]
A variational eigenvalue solver on a photonic quantum processor,
A. Peruzzo, J. McClean, P. Shadbolt, M.-H. Yung, X.-Q. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’brien, “A variational eigenvalue solver on a photonic quantum processor,”Nature communications, vol. 5, no. 1, p. 4213, 2014
2014
-
[41]
The theory of variational hybrid quantum-classical algo- rithms,
J. R. McClean, J. Romero, R. Babbush, and A. Aspuru-Guzik, “The theory of variational hybrid quantum-classical algo- rithms,”New Journal of Physics, vol. 18, no. 2, p. 023023, 2016
2016
-
[42]
Long short-term memory,
S. Hochreiter, “Long short-term memory,”Neural Computation MIT-Press, 1997
1997
-
[43]
Empirical evaluation of gated recurrent neural networks on sequence modeling,
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,”arXiv preprint arXiv:1412.3555, 2014
2014 arXiv
-
[44]
Sequence to sequence learning with neural networks,
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,”Advances in neural infor- mation processing systems, vol. 27, 2014
2014
-
[45]
Cvt: Introducing convolutions to vision transformers,
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang, “Cvt: Introducing convolutions to vision transformers,” inProceedings of the IEEE/CVF international conference on computer vision, pp. 22–31, 2021
2021
-
[46]
Training data-efficient image transformers & distillation through attention,
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” inInternational conference on machine learning, pp. 10347–10357, PMLR, 2021
2021
-
[47]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly,et al., “An image is worth 16x16 words: Transformers for image recognition at scale,”arXiv preprint arXiv:2010.11929, 2020
2010 arXiv
-
[48]
U-net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” inInter- national Conference on Medical image computing and computer-assisted intervention, pp. 234–241, Springer, 2015
2015
-
[49]
nnu-net: a self-configuring method for deep learning-based biomedical image segmentation,
F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, and K. H. Maier-Hein, “nnu-net: a self-configuring method for deep learning-based biomedical image segmentation,”Nature methods, vol. 18, no. 2, pp. 203–211, 2021
2021
-
[50]
Attention u-net: Learning where to look for the pancreas,
O. Oktay, J. Schlemper, L. L. Folgoc, M. Lee, M. Heinrich, K. Misawa, K. Mori, S. McDonagh, N. Y. Hammerla, B. Kainz, et al., “Attention u-net: Learning where to look for the pancreas,”arXiv preprint arXiv:1804.03999, 2018
2018 arXiv
-
[51]
Prescription for experimental determination of the dynamics of a quantum black box,
I. L. Chuang and M. A. Nielsen, “Prescription for experimental determination of the dynamics of a quantum black box,” Journal of Modern Optics, vol. 44, no. 11-12, pp. 2455–2467, 1997
1997
-
[52]
Complete characterization of a quantum process: the two-bit quantum gate,
J. Poyatos, J. I. Cirac, and P. Zoller, “Complete characterization of a quantum process: the two-bit quantum gate,”Physical Review Letters, vol. 78, no. 2, p. 390, 1997
1997
-
[53]
Deep learning on image denoising: An overview,
C. Tian, L. Fei, W. Zheng, Y. Xu, W. Zuo, and C.-W. Lin, “Deep learning on image denoising: An overview,”Neural Networks, vol. 131, pp. 251–275, 2020
2020
-
[54]
Image denoising review: From classical to state-of-the-art approaches,
B. Goyal, A. Dogra, S. Agrawal, B. S. Sohi, and A. Sharma, “Image denoising review: From classical to state-of-the-art approaches,”Information fusion, vol. 55, pp. 220–244, 2020
2020
-
[55]
Dynamic noise aware training for speech enhancement based on deep neural networks.,
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “Dynamic noise aware training for speech enhancement based on deep neural networks.,” inInterspeech, vol. 1, pp. 2670–2674, 2014
2014
-
[56]
Mofusion: A framework for denoising-diffusion-based motion synthesis,
R. Dabral, M. H. Mughal, V. Golyanik, and C. Theobalt, “Mofusion: A framework for denoising-diffusion-based motion synthesis,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9760–9770, 2023
2023
-
[57]
The quantum internet,
H. J. Kimble, “The quantum internet,”Nature, vol. 453, no. 7198, pp. 1023–1030, 2008
2008
-
[58]
Quantum repeaters: From quantum networks to the quantum internet,
K. Azuma, S. E. Economou, D. Elkouss, P. Hilaire, L. Jiang, H.-K. Lo, and I. Tzitrin, “Quantum repeaters: From quantum networks to the quantum internet,”Reviews of Modern Physics, vol. 95, no. 4, p. 045006, 2023
2023
-
[59]
A one-way quantum computer,
R. Raussendorf and H. J. Briegel, “A one-way quantum computer,”Physical review letters, vol. 86, no. 22, p. 5188, 2001
2001
-
[60]
Measurement-based quantum computation on cluster states,
R. Raussendorf, D. E. Browne, and H. J. Briegel, “Measurement-based quantum computation on cluster states,”Physical review A, vol. 68, no. 2, p. 022312, 2003. 15
2003
-
[61]
Information-reality complementarity: The role of measurements and quantum reference frames,
P. R. Dieguez and R. M. Angelo, “Information-reality complementarity: The role of measurements and quantum reference frames,”Phys. Rev. A, vol. 97, no. 2, p. 022107, 2018
2018
-
[62]
Efficient measurement error mitigation with subsystem-balanced pauli twirling,
X.-Y. Xu, C. Ding, and W.-S. Bao, “Efficient measurement error mitigation with subsystem-balanced pauli twirling,” 2025
2025
-
[63]
Quantum fidelity measures for mixed states,
Y.-C. Liang, Y.-H. Yeh, P. E. Mendonça, R. Y. Teh, M. D. Reid, and P. D. Drummond, “Quantum fidelity measures for mixed states,”Reports on Progress in Physics, vol. 82, no. 7, p. 076001, 2019
2019
-
[64]
Neural autoregressive distribution estimation,
B. Uria, M.-A. Côté, K. Gregor, I. Murray, and H. Larochelle, “Neural autoregressive distribution estimation,” 2016
2016
-
[65]
Generating sequences with recurrent neural networks,
A. Graves, “Generating sequences with recurrent neural networks,” 2014
2014
-
[66]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017
2017
-
[67]
Efficient method for computing the maximum-likelihood quantum state from measurements with additive gaussian noise,
J. A. Smolin, J. M. Gambetta, and G. Smith, “Efficient method for computing the maximum-likelihood quantum state from measurements with additive gaussian noise,”Phys. Rev. Lett., vol. 108, p. 070502, 2012
2012
-
[68]
Computing a nearest symmetric positive semidefinite matrix,
N. J. Higham, “Computing a nearest symmetric positive semidefinite matrix,”Linear Algebra and its Applications, vol. 103, pp. 103–118, 1988
1988
-
[69]
Boyd and L
S. Boyd and L. Vandenberghe,Convex Optimization. Cambridge University Press, 2004. Appendix A: Neural-network architectures and training methodology for error mitigation To learn the inverse noisy evolution, we consider three neural-network architectures adapted to different d...
2004
-
[70]
In an LSTM, each cell statecn stores information accumulated up to time stepn and is updated together with the hidden statehn to produce the next pair(cn+1, hn+1)
LSTM ToapproximatetheinversequantumchannelinEq.15wetrainanLSTMnetworktomapnoisyquantumtrajectories back to their original pure states. In an LSTM, each cell statecn stores information accumulated up to time stepn and is updated together with the hidden statehn to produce the n...
-
[71]
This representation is particularly useful when the relevant quantum correlations are spatially non-local in the density-matrix layout
Vision Transformer The ViT [43–45] is more suitable for larger density-matrix representations, which treats the input as a multi- channel image and processes it using self-attention. This representation is particularly useful when the relevant quantum correlations are spatiall...
-
[72]
Whereas the Vision Trans- former emphasizes global self-attention, the U-Net is designed to combine local feature recovery with multiscale contextual reconstruction
Enhanced U-Net As a second image-based model, we employ an enhanced U-Net architecture [46–48]. Whereas the Vision Trans- former emphasizes global self-attention, the U-Net is designed to combine local feature recovery with multiscale contextual reconstruction. This is advanta...
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.