Pith. sign in

REVIEW 3 major objections 5 minor 72 references

Quantum Error Mitigation with Diffusion-Like Models

T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Neural networks trained on synthetic weak-measurement trajectories reconstruct pre-noise quantum states with fidelity above 0.99.

desk verdict Novel weak-measurement diffusion framing for error mitigation, but the headline fidelities may be inflated by channel sharing in the split and there are no baselines. read the letter →

arxiv 2608.05382 v1 pith:5E55FQ3I submitted 2026-08-05 quant-ph

classification quant-ph
keywords quantumerrormitigationweakmeasurementsdiffusionmodelsneuralnetworksdensitymatrixreconstructionlocal-to-globaldecoherencestatetomography
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes that decoherence caused by sequential weak measurements in random Pauli bases can be treated as a quantum diffusion-like forward process, and that neural networks trained on exact synthetic density-matrix trajectories can learn an effective inverse map that recovers the original pre-noise state. The authors report geometric-mean fidelity above 0.99 for a separable six-qubit register and a mean test fidelity of 0.9694 for a five-qubit local-to-global reconstruction task. The practical appeal is that the local-to-global setting relies only on locally accessible information, making the method compatible with noisy and distributed quantum hardware without requiring full state tomography.

What carries the argument

The load-bearing construction is the monitoring quantum channel $M(\rho) = (1-\epsilon)\rho + \epsilon \Phi_A(\rho)$, where $\Phi_A$ is non-selective dephasing in the eigenbasis of a Pauli observable. Composing randomly sampled local copies of this channel along a trajectory produces the forward diffusion-like process, and uniform sampling over $x,y,z$ makes the ensemble-averaged single-qubit step the locally depolarizing channel $D_{1-2\epsilon/3}$. The learned object is the effective inverse map $M^{-1}_{\mathrm{eff}}$, trained with a geometric-mean-fidelity loss and a small L1 regularization term, with outputs projected onto the set of valid density matrices. Architectures are chosen by representation: LSTM for separable registers whose feature count scales linearly with qubit number, and Vision Transformer or enhanced U-Net for full density-matrix representations of entangled registers.

What would settle it

Train the same models on weak-measurement trajectories generated with one interaction strength and state ensemble, then evaluate them on states produced by a different noise model (such as amplitude damping or gate-dependent dephasing reaching the same purity) or on global states whose single-qubit marginals are identical but whose correlations differ; a significant drop in geometric-mean fidelity would show the learned inverse map is tied to the synthetic training channel.

Watch

Extended reading notes

Core claim

The central claim is that a neural network, trained purely on synthetic density matrices produced by a fixed sequential weak-measurement channel, can approximate the inverse of that channel and reconstruct pre-noise states with high fidelity. The forward process is built from random non-selective partial dephasing maps in the $x$, $y$, and $z$ Pauli bases, which on average act as a locally depolarizing channel, while individual trajectories retain basis-dependent structure. The trained models learn an effective data-driven inverse map rather than a physically implementable CPTP inverse channel, and the paper explicitly notes that the local-to-global reconstruction rule is distribution-dependent: one-qubit marginals do not determine a general global density matrix, so the learned map reflects the specific state-preparation and weak-measurement ensemble.

Load-bearing premise

The load-bearing premise is that the noise a real device produces is well approximated by the synthetic sequential weak-measurement channel with a fixed interaction strength, and that the states seen at test time come from the same ensemble used in training; the authors state that the local-to-global reconstruction rule is therefore not a universal map.

Editorial extensions

If this is right

  • For separable registers, local noise channels can be calibrated with a model whose input features grow linearly with qubit count, since each qubit's density matrix is represented by four real parameters and the learned inverse map reaches $F_{\mathrm{GM}} > 0.99$ on a six-qubit example.
  • For entangled registers, the full density-matrix representation is necessary; an LSTM fails with fidelity near 0.5 regardless of dataset size, while Vision Transformer and enhanced U-Net models learn the inverse map once enough training data are supplied.
  • The local-to-global setting, which reconstructs the global state from single-qubit reduced density matrices, is experimentally cheaper than full tomography but is not a universal map from marginals to global states.
  • Because the formal inverse of the depolarizing channel is not completely positive, the learned reconstruction is an effective data-driven inverse applied as classical post-processing rather than a physically implementable inverse channel.
  • The authors view the sequential weak-measurement process as measurement-induced and contextual, making the learned inverse a step toward data-driven characterization of non-unitary dynamics in distributed quantum architectures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next step outside the paper is to test whether the same training procedure transfers to hardware noise: if the real device's effective channel is not close to sequential Pauli-basis dephasing, the learned inverse map should degrade, and retraining on device-calibrated trajectories could restore performance.
  • Because local-to-global reconstruction is distribution-dependent, the method is best understood as a prior-aware estimator rather than a tomographic protocol; it answers what global state is most plausible under the training ensemble, which could be useful for state-preparation verification where the ensemble is known.
  • The ensemble-averaged equivalence to a locally depolarizing channel suggests that deliberately randomizing or twirling measurement bases on hardware could bring real noise closer to the training distribution and improve transfer without changing the architecture.
  • The components, including the geometric-mean-fidelity loss, the density-matrix projection filter, and trajectory-based training with teacher forcing, are reusable as generic post-processing in other learning-based reconstruction pipelines beyond the weak-measurement diffusion picture.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a diffusion-like quantum error mitigation framework in which the forward process is generated by sequential weak measurements in randomly selected Pauli bases, driving the state toward decoherence. Neural networks (LSTM, Vision Transformer, enhanced U-Net) are trained on exact synthetic density-matrix trajectories to learn an effective inverse map that reconstructs the pre-noise state. The method is benchmarked on single-qubit Bloch-vector states, separable multi-qubit registers, entangled registers, and a local-to-global reconstruction task where only reduced density matrices are inputs. The reported results include F_GM > 0.99 for a six-qubit separable register and a mean test fidelity of 0.9694 for a five-qubit local-to-global task.

Significance. If the experimental claims are robust, the work would provide a novel hybrid classical-quantum tool for approximating non-unitary dynamics and mitigating measurement-induced decoherence, with potential applications in distributed and noisy quantum systems. The theoretical framing connecting sequential weak measurements to diffusion processes is conceptually appealing and the paper includes a clear derivation of the averaged depolarizing behavior. The main value rests on the generalization claim that the trained networks learn an effective inverse map for the forward process. However, the current evidence for this generalization is weakened by the absence of channel-disjoint evaluation and the lack of any baseline comparisons, so the significance cannot be fully assessed from the manuscript as written.

major comments (3)
  1. [Sec. IV, Figs. 4 and 5] The data-splitting protocol is not described in a way that excludes a serious confound. Algorithm 1 generates one weak-measurement channel by evolving a single Haar-random state to the purity threshold, and then applies that same recorded channel to many additional random initial states. The paper states an 80/10/10 split over "base states" but never states that all examples generated from one recorded channel are kept in the same partition. If channels are shared between training and test, the model can memorize the inverse of each specific basis sequence (which is identifiable from the trajectory) and report artificially high test fidelity. The central claim of Eq. (10) that the model learns an effective inverse map requires evaluation on channels not seen during training. The authors should report the number of channels, the number of states per channel, and perform a channel-disjoint split, or justify why channel overlap does not inflate the reported mean test fidelity of 0.9694.
  2. [Sec. IV] No baseline comparisons are reported. The paper presents absolute fidelities such as F_GM > 0.99 for a six-qubit separable register, but without comparing to (i) the trivial no-mitigation baseline (fidelity of the final noisy state to the original state), (ii) a standard analytic inversion of the averaged local depolarizing channel, or (iii) established methods such as classical shadows, the reader cannot judge whether the proposed approach provides any practical advantage. For separable states under local depolarizing noise, a simple per-qubit inverse of the averaged channel may already achieve high fidelity, and the LSTM result would then be unsurprising. Adding such baselines is essential to support the claim that the AI models are useful for error mitigation.
  3. [Sec. IV] The paper acknowledges that the local-to-global reconstruction is distribution dependent and that the network learns a rule induced by the specific state-preparation and weak-measurement process. This is an honest and important qualification, but it also means the method is not a general error-mitigation technique; it is a tailored fit to a specific synthetic noise model. The practical relevance depends on whether the trained model can transfer to a real device whose noise is only approximately described by the sequential weak-measurement channel with fixed epsilon. The paper should either provide experiments with a mismatched training/deployment distribution (e.g., different epsilon or different state ensemble) or clearly articulate the conditions under which the method is expected to work. Without this, the claim of "quantum error mitigation" is overly broad.
minor comments (5)
  1. [Sec. IIA] The sentence "Equation 3 assumes this projective-dephasing channel" is confusing because Eq. (3) is the iterative map; the projective-dephasing channel is defined in Eq. (2). Please clarify the cross-reference.
  2. [Sec. IIIB] The line "P_stop = eta(Tr(rho^2) = eta)" is ambiguous; it should read "P_stop is the purity threshold such that the loop stops when Tr(rho^2) <= eta".
  3. [Sec. IV] The caption states "test set (n=18,000 states)" but the text says 1,500 test base states corresponding to 18,000 noisy-clean pairs. Please make the distinction between base states and trajectory pairs consistent.
  4. [Additional Information] The data availability statement says the data are provided within the article, but the actual datasets are not included; only descriptions and aggregated statistics are given. The code is not publicly available due to IP restrictions. For a methodological paper, a public release of code and synthetic-data generation scripts would greatly aid reproducibility; at minimum, the authors should provide the exact hyperparameters and data-generation seeds for all reported experiments.
  5. [Sec. IV] Several reported results (e.g., the six-qubit separable register, the entangled-register scaling curves) are given without error bars or confidence intervals. Only Fig. 5b reports a mean over 10 runs. Error bars or multiple-seed statistics should be provided for all headline fidelities.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the reported reconstruction is a supervised benchmark on synthetic trajectories generated from the paper's own stated weak-measurement model, and the only self-citation is non-load-bearing.

full rationale

The paper's central claim is that neural networks trained on weak-measurement trajectories learn an effective inverse map for the specific synthetic channel ensemble. This is a standard supervised-learning evaluation: the forward process in Eqs. (2) and (11) generates training and test pairs, the loss in Eq. (16) compares the model output to the independently known initial state, and the reported test fidelities are computed on held-out pairs. The target state is not a fitted parameter of the loss, and the test pairs are not identical to the training pairs, so the prediction does not reduce to its inputs by construction. The ideal reconstruction relation in Eq. (10), rho approx M_eff^{-1}(M(rho)), is a definition of the goal, not a proof that the learned map succeeds. The only self-citation is ref. [3] (Tamir and Cohen, with E. Cohen as a co-author), used for the standard von Neumann measurement interaction in Eq. (1); the actual monitoring channel in Eq. (2) is cited to the independent ref. [59], and the numerical results do not depend on ref. [3]. The paper explicitly acknowledges the distribution-dependent nature of the local-to-global task: 'for arbitrary quantum states, one-qubit marginals do not uniquely determine the global density matrix. The network therefore learns a reconstruction rule induced by the specific state-preparation and weak-measurement process, not a universal map.' This honest scoping further supports that the work is a self-contained simulation study rather than a circular derivation. The data-split ambiguity noted in the skeptic materials (whether channels are shared between train and test base states) is a potential evaluation-leakage concern about generalization, but it cannot be exhibited from the paper's equations as a reduction of the claimed prediction to its inputs, so it does not constitute circularity under the stated rules.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The numerical claims rest on the chosen noise model (weak measurement strength, purity threshold), the state-preparation ensemble, and training loss hyperparameters. No new physical entities are proposed.

free parameters (3)
  • measurement strength epsilon = 0.1 to 0.15 (per experiment)
    Sets the per-step dephasing strength in the forward process (Algorithm 1); results were obtained with fixed values and not swept systematically.
  • purity threshold eta (P_stop) = 0.55 or 0.6
    Stops the forward process and fixes the final noise level and sequence length; chosen by the authors.
  • loss weights lambda_F and lambda_1 = 10 and 1e-2
    Weights in the combined fidelity and L1 loss (Eq. 16); chosen as hyperparameters.
assumptions (3)
  • domain assumption The monitoring channel M(ρ) = (1-ε)ρ + ε Φ_A(ρ), with non-selective dephasing Φ_A, describes the effect of a weak measurement with discarded outcomes (Eq. 2).
    Foundation of the forward process; standard weak-measurement model adopted from refs. [3,59].
  • domain assumption The training state distribution consists of Haar-random single-qubit states and, for multi-qubit registers, random product states followed by n-1 random CNOT gates.
    The learned reconstruction is distribution specific; the local-to-global map is only defined on this ensemble.
  • domain assumption The synthetic weak-measurement process is a faithful surrogate for the decoherence that error mitigation is intended to correct.
    Transfers the simulation results to real quantum devices; not validated experimentally.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Quantum Error Mitigation with Diffusion-Like Models." pith.science (2026). https://pith.science/paper/5E55FQ3I

@misc{pith2026260805382,
  author       = {Pith},
  title        = {Pith review of: Quantum Error Mitigation with Diffusion-Like Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5E55FQ3I}},
  note         = {Machine review of arXiv:2608.05382}
}
read the original abstract

Coupling between a quantum system and its environment causes decoherence by transferring information from the system to environmental degrees of freedom. When discretized in time, such interactions can be interpreted as sequences of weak measurements that provide an effective model of noisy quantum dynamics. Motivated by this picture, we propose an AI-assisted error-mitigation framework for quantum diffusion processes generated by sequential local weak measurements. The forward process progressively erases information from the input state through weak measurements performed in randomly selected Pauli bases, producing basis-dependent local dephasing and locally depolarizing dynamics on average. Machine-learning models are trained on exact synthetic density matrices to learn a channel- and distribution-specific denoising map and estimate the corresponding pre-noise state. We benchmark the approach on single-qubit states and separable and entangled multi-qubit registers. We also study distribution-dependent local-to-global reconstruction, in which local reduced density matrices are used to reconstruct the global state. This experimentally motivated setting relies on locally accessible information and is therefore compatible with noisy and distributed quantum systems. More broadly, the framework provides a hybrid classical-quantum approach for approximating non-unitary dynamics and mitigating coherence loss.

Figures

Figures reproduced from arXiv: 2608.05382 by the authors.

Figure 1
Figure 1. FIG. 1: Schematic of the diffusion-like quantum-noise process and the AI-assisted inverse-channel reconstruction. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2: Example of the separable six-qubit denoising task. An input state evolves under a noisy quantum channel [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3: Example of the local-to-global inverse reconstruction task for a five-qubit quantum register, which [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: FIG. 4: Configuration and local-to-global reconstruction results. (a) Example of a noisy local quantum-register [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5: Test fidelity as a function of training-set size for different four-qubit reconstruction tasks and model [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: FIG. 6: Schematic architectures of (a) the Vision Transformer and (b) the enhanced U-Net used in this work. [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

72 extracted references · 57 canonical work pages

  1. [1]

    This representation is useful for building intuition, since the noisy evolution can be visualized directly as a trajectory inside the Bloch sphere

    Single-qubit Bloch-vector dataset We begin with the simplest reconstruction task: the inverse evolution of a single qubit represented by its Bloch vector, ρ= 1 2 (I+⃗ r·⃗ σ),(18) where⃗ ri = (rx i , ry i , rz i )contains three real features per time step. This representation is useful for building intuition, since the noisy evolution can be visualized dir...

  2. [2]

    QERNEL” and the “Interdisciplinary Center for the Theory of Quantum Computing

    Separable multi-qubit dataset We next consider the extension from a single qubit to a register ofnseparable, non-interacting qubits. Each local qubit state is represented by a2×2density matrix, ρ(t,k) i = α(k,t) 1,i β(k,t) 1,i +iβ (k,t) 2,i β(k,t) 1,i −iβ (k,t) 2,i δ(k,t) 1,i ! ,(24) 10 where Hermiticity implies that only four real parameters are required...

  3. [3]

    Coecke and A

    B. Coecke and A. Kissinger,Picturing Quantum Processes: A First Course in Quantum Theory and Diagrammatic Rea- soning. Cambridge: Cambridge University Press, 2017

  4. [4]

    Dynamics of quantum causal structures,

    E. Castro-Ruiz, F. Giacomini, and Č. Brukner, “Dynamics of quantum causal structures,”Physical Review X, vol. 8, no. 1, p. 011047, 2018

  5. [5]

    Introduction to weak measurements and weak values,

    B. Tamir and E. Cohen, “Introduction to weak measurements and weak values,”Quanta, vol. 2, no. 1, pp. 7–17, 2013

  6. [6]

    Decoherence, einselection, and the quantum origins of the classical,

    W. H. Zurek, “Decoherence, einselection, and the quantum origins of the classical,”Reviews of modern physics, vol. 75, no. 3, p. 715, 2003

  7. [7]

    Demonstration of logical qubits and repeated error correction with better-than- physical error rates,

    A. Paetznick, M. P. da Silva, C. Ryan-Anderson, J. M. Bello-Rivas, J. P. Campora III, A. Chernoguzov, J. M. Dreiling, C. Foltz, F. Frachon, J. P. Gaebler,et al., “Demonstration of logical qubits and repeated error correction with better-than- physical error rates,” 2024

  8. [8]

    Quantum error correction below the surface code threshold,

    Google Quantum AI and Collaborators, “Quantum error correction below the surface code threshold,”Nature, vol. 638, pp. 920–926, 2025

Show all 72 references
  1. [9]

    A fault-tolerant neutral-atom architecture for universal quantum computation,

    D. Bluvstein, A. A. Geim, S. H. Li, S. J. Evered, J. P. Bonilla Ataides, G. Baranes, A. Gu, T. Manovitz, M. Xu, M. Kalinowski,et al., “A fault-tolerant neutral-atom architecture for universal quantum computation,”Nature, vol. 649, no. 8095, pp. 39–46, 2026

  2. [10]

    Lattice surgery realized on two distance-three repetition codes with superconducting qubits,

    I. Besedin, M. Kerschbaum, J. Knoll, I. Hesner, L. Bödeker, L. Colmenarez, L. Hofele, N. Lacroix, C. Hellings, F. Swiadek, et al., “Lattice surgery realized on two distance-three repetition codes with superconducting qubits,”Nature physics, vol. 22, no. 2, pp. 189–194, 2026

  3. [11]

    Hardware-efficient quantum error correction via concatenated bosonic qubits,

    H. Putterman, K. Noh, C. T. Hann, G. S. MacCabe, S. Aghaeimeibodi, R. N. Patel, M. Lee, W. M. Jones, H. Moradinejad, R. Rodriguez,et al., “Hardware-efficient quantum error correction via concatenated bosonic qubits,”Nature, vol. 638, no. 8052, pp. 927–934, 2025

  4. [12]

    On the importance of error mitigation for quantum computation,

    D. Aharonov, O. Alberton, I. Arad, Y. Atia, E. Bairey, Z. Brakerski, I. Cohen, O. Golan, I. Gurwich, O. Kenneth, E. Leviatan, N. H. Lindner, R. A. Melcer, A. Meyer, G. Schul, and M. Shutman, “On the importance of error mitigation for quantum computation,” 2025

  5. [13]

    Noisy intermediate-scale quantum algorithms,

    K. Bharti, A. Cervera-Lierta, T. H. Kyaw, T. Haug, S. Alperin-Lea, A. Anand, M. Degroote, H. Heimonen, J. S. Kottmann, T. Menke, W.-K. Mok, S. Sim, L.-C. Kwek, and A. Aspuru-Guzik, “Noisy intermediate-scale quantum algorithms,”Reviews of Modern Physics, vol. 94, p. 015004, 2022

  6. [14]

    M. A. Nielsen and I. L. Chuang,Quantum computation and quantum information. Cambridge university press, 2010

  7. [15]

    On the complementary quantum capacity of the depolarizing channel,

    D. Leung and J. Watrous, “On the complementary quantum capacity of the depolarizing channel,”Quantum, vol. 1, p. 28, 2017

  8. [16]

    The capacity of the quantum depolarizing channel,

    C. King, “The capacity of the quantum depolarizing channel,”IEEE Transactions on Information Theory, vol. 49, no. 1, pp. 221–229, 2003

  9. [17]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”Advances in neural information processing systems, vol. 33, pp. 6840–6851, 2020

  10. [18]

    Improved denoising diffusion probabilistic models,

    A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” inInternational Conference on Machine Learning, pp. 8162–8171, PMLR, 2021

  11. [19]

    Deep unsupervised learning using nonequilibrium thermodynamics,

    J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” inInternational conference on machine learning, pp. 2256–2265, PMLR, 2015

  12. [20]

    Shadow tomography of quantum states,

    S. Aaronson, “Shadow tomography of quantum states,” inProceedings of the 50th annual ACM SIGACT symposium on theory of computing, pp. 325–338, 2018

  13. [21]

    Predicting many properties of a quantum system from very few measurements,

    H.-Y. Huang, R. Kueng, and J. Preskill, “Predicting many properties of a quantum system from very few measurements,” Nature Physics, vol. 16, no. 10, pp. 1050–1057, 2020

  14. [22]

    Classical shadows with noise,

    D. E. Koh and S. Grewal, “Classical shadows with noise,”Quantum, vol. 6, p. 776, 2022

  15. [23]

    Quantum error mitigation,

    Z. Cai, R. Babbush, S. C. Benjamin, S. Endo, W. J. Huggins, Y. Li, J. R. McClean, and T. E. O’Brien, “Quantum error mitigation,”Reviews of Modern Physics, vol. 95, no. 4, p. 045005, 2023

  16. [24]

    Error mitigation extends the computational reach of a noisy quantum processor,

    A. Kandala, K. Temme, A. D. Córcoles, A. Mezzacapo, J. M. Chow, and J. M. Gambetta, “Error mitigation extends the computational reach of a noisy quantum processor,”Nature, vol. 567, no. 7749, pp. 491–495, 2019

  17. [25]

    Learning-based quantum error mitigation,

    A. Strikis, D. Qin, Y. Chen, S. C. Benjamin, and Y. Li, “Learning-based quantum error mitigation,”PRX Quantum, vol. 2, no. 4, p. 040330, 2021

  18. [26]

    Error mitigation with clifford quantum-circuit data,

    P. Czarnik, A. Arrasmith, P. J. Coles, and L. Cincio, “Error mitigation with clifford quantum-circuit data,”Quantum, vol. 5, p. 592, 2021

  19. [27]

    Differentiable quantum architecture search,

    S.-X. Zhang, C.-Y. Hsieh, S. Zhang, and H. Yao, “Differentiable quantum architecture search,”Quantum Sci. Technol., vol. 7, no. 4, p. 045023, 2022. 14

  20. [28]

    Stability of classical shadows under gate-dependent noise,

    R. Brieger, M. Heinrich, I. Roth, and M. Kliesch, “Stability of classical shadows under gate-dependent noise,”Phys. Rev. Lett., vol. 134, p. 090801, Mar 2025

  21. [29]

    Vaqem: A variational ap- proach to quantum error mitigation,

    G. S. Ravi, K. N. Smith, P. Gokhale, A. Mari, N. Earnest, A. Javadi-Abhari, and F. T. Chong, “Vaqem: A variational ap- proach to quantum error mitigation,” in2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pp. 288–303, IEEE, 2022

  22. [30]

    Efficient quantum tomography,

    R. O’Donnell and J. Wright, “Efficient quantum tomography,” inProceedings of the forty-eighth annual ACM symposium on Theory of Computing, pp. 899–912, 2016

  23. [31]

    Quantum resource theories,

    E. Chitambar and G. Gour, “Quantum resource theories,”Rev. Mod. Phys., vol. 91, p. 025001, Apr 2019

  24. [32]

    How the result of a measurement of a component of the spin of a spin-1/2 particle can turn out to be 100,

    Y. Aharonov, D. Z. Albert, and L. Vaidman, “How the result of a measurement of a component of the spin of a spin-1/2 particle can turn out to be 100,”Physical review letters, vol. 60, no. 14, p. 1351, 1988

  25. [33]

    Weak-value amplification of the nonlinear effect of a single photon,

    M. Hallaji, A. Feizpour, G. Dmochowski, J. Sinclair, and A. M. Steinberg, “Weak-value amplification of the nonlinear effect of a single photon,”Nature Physics, vol. 13, no. 6, pp. 540–544, 2017

  26. [34]

    Quantum cryptography with weak measurements,

    J. E. Troupe and J. M. Farinholt, “Quantum cryptography with weak measurements,”arXiv preprint arXiv:1702.04836, 2017

  27. [35]

    Error mitigation for short-depth quantum circuits,

    K. Temme, S. Bravyi, and J. M. Gambetta, “Error mitigation for short-depth quantum circuits,”Physical review letters, vol. 119, no. 18, p. 180509, 2017

  28. [36]

    Quantum error mitigation with artificial neural network,

    C. Kim, K. D. Park, and J.-K. Rhee, “Quantum error mitigation with artificial neural network,”IEEE Access, vol. 8, pp. 188853–188860, 2020

  29. [37]

    Quantum readout error mitigation via deep learning,

    J. Kim, B. Oh, Y. Chong, E. Hwang, and D. K. Park, “Quantum readout error mitigation via deep learning,”New Journal of Physics, vol. 24, no. 7, p. 073009, 2022

  30. [38]

    Reservoir computing approach to quantum state measurement,

    G. Angelatos, S. A. Khan, and H. E. Türeci, “Reservoir computing approach to quantum state measurement,”Physical Review X, vol. 11, no. 4, p. 041062, 2021

  31. [39]

    The variational quantum eigensolver: a review of methods and best practices,

    J. Tilly, H. Chen, S. Cao, D. Picozzi, K. Setia, Y. Li, E. Grant, L. Wossnig, I. Rungger, G. H. Booth,et al., “The variational quantum eigensolver: a review of methods and best practices,”Physics Reports, vol. 986, pp. 1–128, 2022

  32. [40]

    A variational eigenvalue solver on a photonic quantum processor,

    A. Peruzzo, J. McClean, P. Shadbolt, M.-H. Yung, X.-Q. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’brien, “A variational eigenvalue solver on a photonic quantum processor,”Nature communications, vol. 5, no. 1, p. 4213, 2014

  33. [41]

    The theory of variational hybrid quantum-classical algo- rithms,

    J. R. McClean, J. Romero, R. Babbush, and A. Aspuru-Guzik, “The theory of variational hybrid quantum-classical algo- rithms,”New Journal of Physics, vol. 18, no. 2, p. 023023, 2016

  34. [42]

    Long short-term memory,

    S. Hochreiter, “Long short-term memory,”Neural Computation MIT-Press, 1997

  35. [43]

    Empirical evaluation of gated recurrent neural networks on sequence modeling,

    J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,”arXiv preprint arXiv:1412.3555, 2014

  36. [44]

    Sequence to sequence learning with neural networks,

    I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,”Advances in neural infor- mation processing systems, vol. 27, 2014

  37. [45]

    Cvt: Introducing convolutions to vision transformers,

    H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang, “Cvt: Introducing convolutions to vision transformers,” inProceedings of the IEEE/CVF international conference on computer vision, pp. 22–31, 2021

  38. [46]

    Training data-efficient image transformers & distillation through attention,

    H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” inInternational conference on machine learning, pp. 10347–10357, PMLR, 2021

  39. [47]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly,et al., “An image is worth 16x16 words: Transformers for image recognition at scale,”arXiv preprint arXiv:2010.11929, 2020

  40. [48]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” inInter- national Conference on Medical image computing and computer-assisted intervention, pp. 234–241, Springer, 2015

  41. [49]

    nnu-net: a self-configuring method for deep learning-based biomedical image segmentation,

    F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, and K. H. Maier-Hein, “nnu-net: a self-configuring method for deep learning-based biomedical image segmentation,”Nature methods, vol. 18, no. 2, pp. 203–211, 2021

  42. [50]

    Attention u-net: Learning where to look for the pancreas,

    O. Oktay, J. Schlemper, L. L. Folgoc, M. Lee, M. Heinrich, K. Misawa, K. Mori, S. McDonagh, N. Y. Hammerla, B. Kainz, et al., “Attention u-net: Learning where to look for the pancreas,”arXiv preprint arXiv:1804.03999, 2018

  43. [51]

    Prescription for experimental determination of the dynamics of a quantum black box,

    I. L. Chuang and M. A. Nielsen, “Prescription for experimental determination of the dynamics of a quantum black box,” Journal of Modern Optics, vol. 44, no. 11-12, pp. 2455–2467, 1997

  44. [52]

    Complete characterization of a quantum process: the two-bit quantum gate,

    J. Poyatos, J. I. Cirac, and P. Zoller, “Complete characterization of a quantum process: the two-bit quantum gate,”Physical Review Letters, vol. 78, no. 2, p. 390, 1997

  45. [53]

    Deep learning on image denoising: An overview,

    C. Tian, L. Fei, W. Zheng, Y. Xu, W. Zuo, and C.-W. Lin, “Deep learning on image denoising: An overview,”Neural Networks, vol. 131, pp. 251–275, 2020

  46. [54]

    Image denoising review: From classical to state-of-the-art approaches,

    B. Goyal, A. Dogra, S. Agrawal, B. S. Sohi, and A. Sharma, “Image denoising review: From classical to state-of-the-art approaches,”Information fusion, vol. 55, pp. 220–244, 2020

  47. [55]

    Dynamic noise aware training for speech enhancement based on deep neural networks.,

    Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “Dynamic noise aware training for speech enhancement based on deep neural networks.,” inInterspeech, vol. 1, pp. 2670–2674, 2014

  48. [56]

    Mofusion: A framework for denoising-diffusion-based motion synthesis,

    R. Dabral, M. H. Mughal, V. Golyanik, and C. Theobalt, “Mofusion: A framework for denoising-diffusion-based motion synthesis,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9760–9770, 2023

  49. [57]

    The quantum internet,

    H. J. Kimble, “The quantum internet,”Nature, vol. 453, no. 7198, pp. 1023–1030, 2008

  50. [58]

    Quantum repeaters: From quantum networks to the quantum internet,

    K. Azuma, S. E. Economou, D. Elkouss, P. Hilaire, L. Jiang, H.-K. Lo, and I. Tzitrin, “Quantum repeaters: From quantum networks to the quantum internet,”Reviews of Modern Physics, vol. 95, no. 4, p. 045006, 2023

  51. [59]

    A one-way quantum computer,

    R. Raussendorf and H. J. Briegel, “A one-way quantum computer,”Physical review letters, vol. 86, no. 22, p. 5188, 2001

  52. [60]

    Measurement-based quantum computation on cluster states,

    R. Raussendorf, D. E. Browne, and H. J. Briegel, “Measurement-based quantum computation on cluster states,”Physical review A, vol. 68, no. 2, p. 022312, 2003. 15

  53. [61]

    Information-reality complementarity: The role of measurements and quantum reference frames,

    P. R. Dieguez and R. M. Angelo, “Information-reality complementarity: The role of measurements and quantum reference frames,”Phys. Rev. A, vol. 97, no. 2, p. 022107, 2018

  54. [62]

    Efficient measurement error mitigation with subsystem-balanced pauli twirling,

    X.-Y. Xu, C. Ding, and W.-S. Bao, “Efficient measurement error mitigation with subsystem-balanced pauli twirling,” 2025

  55. [63]

    Quantum fidelity measures for mixed states,

    Y.-C. Liang, Y.-H. Yeh, P. E. Mendonça, R. Y. Teh, M. D. Reid, and P. D. Drummond, “Quantum fidelity measures for mixed states,”Reports on Progress in Physics, vol. 82, no. 7, p. 076001, 2019

  56. [64]

    Neural autoregressive distribution estimation,

    B. Uria, M.-A. Côté, K. Gregor, I. Murray, and H. Larochelle, “Neural autoregressive distribution estimation,” 2016

  57. [65]

    Generating sequences with recurrent neural networks,

    A. Graves, “Generating sequences with recurrent neural networks,” 2014

  58. [66]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017

  59. [67]

    Efficient method for computing the maximum-likelihood quantum state from measurements with additive gaussian noise,

    J. A. Smolin, J. M. Gambetta, and G. Smith, “Efficient method for computing the maximum-likelihood quantum state from measurements with additive gaussian noise,”Phys. Rev. Lett., vol. 108, p. 070502, 2012

  60. [68]

    Computing a nearest symmetric positive semidefinite matrix,

    N. J. Higham, “Computing a nearest symmetric positive semidefinite matrix,”Linear Algebra and its Applications, vol. 103, pp. 103–118, 1988

  61. [69]

    Boyd and L

    S. Boyd and L. Vandenberghe,Convex Optimization. Cambridge University Press, 2004. Appendix A: Neural-network architectures and training methodology for error mitigation To learn the inverse noisy evolution, we consider three neural-network architectures adapted to different d...

  62. [70]

    In an LSTM, each cell statecn stores information accumulated up to time stepn and is updated together with the hidden statehn to produce the next pair(cn+1, hn+1)

    LSTM ToapproximatetheinversequantumchannelinEq.15wetrainanLSTMnetworktomapnoisyquantumtrajectories back to their original pure states. In an LSTM, each cell statecn stores information accumulated up to time stepn and is updated together with the hidden statehn to produce the n...

  63. [71]

    This representation is particularly useful when the relevant quantum correlations are spatially non-local in the density-matrix layout

    Vision Transformer The ViT [43–45] is more suitable for larger density-matrix representations, which treats the input as a multi- channel image and processes it using self-attention. This representation is particularly useful when the relevant quantum correlations are spatiall...

  64. [72]

    Whereas the Vision Trans- former emphasizes global self-attention, the U-Net is designed to combine local feature recovery with multiscale contextual reconstruction

    Enhanced U-Net As a second image-based model, we employ an enhanced U-Net architecture [46–48]. Whereas the Vision Trans- former emphasizes global self-attention, the U-Net is designed to combine local feature recovery with multiscale contextual reconstruction. This is advanta...

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.