Pith. sign in

REVIEW 5 major objections 3 minor 2 cited by

The paper claims that quantum error correction can be made adaptive: a bandit-retuned variational unitary tracks time-varying X-to-Z noise, cutting logical infidelity by about 18-fold for qubit codes and 3-fold for qutrit codes when the noi

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Adaptive quantum error correction: multi-agent RL discovers QEC circuits offline; a bandit-controlled variational layer retrains online, cutting logical infidelity about 18x (qubit) and 3x (qutrit) under drifting bit/phase-flip noise at high sampling rates.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection MARL code discovery is a solid, reproducible contribution; the adaptive BRAVE layer is a real idea, but the headline gains only hold for a noise channel the variational ansatz can exactly invert. the 5 major comments →

arxiv 2509.03974 v2 pith:C7XX4ZBN submitted 2025-09-04 quant-ph

Real-time adaptive quantum error correction by model-free multi-agent learning

classification quant-ph
keywords quantum error correctionmulti-agent reinforcement learningvariational unitarybandit algorithmnon-stationary noisequtrit codesKnill-Laflamme conditionsadaptive calibration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that quantum error correction need not be static: a two-level learning scheme can discover codes from scratch and then keep them matched to noise that changes over time. Offline, three reinforcement-learning agents build the encoder, syndrome measurement, and recovery circuits for qubit and qutrit systems, guided only by the Knill-Laflamme orthogonality conditions. Online, a lightweight variational layer—a single unitary applied to the whole register—is periodically retuned by BRAVE, a bandit algorithm that decides when the code needs recalibration. The payoff, if the argument holds, is that learned error-correcting circuits track drifting bit-flip/phase-flip noise: roughly 18-fold lower logical infidelity for qubits and 3-fold for qutrits compared with static QEC, plus a wider range of error probabilities that stay above 99% logical fidelity.

Core claim

Under non-stationary noise E2(t)=sqrt(p)(Z X†)^{α(t)} X, a code optimized for one error basis stops satisfying the Knill-Laflamme conditions as α(t) moves X toward Z. The paper's central claim is that conjugating every stage of an already-learned QEC cycle by a single variational unitary U(θ)=exp(i Σ θ_k λ_k), with θ retuned by a gradient bandit using fidelity as reward, approximately restores those conditions at each time step. Because the stabilizers and recovery operators transform together (S'_i = U S_i U†, E'_s = U E_s U†), the whole cycle co-adapts. Numerical simulations over parameter sweeps show the adaptive cycle outperforms static correction for both qubits and qutrits whenever the

What carries the argument

The load-bearing object is the global variational unitary U(θ)=exp(i Σ_k θ_k λ_k), parametrized by the d²−1 generators of SU(d)—Pauli matrices for qubits, Gell-Mann matrices for qutrits. It is inserted as a calibration layer that simultaneously transforms encoder, stabilizers, and recovery, so all stages of the QEC cycle stay consistent under retraining. The decision of when to retune θ is made by a gradient bandit with a softmax keep/retrain policy and a reset mechanism; retraining runs a simplex-based optimizer on the fidelity reward. This 'discover once, adapt continuously' split carries the argument: offline MARL supplies the code, online BRAVE supplies the tracking.

Load-bearing premise

The noise drift is modeled as one global rotation of a single fixed Pauli error, applied identically to every qudit, so a single global unitary can undo it; real drift that is spatially inhomogeneous or changes the error model itself is outside the paper's evidence.

What would settle it

On a multiqubit device, impose two different α(t) rotations on different subregions while running BRAVE with its single global U(θ): if the logical-fidelity improvement over static QEC drops toward zero as the regions drift out of sync, the single-unitary assumption is the load-bearing limit. Alternatively, run with sampling rate fs below the noise frequency and observe the adaptive gain collapse to the static baseline.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Learned QEC codes can be kept valid under non-stationary noise without retraining the full reinforcement-learning stack.
  • At sampling rates high relative to drift, logical infidelity falls roughly 18-fold for qubits and 3-fold for qutrits versus static codes, and the physical error probability tolerated at 99% logical fidelity increases by Δp=0.095 for qubits and Δp=0.025 for qutrits.
  • Because the Clifford gate set and SU(d) parameterization are defined for general d, the same framework extends beyond qubits and qutrits to arbitrary qudit architectures.
  • Because stabilizers and recovery transform with the encoder, a single variational layer adapts all three QEC stages consistently rather than patching one component.
  • MARL also discovers qutrit codes including a generalized Shor-like code and hybrid-error codes, indicating automated code discovery can reach higher-dimensional codes.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The single global unitary restricts the method to noise drifts that are spatially uniform; per-qubit or per-region drift would likely require local variational parameters, a case the paper does not analyze.
  • The sampling-rate dependence suggests a Nyquist-like limit: when fs drops below roughly twice the noise-drift frequency, retraining decisions become stale and the adaptive gain should vanish; a sweep of fs/ν could map that boundary.
  • The keep/retrain bandit is a generic meta-optimizer; the same fidelity-reward mechanism could be applied to tracking slowly drifting qubit frequencies, gate calibrations, or other time-varying control errors.
  • A direct experimental test would run BRAVE on a tunable transmon with modulated flux noise, comparing logical-fidelity trajectories against static QEC on the same device.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 3 minor

Summary. The paper proposes a two-level framework for quantum error correction under time-varying noise. At the offline level, multi-agent reinforcement learning (MARL) is used to discover complete QEC cycles—encoder, syndrome extraction, and recovery—as explicit Clifford circuits, for both qubits and qutrits, without prescribing a code family. At the online level, the BRAVE layer uses a gradient bandit to decide when to retune a global variational unitary U(θ)=exp(iΣθ_k λ_k) that conjugates the encoder, stabilizers, and recovery, adapting to a time-dependent noise channel. The authors report rediscovery of standard qubit codes, discovery of qutrit codes, and an approximately 18-fold (qubit) and 3-fold (qutrit) improvement in logical fidelity relative to a static baseline under a sinusoidal X↔Z noise rotation, provided the sampling rate is sufficiently high. The main adaptive claim is supported only for a specialized noise model that the variational ansatz can exactly invert, and the quantitative headline figures lack statistical characterization.

Significance. If the 'discover once, adapt continuously' paradigm were demonstrated for realistic non-stationary hardware noise, it would be a valuable contribution to practical QEC. The MARL code-discovery component is a genuine strength: the agents reproduce known codes (bit-flip, phase-flip, 5-qubit, Shor) and extend to qutrit codes, with code and notebook released. However, the online adaptive contribution is currently established only for a noise channel that is a global unitary rotation of a fixed Pauli error, applied identically to every qudit—a channel for which the single global U(θ) ansatz is, by construction, the exact inverse. The paper's quantitative claims also rest on single-trajectory pie charts with no error bars. These issues are load-bearing for the headline 'realistic hardware/non-stationary noise' claim, so the significance of the adaptive part is conditional pending broader noise tests and proper statistics.

major comments (5)
  1. [§Results, Eq. (1) and Supp. Eq. (S9)]
  2. [§Results, Fig. 4 (d)–(g)]
  3. [§Results, Fig. 4 baseline]
  4. [Supp. Algorithm 1; §Methods C]
  5. [Methods Sec. C; Supp. Sec. II 2]
minor comments (3)
  1. [General] Typos and grammar: 'Retrainng' in the BRAVE heading (Methods C), 'Finaly' in the Introduction, 'emcompass' in Conclusions, 'weather' for 'whether' in 'determine weather a QEC code' (Results, Fig. 2 caption area).
  2. [§Results, Fig. 2] The learning curves in Fig. 2(c)–(e) would benefit from labels for the plotted curves; currently the text refers to them as 'the green curve' and 'the red curve' but the colors are not described in the figure itself.
  3. [§Supplemental Materials, Algorithm 1] Algorithm 1 resets the bandit preferences H to [h0, h1] after a retrain, but the text does not state how h0 and h1 are chosen or whether they affect the reported results. A brief note on the sensitivity to these initial preferences would improve reproducibility.

Circularity Check

0 steps flagged

No significant circularity: BRAVE is a closed-loop fidelity optimizer, not a definitional re-derivation; the unitary-rotation noise model is a scope limitation rather than a circular step.

full rationale

Walking the derivation chain: the MARL code-discovery component is self-contained, optimizing Knill-Laflamme conditions (Supp. Eq. S1) and validated by reproducing known codes (bit/phase flip, [[5,1,3]], Shor), so it does not reduce to its inputs. BRAVE is a feedback controller: the gradient bandit receives only fidelity rewards (Methods C; Supp. Algorithm 1) and chooses keep/retrain, while Nelder-Mead updates the variational parameters to maximize that same fidelity. The 18x/3x figures are observed closed-loop simulation results, not quantities forced by definition from the fitted parameters. The genuine limitation is scope: the test channel Eq. (1), E2(t)=sqrt(p)(Z X†)^{alpha(t)}X, is a global unitary conjugation of a fixed Pauli error, and the variational layer U(theta)=exp(i sum theta_k lambda_k) (Supp. Eq. S9) is also a global SU(d) unitary; Methods C conjugates stabilizers and errors by the same U. Thus the optimizer has exactly the degrees of freedom to invert the imposed drift, and Fig. 3e shows it does (Hadamard-like theta). This means the quantitative gains are demonstrated only for this specialized single-parameter, spatially uniform rotation, not for generic multi-axis or inhomogeneous drift. That is an overgeneralization/correctness-claim weakness, not circularity: the paper does not define the improvement as the fit, and the bandit is not given alpha(t). Self-citations (e.g., [89] for unitary interpolation) are background or implementation details, not load-bearing uniqueness theorems. The regret 'treatment' in Supp. II 2 is a differential inequality plus simulations rather than a completed bound, and Algorithm 1 lines 5/9 mention computing/updating a noise model and retraining with current noise, which sits uneasily with the 'model-free' label; these are missing-support and terminology issues for a correctness pass, not circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

The central adaptive claim rests on the assumption that hardware noise drift can be represented as a single-qudit unitary rotation of a fixed Pauli error set. The paper motivates this with a schematic transmon flux-noise argument (Methods D) but the exact equivalence relies on the approximations cos(phi) ~ Z, sin(phi) ~ X and on a single global U(theta). Other choices are domain assumptions standard to stabilizer QEC.

free parameters (4)
  • Variational angles theta_k in U(theta) = exp(i sum theta_k lambda_k) = qubit: (3pi/4, 5pi/4, 5pi/4) at alpha=1 (Hadamard-equivalent); qutrit values not reported
    Optimized online by Nelder-Mead against measured fidelity; these angles implement the code rotation that tracks the noise drift (Fig. 3e).
  • Bandit baseline reward F-bar (fidelity threshold) = 0.99
    Hand-chosen threshold used in the bandit update and in the robustness analysis (Fig. 4c).
  • Bandit learning rate eta = not specified
    Appears in the regret analysis (Supp. II.2) but no numerical value is given for the simulations; the algorithm behavior depends on it.
  • Bandit initial preferences h0, h1 = not specified
    Initial softmax preferences in Algorithm 1, set by hand.
axioms (6)
  • domain assumption The search is restricted to the generalized Clifford gate set {CNOT_d, H_d, S_q}.
    Methods Sec. A Eq. (5); limits discovered codes to stabilizer codes and excludes non-stabilizer codes.
  • domain assumption Noise is discretized into qudit Pauli operators (X, Z, X^2, Z^2, ...) and syndrome measurement projects errors onto Pauli error subspaces.
    Methods Sec. A; assumes Pauli (not amplitude, leakage, or correlated) noise.
  • ad hoc to paper The time-varying noise channel is a one-parameter unitary rotation of a fixed error, E2(t) = sqrt(p)(Z X-dagger)^(alpha(t)) X, applied uniformly to every qudit, so a single global U(theta) can undo it.
    Eq. 1 and Supp. Eq. S9; the whole BRAVE demonstration depends on this matching, acknowledged in Fig. 3(e).
  • domain assumption Fidelity of the recovered logical state is a sufficient reward for the bandit to decide keep/retrain.
    BRAVE algorithm (Supp. Algorithm 1); assumes fidelity statistics are available and informative during normal operation.
  • domain assumption The transmon flux-noise argument justifies the X-Z interpolation using the approximations cos(phi) ~ Z and sin(phi) ~ X.
    Methods Sec. D; these approximations neglect higher-order terms in the transmon eigenbasis.
  • standard math The Knill-Laflamme conditions are necessary and sufficient for correctability of the discrete error set.
    Methods Sec. A Eq. (3); standard QEC theory (Ref 96).

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Real-time adaptive quantum error correction by model-free multi-agent learning." pith.science (2026). https://pith.science/paper/C7XX4ZBN

@misc{pith2026250903974,
  author       = {Pith},
  title        = {Pith review of: Real-time adaptive quantum error correction by model-free multi-agent learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/C7XX4ZBN}},
  note         = {Machine review of arXiv:2509.03974}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Quantum error correction (QEC) is essential for scalable quantum computing, yet existing approaches rely on static assumptions about noise that break down in realistic hardware, where error channels drift over time. We introduce a unified framework that separates QEC into two learning timescales: offline code discovery and online adaptation. Offline, Multi-Agent Reinforcement Learning (MARL) autonomously discovers complete QEC cycles as explicit quantum circuits, with separate agents responsible for encoding, syndrome extraction, and error recovery, and without prescribing a code family or circuit ansatz. Online, a lightweight adaptive layer, termed Bandit Retraining for Adaptive Variational Error Correction (BRAVE), continuously retunes a low-dimensional variational parameterization without retraining the full MARL stack. This yields a "discover once, adapt continuously" strategy that combines the flexibility of learned codes with real-time adaptation to non-stationary noise. At sufficiently high sampling rates relative to the noise drift, our method reduces logical infidelity by roughly 18-fold for qubit codes and 3-fold for qutrit codes compared to static error correction, while substantially extending robustness to noise fluctuations. These results establish a paradigm in which QEC is no longer static but is dynamically optimized for realistic quantum hardware.

Figures

Figures reproduced from arXiv: 2509.03974 by Felix Motzoi, Francesco Preti, Francisco Andr\'es C\'ardenas-L\'opez, Manuel Guatto, Michael Schilling, Tommaso Calarco.

Figure 1
Figure 1. Figure 1: FIG. 1. The two important algorithmic structures developed in our work: (a) the multi-agent RL-based optimization of QEC codes (b) and [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. RL-based optimization of the bit flip (a) and the phase flip (b) code. The circuits in the upper part represent the best solution delivered [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Illustrative plots from Section D and section II: (a)-(b) [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 3
Figure 3. Figure 3: Initially, the parameters are clustered around the pole [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4. Performance comparison of standard and adaptive variational approaches (BRAVE) under different noise and sampling conditions: [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Quantum circuit partition as a maze: emerging percolation transition via path finding

    quant-ph 2026-06 unverdicted novelty 6.0

    Quantum circuit partitioning is formalized as a maze path problem, revealing a percolation phase transition that separates partitionable from non-partitionable regimes when the CNOT-to-qubit ratio is near one.

  2. Simple analytical flux-tuned iSWAP pulses for leakage suppression

    quant-ph 2026-06 unverdicted novelty 5.0

    Φ-DRAG derives modified analytical flux modulation protocols that suppress leakage below 10^{-4} in 15 ns iSWAP gates across varying qubit anharmonicities.

Reference graph

Works this paper leans on

122 extracted references · 51 canonical work pages · cited by 2 Pith papers · 6 internal anchors

  1. [1]

    M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information: 10th Anniversary Edition (Cambridge University Press, 2010)

  2. [2]

    S. J. Devitt, W. J. Munro, and K. Nemoto, Quantum error correction for beginners, Reports on Progress in Physics 76, 076001 (2013)

  3. [3]

    P. W. Shor, Scheme for reducing decoherence in quantum computer memory, Phys. Rev. A52, R2493 (1995)

  4. [4]

    Steane, Multiple-particle interference and quan- tum error correction, Proceedings of the Royal So- ciety of London

    A. Steane, Multiple-particle interference and quan- tum error correction, Proceedings of the Royal So- ciety of London. Series A: Mathematical, Physi- cal and Engineering Sciences 452, 2551 (1996), https://royalsocietypublishing.org/doi/pdf/10.1098/rspa.1996.0136

  5. [5]

    Knill, R

    E. Knill, R. Laflamme, R. Martinez, and C. Negrevergne, Benchmarking quantum computers: The five-qubit error cor- recting code, Phys. Rev. Lett. 86, 5811 (2001)

  6. [6]

    Gottesman, Stabilizer codes and quantum error correction (1997), arXiv:quant-ph/9705052 [quant-ph]

    D. Gottesman, Stabilizer codes and quantum error correction (1997), arXiv:quant-ph/9705052 [quant-ph]

  7. [7]

    Kitaev, Fault-tolerant quantum computation by anyons, An- nals of Physics 303, 2 (2003)

    A. Kitaev, Fault-tolerant quantum computation by anyons, An- nals of Physics 303, 2 (2003)

  8. [8]

    A. G. Fowler, M. Mariantoni, J. M. Martinis, and A. N. Cle- land, Surface codes: Towards practical large-scale quantum computation, Phys. Rev. A 86, 032324 (2012)

  9. [9]

    Acharya, D

    R. Acharya, D. A. Abanin, L. Aghababaie-Beni, et al., Quan- tum error correction below the surface code threshold, Nature 10.1038/s41586-024-08449-y (2024)

  10. [10]

    A. J. Brady, A. Eickbusch, S. Singh, et al. , Ad- vances in bosonic quantum error correction with gottes- man–kitaev–preskill codes: Theory, engineering and applica- tions, Progress in Quantum Electronics 93, 100496 (2024)

  11. [11]

    N. P. Breuckmann and J. N. Eberhardt, Quantum low-density parity-check codes, PRX Quantum 2, 040101 (2021)

  12. [12]

    Preti, T

    F. Preti, T. Calarco, and F. Motzoi, Continuous quantum gate sets and pulse-class meta-optimization, PRX Quantum 3, 040311 (2022)

  13. [13]

    Guatto, G

    M. Guatto, G. A. Susto, and F. Ticozzi, Improving robust- ness of quantum feedback control with reinforcement learn- ing, Phys. Rev. A 110, 012605 (2024)

  14. [14]

    Optimizing optical potentials with physics-inspired learning algorithms

    M. Calzavara, Y . Kuriatnikov, A. Deutschmann-Olek, et al. , Optimizing optical potentials with physics-inspired learning algorithms (2022), arXiv:2210.07776 [cond-mat, physics:physics]

  15. [15]

    Eickbusch, V

    A. Eickbusch, V . Sivak, A. Z. Ding,et al., Fast universal con- trol of an oscillator with weak dispersive coupling to a qubit, Nature Physics , 1 (2022)

  16. [16]

    Dalgaard, F

    M. Dalgaard, F. Motzoi, J. J. Sørensen, and J. Sherson, Global optimization of quantum dynamics with AlphaZero deep ex- ploration, Npj Quantum Inf. 6 (2020)

  17. [17]

    Porotti, A

    R. Porotti, A. Essig, B. Huard, and F. Marquardt, Deep Rein- forcement Learning for Quantum State Preparation with Weak Nonlinear Measurements, Quantum 6, 747 (2022)

  18. [18]

    Nam Nguyen, F

    H. Nam Nguyen, F. Motzoi, M. Metcalf, et al., Reinforcement learning pulses for transmon qubit entangling gates, Mach. Learn. Sci. Technol. 5, 025066 (2024)

  19. [19]

    L. Moro, M. Paris, M. Restelli, and E. Prati, Quantum compil- ing by deep reinforcement learning, Communications Physics 4 (2021)

  20. [20]

    Preti, M

    F. Preti, M. Schilling, S. Jerbi, et al. , Hybrid discrete- continuous compilation of trapped-ion quantum circuits with deep reinforcement learning, Quantum 8, 1343 (2024)

  21. [21]

    F ¨urrutter, G

    F. F ¨urrutter, G. Mu˜noz-Gil, and H. J. Briegel, Quantum circuit synthesis with diffusion models, Nature Machine Intelligence 6, 515 (2024)

  22. [22]

    Z. T. Wang, Q. Chen, Y . Du, et al. , Quantum compiling with reinforcement learning on a superconducting processor (2024), arXiv:2406.12195 [quant-ph]

  23. [23]

    Zhang, P.-L

    Y .-H. Zhang, P.-L. Zheng, Y . Zhang, and D.-L. Deng, Topo- logical quantum compiling with reinforcement learning, Phys. Rev. Lett. 125, 170501 (2020)

  24. [24]

    Preti, T

    F. Preti, T. Calarco, J. M. Torres, and J. Z. Bern ´ad, Optimal two-qubit gates in recurrence protocols of entanglement pu- rification, Phys. Rev. A 106, 022422 (2022)

  25. [25]

    Preti and J

    F. Preti and J. Z. Bern ´ad, Statistical evaluation and optimiza- tion of entanglement purification protocols, Physical Review A 110, 10.1103/physreva.110.022619 (2024)

  26. [26]

    R. Zen, J. Olle, L. Colmenarez, et al., Quantum circuit discov- ery for fault-tolerant logical state preparation with reinforce- ment learning (2024), arXiv:2402.17761 [quant-ph]

  27. [27]

    Puviani, S

    M. Puviani, S. Borah, R. Zen, et al., Boosting the gottesman- kitaev-preskill quantum error correction with non-markovian feedback (2023), arXiv:2312.07391 [quant-ph]

  28. [28]

    C. Cao, C. Zhang, Z. Wu, et al., Quantum variational learning for quantum error-correcting codes, Quantum 6, 828 (2022)

  29. [29]

    Gicev, L

    S. Gicev, L. C. L. Hollenberg, and M. Usman, A scalable and fast artificial neural network syndrome decoder for surface codes, Quantum 7, 1058 (2023)

  30. [30]

    J. Olle, R. Zen, M. Puviani, and F. Marquardt, Simultane- ous discovery of quantum error correction codes and encoders with a noise-aware reinforcement learning agent (2024), arXiv:2311.04750 [quant-ph]

  31. [31]

    J. Olle, O. M. Yevtushenko, and F. Marquardt, Scaling the automated discovery of quantum circuits via reinforcement learning with gadgets (2025), arXiv:2503.11638 [quant-ph]

  32. [32]

    H. P. Nautrup, N. Delfosse, V . Dunjko, et al. , Optimizing Quantum Error Correction Codes with Reinforcement Learn- ing, Quantum 3, 215 (2019)

  33. [33]

    Meyer, C

    N. Meyer, C. Mutschler, A. Maier, and D. D. Scherer, Learning encodings by maximizing state distinguishability: Variational quantum error correction (2025), arXiv:2506.11552 [quant- ph]

  34. [34]

    E. S. Matekole, E. Ye, R. Iyer, and S. Y .-C. Chen, Decod- ing surface codes with deep reinforcement learning and prob- abilistic policy reuse (2022), arXiv:2212.11890 [quant-ph]

  35. [35]

    Lange, P

    M. Lange, P. Havstr ¨om, B. Srivastava, et al., Data-driven de- 13 coding of quantum error correcting codes using graph neu- ral networks, Physical Review Research7, 10.1103/physrevre- search.7.023181 (2025)

  36. [36]

    Sweke, M

    R. Sweke, M. S. Kesselring, E. P. L. van Nieuwenburg, and J. Eisert, Reinforcement learning decoders for fault-tolerant quantum computation, Mach. Learn. Sci. Technol. 2, 025005 (2021)

  37. [37]

    Gottesman, Fault-tolerant quantum computation with higher-dimensional systems, in Quantum Computing and Quantum Communications, edited by C

    D. Gottesman, Fault-tolerant quantum computation with higher-dimensional systems, in Quantum Computing and Quantum Communications, edited by C. P. Williams (Springer Berlin Heidelberg, Berlin, Heidelberg, 1999) pp. 302–313

  38. [38]

    P. J. Low, B. M. White, A. A. Cox, et al., Practical trapped- ion protocols for universal qudit-based quantum computing, Physical Review Research 2, 033128 (2020)

  39. [39]

    Ringbauer, M

    M. Ringbauer, M. Meth, L. Postler, et al., A universal qudit quantum processor with trapped ions, Nature Physics18, 1053 (2022)

  40. [40]

    P. Hrmo, B. Wilhelm, L. Gerster, et al., Native qudit entangle- ment in a trapped ion quantum processor, Nature Communi- cations 14, 2242 (2023)

  41. [41]

    P. J. Low, B. White, and C. Senko, Control and Readout of a 13-level Trapped Ion Qudit (2023), arXiv:2306.03340

  42. [42]

    Gonz ´alez-Cuadra, T

    D. Gonz ´alez-Cuadra, T. V . Zache, J. Carrasco,et al., Hardware Efficient Quantum Simulation of Non-Abelian Gauge Theo- ries with Qudits on Rydberg Platforms, Physical Review Let- ters 129, 160501 (2022)

  43. [43]

    Hussain, G

    R. Hussain, G. Allodi, A. Chiesa, et al. , Coherent ma- nipulation of a molecular ln-based nuclear qudit cou- pled to an electron qubit, Journal of the American Chemical Society 140, 9814 (2018), pMID: 30040890, https://doi.org/10.1021/jacs.8b05934

  44. [44]

    M. Kues, C. Reimer, P. Roztocki, et al., On-chip generation of high-dimensional entangled quantum states and their coherent control, Nature 546, 622 (2017)

  45. [45]

    Erhard, M

    M. Erhard, M. Malik, M. Krenn, and A. Zeilinger, Exper- imental Greenberger–Horne–Zeilinger entanglement beyond qubits, Nature Photonics 12, 759 (2018)

  46. [46]

    Luo, H.-S

    Y .-H. Luo, H.-S. Zhong, M. Erhard, et al. , Quantum Tele- portation in High Dimensions, Physical Review Letters 123, 070505 (2019)

  47. [47]

    E. J. Davis, G. Bentsen, L. Homeier, et al., Photon-Mediated Spin-Exchange Dynamics of Spin-1 Atoms, Physical Review Letters 122, 010405 (2019)

  48. [48]

    Y . Chi, J. Huang, Z. Zhang, et al. , A programmable qudit- based quantum processor, Nature Communications 13, 1166 (2022)

  49. [49]

    M. S. Blok, V . V . Ramasesh, T. Schuster,et al., Quantum In- formation Scrambling on a Superconducting Qutrit Processor, Physical Review X 11, 021010 (2021)

  50. [50]

    P. Liu, R. Wang, J.-N. Zhang, et al., Performing SU ( d ) Op- erations and Rudimentary Algorithms in a Superconducting Transmon Qudit for d = 3 and d = 4, Physical Review X 13, 021028 (2023)

  51. [51]

    Champion, Z

    E. Champion, Z. Wang, R. Parker, and M. Blok, Multi- frequency control and measurement of a spin-7/2 system en- coded in a transmon qudit (2024), arXiv:2405.15857 [quant- ph]

  52. [52]

    Morvan, V

    A. Morvan, V . V . Ramasesh, M. S. Blok, et al., Qutrit Ran- domized Benchmarking, Physical Review Letters126, 210504 (2021)

  53. [53]

    M. A. Yurtalan, J. Shi, M. Kononenko, et al., Implementation of a Walsh-Hadamard Gate in a Superconducting Qutrit, Phys- ical Review Letters 125, 180504 (2020)

  54. [54]

    Kononenko, M

    M. Kononenko, M. A. Yurtalan, S. Ren, et al., Characteriza- tion of control in a superconducting qutrit using randomized benchmarking, Physical Review Research 3, L042007 (2021)

  55. [55]

    Yurtalan, J

    M. Yurtalan, J. Shi, G. Flatt, and A. Lupascu, Characteri- zation of Multilevel Dynamics and Decoherence in a High- Anharmonicity Capacitively Shunted Flux Circuit, Physical Review Applied 16, 054051 (2021)

  56. [56]

    K. Luo, W. Huang, Z. Tao, et al., Experimental Realization of Two Qutrits Gate with Tunable Coupling in Superconducting Circuits, Physical Review Letters 130, 030603 (2023)

  57. [57]

    B. Li, F. C ´ardenas-L´opez, A. Lupascu, and F. Motzoi, Uni- versal pulses for superconducting qudit ladder gates, arXiv preprint arXiv:2412.18339 (2024)

  58. [58]

    Optimal Error Correcting Code For Ternary Quantum Systems

    R. Majumdar and S. Sur-Kolay, Optimal error correcting code for ternary quantum systems (2020), arXiv:1906.11137 [quant-ph]

  59. [59]

    Breuer and F

    H.-P. Breuer and F. Petruccione, The Theory of Open Quan- tum Systems (Oxford University Press, Oxford ; New York, 2002)

  60. [60]

    Etxezarreta Martinez, P

    J. Etxezarreta Martinez, P. Fuentes, A. deMarti iOlius, et al., Multiqubit time-varying quantum channels for nisq-era super- conducting quantum processors, Phys. Rev. Res. 5, 033055 (2023)

  61. [61]

    Enhancing Quantum Circuit Noise Robustness from a Geometric Perspective

    J. Zeng, Y .-J. Hai, H. Liang, and X.-H. Deng, Quantum Cir- cuits Noise Tailoring from a Geometric Perspective (2023), arXiv:2305.06795 [quant-ph]

  62. [62]

    Tan, Multi-agent reinforcement learning: Independent ver- sus cooperative agents, in International Conference on Ma- chine Learning (1997)

    M. Tan, Multi-agent reinforcement learning: Independent ver- sus cooperative agents, in International Conference on Ma- chine Learning (1997)

  63. [63]

    Claus and C

    C. Claus and C. Boutilier, The dynamics of reinforcement learning in cooperative multiagent systems (1998)

  64. [64]

    Narvekar, B

    S. Narvekar, B. Peng, M. Leonetti, et al., Curriculum learning for reinforcement learning domains: A framework and survey (2020), arXiv:2003.04960 [cs.LG]

  65. [66]

    R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed. (The MIT Press, 2018)

  66. [67]

    Raffin, A

    A. Raffin, A. Hill, A. Gleave, et al., Stable-baselines3: Reli- able reinforcement learning implementations, Journal of Ma- chine Learning Research 22, 1 (2021)

  67. [68]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, et al., Proximal policy optimization algorithms (2017), arXiv:1707.06347 [cs.LG]

  68. [69]

    F ¨osel, S

    T. F ¨osel, S. Krastanov, F. Marquardt, and L. Jiang, Efficient cavity control with SNAP gates (2020), arXiv:2004.14256 [quant-ph]

  69. [70]

    F ¨osel, P

    T. F ¨osel, P. Tighineanu, T. Weiss, and F. Marquardt, Reinforce- ment learning with neural networks for quantum feedback, Physical Review X 8, 10.1103/physrevx.8.031084 (2018)

  70. [71]

    Tilma and E

    T. Tilma and E. C. G. Sudarshan, Generalized euler angle parametrization for su(n), Journal of Physics A: Mathemati- cal and General 35, 10467 (2002)

  71. [72]

    J. A. Nelder and R. Mead, A simplex method for function minimization, The Computer Journal 7, 308 (1965), https://academic.oup.com/comjnl/article- pdf/7/4/308/1013182/7-4-308.pdf

  72. [73]

    Caneva, T

    T. Caneva, T. Calarco, and S. Montangero, Chopped random- basis quantum optimization, Physical Review A 84, 022326 (2011)

  73. [74]

    V . V . Sivak, A. Eickbusch, H. Liu,et al., Model-Free Quantum Control with Reinforcement Learning, Physical Review X 12, 011059 (2022), arXiv:2104.14539 [quant-ph]

  74. [75]

    Combes, C

    J. Combes, C. Granade, C. Ferrie, and S. T. Flammia, Logical randomized benchmarking, arXiv preprint arXiv:1702.03688 14 (2017)

  75. [76]

    L. Besson, SMPyBandits: an Open-Source Research Framework for Single and Multi-Players Multi-Arms Bandits (MAB) Algorithms in Python, Online at: github.com/SMPyBandits/SMPyBandits (2018), code at https://github.com/SMPyBandits/SMPyBandits/, documentation at https://smpybandits.github.io/

  76. [77]

    Lavrijsen, A

    W. Lavrijsen, A. Tudor, J. M ¨uller, et al., Classical optimizers for noisy intermediate-scale quantum devices, in 2020 IEEE International Conference on Quantum Computing and Engi- neering (QCE) (2020) pp. 267–277

  77. [78]

    J. Koch, T. M. Yu, J. Gambetta,et al., Charge-insensitive qubit design derived from the cooper pair box, Phys. Rev. A 76, 042319 (2007)

  78. [79]

    Blais, A

    A. Blais, A. L. Grimsmo, S. M. Girvin, and A. Wallraff, Cir- cuit quantum electrodynamics, Rev. Mod. Phys. 93, 025005 (2021)

  79. [80]

    C. D. Bruzewicz, J. Chiaverini, R. McConnell, and J. M. Sage, Trapped-ion quantum computing: Progress and challenges, Applied Physics Reviews 6, 021314 (2019), https://pubs.aip.org/aip/apr/article- pdf/doi/10.1063/1.5088164/19742554/021314 1 online.pdf

  80. [81]

    Foss-Feig, G

    M. Foss-Feig, G. Pagano, A. C. Potter, and N. Y . Yao, Progress in trapped-ion quantum simulation (2024), arXiv:2409.02990 [quant-ph]

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.