Pith. sign in

REVIEW 2 cited by

Decoding surface codes with deep reinforcement learning and probabilistic policy reuse

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.11890 v1 pith:UQCZPLPZ submitted 2022-12-22 quant-ph cs.AIcs.ETcs.LGcs.NE

Decoding surface codes with deep reinforcement learning and probabilistic policy reuse

classification quant-ph cs.AIcs.ETcs.LGcs.NE
keywords quantumdecodinglearningcodesreinforcementsurfacecertainchallenges
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Quantum computing (QC) promises significant advantages on certain hard computational tasks over classical computers. However, current quantum hardware, also known as noisy intermediate-scale quantum computers (NISQ), are still unable to carry out computations faithfully mainly because of the lack of quantum error correction (QEC) capability. A significant amount of theoretical studies have provided various types of QEC codes; one of the notable topological codes is the surface code, and its features, such as the requirement of only nearest-neighboring two-qubit control gates and a large error threshold, make it a leading candidate for scalable quantum computation. Recent developments of machine learning (ML)-based techniques especially the reinforcement learning (RL) methods have been applied to the decoding problem and have already made certain progress. Nevertheless, the device noise pattern may change over time, making trained decoder models ineffective. In this paper, we propose a continual reinforcement learning method to address these decoding challenges. Specifically, we implement double deep Q-learning with probabilistic policy reuse (DDQN-PPR) model to learn surface code decoding strategies for quantum environments with varying noise patterns. Through numerical simulations, we show that the proposed DDQN-PPR model can significantly reduce the computational complexity. Moreover, increasing the number of trained policies can further improve the agent's performance. Our results open a way to build more capable RL agents which can leverage previously gained knowledge to tackle QEC challenges.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Real-time adaptive quantum error correction by model-free multi-agent learning

    quant-ph 2025-09 conditional novelty 7.0

    Adaptive quantum error correction: multi-agent RL discovers QEC circuits offline; a bandit-controlled variational layer retrains online, cutting logical infidelity about 18x (qubit) and 3x (qutrit) under drifting bit/...

  2. Real-time decoding of quantum error correction codes using high-performance computing

    quant-ph 2026-08 conditional novelty 5.0

    An HPC-to-quantum-control interconnect achieves 2.944 µs round-trip latency and CPU-based real-time surface-code decoding up to distance 19 at about 1 µs per round.