Pith. sign in

REVIEW 3 major objections 5 minor 32 references

A distance-4 bivariate bicycle code with a neural decoder keeps a 4-qubit QCNN trainable at 0.1% gate noise where the unprotected circuit fails.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 03:01 UTC pith:KILTVSBQ

load-bearing objection Useful engineering sketch of BB-encoded QCNNs, but the validation claim overreaches because the full protected circuit and a working non-Clifford decoder are never shown. the 3 major comments →

arxiv 2607.05724 v1 pith:KILTVSBQ submitted 2026-07-07 cs.LG quant-ph

Low-Overhead Error-Corrected QCNNs Using Bivariate Bicycle Codes

classification cs.LG quant-ph
keywords quantum convolutional neural networksbivariate bicycle codesquantum error correctionqLDPC codesNISQtransversal encodingneural decoderSPT phase recognition
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Noise on present-day quantum hardware ruins the loss landscape of quantum convolutional neural networks, so that even a small 4-qubit model fails to learn. Surface-code protection would fix the errors but demands far more qubits than near-term machines can spare. This paper shows that a compact [[18,4,4]] bivariate bicycle code, together with a feed-forward neural network that corrects errors between layers, encodes the same QCNN into only 29 physical qubits. Under a realistic Pauli noise model the protected circuit still converges at a physical error rate of 0.001, while the unprotected baseline does not. The result is offered as concrete evidence that constant-overhead quantum error correction can make QCNNs practical before large surface-code machines arrive.

Core claim

A distance-4 bivariate bicycle code with an interleaved feed-forward neural decoder yields a constant-overhead, 29-physical-qubit encoding of a 4-qubit QCNN that retains trainability at physical error rate p=0.001, where the unprotected circuit fails to converge (final loss 0.348 versus approximately 0.65).

What carries the argument

The [[18,4,4]] bivariate bicycle code whose transversal gate set, combined with a syndrome-based feed-forward neural network that predicts Pauli corrections after every convolutional and pooling layer, keeps the logical state inside the codespace without magic-state factories.

Load-bearing premise

The neural decoder, trained only on single-layer Pauli errors, will still correct the non-Clifford variational rotations of the QCNN without systematic over-correction that collapses every prediction to 1.

What would settle it

Run the full 29-qubit protected circuit (or a higher-fidelity classical simulation of it) at p=0.001 with the decoder co-trained on the actual non-Clifford layers; if the training loss still plateaus near 0.35 or the phase-recognition decision boundary collapses, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Near-term processors with a few dozen qubits can host small protected QCNNs without waiting for surface-code factories.
  • The same BB-code plus neural-decoder pattern can be scaled to deeper QCNNs once larger qLDPC codes become available.
  • Interleaved soft decoding removes the microsecond classical-post-processing bottleneck that otherwise idles superconducting qubits between layers.
  • Constant encoding rate removes the quadratic qubit wall that currently separates NISQ devices from fault-tolerant QML.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Co-training the decoder jointly with the QCNN weights, or initializing from a pre-trained model, is likely required before the method works on hardware, because the paper already sees over-correction on non-Clifford gates.
  • The same low-overhead idea could be tested on other shallow variational models (QAOA, VQE) whose depth is still within the reach of distance-4 qLDPC codes.
  • If the bold-driver SPSA schedule continues to collapse at p=0.003 even with perfect correction, gradient estimation itself, not just gate noise, will need redesign for deeper protected QCNNs.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a low-overhead quantum error-correction scheme for 4-qubit QCNNs based on a distance-4 [[18,4,4]] bivariate bicycle (BB) code, combined with an interleaved feed-forward neural network (FFNN) soft decoder. It argues that surface-code overhead is prohibitive for near-term QML, while BB codes offer constant encoding rate and linear distance. Under a circuit-level Pauli noise model, an unprotected 4-qubit QCNN fails to converge at physical error rates p ≥ 0.001, whereas a transversally encoded 11-qubit ansatz retains trainability (final loss ≈ 0.348 at p = 0.001). The authors further describe encoding the QCNN into a 29-physical-qubit layout (11 data + 18 ancilla) with transversal gates and an SPSA-trained FFNN that maps syndromes to Pauli corrections between layers, and they claim this combination validates a step toward practical QCNNs.

Significance. If the full BB-encoded QCNN with a working non-Clifford-aware decoder were shown to restore trainability at NISQ-relevant error rates with only ~29 physical qubits, the result would be a meaningful step for near-term fault-tolerant QML, given the well-known surface-code overhead for non-Clifford gates. The manuscript usefully imports recent BB-code constructions into the QCNN setting, provides a concrete resource comparison against toric/surface codes, and documents that unprotected QCNNs degrade under realistic gate noise while a transversal skeleton retains some learning signal. Those elements are of genuine interest. However, the load-bearing validation of the proposed technique (BB transversal encoding + interleaved FFNN) is not yet demonstrated in the reported experiments, so the significance of the present claims is more architectural and diagnostic than empirical.

major comments (3)
  1. [Abstract, §IV-B, Fig. 3] Abstract, §I-B, and §V claim that the distance-4 BB QEC technique for QCNNs is validated and represents a step toward practical QCNNs. Yet §IV-B explicitly states that the full 29-qubit protected circuit (11 data + 18 ancilla with interleaved FFNN) is never simulated under noise and is left for future research. All training curves that show resilience at p = 0.001 (final loss 0.348) are for the 11-qubit transversal ansatz alone, without the FFNN decoder. The central validation claim therefore rests on a partial architecture that is not the proposed technique.
  2. [§III-B, §IV-B, Fig. 6] The proposed technique is precisely the combination of BB transversal encoding and an interleaved FFNN soft decoder (§III-B–C). Fig. 6 and the accompanying text in §IV-B show that the FFNN, trained only on single-layer Pauli/Clifford error patterns (§III-B, F_train), systematically over-corrects the non-Clifford variational rotations of the QCNN, collapsing predictions toward entropy-1. Because non-Clifford layers are essential to the QCNN ansatz and cannot be implemented transversally in the BB block, this failure mode is load-bearing: the decoder as trained does not support the claimed error-corrected QCNN.
  3. [§IV-C, §V] The resource and trainability conclusions in §IV-C and §V compare the unprotected 4-qubit QCNN against the transversal 11-qubit skeleton and then extrapolate to the full 29-qubit BB+FFNN system. Without noisy end-to-end simulation (or at least a controlled ablation that isolates the FFNN under the same noise model used for Fig. 3), it is not established that the constant-overhead encoding plus decoder restores convergence where the unprotected circuit fails. The manuscript should either (i) provide such results or (ii) substantially narrow the abstract/conclusion claims to what is actually shown: noise resilience of a transversal QCNN skeleton and a separate, Clifford-trained decoder that currently over-corrects non-Clifford layers.
minor comments (5)
  1. [Fig. 3, §IV-B] Fig. 3 caption and body refer to both '4-qubit QCNN and 11-qubit transversal QCNN' and later to panels (a)–(d) at different p; the figure itself is described as a single multi-panel plot but the text sometimes treats (c) and (d) as both p = 0.003 for different architectures. Clarify panel labels and which curve is which.
  2. [Abstract, §I-B, §IV-B] Typos and consistency: 'QCNNS' in the abstract; 'the the application' in §I-B; 'reduces, the effective expressibility' in §IV-B; 'learning loss was lowered' etc. A careful proofread is needed.
  3. [§III-A, §III-B] Eq. (17) defines Loss(φ) as a 0–1 misclassification rate over N = 16 patterns, while the QCNN itself is trained with MSE (Eq. 14). The relationship between the FFNN loss and the QCNN MSE under joint or sequential training is not stated; a short clarification would help.
  4. [§III-C] The qubit-count reduction from 36 to 29 via shared-index mapping (§III-C) is important; a small table or explicit list of the 11 data indices and which logical qubits share them would make the construction easier to reproduce.
  5. [§I-A] Related-work discussion of [16] is useful but could more clearly separate classical simulability of QCNNs from hardware trainability under noise, which is the paper's actual focus.

Circularity Check

0 steps flagged

No circularity: empirical simulation results and literature-sourced BB/QCNN constructions do not reduce by construction to their inputs.

full rationale

The paper's central claims rest on stochastic circuit simulations under a depolarizing noise model (p in {0, 0.0001, 0.001, 0.003}) of an unprotected 4-qubit QCNN versus a transversally encoded 11-qubit ansatz derived from the [[18,4,4]] BB code of Bravyi et al. Training losses, best-so-far curves, and prediction distributions (Figs. 3–6) are obtained by independent finite-difference/SPSA optimization runs; they are not algebraic rearrangements of fitted parameters. The BB check matrices, transversal gate expansion (Eqs. 21–22), and syndrome-extraction architecture are taken verbatim from the cited literature ([12], [15]) rather than re-derived or fitted to the QCNN loss surface. The FFNN decoder is trained on an exhaustive set of single-layer Pauli patterns (Eq. 17) and then applied; its acknowledged over-correction of non-Clifford layers (Fig. 6) is an empirical failure mode, not a definitional identity. No uniqueness theorem, self-citation chain, or ansatz smuggled via prior author work forces the reported resilience at p=0.001. The derivation chain is therefore self-contained against external benchmarks and exhibits zero circular steps.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 1 invented entities

The paper inherits the BB-code construction, the QCNN ansatz and the circuit-level depolarizing noise model from prior literature; the only free choices are the specific [[18,4,4]] parameters, the SPSA hyper-parameters and the decision to train the decoder solely on Pauli errors. No new physical entities are postulated.

free parameters (3)
  • BB polynomial coefficients [a1,a2,a3]=[1,0,1], [b1,b2,b3]=[2,0,2]
    Hand-chosen to realize the [[18,4,4]] code; different generators would change distance and connectivity.
  • SPSA perturbation magnitude c=0.2 rad and bold-driver factors 1.05/0.5
    Optimizer hyper-parameters that control the decoder training trajectory; not derived from first principles.
  • Physical error rates p ∈ {0.0001,0.001,0.003}
    Sweep points chosen by the authors to straddle the reported BB threshold; the qualitative claim depends on the intermediate value.
axioms (3)
  • domain assumption Bivariate bicycle codes admit a depth-7 syndrome extraction circuit with constant encoding rate and linear distance (Bravyi et al. 2024).
    Used throughout §II-D and §III-B as the foundation for the 29-qubit encoding.
  • domain assumption A QCNN realizes an inverse MERA and can classify 1-D SPT phases with O(log N) parameters (Cong et al. 2019).
    Defines the unprotected baseline circuit whose noise sensitivity is measured.
  • domain assumption Circuit-level depolarizing noise with independent single- and two-qubit Pauli channels is a realistic model of NISQ hardware.
    Underpins all numerical claims in §IV.
invented entities (1)
  • CQCNN with interleaved FFNN soft decoder no independent evidence
    purpose: To apply non-Clifford variational layers inside a BB code block without magic-state factories by correcting syndromes between layers.
    The architecture is introduced in §III-C; independent evidence is limited to the partial transversal simulations.

pith-pipeline@v1.1.0-grok45 · 19759 in / 2793 out tokens · 33387 ms · 2026-07-11T03:01:17.708286+00:00 · methodology

0 comments
read the original abstract

Quantum convolutional neural networks (QCNNs) combine the power of quantum computing and classical CNN for computational speedup in classification tasks. However, noise levels on state-of-the-art quantum devices remain too high for practical QCNN execution. In addition, despite the reliable surface code providing a method for error rates below a threshold value, they have a prohibitively large qubit cost. Recently introduced bivariate bicycle (BB) codes are of particular interest for their high error threshold, constant encoding rate, and linear code distance. Through simulation with realistic hardware noise sources, we demonstrate that a 4-qubit unprotected QCNN fails to converge and exhibits a worse learning rate compared to numerical simulations. Addressing both limitations, we propose a distance-4 BB quantum error-correction (QEC) technique for QCNNs. In doing so, we validate that our low-overhead QEC technique for QCNNS represents a step toward practical QCNNs.

Figures

Figures reproduced from arXiv: 2607.05724 by Alejandro Rosales, Animesh Yadav.

Figure 1
Figure 1. Figure 1: (a) Points show a 64 × 64 test set of ground states over h1 and h2 for a Hamiltonian with J = 1. Phase￾boundary points (blue and red diamonds) come from infinite￾size density-matrix renormalization group (DMRG). Colors indicate the circuit’s expectation values found via evolutionary search for N = 4 spins using initial, untrained weights. (An artifact can be seen, possibly because of the low qubit spin sta… view at source ↗
Figure 2
Figure 2. Figure 2: QEC-QCNN circuit. 3) Decoding Algorithm: Consider a [[n, k, d]] BB code and an SM circuit U constructed with Nc syndrome cycles defined in the SM section. Error correction is performed using a classical algorithm that takes a measured error syndrome as input and returns an estimate of the Pauli error that occurred during computation on the data qubits, accounting for all faults provided by the SM circuit. … view at source ↗
Figure 3
Figure 3. Figure 3: Statevector training loss vs iteration for [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Training loss versus iteration (a) the 4-qubit QCNN [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: 4-qubit QCNN prediction of an antiferromagnetic and paramagnetism phase with the trained model. The QCNN (blue) and the QEC-QCNN (orange) are shown. The QEC-QCNN 6 shows that the FFNN trained on Clifford gates over corrects for errors, and Non-Clifford Gates are not apart of the training data, and appear to the FFNN soft decoder as errors. This causes bit and phase flips to the entangled data, causing pred… view at source ↗
Figure 5
Figure 5. Figure 5: The violin plots of QCNN predictions (blue) vs [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

32 extracted references · 32 canonical work pages · 3 internal anchors

  1. [1]

    The vanishing gradient problem during learning recurrent neural nets and problem solutions,

    S. Hochreiter, “The vanishing gradient problem during learning recurrent neural nets and problem solutions,”International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, vol. 6, no. 2, pp. 107–116, 1998

  2. [2]

    On the difficulty of training recurrent neural networks,

    R. Pascanu, T. Mikolov, and Y . Bengio, “On the difficulty of training recurrent neural networks,” inProceedings of the 30th International Conference on Machine Learning, 2013

  3. [3]

    Information-theoretic bounds on quantum advantage in machine learning,

    H.-Y . Huang, R. Kueng, and J. Preskill, “Information-theoretic bounds on quantum advantage in machine learning,”Physical Review Letters, vol. 126, no. 19, p. 190505, 2021

  4. [4]

    Deep learning,

    Y . LeCun, Y . Bengio, and G. Hinton, “Deep learning,”Nature, vol. 521, no. 7553, pp. 436–444, May 2015

  5. [5]

    M. A. Nielsen and I. L. Chuang,Quantum Computation and Quantum Information: 10th Anniversary Edition. Cambridge, U.K.: Cambridge University Press, 2011

  6. [6]

    Hirvensalo,Quantum Computing, 2nd ed

    M. Hirvensalo,Quantum Computing, 2nd ed. Berlin, Germany: Springer, 2001

  7. [7]

    Quantum machine learning,

    J. Biamonteet al., “Quantum machine learning,”Nature, vol. 549, pp. 195–202, 2017

  8. [8]

    Supervised learning with quantum-enhanced feature spaces,

    V . Havl´ıˇceket al., “Supervised learning with quantum-enhanced feature spaces,”Nature, vol. 567, pp. 209–212, 2019

  9. [9]

    Barren plateaus in quantum neural network training landscapes,

    J. R. McCleanet al., “Barren plateaus in quantum neural network training landscapes,”Nature Communications, vol. 9, p. 4812, 2018

  10. [10]

    Building logical qubits in a superconducting quantum computing system,

    J. M. Gambetta, J. M. Chow, and M. Steffen, “Building logical qubits in a superconducting quantum computing system,”npj Quantum Infor- mation, vol. 3, pp. 1–7, 2017

  11. [11]

    Fault-tolerant quantum computation by anyons,

    A. Y . Kitaev, “Fault-tolerant quantum computation by anyons,”Annals of Physics, vol. 303, no. 1, pp. 2–30, 2003

  12. [12]

    High-threshold and low- overhead fault-tolerant quantum memory,

    S. Bravyi, A. W. Cross, J. M. Gambettaet al., “High-threshold and low- overhead fault-tolerant quantum memory,”Nature, vol. 627, pp. 778– 782, 2024

  13. [13]

    Fault-tolerant quantum computation with constant over- head,

    D. Gottesman, “Fault-tolerant quantum computation with constant over- head,”Quantum Information and Computation, vol. 14, no. 15–16, 2014

  14. [14]

    Quantum convolutional neural networks,

    I. Cong, S. Choi, and M. D. Lukin, “Quantum convolutional neural networks,”Nature Physics, vol. 15, no. 12, pp. 1273–1278, Dec. 2019

  15. [15]

    Demonstration of low-overhead quantum error correction codes,

    K. Wang, Z. Lu, C. Zhanget al., “Demonstration of low-overhead quantum error correction codes,”Nature Physics, vol. 22, pp. 308–314, 2026

  16. [16]

    Quantum convolutional neural networks are effec- tively classically simulable,

    P. Bermejoet al., “Quantum convolutional neural networks are effec- tively classically simulable,”PRX Quantum, vol. 7, no. 2, p. 020304, Apr. 2026

  17. [17]

    Convolutional networks for images, speech, and time-series,

    Y . LeCun and Y . Bengio, “Convolutional networks for images, speech, and time-series,” inThe Handbook of Brain Theory and Neural Net- works, 1995

  18. [18]

    Multiscale entanglement renormalization ansatz: Causal- ity and error correction,

    D. Pomarico, “Multiscale entanglement renormalization ansatz: Causal- ity and error correction,”Dynamics, vol. 3, no. 3, pp. 622–635, 2023

  19. [19]

    Infinite size density matrix renormalization group, revisited

    I. P. McCulloch, “Infinite size density matrix renormalization group, revisited,”arXiv preprint arXiv:0804.2509, 2008

  20. [20]

    Multivariate stochastic approximation using a simultaneous perturbation gradient approximation,

    J. C. Spall, “Multivariate stochastic approximation using a simultaneous perturbation gradient approximation,”IEEE Transactions on Automatic Control, vol. 37, no. 3, pp. 332–341, 1992

  21. [21]

    First- and second-order methods for learning: Between steepest descent and Newton’s method,

    R. Battiti, “First- and second-order methods for learning: Between steepest descent and Newton’s method,”Neural Computation, vol. 4, no. 2, pp. 141–166, Mar. 1992

  22. [22]

    Quantum error correction below the surface code threshold,

    Google Quantum AI and Collaborators, “Quantum error correction below the surface code threshold,”Nature, vol. 638, pp. 920–926, 2025

  23. [23]

    Demonstrating real-time and low-latency quantum error correction with superconducting qubits

    L. Cauneet al., “Demonstrating real-time and low-latency quan- tum error correction with superconducting qubits,”arXiv preprint arXiv:2410.05202, 2024

  24. [24]

    Low Latency GNN Accelerator for Quantum Error Correction

    J. Viszlaiet al., “Low latency GNN accelerator for quantum error correction,”arXiv preprint arXiv:2603.22149, 2025

  25. [25]

    Sachdev,Quantum Phase Transitions, 2nd ed

    S. Sachdev,Quantum Phase Transitions, 2nd ed. Cambridge, U.K.: Cambridge University Press, 2011

  26. [26]

    Subsystem codes with high thresh- olds by gauge fixing and reduced qubit overhead,

    O. Higgott and N. P. Breuckmann, “Subsystem codes with high thresh- olds by gauge fixing and reduced qubit overhead,”Physical Review X, vol. 11, no. 3, 2021

  27. [27]

    Magic state distillation: Not as costly as you think,

    D. Litinski, “Magic state distillation: Not as costly as you think,” Quantum, vol. 3, p. 205, Dec. 2019

  28. [28]

    Restrictions on transversal encoded quantum gate sets,

    B. Eastin and E. Knill, “Restrictions on transversal encoded quantum gate sets,”Physical Review Letters, vol. 102, no. 11, p. 110502, 2009

  29. [29]

    IBM quantum delivers on 2022 100×100 performance challenge,

    IBM Quantum, “IBM quantum delivers on 2022 100×100 performance challenge,” IBM Quantum Blog, Nov. 2024, accessed: 2024. [Online]. Available: https://www.ibm.com/quantum/blog/qdc-2024

  30. [30]

    How to factor 2048 bit RSA integers in 8 hours using 20 million noisy qubits,

    C. Gidney and M. Eker ˚a, “How to factor 2048 bit RSA integers in 8 hours using 20 million noisy qubits,”Quantum, vol. 5, p. 433, 2021

  31. [31]

    Even more efficient quantum computations of chemistry through tensor hypercontraction,

    J. Leeet al., “Even more efficient quantum computations of chemistry through tensor hypercontraction,”PRX Quantum, vol. 2, p. 030305, 2021

  32. [32]

    Focus beyond quadratic speedups for error-corrected quan- tum advantage,

    R. Babbush, J. R. McClean, M. Newman, C. Gidney, S. Boixo, and H. Neven, “Focus beyond quadratic speedups for error-corrected quan- tum advantage,”PRX Quantum, vol. 2, p. 010103, 2021