REVIEW 3 major objections 5 minor 32 references
A distance-4 bivariate bicycle code with a neural decoder keeps a 4-qubit QCNN trainable at 0.1% gate noise where the unprotected circuit fails.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 03:01 UTC pith:KILTVSBQ
load-bearing objection Useful engineering sketch of BB-encoded QCNNs, but the validation claim overreaches because the full protected circuit and a working non-Clifford decoder are never shown. the 3 major comments →
Low-Overhead Error-Corrected QCNNs Using Bivariate Bicycle Codes
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A distance-4 bivariate bicycle code with an interleaved feed-forward neural decoder yields a constant-overhead, 29-physical-qubit encoding of a 4-qubit QCNN that retains trainability at physical error rate p=0.001, where the unprotected circuit fails to converge (final loss 0.348 versus approximately 0.65).
What carries the argument
The [[18,4,4]] bivariate bicycle code whose transversal gate set, combined with a syndrome-based feed-forward neural network that predicts Pauli corrections after every convolutional and pooling layer, keeps the logical state inside the codespace without magic-state factories.
Load-bearing premise
The neural decoder, trained only on single-layer Pauli errors, will still correct the non-Clifford variational rotations of the QCNN without systematic over-correction that collapses every prediction to 1.
What would settle it
Run the full 29-qubit protected circuit (or a higher-fidelity classical simulation of it) at p=0.001 with the decoder co-trained on the actual non-Clifford layers; if the training loss still plateaus near 0.35 or the phase-recognition decision boundary collapses, the central claim fails.
If this is right
- Near-term processors with a few dozen qubits can host small protected QCNNs without waiting for surface-code factories.
- The same BB-code plus neural-decoder pattern can be scaled to deeper QCNNs once larger qLDPC codes become available.
- Interleaved soft decoding removes the microsecond classical-post-processing bottleneck that otherwise idles superconducting qubits between layers.
- Constant encoding rate removes the quadratic qubit wall that currently separates NISQ devices from fault-tolerant QML.
Where Pith is reading between the lines
- Co-training the decoder jointly with the QCNN weights, or initializing from a pre-trained model, is likely required before the method works on hardware, because the paper already sees over-correction on non-Clifford gates.
- The same low-overhead idea could be tested on other shallow variational models (QAOA, VQE) whose depth is still within the reach of distance-4 qLDPC codes.
- If the bold-driver SPSA schedule continues to collapse at p=0.003 even with perfect correction, gradient estimation itself, not just gate noise, will need redesign for deeper protected QCNNs.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a low-overhead quantum error-correction scheme for 4-qubit QCNNs based on a distance-4 [[18,4,4]] bivariate bicycle (BB) code, combined with an interleaved feed-forward neural network (FFNN) soft decoder. It argues that surface-code overhead is prohibitive for near-term QML, while BB codes offer constant encoding rate and linear distance. Under a circuit-level Pauli noise model, an unprotected 4-qubit QCNN fails to converge at physical error rates p ≥ 0.001, whereas a transversally encoded 11-qubit ansatz retains trainability (final loss ≈ 0.348 at p = 0.001). The authors further describe encoding the QCNN into a 29-physical-qubit layout (11 data + 18 ancilla) with transversal gates and an SPSA-trained FFNN that maps syndromes to Pauli corrections between layers, and they claim this combination validates a step toward practical QCNNs.
Significance. If the full BB-encoded QCNN with a working non-Clifford-aware decoder were shown to restore trainability at NISQ-relevant error rates with only ~29 physical qubits, the result would be a meaningful step for near-term fault-tolerant QML, given the well-known surface-code overhead for non-Clifford gates. The manuscript usefully imports recent BB-code constructions into the QCNN setting, provides a concrete resource comparison against toric/surface codes, and documents that unprotected QCNNs degrade under realistic gate noise while a transversal skeleton retains some learning signal. Those elements are of genuine interest. However, the load-bearing validation of the proposed technique (BB transversal encoding + interleaved FFNN) is not yet demonstrated in the reported experiments, so the significance of the present claims is more architectural and diagnostic than empirical.
major comments (3)
- [Abstract, §IV-B, Fig. 3] Abstract, §I-B, and §V claim that the distance-4 BB QEC technique for QCNNs is validated and represents a step toward practical QCNNs. Yet §IV-B explicitly states that the full 29-qubit protected circuit (11 data + 18 ancilla with interleaved FFNN) is never simulated under noise and is left for future research. All training curves that show resilience at p = 0.001 (final loss 0.348) are for the 11-qubit transversal ansatz alone, without the FFNN decoder. The central validation claim therefore rests on a partial architecture that is not the proposed technique.
- [§III-B, §IV-B, Fig. 6] The proposed technique is precisely the combination of BB transversal encoding and an interleaved FFNN soft decoder (§III-B–C). Fig. 6 and the accompanying text in §IV-B show that the FFNN, trained only on single-layer Pauli/Clifford error patterns (§III-B, F_train), systematically over-corrects the non-Clifford variational rotations of the QCNN, collapsing predictions toward entropy-1. Because non-Clifford layers are essential to the QCNN ansatz and cannot be implemented transversally in the BB block, this failure mode is load-bearing: the decoder as trained does not support the claimed error-corrected QCNN.
- [§IV-C, §V] The resource and trainability conclusions in §IV-C and §V compare the unprotected 4-qubit QCNN against the transversal 11-qubit skeleton and then extrapolate to the full 29-qubit BB+FFNN system. Without noisy end-to-end simulation (or at least a controlled ablation that isolates the FFNN under the same noise model used for Fig. 3), it is not established that the constant-overhead encoding plus decoder restores convergence where the unprotected circuit fails. The manuscript should either (i) provide such results or (ii) substantially narrow the abstract/conclusion claims to what is actually shown: noise resilience of a transversal QCNN skeleton and a separate, Clifford-trained decoder that currently over-corrects non-Clifford layers.
minor comments (5)
- [Fig. 3, §IV-B] Fig. 3 caption and body refer to both '4-qubit QCNN and 11-qubit transversal QCNN' and later to panels (a)–(d) at different p; the figure itself is described as a single multi-panel plot but the text sometimes treats (c) and (d) as both p = 0.003 for different architectures. Clarify panel labels and which curve is which.
- [Abstract, §I-B, §IV-B] Typos and consistency: 'QCNNS' in the abstract; 'the the application' in §I-B; 'reduces, the effective expressibility' in §IV-B; 'learning loss was lowered' etc. A careful proofread is needed.
- [§III-A, §III-B] Eq. (17) defines Loss(φ) as a 0–1 misclassification rate over N = 16 patterns, while the QCNN itself is trained with MSE (Eq. 14). The relationship between the FFNN loss and the QCNN MSE under joint or sequential training is not stated; a short clarification would help.
- [§III-C] The qubit-count reduction from 36 to 29 via shared-index mapping (§III-C) is important; a small table or explicit list of the 11 data indices and which logical qubits share them would make the construction easier to reproduce.
- [§I-A] Related-work discussion of [16] is useful but could more clearly separate classical simulability of QCNNs from hardware trainability under noise, which is the paper's actual focus.
Circularity Check
No circularity: empirical simulation results and literature-sourced BB/QCNN constructions do not reduce by construction to their inputs.
full rationale
The paper's central claims rest on stochastic circuit simulations under a depolarizing noise model (p in {0, 0.0001, 0.001, 0.003}) of an unprotected 4-qubit QCNN versus a transversally encoded 11-qubit ansatz derived from the [[18,4,4]] BB code of Bravyi et al. Training losses, best-so-far curves, and prediction distributions (Figs. 3–6) are obtained by independent finite-difference/SPSA optimization runs; they are not algebraic rearrangements of fitted parameters. The BB check matrices, transversal gate expansion (Eqs. 21–22), and syndrome-extraction architecture are taken verbatim from the cited literature ([12], [15]) rather than re-derived or fitted to the QCNN loss surface. The FFNN decoder is trained on an exhaustive set of single-layer Pauli patterns (Eq. 17) and then applied; its acknowledged over-correction of non-Clifford layers (Fig. 6) is an empirical failure mode, not a definitional identity. No uniqueness theorem, self-citation chain, or ansatz smuggled via prior author work forces the reported resilience at p=0.001. The derivation chain is therefore self-contained against external benchmarks and exhibits zero circular steps.
Axiom & Free-Parameter Ledger
free parameters (3)
- BB polynomial coefficients [a1,a2,a3]=[1,0,1], [b1,b2,b3]=[2,0,2]
- SPSA perturbation magnitude c=0.2 rad and bold-driver factors 1.05/0.5
- Physical error rates p ∈ {0.0001,0.001,0.003}
axioms (3)
- domain assumption Bivariate bicycle codes admit a depth-7 syndrome extraction circuit with constant encoding rate and linear distance (Bravyi et al. 2024).
- domain assumption A QCNN realizes an inverse MERA and can classify 1-D SPT phases with O(log N) parameters (Cong et al. 2019).
- domain assumption Circuit-level depolarizing noise with independent single- and two-qubit Pauli channels is a realistic model of NISQ hardware.
invented entities (1)
-
CQCNN with interleaved FFNN soft decoder
no independent evidence
read the original abstract
Quantum convolutional neural networks (QCNNs) combine the power of quantum computing and classical CNN for computational speedup in classification tasks. However, noise levels on state-of-the-art quantum devices remain too high for practical QCNN execution. In addition, despite the reliable surface code providing a method for error rates below a threshold value, they have a prohibitively large qubit cost. Recently introduced bivariate bicycle (BB) codes are of particular interest for their high error threshold, constant encoding rate, and linear code distance. Through simulation with realistic hardware noise sources, we demonstrate that a 4-qubit unprotected QCNN fails to converge and exhibits a worse learning rate compared to numerical simulations. Addressing both limitations, we propose a distance-4 BB quantum error-correction (QEC) technique for QCNNs. In doing so, we validate that our low-overhead QEC technique for QCNNS represents a step toward practical QCNNs.
Figures
Reference graph
Works this paper leans on
-
[1]
The vanishing gradient problem during learning recurrent neural nets and problem solutions,
S. Hochreiter, “The vanishing gradient problem during learning recurrent neural nets and problem solutions,”International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, vol. 6, no. 2, pp. 107–116, 1998
work page 1998
-
[2]
On the difficulty of training recurrent neural networks,
R. Pascanu, T. Mikolov, and Y . Bengio, “On the difficulty of training recurrent neural networks,” inProceedings of the 30th International Conference on Machine Learning, 2013
work page 2013
-
[3]
Information-theoretic bounds on quantum advantage in machine learning,
H.-Y . Huang, R. Kueng, and J. Preskill, “Information-theoretic bounds on quantum advantage in machine learning,”Physical Review Letters, vol. 126, no. 19, p. 190505, 2021
work page 2021
-
[4]
Y . LeCun, Y . Bengio, and G. Hinton, “Deep learning,”Nature, vol. 521, no. 7553, pp. 436–444, May 2015
work page 2015
-
[5]
M. A. Nielsen and I. L. Chuang,Quantum Computation and Quantum Information: 10th Anniversary Edition. Cambridge, U.K.: Cambridge University Press, 2011
work page 2011
-
[6]
Hirvensalo,Quantum Computing, 2nd ed
M. Hirvensalo,Quantum Computing, 2nd ed. Berlin, Germany: Springer, 2001
work page 2001
-
[7]
J. Biamonteet al., “Quantum machine learning,”Nature, vol. 549, pp. 195–202, 2017
work page 2017
-
[8]
Supervised learning with quantum-enhanced feature spaces,
V . Havl´ıˇceket al., “Supervised learning with quantum-enhanced feature spaces,”Nature, vol. 567, pp. 209–212, 2019
work page 2019
-
[9]
Barren plateaus in quantum neural network training landscapes,
J. R. McCleanet al., “Barren plateaus in quantum neural network training landscapes,”Nature Communications, vol. 9, p. 4812, 2018
work page 2018
-
[10]
Building logical qubits in a superconducting quantum computing system,
J. M. Gambetta, J. M. Chow, and M. Steffen, “Building logical qubits in a superconducting quantum computing system,”npj Quantum Infor- mation, vol. 3, pp. 1–7, 2017
work page 2017
-
[11]
Fault-tolerant quantum computation by anyons,
A. Y . Kitaev, “Fault-tolerant quantum computation by anyons,”Annals of Physics, vol. 303, no. 1, pp. 2–30, 2003
work page 2003
-
[12]
High-threshold and low- overhead fault-tolerant quantum memory,
S. Bravyi, A. W. Cross, J. M. Gambettaet al., “High-threshold and low- overhead fault-tolerant quantum memory,”Nature, vol. 627, pp. 778– 782, 2024
work page 2024
-
[13]
Fault-tolerant quantum computation with constant over- head,
D. Gottesman, “Fault-tolerant quantum computation with constant over- head,”Quantum Information and Computation, vol. 14, no. 15–16, 2014
work page 2014
-
[14]
Quantum convolutional neural networks,
I. Cong, S. Choi, and M. D. Lukin, “Quantum convolutional neural networks,”Nature Physics, vol. 15, no. 12, pp. 1273–1278, Dec. 2019
work page 2019
-
[15]
Demonstration of low-overhead quantum error correction codes,
K. Wang, Z. Lu, C. Zhanget al., “Demonstration of low-overhead quantum error correction codes,”Nature Physics, vol. 22, pp. 308–314, 2026
work page 2026
-
[16]
Quantum convolutional neural networks are effec- tively classically simulable,
P. Bermejoet al., “Quantum convolutional neural networks are effec- tively classically simulable,”PRX Quantum, vol. 7, no. 2, p. 020304, Apr. 2026
work page 2026
-
[17]
Convolutional networks for images, speech, and time-series,
Y . LeCun and Y . Bengio, “Convolutional networks for images, speech, and time-series,” inThe Handbook of Brain Theory and Neural Net- works, 1995
work page 1995
-
[18]
Multiscale entanglement renormalization ansatz: Causal- ity and error correction,
D. Pomarico, “Multiscale entanglement renormalization ansatz: Causal- ity and error correction,”Dynamics, vol. 3, no. 3, pp. 622–635, 2023
work page 2023
-
[19]
Infinite size density matrix renormalization group, revisited
I. P. McCulloch, “Infinite size density matrix renormalization group, revisited,”arXiv preprint arXiv:0804.2509, 2008
work page internal anchor Pith review Pith/arXiv arXiv 2008
-
[20]
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation,
J. C. Spall, “Multivariate stochastic approximation using a simultaneous perturbation gradient approximation,”IEEE Transactions on Automatic Control, vol. 37, no. 3, pp. 332–341, 1992
work page 1992
-
[21]
First- and second-order methods for learning: Between steepest descent and Newton’s method,
R. Battiti, “First- and second-order methods for learning: Between steepest descent and Newton’s method,”Neural Computation, vol. 4, no. 2, pp. 141–166, Mar. 1992
work page 1992
-
[22]
Quantum error correction below the surface code threshold,
Google Quantum AI and Collaborators, “Quantum error correction below the surface code threshold,”Nature, vol. 638, pp. 920–926, 2025
work page 2025
-
[23]
Demonstrating real-time and low-latency quantum error correction with superconducting qubits
L. Cauneet al., “Demonstrating real-time and low-latency quan- tum error correction with superconducting qubits,”arXiv preprint arXiv:2410.05202, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[24]
Low Latency GNN Accelerator for Quantum Error Correction
J. Viszlaiet al., “Low latency GNN accelerator for quantum error correction,”arXiv preprint arXiv:2603.22149, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[25]
Sachdev,Quantum Phase Transitions, 2nd ed
S. Sachdev,Quantum Phase Transitions, 2nd ed. Cambridge, U.K.: Cambridge University Press, 2011
work page 2011
-
[26]
Subsystem codes with high thresh- olds by gauge fixing and reduced qubit overhead,
O. Higgott and N. P. Breuckmann, “Subsystem codes with high thresh- olds by gauge fixing and reduced qubit overhead,”Physical Review X, vol. 11, no. 3, 2021
work page 2021
-
[27]
Magic state distillation: Not as costly as you think,
D. Litinski, “Magic state distillation: Not as costly as you think,” Quantum, vol. 3, p. 205, Dec. 2019
work page 2019
-
[28]
Restrictions on transversal encoded quantum gate sets,
B. Eastin and E. Knill, “Restrictions on transversal encoded quantum gate sets,”Physical Review Letters, vol. 102, no. 11, p. 110502, 2009
work page 2009
-
[29]
IBM quantum delivers on 2022 100×100 performance challenge,
IBM Quantum, “IBM quantum delivers on 2022 100×100 performance challenge,” IBM Quantum Blog, Nov. 2024, accessed: 2024. [Online]. Available: https://www.ibm.com/quantum/blog/qdc-2024
work page 2022
-
[30]
How to factor 2048 bit RSA integers in 8 hours using 20 million noisy qubits,
C. Gidney and M. Eker ˚a, “How to factor 2048 bit RSA integers in 8 hours using 20 million noisy qubits,”Quantum, vol. 5, p. 433, 2021
work page 2048
-
[31]
Even more efficient quantum computations of chemistry through tensor hypercontraction,
J. Leeet al., “Even more efficient quantum computations of chemistry through tensor hypercontraction,”PRX Quantum, vol. 2, p. 030305, 2021
work page 2021
-
[32]
Focus beyond quadratic speedups for error-corrected quan- tum advantage,
R. Babbush, J. R. McClean, M. Newman, C. Gidney, S. Boixo, and H. Neven, “Focus beyond quadratic speedups for error-corrected quan- tum advantage,”PRX Quantum, vol. 2, p. 010103, 2021
work page 2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.