REVIEW 2 major objections 5 minor 37 references
On a fixed four-qubit ZZ kernel, all three hardware execution configurations preserved the statevector geometry substantially (CKA 0.933–0.989), with gate twirling the most faithful; fidelity and label alignment nevertheless reversed, and t
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 09:59 UTC pith:AQHO7WLZ
load-bearing objection A careful, honest single-backend diagnostic of ZZ4 kernel geometry survival; the twirling ordering is real only as a description of three jobs, not as mitigation efficacy. the 2 major comments →
Statevector-Referenced Geometry Survival of a Four-Qubit ZZ Quantum Kernel on IBM Quantum Hardware: A Fixed-Subset Diagnostic Across Three Execution Configurations
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the statevector ZZ4 kernel geometry can be reconstructed on real quantum hardware to a substantial but incomplete descriptive degree, and that the three execution configurations differ in a specific, quantifiable way: gate twirling produced the most faithful reconstruction (full-matrix CKA 0.989, off-diagonal RMSE 0.0427), while dynamical decoupling alone did not separate from baseline. The paper further claims that the centered kernel–target alignment ordering is reversed from the fidelity ordering and that this uplift is normalization-associated: hardware compresses the centered kernel norm more than the label-directed numerator, so the alignment ratio rises witho
What carries the argument
The argument rides on the centered kernel alignment (CKA) functional—a centered, Frobenius-normalized inner product between Gram matrices that is invariant to affine maps K ↦ aK + b·11ᵀ. Because pure depolarizing contraction of the kernel is affine, CKA isolates non-affine distortion; the same functional evaluated against the label Gram matrix becomes centered kernel–target alignment (KTA), letting the paper attribute the hardware KTA uplift to norm compression rather than signal. The finite-shot reference-scale decomposition (binomial plug-in at 1024 shots) supplies the scale against which residual hardware distortion is shown to dominate.
Load-bearing premise
The three configurations each ran once, sequentially, on a single device, so the observed ordering may reflect drift or calibration change between jobs rather than the mitigation options themselves—a confounding the paper itself explicitly acknowledges.
What would settle it
Interleave and repeat the baseline, dynamical-decoupling, and gate-twirling jobs across several days on the same device: if the twirling-over-baseline advantage in CKA and RMSE does not persist under interleaved replication, the ordering is job-to-job drift, not mitigation efficacy. Likewise, if a hardware kernel whose centered norm is renormalized to the statevector value still shows the KTA uplift, the normalization-artifact reading would be wrong.
If this is right
- If correct, selecting gate-twirling runtime options can improve quantum-kernel geometry survival on similar devices without changing the feature map or circuits.
- Hardware quantum machine-learning studies should report both implementation fidelity and task relevance, since high fidelity can coexist with label alignment at or below chance.
- Positive kernel–target alignment uplifts observed on hardware should be compared against label-permutation and finite-shot references before being interpreted as signal.
- Residual hardware distortion, not finite-shot sampling, limits kernel fidelity even under the best configuration at 1024 shots; a projected 4096-shot budget would not change that conclusion.
- The configuration that best preserved the intended geometry is the configuration whose label alignment stayed closest to the (chance-level) statevector reference, so restoring the ideal map cannot by itself restore task relevance.
Where Pith is reading between the lines
- The paper's twist—that the most faithful configuration shows the least label alignment—suggests that pursuing fidelity alone can move a kernel away from task-relevant structure when the ideal feature map itself is misaligned with the labels; this is an editorial inference, not the paper's claim.
- A concrete testable extension: interleaving and replicating the three configurations on the same backend across multiple days would tell whether the twirling-over-baseline ordering survives calibration drift, converting the descriptive ordering into a causal one.
- The normalization-artifact reading predicts that any noisy kernel with strongly compressed centered norm will show inflated centered KTA; this could be tested on other feature maps by renormalizing the hardware kernel's centered norm to the statevector value and checking whether the uplift disappears.
- The finite-shot decomposition suggests that shot-budget increases only matter in the low-distortion regime, so future pilots seeking to separate twirling from baseline should prioritize interleaved replication over simply raising the shot count.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a fixed-subset diagnostic of how faithfully a frozen four-qubit ZZ feature-map kernel (ZZ4) is reconstructed on IBM Quantum hardware (ibm_fez) relative to an exact statevector reference, across three execution configurations: baseline, dynamical decoupling alone, and gate twirling alone. Each configuration was run as a single non-interleaved job at 1024 shots per circuit on a frozen N=24 subset of indoor-air-quality windows. The authors report that all three reconstructed Gram matrices are complete, finite, positive-semidefinite, and preserve the centered statevector geometry to a substantial but incomplete degree (full-matrix CKA 0.933–0.989). Gate twirling yields the best observed geometry-preservation on every reported axis and is the only configuration with jackknife-resolved improvement over baseline for the persisted Spearman, MAE, and full-matrix CKA diagnostics; dynamical decoupling alone is not separated from baseline. A finite-shot reference decomposition indicates that residual hardware distortion, not finite sampling, dominates the off-diagonal RMSE. The statevector kernel shows no above-chance label alignment, and the hardware centered-KTA uplift is interpreted as a normalization-associated distortion property rather than captured label signal. The paper is explicitly scoped as a descriptive single-backend, single-job pilot and makes no quantum-advantage or predictive-performance claim.
Significance. If taken at face value, the manuscript makes a useful, carefully bounded contribution: it separates implementation fidelity from task relevance, provides a reproducible artifact-level package, and demonstrates a disciplined workflow for statevector-referenced hardware-kernel diagnostics. The paper is unusually transparent: the frozen subset is locked, the statevector reference is exact, the hardware kernels are directly measured, no parameters are fitted to the data, the revision-added diagnostics are explicitly disclosed as post hoc, and the CKA affine-invariance argument is internally consistent. The central CKA/KTA tension is a genuinely instructive caution for quantum-machine-learning hardware reporting. The significance is modest because the evidence rests on one small subset, one backend, and single non-interleaved jobs; nevertheless, the methodology and the separation of fidelity from alignment are valuable.
major comments (2)
- [§3.1, §5, Table 5] The headline configuration ordering (M2 > M1 > M0) is computed from three sequential, non-interleaved, unreplicated backend-mode jobs on ibm_fez submitted within a 13.3-minute window. The leave-one-window-out jackknife in Table 5 resamples windows within each job and therefore cannot separate the execution-configuration effect from job-to-job temporal drift, calibration changes, or queue/backend state. The manuscript explicitly concedes this in §5 ('The jobs were neither interleaved nor replicated...'), but the abstract, RQ2, and parts of Section 4.2 still frame the result as 'gate twirling' being most faithful rather than as 'the M2 job on ibm_fez on 2026-05-09' being most faithful. Because every load-bearing M2-vs-M0 contrast is within-job, the ordering is not evidence about the mitigation option per se. Please reframe the headline claims to job-level language throughout (title, abstra
- [§2.5, §3.4.1, Table 8] The H2/M2 results pool 1024 shots across 16 Pauli-twirling randomizations, and per-randomization counts were not persisted. The analysis treats H2 as a single binomial estimator with S=1024 (Eq. 10) and uses that scale in the shot-noise decomposition of Table 8. For twirling, the 16 randomizations are a deliberate averaging procedure; their between-randomization dispersion is not equivalent to binomial shot noise and is unquantified. This makes the H2 RMSE, the H2 jackknife standard errors, and the M2-vs-M0 contrasts in Table 5 not shot-matched to M0/M1. The manuscript acknowledges this for the shot-noise reference, but it does not carry the caveat into the central M2-M0 comparison. Please either recover and report per-randomization dispersion, or explicitly state that the M2-M0 comparison is not shot-matched and soften the 'window-resolved' language for H2 contrasts, which cannot be ful
minor comments (5)
- [§2.2 heading] The heading 'F rozen subset' contains an errant space/ligature; fix to 'Frozen subset'.
- [Throughout] The unusual ligatures in words such as 'efficacy' and 'sufficient' render inconsistently; use standard ASCII spelling ('efficacy', 'sufficient') in the final manuscript.
- [Table 3] The column header 'Pair/PUB entries' is ambiguous; since the inventory is 300 unordered pairs and one PUB per pair, consider labeling it 'PUBs (one per pair)' for clarity.
- [§3.3.1 and Supplementary Table S2.5] The diagonal-excluded U-centered CKA contrast for M2-M0 falls below the narrative cut (z_desc = 1.95), while the full-matrix CKA contrast is above it. The main text mentions this, but the abstract-level summary should also state that the full-matrix CKA separation is partly diagonal-dependent, to avoid overstating that diagnostic as the strongest evidence.
- [Figure 2 caption] Panel (d) uses a grey interval for the statevector-only permutation reference; the caption explains it, but a direct legend inside the panel would help the reader parse the null mean vs q95 without referring to the caption text.
Circularity Check
No significant circularity: the hardware kernels are directly measured, the statevector reference is exact, and no fitted parameter or self-citation drives the results.
full rationale
The paper's derivation chain is self-contained and measurement-grounded. The ZZ4 feature-map hyperparameters are fixed before execution (Section 2.3: 'The configuration ... was fixed before hardware execution and was not re-tuned after observing statevector or IBM hardware outputs'), the statevector reference is the exact squared-fidelity kernel, and each hardware kernel entry is the directly observed all-zero count divided by 1024 shots (Eq. 8). No parameter is fitted to the hardware data; CKA, centered KTA, effective rank, RMSE, and the shot-reference decomposition are deterministic functionals of the measured matrices and the exact reference. The label-permutation null and binomial finite-shot reference are parameter-free diagnostics, not fitted models. Cited works are external and none is load-bearing: the ZZ feature map and metrics come from standard literature by other authors, and the paper does not invoke a uniqueness theorem or an author-imported ansatz. The one near-circular feature is the directional-expectation record E1–E4, but the paper explicitly discloses that it 'is not an independently timestamped pre-execution registration' (Section 2.13) and that the revision diagnostics 'are not pre-specified, and none modifies a frozen artifact' (Section 2.13); these are transparency statements, not fitted inputs renamed as predictions. The main validity threat is the non-interleaved single-job execution, which the paper itself flags in Section 5: 'The jobs were neither interleaved nor replicated, so the configuration effect is confounded with job-to-job variation and calibration drift.' This weakens the causal reading of the descriptive ordering but does not make any reported quantity reduce to its own input. Accordingly, there is no circular step and the appropriate score is 0.
Axiom & Free-Parameter Ledger
axioms (6)
- standard math In the noiseless limit, the all-zero probability of the compute–uncompute circuit equals the squared overlap of the two encoded states (Eq. 7).
- standard math Centered CKA/KTA is invariant to affine maps K ↦ aK + b11^T with a>0, so a pure depolarizing contraction cannot create centered-KTA uplift.
- domain assumption The statevector kernel computed on a noiseless simulator is the 'intended geometry' against which hardware kernels are judged.
- domain assumption The three jobs differ only in the Sampler-level runtime options and are comparable despite sequential execution.
- ad hoc to paper The narrative resolution convention |z_desc| ≥ 2 is an interpretative threshold, not a statistical-significance cut.
- standard math The binomial plug-in and the global reference 1/sqrt(2S) are treated as finite-shot magnitude references, not uncertainty models.
read the original abstract
Quantum-kernel methods encode a dataset's geometry in a Gram matrix, so learning claims on hardware kernels assume the intended geometry survives execution. We measure that survival for one frozen four-qubit ZZ feature-map kernel on $N=24$ real indoor air-quality windows, reconstructed on ibm_fez (1024 shots per circuit) under baseline, dynamical decoupling alone, and gate twirling alone, each a single non-interleaved job. Every configuration returned a complete, finite, positive-semidefinite Gram matrix and preserved the centered statevector geometry to a substantial but incomplete descriptive degree (full-matrix centered kernel alignment, CKA, 0.933-0.989). Gate twirling was most faithful on every reported geometry axis, with the only jackknife-resolved improvement over baseline (persisted Spearman, mean absolute error, and full-matrix CKA diagnostics); dynamical decoupling alone was not separated from baseline at the frozen-window scale. Residual hardware distortion, not finite sampling, dominates the discrepancy. Yet fidelity and label alignment were reversed: the most faithful configuration had the lowest centered kernel-target alignment, which sits at or below label-permutation references for statevector and hardware alike. We read the small hardware uplift as a normalization property of the non-affine distortion, not captured signal. These are descriptive results for single jobs on one backend, not causal mitigation-efficacy estimates; no quantum-advantage, hardware-classifier-superiority, or forecasting claim is made. Implementation fidelity and task relevance are distinct axes; hardware quantum machine-learning studies should report both.
Figures
Reference graph
Works this paper leans on
-
[1]
L. Morawska, P. K. Thai, X. Liu, et al., “Applications of low-cost sensing technologies for air quality monitoring and exposure assessment: how far have they gone?” Environment International , vol. 116, pp. 286–299, 2018, doi: 10.1016/j.envint.2018.04.018
-
[2]
Review of the performance of low-cost sensors for air quality monitoring,
F. Karagulian, M. Barbiere, A. Kotsev, et al. , “Review of the performance of low-cost sensors for air quality monitoring,” Atmosphere, vol. 10, no. 9, p. 506, 2019, doi: 10.3390/atmos10090506
-
[3]
Supervised learning with quantum-enhanced feature spaces,
V. Havlíček et al. , “Supervised learning with quantum-enhanced feature spaces,” Nature, vol. 567, pp. 209–212, 2019, doi: 10.1038/s41586-019-0980-2
-
[4]
Quantum machine learning in feature Hilbert spaces,
M. Schuld and N. Killoran, “Quantum machine learning in feature Hilbert spaces,” Physical Review Letters , vol. 122, p. 040504, 2019, doi: 10.1103/PhysRevLett.122.040504
-
[5]
Supervised quantum machine learning models are kernel methods
M. Schuld, “Supervised quantum machine learning models are kernel methods. ” 2021. doi: 10.48550/arXiv.2101.11020
-
[6]
Exponential concentration in quantum kernel methods,
S. Thanasilp, S. Wang, M. Cerezo, and Z. Holmes, “Exponential concentration in quantum kernel methods,” Nature Communications, vol. 15, p. 5200, 2024, doi: 10.1038/s41467-024-49287-w
-
[7]
Noisy quantum kernel machines,
V. Heyraud, Z. Li, Z. Denis, A. Le Boité, and C. Ciuti, “Noisy quantum kernel machines,” Physical Review A , vol. 106, p. 052421, 2022, doi: 10.1103/PhysRevA.106.052421
-
[8]
Towards understanding the power of quantum kernels in the NISQ era,
X. Wang, Y. Du, Y. Luo, and D. Tao, “Towards understanding the power of quantum kernels in the NISQ era,” Quantum, vol. 5, p. 531, 2021, doi: 10.22331/q-2021-08-30-531
-
[9]
Machine learning of high dimensional data on a noisy quantum processor,
E. Peters et al. , “Machine learning of high dimensional data on a noisy quantum processor,” npj Quantum Information, vol. 7, p. 161, 2021, doi: 10.1038/s41534-021-00498-9
-
[10]
Training quantum embedding kernels on near-term quantum computers,
T. Hubregtsen, D. Wierichs, E. Gil-Fuster, P.-J. H. S. Derks, P. K. Faehrmann, and J. J. Meyer, “Training quantum embedding kernels on near-term quantum computers,” Physical Review A , vol. 106, p. 042431, 2022, doi: 10.1103/PhysRevA.106.042431. 35
-
[11]
Covariant quantum kernels for data with group structure,
J. R. Glick et al. , “Covariant quantum kernels for data with group structure,” Nature Physics , vol. 20, pp. 479–483, 2024, doi: 10.1038/s41567-023-02340-9
-
[12]
Power of data in quantum machine learning,
H.-Y. Huang et al. , “Power of data in quantum machine learning,” Nature Communications, vol. 12, p. 2631, 2021, doi: 10.1038/s41467-021-22539-9
-
[13]
Quantum kernel methods under scrutiny: a benchmarking study,
J. Schnabel and M. Roth, “Quantum kernel methods under scrutiny: a benchmarking study,” Quantum Machine Intelligence, vol. 7, p. 58, 2025, doi: 10.1007/s42484-025-00273-5
-
[14]
On the expressivity of embedding quantum kernels,
E. Gil-Fuster, J. Eisert, and V. Dunjko, “On the expressivity of embedding quantum kernels,” Machine Learning: Science and Technology, vol. 5, p. 025003, 2024, doi: 10.1088/2632-2153/ad2f51
-
[15]
The inductive bias of quantum kernels,
J. M. Kübler, S. Buchholz, and B. Schölkopf, “The inductive bias of quantum kernels,” in Advances in neural information processing systems (NeurIPS) , 2021, pp. 12661–12673
2021
-
[16]
Y. Ji and I. Polian, “Synergistic dynamical decoupling and circuit design for enhanced algorithm performance on near-term quantum devices,” Entropy, vol. 26, no. 7, p. 586, 2024, doi: 10.3390/e26070586
-
[17]
Active readout-error mitigation,
R. Hicks, B. Kobrin, C. W. Bauer, and B. Nachman, “Active readout-error mitigation,” Physical Review A , vol. 105, p. 012419, 2022, doi: 10.1103/PhysRevA.105.012419
-
[18]
Qubit readout error mitigation with bit-flip averaging,
A. W. R. Smith, K. E. Khosla, C. N. Self, and M. S. Kim, “Qubit readout error mitigation with bit-flip averaging,” Science Advances, vol. 7, no. 47, p. eabi8009, 2021, doi: 10.1126/sciadv.abi8009
-
[19]
Z. Cai et al. , “Quantum error mitigation,” Reviews of Modern Physics , vol. 95, p. 045005, 2023, doi: 10.1103/RevModPhys.95.045005
-
[20]
S. Kakavand, C. Strohmeyer, and M. Schlotter, “Benchmarking quantum kernel support vector machines against classical baselines on tabular data: a rigorous empirical study with hardware validation. ” 2026. doi: 10.48550/arXiv.2604.18837
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2604.18837 2026
-
[21]
Challenges and opportunities in quantum machine learning,
M. Cerezo, G. Verdon, H.-Y. Huang, Ł. Cincio, and P. J. Coles, “Challenges and opportunities in quantum machine learning,” Nature Computational Science , vol. 2, pp. 567–576, 2022, doi: 10.1038/s43588-022-00311-3
-
[22]
Dynamical decoupling of open quantum systems,
L. Viola, E. Knill, and S. Lloyd, “Dynamical decoupling of open quantum systems,” Physical Review Letters , vol. 82, pp. 2417–2421, 1999, doi: 10.1103/PhysRevLett.82.2417
-
[23]
Dynamical decoupling for superconducting qubits: a performance survey,
N. Ezzell, B. Pokharel, L. Tewala, G. Quiroz, and D. A. Lidar, “Dynamical decoupling for superconducting qubits: a performance survey,” Physical Review Applied , vol. 20, p. 064027, 2023, doi: 10.1103/PhysRevAp- plied.20.064027
-
[24]
Noise tailoring for scalable quantum computation via randomized compiling,
J. J. Wallman and J. Emerson, “Noise tailoring for scalable quantum computation via randomized compiling,” Physical Review A , vol. 94, p. 052325, 2016, doi: 10.1103/PhysRevA.94.052325
-
[25]
Error mitigation for short-depth quantum circuits,
K. Temme, S. Bravyi, and J. M. Gambetta, “Error mitigation for short-depth quantum circuits,” Physical Review Letters, vol. 119, p. 180509, 2017, doi: 10.1103/PhysRevLett.119.180509
-
[26]
Benchmarking MedMNIST dataset on real quantum hardware,
G. Singh, H. Jin, and K. M. Merz Jr., “Benchmarking MedMNIST dataset on real quantum hardware,” Scientific Reports, vol. 16, p. 9017, Feb. 2026, doi: 10.1038/s41598-026-35605-3
-
[27]
Similarity of neural network representations revisited,
S. Kornblith, M. Norouzi, H. Lee, and G. Hinton, “Similarity of neural network representations revisited,” in Proceedings of the 36th international conference on machine learning (ICML) , in PMLR, vol. 97. 2019, pp. 3519–3529. A vailable: https://proceedings.mlr.press/v97/kornblith19a.html
2019
-
[28]
Algorithms for learning kernels based on centered alignment,
C. Cortes, M. Mohri, and A. Rostamizadeh, “Algorithms for learning kernels based on centered alignment,” Journal of Machine Learning Research , vol. 13, pp. 795–828, 2012, A vailable: https://www.jmlr.org/papers/ v13/cortes12a.html
2012
-
[29]
The effective rank: a measure of effective dimensionality,
O. Roy and M. Vetterli, “The effective rank: a measure of effective dimensionality,” in 2007 15th european signal processing conference (EUSIPCO) , 2007, pp. 606–610. doi: 10.5281/zenodo.40328
-
[30]
Computing a nearest symmetric positive semidefinite matrix,
N. J. Higham, “Computing a nearest symmetric positive semidefinite matrix,” Linear Algebra and its Applica- tions, vol. 103, pp. 103–118, 1988, doi: 10.1016/0024-3795(88)90223-6
-
[31]
S. Rza, “Beyond accuracy: a kernel-level comparative analysis of quantum and classical support vector ma- chines. ” 2026. doi: 10.21203/rs.3.rs-9057443/v1
-
[32]
Shot-frugal and robust quantum kernel classi- fiers
A. Shastry, A. Jayakumar, A. D. Patel, and C. Bhattacharyya, “Shot-frugal and robust quantum kernel classi- fiers. ” 2022. doi: 10.48550/arXiv.2210.06971
-
[33]
A. Javadi-Abhari et al. , “Quantum computing with Qiskit. ” 2024. doi: 10.48550/arXiv.2405.08810
-
[34]
Feature selection via dependence maximization,
L. Song, A. Smola, A. Gretton, J. Bedo, and K. Borgwardt, “Feature selection via dependence maximization,” Journal of Machine Learning Research , vol. 13, pp. 1393–1434, 2012
2012
-
[35]
The jackknife estimate of variance,
B. Efron and C. Stein, “The jackknife estimate of variance,” The Annals of Statistics , vol. 9, no. 3, pp. 586–596, 1981, doi: 10.1214/aos/1176345462. 36
arXiv 1981
-
[36]
M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein, “SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability,” in Advances in neural information processing systems , 2017. doi: 10.48550/arXiv.1706.05806
-
[37]
Grounding representation similarity with statistical testing,
F. Ding, J.-S. Denain, and J. Steinhardt, “Grounding representation similarity with statistical testing,” in Advances in neural information processing systems , 2021. A vailable: https://arxiv.org/abs/2108.01661 37
Pith/arXiv arXiv 2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.