REVIEW 2 major objections 6 minor 1 cited by
This paper proves that for any bounded-reward task, a classical strategy with benchmark dependence η can exceed the optimal dependence-free classical score by at most η, so an observed quantum score gap directly lower-bounds the benchmark-c
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 06:43 UTC pith:RSPLTPE6
load-bearing objection The universal relaxed-advantage inequality is correct and the framing is useful; the non-Bell application's quantitative claim rests on a restricted SVM baseline, not the theorem's S_cl. the 2 major comments →
Loophole-Robust Certification of Quantum Advantage
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that benchmark dependence—allowing the hidden variable of a classical strategy to be distributed differently for different task instances—raises the optimal classical score by at most the dependence strength itself. Formally, S_η ≤ min{1, S_cl + η} for any bounded-reward task, and the linear dependence on η is tightly saturated by an explicit family. For repeated product tasks where each round's hidden variable is independent and depends only on that round's instance, the bound sharpens to S_η^{(n)} ≤ (ω_c + η)^n, which is also tight for the hidden-instance family. Inverting these bounds converts a measured quantum–classical score gap into a minimum required benchmark de
What carries the argument
The central object is the benchmark-dependence parameter η_bm = sup_z d_TV(p(λ|z), p_0(λ)), the largest total-variation distance between the hidden variable's law conditioned on a benchmark instance and its marginal law over instances. The argument rests on the standard variational fact that a [0,1]-valued reward function can change its expectation by at most the total-variation distance between two distributions; applying this instance-wise and averaging over the benchmark yields the additive ceiling, and optimizing over local response functions converts it into a bound on the classical value. The multiplicative product-task bound uses the same per-round estimate under the roundwise-indepen
Load-bearing premise
The sharp product-task certificates rest on the roundwise model in which each round's hidden variable is independent across rounds and depends only on that round's instance; if a surrogate may use a single hidden variable correlated with the entire context sequence, the multiplicative bounds do not apply—the paper flags this in Sec. III.A—and the hardware 'cycle-product' certificate is a reconstruction from context-aggregated data, with one formula referenced only as 'Eq. (??
What would settle it
Run a repeated Bell product task with shot-ordered data and allow a classical simulator to use one hidden variable correlated with the whole context sequence; if it reproduces the observed all-round success rate with per-round total-variation dependence below Q_n^{1/n} – ω_c, the multiplicative certificate fails. The paper's own admission that inter-round hidden-state correlations invalidate the factorization points directly to this test.
If this is right
- Any observed quantum score gap S > S_cl becomes a quantitative certificate: no benchmark-dependent classical surrogate can explain the gap unless its dependence is at least S – S_cl.
- For repeated product tasks scored by winning every round, the required per-round dependence is Q_n^{1/n} – ω_c, which is stricter than the additive average-score threshold, so product-style benchmarks give stronger loophole certification.
- Finite-sample versions provide one-sided confidence lower bounds on the required dependence from raw counts, so a hardware demonstration can be reported as 'with 95% confidence, dependence at least X'.
- For pseudo-telepathy Bell games, the threshold to fake a perfect product score equals the classical-value gap: 1/9 for the nine-context magic-square game and 1/4 for the three-party parity game; for the two-party Bell game the threshold to the quantum optimum is about 0.104.
- The same framework audits non-Bell quantum-kernel benchmarks: construction-side variables with measured label dependence above the required threshold support classical shortcut classifiers with perfect accuracy, so the apparent advantage collapses.
Where Pith is reading between the lines
- Editorial inference: this certificate could become a standard reporting item for quantum-advantage claims—alongside a score gap, authors would report the minimum benchmark dependence any classical surrogate needs to close it.
- Editorial inference: worst-case total-variation dependence and average mutual-information dependence can disagree sharply when leakage is concentrated on rare instances, so a serious audit should report both rather than either alone.
- Editorial inference: the framework suggests a direct experimental protocol—measure or bound p(λ|z) for every candidate shortcut variable in a benchmark and compare each to η_req; if none exceeds the threshold, the advantage claim is much harder to undermine.
- Editorial inference: the same additive bound applies to any bounded-reward benchmark, so random-circuit-sampling or other sampling advantages could in principle be given the same treatment by identifying what plays the role of the benchmark instance and the hidden side resource.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces 'benchmark dependence' as a task-level generalization of measurement dependence, quantifying how much a classical strategy's hidden variable may depend on the benchmark instance. The central result is a universal relaxed inequality: for any bounded-reward distributed task, the optimal classical score under benchmark dependence of strength at most η satisfies S_η ≤ min{1, S_cl + η}. The authors prove this bound, construct a family of tasks saturating it, and extend it to repeated product tasks, finite-sample certification, mutual-information constraints, multipartite tasks, and correlated hidden trajectories. They apply the framework to IBM hardware data for CHSH, Mermin–GHZ, and magic-square games, and to a non-Bell quantum-kernel benchmark with an explicit audit of construction-side variables.
Significance. The core inequality is elementary but its formulation as a resource bound for quantum-advantage certification is genuinely useful: it converts an observed score gap into a quantitative lower bound on the benchmark-correlated classical information needed to explain it. The saturation construction shows the linear dependence is tight, and the paper is careful about statistical confidence (Clopper–Pearson with Bonferroni allocation) and about distinguishing exact certificates from readout-mitigated sensitivity estimates. The central proof is correct and parameter-free. However, the non-Bell application contains a load-bearing error: it equates the score of a single SVM with the optimal classical value S_cl, which invalidates the reported quantitative certificate in that section.
major comments (2)
- [Sec. III.I, Eq. (103)] The value S_cl=0.525 is the held-out accuracy of a particular polynomial-kernel SVM, not the optimal benchmark-independent classical value defined in Eq. (3). In the ad hoc task, the label y is a deterministic function of u through Eq. (91) and the gap Δ=0.25, so an unrestricted classical predictor achieves perfect accuracy on the support of the test distribution. Thus S_cl=1, the score gap is S_Q − S_cl = −0.1, and the claimed η_req=0.375 does not follow from Eq. (14). The subsequent comparison Rheta_λ(Y)=0.5 > 0.375 is therefore not a valid necessary-resource audit.
- [Sec. III.I, Table III and Conclusion] Table III and the conclusion state that the apparent quantum-kernel excess is explainable by benchmark-construction side information and requires η_req=0.375. Because the true S_cl is 1, there is no excess over the universal bound; the shortcut classifiers achieving perfect accuracy merely match what an unrestricted classical model without λ already achieves. The example should either compute the true S_cl (which eliminates the claimed gap) or explicitly restrict the classical model class and state that the certificate is relative to that restricted class, not to the universal S_cl of Sec. II.B.
minor comments (6)
- [Sec. III.G, Eq. (??)] There is a broken reference: 'in the sense of Eq. (??)' should be replaced with the correct equation number for the direct block-success analysis.
- [Sec. III.A, Eq. (34)] The phrase 'we posit that' is unnecessary because a proof follows; change to 'we state' or 'we prove'.
- [Sec. III.G] The 'conservative cycle-product statistic' terminology may mislead: the product of per-context Clopper–Pearson lower limits is a simultaneous lower bound on the product of per-context win probabilities, which equals the block success probability only under roundwise independence. Please clarify that the certificate is for the product statistic under the roundwise model.
- [Sec. III.I, Eq. (95)] The superscript Y in Rheta_λ(Y) is not defined in the notation; define Y as the class-label variable.
- [Sec. III.I, near Eq. (104)] There is a typographical/formatting issue: 'LetS (λ) cl denote...' should read 'Let S_cl^{(λ)} denote...'.
- [Table III] The entry 'Loophole-free classical comparator' for the SVM score 0.525 is misleading because this is not the optimal loophole-free value; it is a specific classical model. Please rename to 'Classical SVM comparator (not optimal)'.
Circularity Check
No circular derivation: the universal relaxed inequality and its saturation are proved from stated definitions; the non-Bell and hardware applications contain non-circular application caveats.
full rationale
The central derivation is self-contained. Eq. (14) follows from a direct total-variation estimate (Eqs. (15)-(19)) applied to the definitions of S_η and S_cl; no term in the proof presupposes the conclusion. The saturation construction (Eqs. (23)-(28)) explicitly builds a benchmark-dependent model with d_TV = η and score 1/m + η, independently verifying tightness. The multiplicative roundwise bound (34)-(36) factorizes independent rounds and applies the one-round bound per factor; the paper itself flags the roundwise-independence limitation and supplies the additive pathwise bound (39) for correlated rounds. Finite-sample and hardware certificates (54)-(59), (77)-(80) invert the proven inequalities using Clopper-Pearson/Bonferroni bounds; these are standard statistical lower bounds, not fitted parameters disguised as predictions. The non-Bell application does contain a non-circular but important application error: Eq. (103) sets S_cl to the held-out accuracy of the best of three SVMs (0.525), whereas Eq. (3) defines S_cl as the supremum over all benchmark-independent classical models. Since labels are deterministic functions of g(u), an unrestricted classical model could score far above 0.525, so the quoted η_req = 0.375 is not a valid consequence of Eq. (14). This is a correctness/substitution flaw, not a circular definition: the paper does not use the conclusion to prove the bound. The audit of sign g(u) is intentionally tautological (η_hat = 0.5 follows from the label-generation rule) and is disclosed as such; shortcut classifiers are tested operationally. No self-citation chain or imported uniqueness theorem is load-bearing. Overall circularity: none.
Axiom & Free-Parameter Ledger
axioms (4)
- standard math Total-variation stability of bounded expectations: for 0 ≤ f ≤ 1, |∑_λ (P(λ)−Q(λ)) f(λ)| ≤ d_TV(P,Q).
- domain assumption Classical strategies have the local hidden-variable form p(a,b|z)=∑_λ p(λ|z) p(a|x(z),λ) p(b|y(z),λ).
- domain assumption The benchmark distribution q(z) is known and fixed.
- domain assumption For product tasks, the roundwise model assumes hidden variables are independent across rounds and each depends only on its round's instance (Eq. 32).
invented entities (1)
-
Benchmark dependence η_bm
independent evidence
read the original abstract
Claims of quantum advantage should remain robust even when classical strategies have access to side information correlated with the benchmark under evaluation, just as Bell certification must account for measurement dependence. We formalize such correlations as benchmark dependence, a task-level generalization of measurement dependence. For every bounded-reward task, we show that the optimal benchmark-dependent classical score obeys $S_\eta\leq\min\{1,S_{\mathrm{cl}}+\eta\}$, and construct a family of tasks that saturates this bound, showing that the linear dependence on $\eta$ is tight without further assumptions. For repeated product tasks with roundwise dependence, we obtain the stronger multiplicative bound $S_\eta^{(n)}\leq(\omega_{\mathrm c}+\eta)^n$, and extend the framework to finite-sample data, mutual-information constraints, multipartite tasks, and correlations distributed along a causal path. Applying these results to aggregated IBM hardware data, we obtain positive raw-count cycle-product certificates of 0.0812 for CHSH and 0.2178 for Mermin--GHZ, while the nine-context magic-square construction remains uncertified; readout-mitigated values are reported separately as sensitivity estimates. We also analyze a non-Bell quantum-kernel benchmark, where a label-construction variable has measured conditional dependence $\widehat{\eta}_{\lambda}^{(Y)}=0.5$, above the threshold $\eta_{\mathrm{req}}=0.375$, required to close the reported score gap, and yields perfect classical classification. The framework therefore converts a quantum--classical score separation into a quantitative lower bound on the benchmark-correlated classical information required to explain the score separation.
Figures
Forward citations
Cited by 1 Pith paper
-
Exact minimum measurement dependence for faithful local deterministic models of multipartite GHZ-Mermin correlations
For faithful local deterministic models of n-partite GHZ–Mermin correlations, the exact minimum surrendered measurement independence for n=3–13 is R/[2(R+1)] with R=2^floor((n-1)/2), conjectured to hold for all n.
Reference graph
Works this paper leans on
-
[1]
A. W. Harrow and A. Montanaro, Nature549, 203 (2017)
2017
-
[2]
Aaronson and L
S. Aaronson and L. Chen, in32nd Computational Com- plexity Conference, Leibniz International Proceedings in Informatics, Vol. 79, edited by R. O’Donnell (Schloss Dagstuhl–Leibniz-Zentrum für Informatik, Dagstuhl, Germany, 2017) pp. 22:1–22:67
2017
-
[3]
Preskill, arXiv preprint arXiv:1203.5813 (2012)
J. Preskill, arXiv preprint arXiv:1203.5813 (2012)
Pith/arXiv arXiv 2012
-
[4]
Aruteet al., Nature574, 505 (2019)
F. Aruteet al., Nature574, 505 (2019)
2019
-
[5]
Zhong, H
H.-S. Zhong, H. Wang, Y.-H. Deng, M.-C. Chen, L.-C. Peng, Y.-H. Luo, J. Qin, D. Wu, X. Ding, Y. Hu,et al., Science370, 1460 (2020)
2020
-
[6]
Wu, W.-S
Y. Wu, W.-S. Bao, S. Cao, F. Chen, M.-C. Chen, X. Chen, T.-H. Chung, H. Deng, Y. Du, D. Fan,et al., Physical review letters127, 180501 (2021)
2021
-
[7]
Vincent, J
L.S.Madsen, F.Laudenbach, M.F.Askarani, F.Rortais, T. Vincent, J. F. Bulmer, F. M. Miatto, L. Neuhaus, L. G. Helt, M. J. Collins,et al., Nature606, 75 (2022)
2022
-
[8]
C. Huang, F. Zhang, M. Newman, J. Cai, X. Gao, Z. Tian, J. Wu, H. Xu, H. Yu, B. Yuan,et al., arXiv preprint arXiv:2005.06787 (2020)
Pith/arXiv arXiv 2005
-
[9]
Troyer, and P
A.J.Daley, I.Bloch, C.Kokail, S.Flannigan, N.Pearson, M. Troyer, and P. Zoller, Nature607, 667 (2022)
2022
-
[10]
Hibat-Allah, M
M. Hibat-Allah, M. Mauri, J. Carrasquilla, and A.Perdomo-Ortiz,CommunicationsPhysics7,68(2024)
2024
-
[11]
C. W. Helstrom, Journal of statistical physics1, 231 (1969)
1969
-
[12]
Cleve, P
R. Cleve, P. Høyer, B. Toner, and J. Watrous, inPro- ceedings of the 19th IEEE Annual Conference on Com- putational Complexity(IEEE, 2004) pp. 236–249
2004
-
[13]
H. Buhrman, R. Cleve, S. Massar, and R. De Wolf, arXiv preprint arXiv:0907.3584 (2009)
Pith/arXiv arXiv 2009
-
[14]
Brunner, D
N. Brunner, D. Cavalcanti, S. Pironio, V. Scarani, and S. Wehner, Reviews of Modern Physics86, 419 (2014)
2014
-
[15]
J. S. Bell, Physics Physique Fizika1, 195 (1964)
1964
-
[16]
Y. Liu, S. Arunachalam, and K. Temme, Nature Physics 17, 1013 (2021)
2021
-
[17]
Huang, R
H.-Y. Huang, R. Kueng, and J. Preskill, Nature Physics 16, 1050 (2020)
2020
-
[18]
Huanget al., Science376, 1182 (2022)
H.-Y. Huanget al., Science376, 1182 (2022)
2022
-
[19]
Buhrman, R
H. Buhrman, R. Cleve, and A. Wigderson, inProceedings of the thirtieth annual ACM symposium on Theory of computing(1998) pp. 63–68
1998
-
[20]
Zhao and D.-L
H. Zhao and D.-L. Deng, npj Quantum Information11, 127 (2025)
2025
-
[21]
Havlíček, A
V. Havlíček, A. D. Córcoles, K. Temme, A. W. Harrow, A. Kandala, J. M. Chow, and J. M. Gambetta, Nature 567, 209 (2019)
2019
-
[22]
Huang, M
H.-Y. Huang, M. Broughton, M. Mohseni, R. Babbush, S. Boixo, H. Neven, and J. R. McClean, Nature Commu- nications12, 2631 (2021)
2021
-
[23]
Cerezo, G
M. Cerezo, G. Verdon, H.-Y. Huang, L. Cincio, and P. J. Coles, Nature Computational Science2, 567 (2022)
2022
-
[24]
Schuld and N
M. Schuld and N. Killoran, PRX Quantum3, 030101 (2022)
2022
-
[25]
Kaufman, S
S. Kaufman, S. Rosset, C. Perlich, and O. Stitelman, ACM Transactions on Knowledge Discovery from Data 6, 15:1 (2012)
2012
-
[26]
Geirhos, J.-H
R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann, Nature Machine Intelligence2, 665 (2020)
2020
-
[28]
Hensenet al., Nature526, 682 (2015)
B. Hensenet al., Nature526, 682 (2015)
2015
-
[29]
Giustinaet al., Physical Review Letters115, 250401 (2015)
M. Giustinaet al., Physical Review Letters115, 250401 (2015)
2015
-
[30]
L. K. Shalmet al., Physical Review Letters115, 250402 (2015)
2015
-
[31]
M. J. W. Hall, Physical Review Letters105, 250404 (2010)
2010
-
[32]
M. J. W. Hall, Physical Review A84, 022102 (2011)
2011
-
[33]
L. P. Thinh, L. Sheridan, and V. Scarani, Physical Re- view A87, 062121 (2013)
2013
-
[34]
G. Pütz, D. Rosset, T. J. Barnea, Y.-C. Liang, and N. Gisin, Physical Review Letters113, 190402 (2014)
2014
-
[35]
Pütz and N
G. Pütz and N. Gisin, New Journal of Physics18, 055006 (2016)
2016
-
[37]
P. J. Huber, The Annals of Mathematical Statistics35, 73 (1964)
1964
-
[38]
Le Cam, The Annals of Mathematical Statistics35, 1419 (1964)
L. Le Cam, The Annals of Mathematical Statistics35, 1419 (1964)
1964
-
[39]
Le Cam and G
L. Le Cam and G. L. Yang,Asymptotics in Statistics: Some Basic Concepts, 2nd ed. (Springer, New York, 2000)
2000
-
[41]
Peres, Physics Letters A151, 107 (1990)
A. Peres, Physics Letters A151, 107 (1990)
1990
-
[42]
Santha and U
M. Santha and U. V. Vazirani, Journal of Computer and System Sciences33, 75 (1986)
1986
-
[43]
R. Douc, É. Moulines, and J. S. Rosenthal, The Annals of Applied Probability14, 1643 (2004)
2004
-
[44]
A. Y. Mitrophanov, Journal of Applied Probability42, 1003 (2005)
2005
-
[45]
Hairer and J
M. Hairer and J. C. Mattingly, inSeminar on Stochastic Analysis, Random Fields and Applications VI, Progress in Probability, Vol. 63, edited by R. C. Dalang, M. Dozzi, and F. Russo (Springer, Basel, 2011) pp. 109–117
2011
-
[46]
Rudolf and N
D. Rudolf and N. Schweizer, Bernoulli24, 2610 (2018)
2018
-
[47]
C. J. Clopper and E. S. Pearson, Biometrika26, 404 (1934)
1934
-
[48]
L. D. Brown, T. T. Cai, and A. DasGupta, Statistical Science16, 101 (2001)
2001
-
[49]
T. M. Cover and J. A. Thomas,Elements of Informa- tion Theory, 2nd ed. (Wiley-Interscience, Hoboken, NJ, 2006)
2006
-
[50]
M. L. Almeida, J.-D. Bancal, N. Brunner, A. Acín, N. Gisin, and S. Pironio, Physical Review Letters104, 230404 (2010)
2010
-
[51]
Brassard, A
G. Brassard, A. Broadbent, and A. Tapp, Quantum In- formation and Computation5, 538 (2005)
2005
-
[52]
G. Homa, A. Bodor, and J. Z. Bernád, Journal of Physics A: Mathematical and Theoretical59, 065305 (2026)
2026
-
[53]
Pawela, P
Ł. Pawela, P. Gawron, Z. Puchała, and J. Sładkowski, PLOS ONE8, e64694 (2013). 18
2013
-
[54]
Kelleher, M
C. Kelleher, M. Roomy, and F. Holweck, Quantum Infor- mation Processing23, 187 (2024)
2024
-
[55]
Brassard, A
G. Brassard, A. Broadbent, and A. Tapp, Foundations of Physics35, 1877 (2005)
2005
-
[56]
D. M. Greenberger, M. A. Horne, A. Shimony, and A. Zeilinger, American Journal of Physics58, 1131 (1990)
1990
-
[57]
N. D. Mermin, American Journal of Physics58, 731 (1990)
1990
-
[58]
N. D. Mermin, Physical Review Letters65, 3373 (1990)
1990
-
[59]
J. F. Clauser, M. A. Horne, A. Shimony, and R. A. Holt, Physical Review Letters23, 880 (1969)
1969
-
[60]
A. S. Friedman, A. H. Guth, M. J. W. Hall, D. I. Kaiser, and J. Gallicchio, Physical Review A99, 012121 (2019)
2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.