{"id":"405b1166-f0b3-48f6-af74-5f76ff9cd7e5","arxiv_id":"2607.23574","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Finite-setting generalized Mermin inequalities lower classical bounds at fixed quantum value, yielding exponentially growing Bell ratios and nonlocality depth 14 on 80-qubit GHZ states.","lead":"Researchers generalized Mermin Bell inequalities so more measurement settings tighten classical bounds without changing the ideal quantum score, then used them to benchmark GHZ states of up to 80 superconducting qubits. The method gives stronger nonlocality ratios and deeper certified nonlocality from the same noisy states, using only raw correlators.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"Reader's flagged soft spot is real but narrower than stated: the headline depth-14 (m=8) and Bell-ratio claims sit inside the proven bound regime; only the m=16/32 depth entries in Fig. 4a/Table S4 rest on unproven local bounds C_{16,7} and C_{32,7}.","rationale":"I agree with the reader that bound applicability at edge (m,n) is the correct place to probe, and sharpen it. Verifying each reported point against the stated thresholds: the strongest_claim's headline numbers are all inside proven or explicitly verified regimes — C_{8,7} (depth 14 at n=80) satisfies n=7 ≥ n_odd_8=6; the ratio fit D^exp_{8,n}∝(1.5168)^n uses n∈{4,8,16,28,43,60,80}, all even ≥ n_even_8=4 or odd ≥ n_odd_8=6; the m=2 comparison uses exactly known Mermin bounds; the n=80 ratio scan at m=16,32 uses even-n bounds covered by Prop. 1 (n_even_16=8, n_even_32=24 ≤ 80). The genuine gap is confined to the depth labels for m=16 and m=32 at n=80, which need C_{16,7} and C_{32,7} — outside Prop. 2's proven range (n_odd_16=9, n_odd_32=25) and not covered by the §SIII.C.1 verifications, which handle only m=8 odd cases. This qualifies, but does not overturn, the fixed-size \"increasing m helps\" claim, whose Bell-ratio component remains fully proven. Secondary observations: the depth-14 (m=8) certification is statistically marginal (margin 0.0143, log10 pEB = −3.659 vs the stated −3 criterion), but the rejection rule is pre-stated and transparent, so this is fragility rather than error; the nested depth-null structure avoids a multiplicity problem. The theory proofs are well-structured with checkable intermediate identities (product form Eq. S21, Lemmas 3–4), the proofs explicitly flag their non-tight thresholds, and the correlation-only data reporting is appropriately conservative (settings-level N, not shot count). Net: central claim holds up; recommend ACCEPT unchanged, with the (16,7)/(32,7) verification as a cheap, decisive follow-up.","tokens_in":32182,"tokens_out":12129,"duration_ms":403016,"concrete_test":"Numerically determine the exact local bounds at (m,n)=(16,7) and (32,7). First brute-force the single-party norm maximization max_a‖v_m(a)‖_6^6 over all 2^16 (trivial) and 2^32 (feasible via FFT evaluation of v(a), Eq. S32) deterministic strategies to establish the even-n=6 bound; then evaluate the odd-n comparison inequalities (S68)–(S73) directly at n=7, or enumerate the aligned half-plane strategy family (Eq. S48, only m strategies) over relative offsets r′ (Eq. S67). If the true local value exceeds 0.0853 (m=16) or 0.0849 (m=32), the corresponding depth-14 entries in Table S4/Fig. 4a drop to depth 13; otherwise the conjectured bounds are confirmed at exactly the points used.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Depth certification tests ⟨M⟩ against C_{m,⌈n/(k−1)⌉} (Cor. 2). At n=80, depth k=14 requires the 7-party local bound C_{m,7}. Checking the proven validity ranges: for m=8, n_odd_8 = max{n_even_8+1=5, n_odd,1_8=4, n_odd,2_8=6} = 6 ≤ 7, so C_{8,7}=0.0869 is covered by Proposition 2 — the central depth-14-vs-10 claim is safe. For m=4, n_odd_4=4, also covered; m=2 is standard Mermin. But for m=16, n_odd_16 = max{9,5,7} = 9 > 7, and for m=32, n_odd_32 = 25 ≫ 7. The thresholds C_{16,7}=0.0853 and C_{32,7}=0.0849 used in Table S4 to certify depth 14 for m=16 and m=32 are therefore outside Proposition 2's proven range. §SIII.C.1 closes gaps only for m=8 (n=3 numerically, n=5 by a sketched comparison) and for even n at m∈{4,16,32}; it does not verify odd l=7 at m=16,32. Note the odd-n proof at n=7 also needs the even-n bound at n=6, which is below n_even_16=8 and n_even_32=24. The paper conjectures the formulas hold for all n and found no counterexample, so the risk is low, but the \"increasing m strengthens depth certification\" trend at fixed n=80 for m>8 currently rests on conjectured bounds. The even-n Bell-ratio scan at n=80 for m=16,32 is unaffected (Prop. 1 covers n≥8 and n≥24 respectively), and the exponential-ratio fit uses (m,n) pairs all inside proven ranges.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is simple: they give a practical extra dial m on GHZ Bell tests. For m a power of two the ideal normalized quantum value stays 1 while the local and grouping bounds drop, so ratios and nonlocality depth improve without better state prep. On Zuchongzhi they hold the 80-qubit GHZ fixed, vary only the equatorial analysis, sample correlators, and get D_exp_8,n scaling like (1.52)^n plus depth 14 for m=8 versus 10 for m=2—all from raw correlators and analytic bounds, no mitigation.\n\nWhat is actually new is the organized finite-setting normalized family with closed-form local bounds (Props. 1–2) and the grouping-depth corollary, plus the large-scale experiment that isolates the bound-tightening mechanism. Continuum and multi-setting GHZ inequalities already existed; they cite Żukowski, Nagata et al., extended parity games, and Pál–Vértesi. The incremental theory value is the usable finite-m formulas and depth bounds you can sample without exponential enumeration. The SM is careful: product form, Hölder, conjugate strategies, norm interpolation, explicit n thresholds, and they check the (m,n) pairs they use. EB p-values at the settings level are the right conservative choice. Loopholes are disclosed honestly; they do not claim DI.\n\nSoft spot, in proportion: the headline depth-14 (m=8) and the exponential-ratio fit sit inside the proven regime. The m=16 and m=32 depth entries at n=80 lean on C_m,7 outside the proven odd-n range (they conjecture the formulas hold everywhere and report no counterexample). That is a narrow gap, not a load-bearing collapse—the even-n ratio scan and the m=2 vs m=8 depth comparison are fine. Theory novelty is real but not revolutionary; the package plus the 80-qubit raw-correlator demo is what carries it.\n\nThis is for people who do multipartite Bell benchmarking or large GHZ characterization on NISQ hardware. Math, data, and citations look solid. I would engage, send it to referees, and cite it if I were writing on GHZ Bell tests or correlation-only processor benchmarks.","headline":"Useful finite-setting Mermin package plus real 80-qubit data: m tightens classical bounds and deepens certification on the same noisy GHZ states.","tokens_in":34035,"tokens_out":572,"would_cite":true,"duration_ms":20267,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Raising the number of measurement settings in generalized Mermin inequalities tightens classical bounds and sharpens Bell certification of large noisy GHZ states.","keywords":["multipartite Bell inequalities","generalized Mermin inequality","GHZ states","nonlocality depth","Bell benchmarking","superconducting qubits","noise robustness"],"falsifier":"Fix the same noisy n-qubit GHZ preparation and sampling budget, raise m through powers of two, and check whether the measured operator stays near the m=2 value while the experimental Bell ratio and grouping-model depth both increase exactly as the analytic bounds predict; a failure of either the ratio growth or the depth ordering falsifies the central claim.","tokens_in":33664,"feed_emoji":"⚛️","tokens_out":880,"duration_ms":24903,"temperature":0.7,"pith_summary":"Multipartite Bell tests can benchmark quantum processors from correlators alone, but noise kills many-body signals and ordinary Bell expressions explode in size. This paper introduces a finite-setting family of generalized Mermin inequalities tailored to GHZ states, treating the local setting count m as a second certification knob alongside system size n. For powers-of-two m, the ideal normalized quantum value stays 1 while the relevant classical and grouping-model bounds fall, so Bell-violation ratios and nonlocality-depth claims strengthen for the same noisy state. On a superconducting processor the authors prepare GHZ states up to 80 qubits, estimate the operators by randomized sampling of correlators, and report exponentially growing ratios plus a certified nonlocality depth of 14 at n=80 for m=8 versus 10 for the standard two-setting test—without readout correction, tomography, or model-based mitigation.","feed_headline":"More settings, stronger Bell tests on 80-qubit GHZ states","feed_subtitle":"Raising measurement choices tightens classical bounds and certifies deeper nonlocality from the same noisy states","key_machinery":"The normalized generalized Mermin operator Mm,n, which averages signed products of m coplanar equatorial observables only over setting vectors whose indices sum to a multiple of m; analytic local and grouping-model bounds for this operator (closed forms for powers-of-two m) carry the certification.","core_discovery":"For powers-of-two setting numbers m, the normalized generalized Mermin operator has ideal GHZ value 1 while its local and k-producible classical bounds decrease with m, producing larger Bell ratios and deeper nonlocality-depth certification from essentially the same measured correlators; experiment on up to 80-qubit GHZ states confirms exponentially growing ratios and stronger depth claims as m increases.","pith_inferences":["If the noise-robustness base continues to improve toward the continuous-setting limit ~2/π, moderate-m tests may already capture most of the available certification gain on present devices.","The construction suggests a design pattern for other stabilizer states: enlarge the equatorial setting set to suppress classical bounds while freezing the ideal quantum value.","Closing locality and freedom-of-choice loopholes on a future architecture would convert the same operators into a scalable device-independent depth witness rather than a correlation-only benchmark."],"forward_implications":["Bell benchmarks of large GHZ states can be strengthened by changing only the measurement layer, without better state preparation.","Nonlocality-depth claims on NISQ hardware become tighter once m is treated as a free certification parameter.","Randomized sampling of the finite-setting operator makes direct Bell-operator estimation scalable past exhaustive correlator lists.","The same analytic bounds supply a correlation-only figure of merit portable across hardware platforms that can prepare GHZ states."],"fun_headline_variants":["More settings tighten classical bounds in 80-qubit GHZ Bell tests","Generalized Mermin family certifies depth-14 nonlocality on 80 qubits","Raising m boosts Bell ratios for noisy large-scale GHZ states","Powers-of-two settings strengthen Mermin benchmarks up to 80 qubits","Finite-setting Mermin inequalities yield deeper GHZ nonlocality claims"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The closed-form classical and grouping bounds used for every reported ratio and depth claim must be valid at the experimental pairs (m, n); if those formulas do not apply or are loose at a claimed point, the certified depths and scaling advantage are overstated.","fun_headline_variants_meta":{"raw":{"variants":["More settings tighten classical bounds in 80-qubit GHZ Bell tests","Generalized Mermin family certifies depth-14 nonlocality on 80 qubits","Raising m boosts Bell ratios for noisy large-scale GHZ states","Powers-of-two settings strengthen Mermin benchmarks up to 80 qubits","Finite-setting Mermin inequalities yield deeper GHZ nonlocality claims"]},"model":"grok-4.5","effort":"low","cost_usd":0.004043,"raw_usage":{"total_tokens":1285,"prompt_tokens":808,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":40428000,"prompt_tokens_details":{"text_tokens":808,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":395,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":808,"tokens_out":82,"duration_ms":7790,"temperature":1.0,"reasoning_tokens":395,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T18:26:43.781369+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Fix the same noisy n-qubit GHZ preparation and sampling budget, raise m through powers of two, and check whether the measured operator stays near the m=2 value while the experimental Bell ratio and grouping-model depth both increase exactly as the analytic bounds predict; a failure of either the ratio growth or the depth ordering falsifies the central claim.","supporting_citations":[],"review_version":1}