{"id":"242d8868-8ae5-4c3c-8c6c-8cb32b5e8e64","arxiv_id":"2504.18429","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"On IBM's ibm_quebec processor, dynamic circuits yield higher CHSH correlation values than SWAP-based CNOTs beyond 10 qubits, while a post-selected version keeps values above the classical bound up to 13 qubits.","lead":"This paper compares three ways to entangle distant qubits on an IBM 127-qubit quantum computer, using the CHSH test to measure entanglement quality. Beyond about 10 qubits, the dynamic feedforward approach beats the simple SWAP approach, though a post-processing variant gives the best numbers overall.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The zero-parity post-selection filter error-detects rather than emulating ideal feedforward, so the |S|>2-up-to-13-qubits claim and the feedforward-overhead gap may be inflated; a corrected-all-shots reanalysis would settle it.","rationale":"I read this as an honest experimental benchmark with real independent strengths: all three CNOT methods are run on the identical physical qubit chain, with the same M3 readout mitigation and dynamical-decoupling settings, and the paper openly discloses the locality loophole, the post-selection criterion, and the non-scalability caveat. The reader's weakest-assumption choice is correct: the post-processing filter's unbiasedness is the linchpin for the post-processed |S|>2 claim and for the interpretation of the dynamic-vs-post-processed gap as feedforward overhead. My analysis gives that assumption a concrete failure mechanism (the parity filter functions as an error-detecting syndrome check) and a cheap decisive test. The concern is not yet established in magnitude, so it warrants conditional acceptance with a required branch-resolved analysis rather than rejection. Two secondary issues reinforce the same condition: the crossover between dynamic and unitary circuits at ~11 qubits rests on differences of about 0.1 in |S| with overlapping standard deviations and no significance test, and the stated exponential decay of the retained fraction sits uneasily with a parity-based filter (ideal retention ~1/4), so retained-shot counts should be reported. Both are addressable in revision. The dynamic-vs-unitary crossover itself is independent of the post-selection filter, so that part of the abstract survives this concern; the title-level 'CHSH violations' attribution, however, is carried largely by the post-selected data and would need to be framed as conditional-state violations unless the corrected-all-shots check passes. I agree with the reader's CONDITIONAL verdict; my stress-test does not move it.","tokens_in":10697,"tokens_out":23665,"duration_ms":222990,"concrete_test":"Using the stored post-processed bitstrings, reprocess all shots without discarding: assign to each shot the Pauli correction dictated by its ancilla parities (bit-flip for x-parity=1 in X-basis measurements, sign-flip for z-parity=1 in Z-basis measurements, composed where both apply), apply those flips/signs to the data-qubit outcomes, and recompute S over all 10,000 shots at each distance. Compare the corrected-all-shots S to the zero-parity-only S for distances 11-15. If they agree within error bars, the filter is unbiased and the post-processed claims stand; if corrected-all S drops below 2 where the zero-parity S is above 2, the |S|>2 claim and the feedforward-overhead gap are inflated by post-selective error detection. Also report the retained-shot fraction per distance to check the exponential-decay statement and the statistical power of the 13-qubit point.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In Section III.B.3 the paper retains only shots with zero parity in both ancilla registers, saying this 'isolates the experimental outcomes corresponding to the ideal case where no active correction gates would have been necessary.' The load-bearing premise is that this filter is an unbiased proxy for an ideal-feedforward dynamic circuit. The premise is insecure: in the LAQCC protocol the ancilla parities are syndrome bits of an error-detecting code. A Pauli error on a data qubit that anticommutes with a mid-circuit measurement flips a syndrome bit; the dynamic circuit keeps that shot and applies a correction appropriate to the flipped syndrome, while the post-processing filter discards it. The zero-parity branch is therefore not representative of the syndrome-averaged state: a class of data-qubit errors is removed, which systematically raises the post-selected CHSH value relative to what a real dynamic circuit (even with perfect feedforward) would produce. Hence the paper's post-processed |S|>2 values up to 13 qubits describe a heralded conditional state, not the unconditional dynamic output, and the dynamic-vs-post-processed gap (Section V.A) conflates feedforward overhead with an error-detection benefit. Also, Section IV.A claims the retained fraction decays exponentially with ancilla count, but a parity-based filter retains about 1/4 of shots in the ideal case; if the true filter instead requires a per-ancilla pattern, 10,000 shots per observable cannot support the reported statistics at 13 qubits. The paper reports neither retained fractions nor significance tests, so this filter mechanism should be resolved with a branch-resolved analysis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an experimental comparison of three ways to implement a long-range CNOT for Bell-state preparation on the 127-qubit IBM Quantum Eagle processor ibm_quebec: a unitary SWAP-based implementation, a dynamic circuit with mid-circuit measurements and classical feedforward, and a post-processed version that omits the feedforward corrections. Entanglement quality is quantified by the maximum CHSH parameter max(|S|) obtained from a phase sweep. The central claims are that dynamic circuits preserve distance-dependent entanglement better than unitary circuits beyond roughly 10 qubits, that the post-processed implementation gives |S| > 2 up to 13 qubits, and that the gap between dynamic and post-processed results quantifies the overhead of mid-circuit measurement and feedforward on current hardware. The authors acknowledge the locality loophole and state that the post-processing approach is not scalable.","tokens_in":10913,"tokens_out":6234,"duration_ms":63852,"significance":"If the central comparison were sound, this would be a useful experimental benchmark: it uses a direct, parameter-free CHSH metric on a programmable general-purpose processor, spans a wide range of qubit separations, and complements prior fidelity-based studies of dynamic circuits. The paper is transparent about the locality loophole and the non-scalability of the post-processing route, and it reports raw-shot statistics with no fitted free parameters. However, the main quantitative claims currently rest on a post-selection procedure that is not an unbiased proxy for ideal feedforward, and the boundary claims (|S| > 2 at 13 qubits, dynamic-over-unitary advantage beyond 10 qubits) are not supported with statistical significance. These issues are load-bearing, so the manuscript needs major revision before the stated conclusions can be accepted.","major_comments":[{"comment":"The post-processing filter is not an unbiased emulation of ideal feedforward. The ancilla parities are syndrome bits of the LAQCC protocol; retaining only shots with zero XOR in both ancilla registers discards shots in which a data-qubit error would have triggered a feedforward correction. The retained subensemble is therefore a heralded conditional state, not the unconditional output of a dynamic circuit with ideal feedforward. This is analogous to a detection-loophole selection in a Bell test: the CHSH value of the retained subensemble can exceed 2 even when the unconditional correlations are classical. Consequently, the statement that the post-processed approach 'yields the highest CHSH values' and the interpretation in Section V.A that the dynamic-versus-post-processed gap measures feedforward overhead are not justified. The authors should reanalyze all shots by classically propagating the measured parities as Pauli-frame corrections to the final data-qubit outcomes, or explicitly present the post-selected results as conditional/error-detected and refrain from using them as a proxy for ideal dynamic circuits.","section":"Section III.B.3 and Fig. 2(d)"},{"comment":"The boundary claims are not supported statistically. At 13 qubits the reported mean max(|S|) is approximately 2.01, which is within the noise of the classical bound 2, and the 12-qubit dynamic-versus-unitary difference is approximately 0.12 (1.39 versus 1.27). No confidence intervals, standard errors, or hypothesis tests are reported for the n = 20 repetitions, so statements such as 'demonstrating improved distance-dependent entanglement preservation' and '|S| > 2 up to 13 qubits' are not quantitatively established. The authors should report per-distance means with confidence intervals and perform a significance test for (i) max(|S|) > 2 at each distance and (ii) dynamic max(|S|) > unitary max(|S|) beyond 10 qubits.","section":"Section IV.B and Fig. 3(a)"},{"comment":"The claim that the retained fraction 'decreases exponentially with the number of ancillary qubits' is inconsistent with the parity-based filter defined in Section III.B.3 and Fig. 2(d). If the filter is the overall XOR of the z-register and the overall XOR of the x-register, then in the ideal case each parity is 0 with probability 1/2, so the retention probability is about 1/4 independent of distance. If the actual filter instead requires a particular per-ancilla pattern, that stronger post-selection must be stated explicitly, and its effect on the reported shot statistics and on the CHSH estimate must be quantified. As written, the text cannot support the claimed exponential decay or the assertion that 10,000 shots per observable remain sufficient at the largest distances.","section":"Section IV.A"},{"comment":"The cross-over claim that dynamic circuits outperform unitary circuits 'beyond 10 qubits' is weakened by the observation that at those distances both implementations are far below the classical bound (for example, dynamic max(|S|) is about 1.43 at 11 qubits and about 1.39 at 12 qubits). Reporting this as 'improved distance-dependent entanglement preservation' requires showing that the small differences are not due to calibration drift, crosstalk, or other distance-correlated hardware effects. A per-distance calibration and drift characterization, or at least a discussion of how the 20 repetitions were distributed in time, is needed to make this claim credible.","section":"Section IV.B and Abstract"}],"minor_comments":[{"comment":"The introduction states that the unitary approach maintains |S| > 2 up to 7 qubits, while Section IV.B says 'below the classical bound for distances beyond approximately 6 qubits'; these threshold statements should be made consistent.","section":"Section I"},{"comment":"There is a typo, 'Circtuits', in the introduction; it should read 'Circuits'.","section":"Section I"},{"comment":"The horizontal axis of Fig. 3(a) appears to be a path length along selected hardware qubits rather than physical distance; the text should define 'qubit distance' explicitly and explain the non-uniform gaps in the axis labels.","section":"Fig. 3(a)"},{"comment":"The use of M3 readout mitigation is described only for final measurement outcomes; it would be helpful to state explicitly whether mid-circuit measurement errors on ancillary qubits are also corrected or characterized, since these dominate the dynamic-circuit comparison.","section":"Section III.C.2"},{"comment":"The sentence 'To ensure statistical robustness, the entire experiment is repeated multiple times' and the later statement 'calculated from statistics over n = 20 sampled repetitions' are vague: the authors should report whether the 20 repetitions are independent runs on different calibration dates or repeated submissions within a single calibration window.","section":"Section IV.A"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this is a useful experimental dataset with a shaky interpretive frame. The authors measure CHSH violations for Bell states created through three CNOT implementations—SWAP-based unitary, dynamic (mid-circuit measurement + feedforward), and post-processed—across qubit separations up to 15 on ibm_quebec. As far as I can tell, that specific distance-dependent comparison is new, and the raw data are a reasonable contribution.\n\nWhat they do well: the experiment is direct, no fitted parameters, and the comparison is fair on a fixed chain of physical qubits. They are candid about the locality loophole and about post-processing not scaling. The observation that dynamic circuits degrade more slowly than SWAP-based unitary at long distances aligns with earlier fidelity studies, which lends plausibility.\n\nThe soft spots are mostly in the interpretation. The biggest one: the post-selected branch used to represent 'no corrections needed' is not an unbiased proxy for an ideal-feedforward dynamic circuit. A nonzero syndrome can come from a real Pauli error that the feedforward would have corrected; the post-processing filter discards those shots. That removes errors the dynamic circuit would have survived, so the post-processed |S| values are selectively cleaned, not a measurement of feedforward overhead. The gap between post-processed and dynamic results therefore conflates feedforward cost with an error-detection benefit. That needs a branch-resolved analysis or a correction-all-shots re-evaluation.\n\nSecond, the paper states the retained fraction decays exponentially with ancilla count, but a parity-only XOR filter should retain about a quarter of shots independent of chain length. Either the filter is more restrictive than described or that statement is wrong; either way the paper should report actual retained counts, especially at 13 qubits where only 10,000 shots per observable are used.\n\nThird, the boundary claims lack statistics. Mean max(|S|) ≈ 2.01 at 13 qubits is right on the classical threshold, and the crossover around 11 qubits has no significance test. The paper should report standard errors or at least bootstrapped intervals.\n\nFor whom: people benchmarking dynamic circuits on IBM hardware, and compiler folks thinking about teleportation routing. It deserves a serious referee because the dataset is reusable and the flaw is fixable. I would send it out, with the expectation that the authors either reanalyze the post-selected data or soften the claims. Cite? I'd cite the dataset if I were working on dynamic circuit benchmarking.\n\nCheers.","headline":"Useful distance-dependent CHSH dataset for dynamic vs unitary CNOTs, but the post-selection filter inflates the post-processed claim and the boundary results lack statistical support.","tokens_in":11527,"tokens_out":5051,"would_cite":true,"duration_ms":50260,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On chains beyond ten qubits, dynamic circuits preserve more non-classical correlation than SWAP-based CNOTs.","keywords":["CHSH inequality","Bell state","dynamic circuits","LAQCC","entanglement preservation","mid-circuit measurement","NISQ","superconducting qubits"],"falsifier":"Compute $S$ separately for each ancilla-XOR outcome bin and check whether it varies across bins; if the zero-XOR subspace is not representative, the per-bin values will differ beyond statistical error, while an unbiased filter would show bin-independent $S$.","tokens_in":10436,"feed_emoji":"⚛️","tokens_out":6794,"duration_ms":61311,"temperature":0.7,"pith_summary":"This paper reports an experimental comparison of three ways to generate a Bell state between two distant qubits on a 127-qubit superconducting processor: a unitary CNOT implemented by moving one qubit through SWAP gates, a dynamic circuit that uses mid-circuit measurements and classical feedforward to implement the CNOT remotely, and a post-processed version of the dynamic circuit that keeps only shots where the intervening ancillas would not have needed corrections. Using the CHSH parameter $S$ as a distance-dependent measure of entanglement quality, the authors find that the unitary approach starts near $|S|\\approx 2.64$ but falls below the classical bound $|S|=2$ beyond about six qubits, while the dynamic approach decays more slowly and yields higher $|S|$ than the unitary one for distances beyond roughly ten qubits. The post-processed approach gives the highest values throughout, staying above $|S|=2$ up to 13 qubits. A sympathetic reader would take this as evidence that dynamic-circuit-style routing can preserve long-range non-classical correlations better than SWAP chains, and that on current hardware the remaining bottleneck is the cost of mid-circuit measurement and feedforward rather than the dynamic-circuit construction itself.","feed_headline":"Past 10 qubits, dynamic CNOTs beat SWAP-based routing","feed_subtitle":"On a 127-qubit chip, CHSH tests show dynamic circuits keep non-classical correlations farther; post-processing tops the classical bound at…","key_machinery":"The load-bearing mechanism is the LAQCC-style dynamic CNOT: a chain of Bell pairs prepared among ancilla qubits, mid-circuit measurements on two ancilla registers $(z,x)$, and classical feedforward of the measured bits to decide whether $X$ or $Z$ corrections are applied. The post-processed variant deletes the feedforward and instead filters the final measurements, keeping shots where the XOR of the $z$ register and the XOR of the $x$ register are both zero. The CHSH parameter, computed from measured expectation values of $ZZ$, $ZX$, $XZ$, and $XX$ across a sweep of the rotation angle $\\phi$, serves as the metric that quantifies how much non-classical correlation survives at each qubit separation.","core_discovery":"The central discovery is that dynamic circuits mitigate the distance-dependent degradation of entanglement more effectively than unitary SWAP-based routing on this hardware, even though they underperform at short distances. At a 12-qubit separation the maximum CHSH value of the dynamic implementation is about 0.12 higher than the unitary implementation ($\\approx 1.39$ versus $\\approx 1.27$), while the post-processed value sits near the classical bound at $\\approx 2.01$. The post-processed implementation, which removes the feedforward and mid-circuit measurement overhead by filtering on the ancillary registers, achieves the highest max($|S|$) values at every measured distance and maintains violations ($|S|>2$) up to 13 qubits. The authors interpret the gap between the dynamic and post-processed curves as a quantitative measure of the feedforward and measurement overhead that currently prevents dynamic circuits from reaching their theoretical potential.","pith_inferences":["If the zero-XOR filter is unbiased, the dynamic-versus-post-processed gap isolates the combined cost of mid-circuit measurement errors, feedforward latency, and conditional-gate errors; repeating the measurement on a processor with faster feedforward and checking whether the gap shrinks would test this decomposition.","The same experimental protocol could be applied to other connectivity-limited architectures, using only calibration data to predict the distance at which dynamic routing overtakes SWAP-based routing.","Because post-processing discards shots exponentially with the number of ancillas, the post-processed advantage likely degrades at larger separations than the 15-qubit range tested, so extrapolating beyond 13 qubits is unsafe."],"forward_implications":["Algorithm designers on connectivity-limited hardware can expect dynamic routing to become the better choice for long-distance entangling gates, with a crossover on this processor around 10–11 qubits.","Faster mid-circuit measurement and classical feedforward should move that crossover to shorter distances and lift the dynamic curve toward the post-processed curve.","The post-processed curve is a practical upper bound for what a noiseless dynamic circuit on this hardware could achieve, so the gap between the two curves quantifies the feedforward overhead.","CHSH violations can serve as a complementary benchmark to fidelity for evaluating entanglement-preserving routing strategies on near-term devices."],"supporting_citations":[{"why":"Defines the LAQCC formalism that the dynamic CNOT implementation follows.","marker":"[5]"},{"why":"Supplies the dynamic CNOT construction and the earlier fidelity-based evidence this work extends to CHSH violations.","marker":"[7]"},{"why":"Defines the CHSH inequality and the classical bound $|S|\\le 2$ used as the benchmark.","marker":"[10]"},{"why":"Provides the base CHSH circuit and the measurement procedure for computing $S$ from Pauli expectation values.","marker":"[19]"},{"why":"Supplies the readout error-mitigation technique used to correct the raw measurement counts.","marker":"[23]"},{"why":"Supplies the dynamical decoupling technique applied to reduce idle-time decoherence in all three implementations.","marker":"[24]"}],"fun_headline_variants":["Dynamic CNOTs beat SWAP routing past 10 qubits","CHSH tests show dynamic circuits win at long range","On 127 qubits, dynamic circuits beat routing in CHSH","Post-processed CNOTs lead CHSH, dynamic beats unitary","Beyond 10 qubits, dynamic circuits outperform unitary in CHSH"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that keeping only shots where the XOR of the ancilla $z$-register and the XOR of the ancilla $x$-register are both zero gives an unbiased estimate of what an ideal dynamic circuit with perfect feedforward would produce; if those discarded shots are systematically different, the post-processed CHSH values are inflated and the reported feedforward overhead is overstated.","fun_headline_variants_meta":{"raw":{"variants":["Dynamic CNOTs beat SWAP routing past 10 qubits","CHSH tests show dynamic circuits win at long range","On 127 qubits, dynamic circuits beat routing in CHSH","Post-processed CNOTs lead CHSH, dynamic beats unitary","Beyond 10 qubits, dynamic circuits outperform unitary in CHSH"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000655,"raw_usage":{"total_tokens":2983,"prompt_tokens":909,"completion_tokens":2074,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":1986}},"tokens_in":525,"tokens_out":2074,"duration_ms":16013,"temperature":1.0,"reasoning_tokens":1986,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:16:38.894541+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute $S$ separately for each ancilla-XOR outcome bin and check whether it varies across bins; if the zero-XOR subspace is not representative, the per-bin values will differ beyond statistical error, while an unbiased filter would show bin-independent $S$.","supporting_citations":[{"cited_title":"State preparation by shallow circuits using feed forward,","cited_arxiv_id":null,"evidence_quote":"Defines the LAQCC formalism that the dynamic CNOT implementation follows."},{"cited_title":"Proposed Experiment to Test Local Hidden-Variable Theories,","cited_arxiv_id":null,"evidence_quote":"Defines the CHSH inequality and the classical bound $|S|\\le 2$ used as the benchmark."},{"cited_title":"CHSH Inequality | IBM Quantum Learning","cited_arxiv_id":null,"evidence_quote":"Provides the base CHSH circuit and the measurement procedure for computing $S$ from Pauli expectation values."}],"review_version":1}