{"id":"9921df6c-4087-47b8-9540-6e26dbf3d823","arxiv_id":"2607.14260","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Mathematically equivalent depolarizing-noise implementations produce markedly different results when emulating an entanglement distribution network on real quantum hardware.","lead":"Using a quantum computer, the authors tested three equivalent ways to add realistic noise to simulated entanglement-distribution networks and found the results diverge sharply on real hardware. The work warns experimenters that mathematically equal noise models can behave very differently when actually run, so implementation choice matters.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Independent calibration snapshots per depolarization method confound the hardware comparison; transient noise, not implementation, may drive the reported 'profound differences.'","rationale":"The reader correctly identifies the calibration-snapshot confound as the key threat to the central claim. Section 3.1 and Section 4.1 explicitly acknowledge independent calibration snapshots between methods, and the conclusion admits transient hardware noise. The strongest form of the claim—'profound differences' between mathematically equivalent implementations—requires that differences in Fig. 4 be caused by the implementation, not by when the job ran. A fair test would hold calibration fixed or quantify its drift. The paper's own admission that Pauli no longer implements a true depolarizing channel after transpilation also means the hardware comparison changes two variables at once (implementation and calibration). My proposed interleaved-snapshot check would settle whether the effect is real. Because the issue is addressable and the paper is otherwise transparent, the reader's CONDITIONAL verdict remains appropriate; no verdict adjustment is needed.","tokens_in":6891,"tokens_out":3787,"duration_ms":42665,"concrete_test":"Re-run the hardware comparison (ideal vs QPD vs Pauli at 99.6% fidelity) with all methods interleaved within a single calibration snapshot, or execute each method back-to-back and repeat over several calibration snapshots. Compute the variance in edge fidelities attributable to calibration cycle (same method, different snapshots) versus method (same snapshot, different methods). If the method-related variance is not significantly larger than the calibration-related variance, the central claim is weakened. Also report which calibration snapshot/date each point in Fig. 4b corresponds to.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central experimental claim—that mathematically equivalent depolarizing implementations diverge on hardware—requires that observed differences be attributable to the depolarization method. Section 3.1 and Section 4.1 state separately: 'separate runs for different depolarization methods each retrieve an independent calibration snapshot, so the noise model may differ slightly between methods,' and Section 4.1 explains even a depolarized run outperforming the ideal case as due to fresh calibration snapshots. The conclusion then admits 'the results here are influenced by transient hardware noise.' Because QPD and Pauli runs are executed at different times with different calibration data, the between-method spread in Fig. 4 could be dominated by calibration drift rather than by the depolarization implementation. This is especially dangerous for the headline QPD-vs-Pauli comparison: the Pauli method is also admitted in Section 4.2 to no longer implement a true depolarizing channel after transpilation, so two variables change at once. Without controlling calibration, the 'profound differences' claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper demonstrates a framework for emulating a network-distributed entanglement link on a single quantum processor. It realizes a virtual CZ gate and a ring graph-state edge using a cut Bell pair, following Carrera Vazquez et al., and compares three mathematically equivalent depolarizing-channel implementations (Stinespring unitary, random Pauli, and quasi-probability decomposition). The methods are benchmarked in simulation on an IBM Torino noise model, with the Pauli and QPD variants also run on IQM Emerald hardware. The paper sweeps the input Bell-pair fidelity and the classical communication delay, reporting lower-bound graph-state fidelities from entanglement witnesses. The central claim is that, despite mathematical equivalence of the noise models, hardware constraints cause 'profound differences' in results, with QPD depolarization performing best, while Pauli depolarization degrades significantly after transpilation. The paper also reports that on hardware, successful graph-state creation requires network-distributed fidelity of at least 90%, and that classical delays only become detrimental at tens of kilometers.","tokens_in":7136,"tokens_out":2935,"duration_ms":34846,"significance":"If the central claim were fully established, the paper would provide a useful, cautionary message for distributed-quantum-computing emulation: the choice of how to implement an abstract noise model can materially change conclusions drawn from hardware experiments, and QPD offers a low-gate-overhead alternative. The paper is also valuable for its concrete 60/84-circuit measurement workflow, its use of witness-based fidelity lower bounds, and its transparent admission of limitations. However, the main comparative claim is currently underdetermined by the reported data because of calibration-snapshot confounding and because the Pauli implementation is admitted not to realize a depolarizing channel after transpilation. These issues are load-bearing: without controlling them, the 'profound differences' headline reduces to a comparison of uncompiled vs. compiled operations under differing hardware conditions. The delay-sweep conclusion is also not reproducible because the thermal-relaxation parameters are not reported. The paper's significance is therefore conditional on additional controlled experiments or substantially weakened claims.","major_comments":[{"comment":"The comparison of depolarization methods is confounded by the use of independent calibration snapshots. Sec. 3.1 states that 'separate runs for different depolarization methods each retrieve an independent calibration snapshot,' and Sec. 4.1 repeats that a depolarized run can outperform ideal because each run receives 'a fresh calibration snapshot.' The conclusion then admits 'the results here are influenced by transient hardware noise.' Since the QPD and Pauli runs are executed at different times, the between-method spread in Fig. 4 could be dominated by calibration drift rather than by the depolarization implementation. This directly undermines the central 'profound differences' claim. A control experiment with interleaved runs under a single calibration snapshot, or a paired statistical comparison with reported calibration data, is needed.","section":"Sec. 3.1 and Sec. 4.1"},{"comment":"The Pauli 'depolarizing' implementation on hardware is no longer a depolarizing channel after transpilation. Sec. 4.2 explains that the required Hadamard operations create unequal gate counts and accumulated error across Pauli operators, so 'the implemented operation therefore no longer represents a true depolarizing channel.' Thus Fig. 4 does not compare two mathematically equivalent noise models: it compares QPD depolarization with an arbitrarily compiled Pauli circuit that has a different effective channel. The abstract's claim that 'mathematically equivalent' implementations diverge is therefore not supported by the hardware data as presented. Either the Pauli implementation should be compiled in a way that preserves the depolarizing structure, or the implemented channel should be characterized (e.g., by process tomography) and the claim reframed.","section":"Sec. 4.2"},{"comment":"The classical-communication-delay sweep is not reproducible because the thermal-relaxation parameters T1 and T2 used in the simulation are not reported. Sec. 3.1 says that classical delays are modeled 'using thermal relaxation' and Sec. 4.4 claims delays become detrimental only at 'tens of kilometers.' This threshold depends exponentially on the assumed T1/T2 values and on the per-gate timing model. Without these parameters, the quantitative conclusion—and even the order-of-magnitude threshold—cannot be verified or compared with other hardware models. Please report the T1/T2 values, the gate-duration model, and the mapping from delay to thermal-relaxation error.","section":"Sec. 4.4 and Sec. 3.1"}],"minor_comments":[{"comment":"The text says that on hardware 'successful graph-state creation' occurs at network-distributed fidelities of 0.9 and 1, but Fig. 5b apparently shows only data points at those values. Clarify whether 1.0 is a physically meaningful input or a normalization point, and whether the statement refers to the lowest successful point or to all points above threshold.","section":"Sec. 4.3 / Fig. 5"},{"comment":"The relationship between 'Runs per Observable' in Table 2 (0 for ideal/Pauli, 2 for QPD) and the total circuit counts (60 vs. 84) is not immediately clear. Define the counting convention explicitly, including how the 12 observables combine with QPD sampling terms.","section":"Table 2 / Sec. 3"},{"comment":"The simulation uses an IBM Torino noise model while the hardware experiments use IQM Emerald. The architectural mismatch is acknowledged, but the paper would benefit from a table comparing gate sets, error rates, and connectivity for the two platforms, since the central claims depend on hardware-specific behavior.","section":"Sec. 3.2 / Sec. 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper leans heavily on the first author's master's thesis [13] for the depolarization circuit details and the witness formalism. For a journal publication, the essential details should be self-contained or the thesis should be publicly accessible. The calibration-snapshot issue is not merely a presentation weakness; it is the main threat to the paper's central claim. The authors' own conclusion ('results here are influenced by transient hardware noise') is an explicit acknowledgment that the headline comparison is not yet controlled. If the authors can add interleaved calibration-controlled runs or temper the claim to 'with our specific calibration snapshots, the implementations gave different results,' the paper could be suitable after revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper has a real and useful observation—that on actual hardware you cannot treat three 'equivalent' depolarizing implementations as interchangeable—but the headline 'profound differences' claim is weaker than it looks because the comparison runs were taken under different calibration snapshots and one of the methods stops being a depolarizing channel after transpilation.\n\nWhat's actually new: this is the first side-by-side evaluation I know of that puts QPD, Pauli, and (in simulation) unitary depolarization through the same cut-Bell-pair graph-state emulation pipeline. The measured result that IQM Emerald needs network-distributed fidelity ≥90% to produce a ring graph state above the 0.5 witness threshold is a concrete data point for anyone building near-term DQC emulations. The delay sweep result—no appreciable degradation until ~1000 ns, failure only at tens of km—is also useful for data-center-scale planning.\n\nWhat the paper does well: the experimental design is honestly documented. They tell you that each method gets a fresh calibration snapshot, that the Pauli transpilation no longer realizes a true depolarizing channel, and that transient hardware noise influences results. That transparency is worth preserving.\n\nSoft spots, in order: (1) The central comparison is confounded. Separate calibration snapshots per method mean drift, not implementation, could drive the QPD-vs-Pauli spread in Fig. 4. The authors themselves point this out in Sections 3.1 and 4.1 and in the conclusion, so the abstract's 'profound differences' is not established. (2) The Pauli method after transpilation is not a depolarizing channel; comparing it to QPD is therefore comparing a broken implementation to a clean one. (3) QPD performing best is partly by construction since it introduces no gates; that's fine as an implementation note, but it isn't a discovery about the noise models. (4) No code/data released, and thermal relaxation parameters for the delay sweep aren't reported, so the thresholds are hard to reproduce.\n\nNone of these are fatal to the paper's usefulness as a cautionary empirical study. The right fix is to run methods with the same calibration snapshot or to explicitly sweep calibration, and to release the artifacts. As written, it deserves a serious referee, mainly because the confound is addressable and the measurement is freshly useful.\n\nI'd want to see the revision before citing it.","headline":"Useful empirical caution about emulating network noise on hardware, but the 'profound differences' claim is undercut by calibration drift and one broken implementation; worth refereeing if the confound is addressed.","tokens_in":7637,"tokens_out":2605,"would_cite":false,"duration_ms":29497,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68","81P40","81P45"],"pacs":["03.67.-a","03.67.Hk"],"model":"deepseek-v4-flash","headline":"Mathematically identical ways of modeling noise in a quantum-network emulation produce different results on real hardware, and the quasi-probability implementation is the only one that stays close to ideal.","keywords":["distributed quantum computing","entanglement distribution","cut Bell pair","quasi-probability decomposition","depolarizing channel","graph states","teleportation","classical communication delay"],"falsifier":"Run all three depolarization implementations interleaved within a single calibration snapshot (or averaged over many calibration cycles) on the same hardware and check whether the quasi-probability-vs-Pauli fidelity gap persists; if the gap disappears or shrinks below shot noise, the paper's 'profound differences' conclusion is an artifact of calibration timing rather than the method.","tokens_in":6791,"feed_emoji":"⚛️","tokens_out":5672,"duration_ms":60483,"temperature":0.7,"pith_summary":"This paper uses a single quantum computer to emulate a distributed quantum network: it builds a ring graph state in which one edge is created through a teleported cut Bell pair instead of a physical connection. The authors model imperfect entanglement sources with three depolarizing channels that are mathematically equivalent (a Stinespring-dilated unitary, random Pauli errors, and a quasi-probability decomposition) and compare them in simulation and on hardware. The central finding is that the equivalence holds only in theory: hardware constraints flip the comparison, with the quasi-probability method performing best because it adds no gates, while Pauli-error depolarization fails to act as a true depolarizing channel. The paper also reports that successful graph-state creation on hardware requires pre-distributed entanglement fidelity of at least 90%, and that classical communication latency only becomes detrimental at tens of kilometers. These results matter because they tell algorithm designers which noise-emulation choices are trustworthy when developing distributed quantum computing protocols on today's devices.","feed_headline":"Equivalent noise models diverge on real quantum hardware","feed_subtitle":"The gate-free quasi-probability method stays true to the model; Pauli errors break down in practice.","key_machinery":"The load-bearing object is the cut Bell pair: a Bell state |Φ+⟩ written as a quasi-probability distribution (QPD) over separable states ρ+_k and ρ−_k with weights summing to one but taking negative values. Each cut pair is used to implement a teleportation-based virtual controlled-Z gate across a previously disconnected edge of a six-qubit linear graph state, turning it into a ring. Depolarizing noise is applied to the cut Bell pair in three equivalent forms—a unitary Stinespring dilation, randomly applied Pauli gates, and the QPD itself—and the witness-based fidelities of the resulting graph state are measured using 12 observable circuits combined over QPD terms (60–84 runs total). The mach","core_discovery":"On its own terms, the paper establishes that the three depolarizing implementations are not interchangeable on a physical quantum processor. Despite being mathematically equivalent channels, the unitary, Pauli, and quasi-probability versions diverge because hardware transpilation adds different numbers of gates, changes qubit maps, and introduces calibration drift between runs. The quasi-probability decomposition, which adds no gates, consistently matches simulation and the ideal no-noise case, whereas the Pauli implementation requires extra Hadamard gates that break the depolarizing property, producing fidelities below the entanglement threshold. The paper directly claims that 'the more err","pith_inferences":["If the paper's hardware-vs-simulation divergence is general, then any noise-model comparison in quantum-network emulation should be conducted under a single calibration snapshot or averaged over many, since the paper's own runs used separate snapshots that could bias the 'profound differences' conclusion.","The discrepancy between the Pauli and quasi-probability methods on hardware could be repurposed as a probe of transpilation-induced error: the extra gate depth is measurable, and process tomography of the effective channel would test whether it is truly non-depolarizing.","A natural testable extension is to repeat the fidelity sweep with standard error-mitigation techniques; if the 90% threshold drops, the requirement is a hardware-noise artifact rather than a fundamental property of teleportation-based graph-state construction.","The paper's latency results suggest that classical-communication delays are not the bottleneck for data-center-scale distributed quantum computing; the more pressing constraint is the quality of pre-distributed entanglement, shifting research attention to source fidelity and purification."],"forward_implications":["If the results hold, quantum-network emulation on current hardware should use gate-free depolarization (QPD) to avoid implementation artifacts.","Network-distributed entanglement fidelity must be ≥90% on near-term hardware for a usable teleportation-built graph state, versus only 60% in an ideal-noise simulation.","Data-center-scale interconnects (meters of fiber, ≤100 ns) are safe from classical-communication-induced dephasing; effects only appear at hundreds of meters and break down at tens of kilometers.","The directional asymmetry of the virtual CZ (control qubit 5, target qubit 0) consistently degrades the target edge more, a feature future distributed-graph-state protocols must budget for.","The framework extends a single-QPU cut-Bell-pair construction to a hardware benchmark, establishing a workflow for testing future network-enabled distributed quantum computing algorithms."],"fun_headline_variants":["Equivalent noise models split on quantum hardware","Mathematically identical error channels diverge on real quantum chips","Quasi-probability keeps noise models true on real hardware","Hardware shows equivalent depolarizing models are not equal","Matched noise models fail to match on quantum processors"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central comparison assumes that differences between the three depolarization methods are not caused by calibration drift between the separate hardware runs, even though the paper acknowledges 'the noise model may differ slightly between methods.'","fun_headline_variants_meta":{"raw":{"variants":["Equivalent noise models split on quantum hardware","Mathematically identical error channels diverge on real quantum chips","Quasi-probability keeps noise models true on real hardware","Hardware shows equivalent depolarizing models are not equal","Matched noise models fail to match on quantum processors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000562,"raw_usage":{"total_tokens":2448,"prompt_tokens":628,"completion_tokens":1820,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":372,"completion_tokens_details":{"reasoning_tokens":1758}},"tokens_in":372,"tokens_out":1820,"duration_ms":14807,"temperature":1.0,"reasoning_tokens":1758,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:37:15.173358+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run all three depolarization implementations interleaved within a single calibration snapshot (or averaged over many calibration cycles) on the same hardware and check whether the quasi-probability-vs-Pauli fidelity gap persists; if the gap disappears or shrinks below shot noise, the paper's 'profound differences' conclusion is an artifact of calibration timing rather than the method.","supporting_citations":[],"review_version":1}