{"id":"57c53f4f-37af-401b-87dc-dc0a7ccd6e5c","arxiv_id":"2411.19308","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A per-qubit-pair pulse profiling and parallel calibration protocol lowers two-qubit gate errors and doubles quantum volume on 127-qubit IBM processors.","lead":"Quantum computers use microwave pulses to perform two-qubit operations, and this paper shows that picking a different pulse shape for each pair of qubits, instead of using one shape everywhere, makes IBM's 127-qubit machines more accurate. The authors report roughly half the two-qubit gate error, double the quantum volume, and shorter calibration time on real hardware.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 1.84x median-error claim is not anchored to the deployed policy assignments; the only policy-level data show representative generalization missing the optimal waveform on 5/21 or 2/21 pairs, and the abstract/body medians disagree.","rationale":"The reader's verdict is CONDITIONAL, and this stress-test identifies a concrete condition that must be met before the central claim can be accepted as stated: the 1.84x median improvement must be shown to hold under the actual policy-based waveform assignments, not under the idealized per-pair optimal selection. The reader's weakest assumption is the same one: representative generalization is imperfect, so 'optimal profiling for each qubit pair' is not actually achieved on all pairs. The paper's own Figure 12 admits failures on 5/21 and 2/21 pairs, which makes it essential to report the aggregate median under the deployed policies. The additional mismatch between the abstract median (0.006) and the body median (4.4e-3) reinforces that the headline number is not currently reproducible from the paper's text. Because the concern is about the anchoring and fair comparison of the reported improvement rather than about the existence of real device-level gains, the appropriate resolution is to keep the conditional verdict and require either a reanalysis of existing data or a clear statement of the exact population and protocol behind each median. No change of verdict is recommended beyond the reader's conditional acceptance.","tokens_in":18098,"tokens_out":7699,"duration_ms":73375,"concrete_test":"Using the raw IRB data behind Figure 12, compute for each of the 21 pairs the error of the waveform that each policy would actually assign, pair it with the default Echoed CR error for the same pair measured in the same calibration window, and calculate the paired median ratio default/policy. Check whether that ratio is 1.84x and whether the resulting policy median is 4.4e-3 (or 0.006, as in the abstract). If the policy-assignment ratio is materially below 1.84x or the medians cannot be reproduced, the central Section V-B claim is not supported by the provided data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central assertion in Section V-B is that after applying the full calibration process, the median two-qubit gate error is 4.4e-3, a 1.84x improvement over IBM's default pulse configuration. For this to support the paper's headline, the comparison must be between the waveform actually assigned by the proposed policies and the default Echoed CR waveform, measured on the same pairs in the same time window. The paper never reports that comparison. The only gate-level data with explicit policy assignments are the 21 pairs in Figure 12, where the paper's own results show Brute-force Clustering selects a non-optimal waveform on 5 of 21 pairs and Topology-oriented Representative on 2 of 21 pairs. Moreover, the Section V-B median is not tied to a stated population, policy, or measurement protocol, and the abstract reports a different optimized median (0.006) than the body (4.4e-3). If the 1.84x factor was computed from the individually optimized 'optimal' assignments used to construct Figure 12, rather than from the deployed representative-based policy assignments, then the central fidelity improvement is overstated even if the device-level metrics are real.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a fine-grained calibration protocol for two-qubit ECR gates on IBM heavy-hex superconducting processors. It enlarges the pulse candidate set to three waveforms (echoed CR, multi-derivative DRAG, direct CR), introduces three policies for profiling each qubit pair's optimal waveform (Brute-force Clustering, Topology-oriented Representative, Hardware-oriented Policy), and adds a graph-based parallel calibration scheme. The protocol is evaluated on 127-qubit machines (ibm_rensselaer, ibm_nazca, ibm_strasbourg) with gate-level IRB measurements, calibration-time measurements, device-level QV and EPLG benchmarks, and application-level circuits. The headline claims are a 1.84x reduction in median two-qubit gate error, a doubling of quantum volume, up to 2.3x reduction in EPLG, and that the minimum error falls below a QEC threshold.","tokens_in":18447,"tokens_out":3795,"duration_ms":30260,"significance":"If the central claims hold, this is a practically useful engineering contribution: it demonstrates per-qubit-pair pulse profiling with multiple waveform types on real IBM hardware, and it backs the claims with machine-measured gate-level and device-level benchmarks. The paper deserves credit for reporting IRB measurements, QV and EPLG results in Table I, and for including calibration-time measurements in Figure 14. It also honestly discloses in Figure 12 that the representative-based policies do not always select the optimal waveform. The main weaknesses are that the headline median-improvement claim is not anchored to the deployed policy assignments, and the QEC-threshold conclusion is based on the minimum rather than the typical error rate.","major_comments":[{"comment":"The abstract states an optimized median error of 0.006, while Section V-B states the median is reduced to 4.4e-3; these numbers are inconsistent and should be reconciled. More importantly, the Section V-B claim that the median is reduced to 4.4e-3, representing a 1.84x improvement over IBM's default pulse configuration, is not accompanied by the measurement protocol: the reader is not told which qubit pairs were included, whether the calibrated median was computed under the actual policy assignments or under the individually optimized waveforms used to construct Figure 12, or how the default-window comparison was controlled for drift. Please specify the population, the deployed waveform assignment, and the exact comparison that produced the 1.84x factor.","section":"Section V-B and Abstract"},{"comment":"The statement that the quantum error correction code has entered the region where errors are suppressed relies on the minimum achieved error of 1.3e-3 being below the 3e-3 threshold from reference [4]. A single pair below threshold does not support a claim that the code operates below threshold; the relevant quantity for a threshold argument is the typical or worst-case physical error rate of the gates used by the code, and the reported median of 4.4e-3 is above 3e-3. Additionally, the applicability of the specific threshold value from [4] to IRB-measured ECR error rates on this hardware is not justified. The threshold claim should be removed or heavily qualified.","section":"Section V-B (QEC threshold claim)"},{"comment":"The protocol's generalization assumption—that a waveform optimized on a representative pair remains optimal for other pairs in the same cluster or heavy-hex unit-cell position—is violated on a nontrivial fraction of the tested pairs: Figure 12a shows Brute-force Clustering missing the optimal waveform on about 5 of 21 pairs, and Figure 12b shows Topology-oriented Representative missing it on about 2 of 21 pairs. This contradicts the Section V-B statement that the fine-tuned protocol 'can achieve almost the optimal error rate on all qubit pairs.' Please quantify the fidelity penalty from these misassignments and state explicitly whether the headline median error and device-level metrics include these penalties.","section":"Section IV-B2, IV-B3, and Figure 12"}],"minor_comments":[{"comment":"The word 'medium' is used where 'median' is meant (Abstract and Section V-B); the abstract also contains the typo 'calibraton' and 'server' should be 'serve.'","section":"Abstract and Section V-B"},{"comment":"The abstract reports a minimum gate error of 0.001, while Section V-B reports a minimum of 1.3e-3; these should be made consistent.","section":"Abstract vs Section V-B"},{"comment":"The phrase 'the total pulse duration is reduced by a factor of 1.26' is ambiguous: it could mean duration becomes 1/1.26 of the original or becomes 1.26 times the original; please clarify.","section":"Section V-B"},{"comment":"The sentence 'However, QPT's bad scalability and it's not applicable to larger systems' is grammatically incomplete; it should be rephrased.","section":"Section II-A"},{"comment":"The normalization used in Figure 13 is not fully defined; the text explains the optimal calibration cost but does not specify how 'optimal' values are determined for gate error and gate duration, nor how the normalized sums are computed.","section":"Figure 13"},{"comment":"The description 'all single paths with a length of one are selected from the coupling map' is unclear and should be rephrased to 'all edges of the coupling graph are selected' if that is the intended meaning.","section":"Section IV-C"},{"comment":"The claim of 'the first large-scale implementation of multi-derivative DRAG and direct CR operations on real quantum machines' is strong; since reference [24] reports experimental results for multi-derivative DRAG, the novelty claim should be qualified to avoid overstatement.","section":"Section I (Contributions)"}],"recommendation":"major_revision","confidential_remarks":"The paper contains genuine empirical results on real hardware and addresses a relevant calibration problem. The main risk is that the central quantitative claims are currently presented in a way that could mislead readers: the median-improvement figure is not methodologically anchored to the policy that was actually deployed, and the QEC-threshold conclusion is based on the minimum error. These issues are fixable within the manuscript's scope by adding the missing measurement protocol and softening the threshold claim, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline first: this is a genuinely useful engineering paper, with real hardware results on 127-qubit IBM processors (QV 128→256, EPLG roughly halved, consistent application-level gains). The genuinely new thing is the integrated protocol—representative-based waveform profiling across three pulse candidates (echoed CR, multi-derivative DRAG, direct CR) plus a graph-parallel calibration schedule. The pulse shapes themselves are known, and the paper properly cites [24] for multi-derivative DRAG; the contribution is the profiling policies and the first large-scale deployment on real machines. That is worth a serious look.\n\nCredit: the device-level and application-level measurements are field data, not extrapolation. The parallel calibration design is clean, and the reported 7.9x real speedup (25x ideal) includes an honest caveat about hardware limits forcing subgraph splitting. Figure 12 is also useful: it shows the authors know their representative policies do not always pick the optimal waveform—5/21 misses for brute-force clustering, 2/21 for topology-oriented—which is the kind of self-critical data that builds trust.\n\nThe soft spots are real but fixable. The headline gate-error claim is under-anchored: the abstract says median 0.006, the body says 4.4e-3 (and calls it 'medium'). The 1.84x improvement is not tied to a stated qubit-pair population, the actual waveform assignments deployed, or a measurement window; it is not the same as the Figure 12 policy-level comparison. The QEC threshold conclusion is overstated—they cite their minimum error (1.3e-3) below the 3e-3 threshold, but the median sits above it, so 'errors are suppressed' is too strong. Missing code/data (pulse templates, per-pair error tables, calibration scripts) hurts reproducibility. These are reportable issues, not a fatal flaw.\n\nOverall: the central engineering claim—that per-pair profiling plus parallel calibration improves real-device performance—probably holds, but the paper as written does not pin down the headline numbers tightly enough to trust the exact factors. A careful revision with explicit statistics, a same-pairs before/after comparison, and a corrected QEC claim would make this a solid systems paper. I'd send it to peer review, and I'd expect conditional acceptance after major revision. For my own work I'd cite it once the data are available.","headline":"Real-machine calibration protocol with three pulse waveforms and three profiling policies shows real improvements on 127-qubit IBM devices, but the headline median-error claim is under-specified and the QEC threshold conclusion is overstated.","tokens_in":18872,"tokens_out":3236,"would_cite":true,"duration_ms":29402,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68","81-05"],"pacs":["03.67.-a","85.25.-j"],"model":"deepseek-v4-flash","headline":"This paper establishes that profiling each qubit pair with its own optimal two-qubit pulse waveform—selected from three candidates—reduces median gate error by 1.84x and doubles quantum volume on 127-qubit IBM processors.","keywords":["pulse profiling","cross-resonance gate","multi-derivative DRAG","direct CR","heavy-hex topology","parallel calibration","quantum volume","two-qubit gate error"],"falsifier":"Measure all three candidate waveforms with interleaved randomized benchmarking on every pair of a 127-qubit heavy-hex processor and compare the policy-assigned waveform to the measured best waveform. If the assigned waveform is not the lowest-error one for most pairs, the claim of per-pair optimal profiling fails; the paper's own Figure 12 already reports misses on 5 of 21 pairs with Brute-force Clustering and 2 of 21 with Topology-oriented Representative.","tokens_in":17943,"feed_emoji":"⚛️","tokens_out":10490,"duration_ms":83683,"temperature":0.7,"pith_summary":"This paper tries to establish that two-qubit gate calibration on large superconducting processors should be done per qubit pair, not with one pulse shape for all pairs. It enlarges the candidate set for the echoed cross-resonance (ECR) gate to three waveforms—echoed CR, multi-derivative DRAG, and direct CR—and assigns each pair the waveform best suited to its physical properties, heavy-hex topology position, and coherence limits. On 127-qubit IBM processors the protocol reports a median two-qubit gate error of $4.4\\times 10^{-3}$, a 1.84x improvement over the default pulse configuration, a doubling of quantum volume from 128 to 256, and up to a 2.3x reduction in error per layered gate. The payoff of the claim is that existing hardware can be pushed closer to the error threshold where quantum error correction begins to suppress errors, without waiting for new chip designs.","feed_headline":"Per-qubit pulse profiling cuts two-qubit gate error 1.84x","feed_subtitle":"Choosing one of three pulse shapes per pair doubles quantum volume on 127-qubit processors.","key_machinery":"The load-bearing object is the echoed cross-resonance (ECR) two-qubit gate, realized by any of three microwave pulse waveforms: the standard echoed CR pulse, a multi-derivative DRAG pulse that suppresses multiple transition errors through recursive derivative corrections, and a direct CR pulse that is shorter but costlier to calibrate. The profiling step maps each qubit pair to its preferred waveform using either physics-based clustering, heavy-hex unit-cell position, or hardware knowledge such as frequency detuning and coherence times. The parallel-calibration step partitions the coupling graph $\\mathcal{G}$ into calibration subgraphs in which concurrently calibrated edges are separated by a minimum graph distance of two, enabling up to 38 pairs to be calibrated at once on the 127-qubit heavy-hex layout. Together, waveform assignment plus subgraph-parallel calibration is the mechanism that converts per-pair optimization into processor-wide gate-error reduction.","core_discovery":"The central claim is that per-qubit-pair pulse profiling—selecting, for each coupled pair, one of three calibrated ECR waveforms—delivers fidelity gains that survive at the device level, not just on individually tuned pairs. The paper demonstrates this with three profiling policies: a clustering policy grouping pairs by frequency detuning, coupling strength, and anharmonicity; a topology policy exploiting the repeating unit-cell structure of the heavy-hex lattice; and a hardware-oriented policy that also accounts for qubit-qubit detuning ranges and $T_1$/$T_2$ decoherence times. After calibration, the median two-qubit gate error on the real machine drops to $4.4\\times 10^{-3}$, a 1.84x improvement relative to the vendor default, while the minimum error reaches $1.3\\times 10^{-3}$. On two 127-qubit processors the quantum volume doubles from 128 to 256 and the error per layered gate falls by factors of 2.3 and 1.99. The paper reads these results as evidence that the protocol captures hardware-specific differences that uniform calibration misses, and that this is an immediate step toward operating below the error-correction threshold.","pith_inferences":["The paper's own benchmarking shows the representative-based profiling is not truly per-pair optimal: 5 of 21 pairs receive a non-optimal waveform under Brute-force Clustering and 2 of 21 under Topology-oriented Representative. A hybrid protocol that spot-checks a second waveform on outlier pairs would likely close most of that gap.","The three waveforms trade fidelity against calibration cost and duration, so the optimal assignment will drift as qubits age; an online re-profiling step that periodically re-measures a small sample of pairs could maintain the reported gains over time.","The profiling logic is tied to cross-resonance gates on fixed-frequency superconducting qubits with heavy-hex layouts; porting it to tunable qubits or other two-qubit gate families would require re-deriving the feature set and the waveform candidates, though the subgraph-parallelization idea transfers directly.","The parallel-calibration speedup is bounded by hardware limits on concurrent arbitrary waveforms (the paper splits subgraphs larger than 20 edges into groups of no more than 10), so the ideal 25x reduction would need controller hardware that handles many custom pulses simultaneously."],"forward_implications":["If the reported numbers hold, a calibration pass that selects among three pulse waveforms can push 127-qubit processors from quantum volume 128 to 256 without any change in chip fabrication.","The minimum two-qubit gate error of $1.3\\times 10^{-3}$ lies below the $3\\times 10^{-3}$ threshold the paper cites for error correction on the heavy-hex lattice, so the protocol moves surface-code operation closer to the error-suppression regime.","Because calibration subgraphs allow up to 38 pairs to be calibrated concurrently, routine recalibration becomes affordable: the paper reports a 7.9x reduction in calibration wall-clock time in practice and up to 25x under ideal hardware, meaning drift can be countered more often.","Application-level benchmarks on standard quantum circuits all show lower error rates and higher fidelities after calibration, with a maximum fidelity increase of 16%, indicating the gain reaches compiled user programs.","Shorter direct CR pulses reduce total pulse duration by a factor of 1.26 for pairs with fabrication defects or short coherence times, extending how many sequential entangling gates fit within the decoherence limit."],"supporting_citations":[{"why":"supplies the multi-derivative DRAG pulse-shaping method used as one of the three waveform candidates and the source of the recursive derivative corrections.","marker":"[24]"},{"why":"supplies the direct cross-resonance gate method that the direct CR waveform candidate is built on.","marker":"[8]"},{"why":"provides the pulse-level programming and calibration library used to implement the echoed CR waveform on real hardware.","marker":"[2]"},{"why":"supplies the Hamiltonian tomography procedure used to measure and cancel unwanted terms during CR calibration.","marker":"[41]"},{"why":"provides interleaved randomized benchmarking, the method used to measure the reported two-qubit gate error rates.","marker":"[28]"},{"why":"supplies the two-qubit gate error threshold ($3\\times 10^{-3}$) the paper uses to argue the calibrated error rate enters the error-correction regime.","marker":"[4]"},{"why":"documents the repeating unit-cell structure of the heavy-hex lattice that the Topology-oriented Representative policy relies on.","marker":"[34]"}],"fun_headline_variants":["Per-pair pulse choice cuts gate error 1.84x","Three pulse shapes per pair double quantum volume","Fine-grained pulse profiling: 1.84x error cut, 2x volume","Optimal per-pair pulses slash error, double volume","Choose one of three waveforms per qubit pair: 2x volume"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The protocol assumes that a pulse calibrated on one representative qubit pair remains the optimal choice for every other pair assigned to the same cluster or the same heavy-hex unit-cell position.","fun_headline_variants_meta":{"raw":{"variants":["Per-pair pulse choice cuts gate error 1.84x","Three pulse shapes per pair double quantum volume","Fine-grained pulse profiling: 1.84x error cut, 2x volume","Optimal per-pair pulses slash error, double volume","Choose one of three waveforms per qubit pair: 2x volume"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1432,"prompt_tokens":979,"completion_tokens":453,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":364}},"tokens_in":595,"tokens_out":453,"duration_ms":4369,"temperature":1.0,"reasoning_tokens":364,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:19:07.539347+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure all three candidate waveforms with interleaved randomized benchmarking on every pair of a 127-qubit heavy-hex processor and compare the policy-assigned waveform to the measured best waveform. If the assigned waveform is not the lowest-error one for most pairs, the claim of per-pair optimal profiling fails; the paper's own Figure 12 already reports misses on 5 of 21 pairs with Brute-force Clustering and 2 of 21 with Topology-oriented Representative.","supporting_citations":[{"cited_title":"Experimental error suppression in Cross-Resonance gates via multi-derivative pulse shaping","cited_arxiv_id":null,"evidence_quote":"supplies the multi-derivative DRAG pulse-shaping method used as one of the three waveform candidates and the source of the recursive derivative corrections."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the direct cross-resonance gate method that the direct CR waveform candidate is built on."},{"cited_title":"Qiskit pulse: programming quantum computers through the cloud with pulses","cited_arxiv_id":null,"evidence_quote":"provides the pulse-level programming and calibration library used to implement the echoed CR waveform on real hardware."},{"cited_title":"Chow, and Jay M","cited_arxiv_id":null,"evidence_quote":"supplies the Hamiltonian tomography procedure used to measure and cancel unwanted terms during CR calibration."},{"cited_title":"Ketchen, and M","cited_arxiv_id":null,"evidence_quote":"provides interleaved randomized benchmarking, the method used to measure the reported two-qubit gate error rates."},{"cited_title":"The IBM Quantum heavy hex lattice | IBM Quantum Computing Blog","cited_arxiv_id":null,"evidence_quote":"documents the repeating unit-cell structure of the heavy-hex lattice that the Topology-oriented Representative policy relies on."}],"review_version":1}