{"id":"a74f67f9-f035-476b-801a-7426e83751b3","arxiv_id":"2607.27111","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Duffing-derived gate corrections that look excellent inside the Duffing model transfer poorly—and can even worsen error—when applied to an independently calibrated diagonalized-transmon baseline.","lead":"Corrections designed with the common Duffing model of a transmon qubit often fail when checked against a more faithful diagonalized model, especially for fast gates. The result warns that open-loop pulse design is only as good as the driven Hamiltonian used to predict multilevel errors.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The internal hybrid result is well supported, but its practical significance rests on treating a four-level RWA diagonalized-transmon model as the device proxy without checking it against the full lab-frame circuit dynamics.","rationale":"I agree with the Reader’s identification of the weakest assumption. Read narrowly, the paper’s strongest claim is a controlled model-comparison result: the same Magnus construction gives effective self-consistent corrections in either model, while Duffing-derived correction fields transfer poorly to an independently calibrated H_diag baseline and can increase fast-gate error. The design is reasonably careful: model-specific n_0,1 calibration removes a trivial Rabi-rate mismatch, the hybrid is compared at the same physical gate time rather than the same dimensionless time, and the four-level truncation is checked for both uncorrected and corrected evolution. The AC Stark comparison also provides a plausible mechanism rather than the sole evidence.\n\nWhat remains missing is validation of H_diag against the less-restricted circuit Hamiltonian from which it was derived. Figs. 2 and 6 are useful convergence tests, but they are convergence within an already RWA, adjacent-coupling, ideal-drive model. Since the practical warning concerns very small coherent residuals in the fast-gate regime, common-mode omitted terms are potentially decision-relevant. This does not overturn the paper’s internal numerical result, and the manuscript explicitly frames its conclusions as coherent model dependence under an ideal waveform. It does, however, keep the contribution at theory/numerics level rather than hardware guidance. The Reader’s CONDITIONAL verdict and MODERATE confidence therefore remain appropriate; no adjustment is warranted.","tokens_in":24734,"tokens_out":4704,"duration_ms":209449,"concrete_test":"Repeat the fast-gate cases, especially E_J/E_C=30 and |alpha_2|t_f≈5.74, by propagating the lab-frame driven cosine Hamiltonian of Eq. (3) without RWA or an adjacent-matrix-element restriction, retaining enough transmon levels for convergence while keeping V(t) ideal to isolate Hamiltonian effects. Independently calibrate the baseline in this full model, then evaluate the uncorrected baseline, the transferred Duffing correction, and the full-model self-consistent correction. If the hybrid remains no better than—or worse than—the full-model baseline with comparable error differences, the concern does not land; if the ordering changes or the discrepancies shift substantially, H_diag is not a sufficient proxy for the practical conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The main soft spot is the ground-truth status of H_diag, not an internal contradiction in Fig. 4. In Sec. II A, the “more faithful” Hamiltonian already makes several common-mode approximations: rotating-wave approximation, retention of only adjacent charge matrix elements, truncation to four levels, an ideal distortion-free voltage map, and purely coherent evolution. The convergence studies in Figs. 2 and 6 vary the number of levels within this same rotating-frame effective representation. They therefore show that four levels are enough for the retained dynamics, but cannot detect errors shared by the four-, eight-, and ten-level models.\n\nThis matters because the hybrid failure is a small-residual-error phenomenon driven by differences in leakage phases, AC Stark accumulation, and higher-level matrix elements. Counter-rotating and nonadjacent-drive terms are argued to be parametrically small, but the paper does not compare their contributions with the corrected error floors, which reach roughly 10^-5 to 10^-7. A contribution that is small relative to the uncorrected leakage error could still materially change the corrected error generator or the baseline-versus-hybrid ordering in the fast-gate regime. Thus the numerical claim about transfer from Duffing to H_diag remains convincing, while the stronger inference that this demonstrates failure for the physical driven transmon depends on unvalidated omissions from H_diag. The authors largely state this scope limitation themselves, so the concern supports conditionality rather than rejection.","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript studies how the choice of effective Hamiltonian affects open-loop, model-based pulse design for single-qubit gates in a capacitively driven transmon. Two four-level models are compared: a Hamiltonian obtained by diagonalizing the static transmon and expressing the charge drive in the dressed eigenbasis, and the standard Duffing (Kerr-oscillator) approximation. The same Magnus-based constrained-control construction (from the authors' prior work) is applied to both models, with the baseline pulse calibrated model-specifically through n0,1 to remove the trivial Rabi-rate mismatch. The central result (Fig. 4) is a hybrid test: Duffing-derived correction quadratures and constant detuning, added to an independently calibrated diagonalized-transmon baseline and evaluated at fixed physical gate time, underperform the self-consistent correction and at short gate times can increase the error over the uncorrected baseline. A relative mismatch in the accumulated AC Stark phase (Fig. 5), growing as EJ/EC decreases and tf shortens, is presented as an illustrative diagnostic. A sequential-transition estimate (Eq. 21, App. A, validated in Fig. 9) plus four- vs ten-level checks (Figs. 2, 6) support the truncation. Finally, three- vs four-level truncations change which correction strategy appears sufficient (Fig. 7), and a nonlinear construction yields a distinct, improved control solution (Fig. 8).","tokens_in":25117,"tokens_out":10593,"duration_ms":861086,"significance":"If the results hold, the paper makes a useful, quantitatively documented point for the superconducting-circuit control community: accuracy of an effective Hamiltonian for static low-energy quantities does not certify its predictions for driven error generators, and the discrepancy becomes acute precisely in the fast-gate regime the field is pursuing. The comparison design is careful and worth explicit credit: identical correction machinery in both models, model-specific n0,1 calibration that removes a trivial confound, and hybrid evaluation at fixed physical tf rather than fixed |α2|tf. The analytic truncation criterion (Eq. 21) is validated against direct propagation (Fig. 9), and the EJ/EC and gate-time trends in Figs. 4–5 constitute falsifiable predictions. The Sec. V–VI demonstration that Hilbert-space truncation determines which correction framework appears sufficient — and that a nonlinear construction finds a genuinely distinct pulse with the same controls — is a constructive methodological contribution beyond the headline negative result.","major_comments":[{"comment":"The RWA justification compares counter-rotating amplitudes (~1/(ωd tf)) to the *uncorrected* leakage error (~1/(|α2| tf)). But the paper's central quantity in Fig. 4(b) is the corrected error floor, ~10^-5–10^-7, and both models share the RWA and the adjacent-only charge-matrix-element truncation, so the convergence tests of Figs. 2 and 6 are blind to these common-mode omissions by construction. A contribution negligible against uncorrected leakage need not be negligible against the corrected floor, and it could in principle reorder baseline/hybrid/self-consistent curves at short tf. Concrete test: at the shortest |α2| tf and EJ/EC=30, propagate Eq. (6) in the lab frame (or a rotating frame retaining counter-rotating and nonadjacent terms, N=8–10) for the baseline, self-consistent, and hybrid protocols, and report whether the ordering and floors survive. Alternatively, provide perturbati","section":"§II.A.1 (RWA paragraph after Eq. (10)) and Fig. 4(b)"},{"comment":"The hybrid construction recalibrates n0,1 on the evaluation model but transfers Δ_Duff 'without further adjustment.' Calibrating the drive frequency against the (dressed) qubit transition is as routine experimentally as the amplitude calibration the paper does allow, so the asymmetry needs justification. This matters quantitatively: Fig. 5 identifies an accumulated-phase mismatch as a contributor, and a single-scalar Δ recalibration could absorb part of it, potentially changing the fast-gate regime where the hybrid 'can fail to improve... and can even increase' error. Please either (a) show the hybrid curve with Δ recalibrated by one scalar parameter on H_diag, demonstrating the fast-gate failure persists, or (b) argue explicitly why frequency recalibration falls outside the paper's definition of predictive design while amplitude calibration does not.","section":"§III.A, Eq. (36) and surrounding text"},{"comment":"The nonlinear framework is motivated by the short-gate-time saturation of the linear strategy (Fig. 7(b): 'especially at short gate times'), yet Fig. 8(a) shows the fourth-order nonlinear advantage (~1 order of magnitude) only for |α2| tf ≳ 7, with linear and nonlinear errors 'comparable' at shorter times. The claimed 'systematic route' to the shortest-time regime via higher-order nonlinear corrections is an expectation, not a demonstrated result. Please reconcile: either include a higher-order (e.g., sixth-order) nonlinear data point at short tf, or revise the §V framing so the demonstrated regime and the anticipated regime are clearly distinguished. This is load-bearing for the paper's second result, not the first.","section":"§V–VI, Figs. 7(b) and 8(a)"}],"minor_comments":[{"comment":"Neither caption states EJ/EC for the plotted traces (Fig. 5 varies it; these figures presumably fix one value). Also give the physical tf range in ns for context.","section":"Figs. 4, 7, 8 captions"},{"comment":"Report absolute device parameters (ωT/2π, EC/h, and the drive frequency). The RWA hierarchy |α2|/ωd ≪ 1 invoked in §II.A.1 cannot be checked from the dimensionless quantities alone, and experimental relevance (which tf in ns corresponds to |α2| tf = 5.74) requires them.","section":"§II and Fig. 4"},{"comment":"State the Fourier basis size and number of free coefficients used for gx, gy, and comment on conditioning/solution selection for the algebraic system from Eq. (35) (uniqueness, minimal-norm choice if underdetermined). A code/data availability statement would substantially strengthen reproducibility of the 10^-7-scale error floors.","section":"§III, after Eq. (35)"},{"comment":"The second matrix element is written ⟨0|Σ Ω(tf)|0⟩ with the Magnus-order subscript l missing on Ω. Also state which order m is used for the Fig. 5 curves ('up to fourth order' — is m=4 for all traces?).","section":"Eq. (37)"},{"comment":"Check the parenthesization of the n1,2 sector: as typeset, the grouping of the cos[θ/2](...) + sin[θ/2](...) terms is ambiguous about which σx/σy factors belong to which trigonometric prefactor.","section":"Eq. (15)"},{"comment":"§III restricts the detuning to be constant, and §VI states the nonlinear scheme uses 'the same constant detuning control,' yet panel (d) plots Δ/ωC against time. If each trace is a constant, say so in the caption; if the nonlinear construction relaxes the constant-detuning constraint, reconcile with §III.","section":"Fig. 8(d)"},{"comment":"The hybrid failure is closely related to the known model sensitivity of DRAG coefficients and optimal detunings in transmons (e.g., higher-derivative and FAST DRAG calibrations in Ref. [6]). A short paragraph connecting the present mechanism to that literature would help experimental readers place the result.","section":"§I or §III.A"},{"comment":"Consider noting explicitly that this is the √X-type gate used as the benchmark in Refs. [21, 25], to orient readers before Eq. (22)–(23).","section":"Eq. (23), 'complex Hadamard gate'"}],"recommendation":"major_revision","confidential_remarks":"The control methodology is taken from the authors' earlier Magnus-based papers [21, 25]; the novelty here is the controlled model-comparison/hybrid-transfer construction and the truncation-level analysis of Sec. V–VI, which I judge sufficient for this journal. The requested revisions are inexpensive given the authors' existing numerical infrastructure (they already propagate ten-level models; the lab-frame spot check and the one-parameter Δ recalibration are modest extensions), so major revision rather than rejection is appropriate. The manuscript is careful about its own scope in several places; the revision mainly asks the authors to make one quantitative step at the point where it matters most — the corrected error floors — rather than relying on the parametric hierarchy alone."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful result here is the hybrid protocol. Same Magnus correction machinery on both sides, baseline calibrated separately via each model’s n0,1, and only the Duffing gx, gy, and constant Δ transferred onto the diagonalized-transmon baseline. Inside Duffing the corrections crush the error; on H_diag they are much weaker and at short gates can sit at or above the uncorrected baseline (Fig. 4). That isolates multilevel error-generator mismatch rather than a trivial Rabi-rate offset. The AC Stark phase diagnostic (Fig. 5) and the three- vs four-level linear/nonlinear story (Figs. 7–8) are solid supporting pieces.\n\nWhat is actually new is not “models matter for open-loop control”—that is already in the citations—but the controlled transfer test, the quantified dynamical diagnostic, and the explicit link from truncation to which correction ansatz is even viable. Numerics are internally consistent: truncation checks for uncorrected and corrected dynamics, sequential-transition estimate backed by numerics in the appendix, and visibly different nonlinear pulse shapes. Math and citation pattern look fine; prior Magnus papers are tools, not the claim.\n\nSoft spot, in proportion: H_diag is labeled “more faithful” but still RWA, adjacent-only drive, four-level, ideal V(t), coherent only. Figs. 2 and 6 only vary level count inside that same rotating-frame object, so they cannot catch common-mode omissions. Corrected floors go to ~10^-5–10^-7; counter-rotating or nonadjacent pieces that are small vs uncorrected leakage could still reorder hybrid vs baseline. The authors mostly own this scope, so it conditions the hardware warning rather than killing the paper. No experiment, no code—moderate confidence is right.\n\nFor people who design open-loop pulses or pick effective models for transmons, this is worth reading. I would send it to referees. Engage if you care about model choice vs control ansatz; skip if you only want device-level validation.","headline":"Clean hybrid-transfer result: Duffing corrections that look great in-model can fail on a diag-calibrated baseline, with the main caveat that H_diag is still an idealized proxy.","tokens_in":25734,"tokens_out":521,"would_cite":true,"duration_ms":16380,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Correction pulses that work inside the Duffing model can fail when the drive is the real dressed charge operator of a transmon.","keywords":["transmon","quantum control","Duffing approximation","effective Hamiltonian","leakage","AC Stark shift","Magnus expansion","single-qubit gates"],"falsifier":"On a device with known EJ/EC, apply the hybrid protocol (Duffing-derived corrections on a device-calibrated baseline) at short dimensionless gate times and compare measured average gate error to both the uncorrected baseline and a self-consistent correction designed in the diagonalized model; if the hybrid systematically improves as much as the self-consistent pulse, the claimed dynamical mismatch is not operative.","tokens_in":25529,"feed_emoji":"⚡","tokens_out":867,"duration_ms":17431,"temperature":0.7,"pith_summary":"High-fidelity open-loop gates need a model that predicts the errors that actually accumulate under the drive, not only the undriven spectrum. This paper compares the standard Duffing (weakly anharmonic oscillator) Hamiltonian with a more faithful low-energy description obtained by diagonalizing the static transmon and expressing the physical charge drive in that eigenbasis. The same Magnus-based correction construction is applied to both. Corrections that sharply cut error inside the Duffing model transfer poorly onto an independently calibrated diagonalized-transmon baseline; at short gate times they can even raise error above the uncorrected baseline. Small, individually justified differences in level spacings and drive matrix elements compound into different dynamical error generators, with AC Stark phase mismatch as one clear diagnostic. The model also dictates which control framework is adequate: a three-level truncation hides channels that a four-level model exposes, motivating a nonlinear extension that still uses only the same laboratory controls.","feed_headline":"Duffing corrections can fail on a real transmon drive","feed_subtitle":"Small spectrum and charge-matrix mismatches compound into wrong multilevel pulses at short gates","key_machinery":"Hybrid transfer test: Duffing-derived correction envelopes and detuning are combined with a baseline pulse calibrated only in the diagonalized-transmon model (via its own n0,1), then evaluated under the diagonalized-transmon Hamiltonian at fixed physical gate time. The test isolates multilevel model mismatch from trivial Rabi-rate mismatch.","core_discovery":"Using identical Magnus-based correction construction, Duffing-derived correction quadratures and constant detuning substantially reduce average gate error inside the Duffing model, but when added to an independently calibrated diagonalized-transmon baseline they remain far less effective and, in the fast-gate regime, can fail to improve—or can increase—error relative to the uncorrected diagonalized-transmon baseline.","pith_inferences":["Closed-loop or data-driven recalibration may partly mask the hybrid failure, so the practical risk is largest for purely open-loop or first-shot pulse deployment.","The same compounding of small spectrum and matrix-element errors should appear in other weakly anharmonic platforms whenever the control operator is replaced by a harmonic proxy.","Including measured transfer functions or decoherence would likely widen, not close, the gap between Duffing-designed and dressed-charge-designed corrections."],"forward_implications":["Pulse libraries designed only in the Duffing approximation cannot be assumed to transfer to hardware without re-deriving multilevel corrections in a dressed-charge model.","Model selection for control must be judged by driven error generators (e.g., AC Stark phase), not only static spectra or matrix elements.","Hilbert-space truncation is part of control design: too few levels can make a linear correction strategy look sufficient when a fuller model shows it is not.","The same physical controls (two quadratures plus constant detuning) can reach a distinct, better solution once higher-order channels are kept and a nonlinear Magnus strategy is used.","As EJ/EC decreases, the cost of using the Duffing model for fast-gate design rises."],"fun_headline_variants":["Duffing fixes cut error in-model but fail on diagonalized transmons","Fast gates: Duffing corrections can raise error on real transmon drive","Spectrum and charge-matrix mismatch spoils transferred Duffing pulses","Same Magnus corrections diverge once the transmon basis is diagonalized","AC Stark phase tracks why Duffing-derived gates miss full-model dynamics"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That a four-level, rotating-wave, distortion-free diagonalized-transmon Hamiltonian is faithful enough to the real driven device that hybrid failure under this model means Duffing-based predictive design would fail in practice.","fun_headline_variants_meta":{"raw":{"variants":["Duffing fixes cut error in-model but fail on diagonalized transmons","Fast gates: Duffing corrections can raise error on real transmon drive","Spectrum and charge-matrix mismatch spoils transferred Duffing pulses","Same Magnus corrections diverge once the transmon basis is diagonalized","AC Stark phase tracks why Duffing-derived gates miss full-model dynamics"]},"model":"grok-4.5","effort":"low","cost_usd":0.00501,"raw_usage":{"total_tokens":1382,"prompt_tokens":766,"num_sources_used":0,"completion_tokens":96,"cost_in_usd_ticks":50104000,"prompt_tokens_details":{"text_tokens":766,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":520,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":766,"tokens_out":96,"duration_ms":8389,"temperature":1.0,"reasoning_tokens":520,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T11:08:20.674749+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a device with known EJ/EC, apply the hybrid protocol (Duffing-derived corrections on a device-calibrated baseline) at short dimensionless gate times and compare measured average gate error to both the uncorrected baseline and a self-consistent correction designed in the diagonalized model; if the hybrid systematically improves as much as the self-consistent pulse, the claimed dynamical mismatch is not operative.","supporting_citations":[],"review_version":1}