{"id":"701b137c-8857-4f24-be7e-ea44ebc4bbdb","arxiv_id":"2607.28364","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"A single FC1-calibrated noise model predicts staggered magnetization and post-quench population dynamics on three Pasqal neutral atom QPUs within Monte Carlo uncertainty envelopes.","lead":"This paper introduces a noise-aware emulation framework for neutral atom analog quantum computers, and shows that one noise model calibrated on a single Pasqal device reproduces annealing and quench measurements on three devices. A smart generalist might read it because it offers a practical route to trusting analog quantum processors whose outputs cannot always be checked classically.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"A posteriori detuning offset fitted to FC1 (Sec. VI.C) makes the 'single noise model, calibrated on one device' claim partially circular; cross-device transfer of this fitted parameter is the key unvalidated step.","rationale":"The reader's weakest_assumption focuses on MPS convergence of connected correlations in the strong-interaction regime (Sec. VI.D, App. E). That is a real limitation, but it affects mainly the detailed noise attribution for FC1 correlations (Fig. 5), not the cross-device validation, which relies on the staggered magnetization (converged at χ=128, Fig. 7) and the mean population (converged at χ=512 for both regimes, Fig. 8). The MPS truncation therefore does not directly threaten the headline cross-device claim. The a posteriori detuning offset in Sec. VI.C is more load-bearing because it concerns the parameter-free status of the model that is central to the claim. The offset is an extra fitted parameter; without a sensitivity analysis or a demonstration that the cross-device agreement holds for a zero offset, the claim 'calibrated on one device' is overstated. Since the paper itself acknowledges the offset is not known a priori, and the reader already marks the paper CONDITIONAL, my read supports an unchanged CONDITIONAL verdict.","tokens_in":26235,"tokens_out":4700,"duration_ms":41303,"concrete_test":"Re-run the Fig. 1(c,d) and Fig. 4 post-quench comparisons for FC1, SA1, and Ruby with the detuning offset fixed to zero (or to the Table II calibration uncertainty of +0.1 MHz) while keeping all other noise parameters fixed; count experimental points outside the 2.5-97.5% trajectory envelopes. If the cross-device agreement is preserved, the fitted offset is not load-bearing; if it is lost, the central claim depends on a parameter that was not predicted a priori.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. VI.C introduces a systematic detuning offset Δδ/(2π) = -0.2 MHz to the noise model, explicitly stated to be 'inferred a posteriori from the comparison itself' to remove residual early-time discrepancies in FC1 post-quench data (Fig. 4). This offset magnitude is twice the calibration precision quoted in Table II (≈100 kHz), so it is a genuine extra degree of freedom, not an uncertainty already contained in the model. The central claim of the paper (abstract, intro, conclusion) is that a single noise model calibrated on one device reproduces the behavior of all three QPUs without device-specific refitting. If the calibration includes a parameter chosen to match one of the validation datasets, the clean predictive claim is weakened: the FC1 agreement is partly a fit, not a prediction. The deliberately programmed +0.67 MHz offset is a good consistency check, but it tests the physical mechanism, not the uniqueness or necessity of the -0.2 MHz value. The paper does not show that SA1 and Ruby data agree with the model without this offset, nor does it quantify how much of the cross-device agreement depends on this single fitted value. If the same offset is applied to all devices, the transfer is still informative, but the claim should be restated as 'a single model with one parameter fitted to FC1 data transfers to the other devices.' Without that qualification, the evidence overstates the degree to which the noise model is parameter-free and predictive.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces a noise-aware emulation framework for neutral-atom analog quantum processing units, implemented in the open-source Pulser/emu-mps toolchain. The framework propagates thermal, laser, SPAM, decoherence, and systematic-bias noise through Monte Carlo wavefunction trajectories and predicts observables for two benchmark protocols: antiferromagnetic quantum annealing and post-quench Rydberg population dynamics. The authors validate the model against measurements on three Pasqal QPUs (FC1, SA1, Ruby) and claim that a single noise model, calibrated on FC1, reproduces the behavior of all three devices without device-specific refitting. They also decompose the noise budget by mechanism, analyze residual systematic detuning offsets, and discuss the convergence limits of the matrix-product-state emulations.","tokens_in":26756,"tokens_out":3601,"duration_ms":34221,"significance":"If the central transfer claim holds, the paper is valuable: it provides a concrete, open-source path from microscopic hardware parameters to quantitative predictions of observables on analog neutral-atom devices, and it explicitly tests cross-device reproducibility, which is often only assumed. The manuscript is unusually careful in several respects: it distinguishes Monte Carlo estimation error from trajectory-to-trajectory variability, reports a detailed noise parameter table with physically motivated values, checks bond-dimension convergence in Appendix E, and openly acknowledges residual discrepancies and the post hoc nature of the detuning offset. These practices raise the standard for validation studies in this area. The stress-test concern about the a posteriori detuning offset is real and load-bearing, and the convergence caveat for the strong-interaction correlation data is also central to the validation claim.","major_comments":[{"comment":"The residual detuning offset δ0/(2π) = −0.2 MHz is explicitly stated to be 'inferred a posteriori from the comparison itself,' and its magnitude is twice the calibration precision quoted in Table II (≈100 kHz). This is an additional free parameter selected on FC1 post-quench data, so the abstract, introduction, and conclusion claim that a single noise model was 'calibrated on one device' overstates the predictive content of the FC1 agreement. The paper should either restate the claim as 'one model with one parameter fitted to FC1 transfers to SA1 and Ruby,' or show quantitatively that the SA1/Ruby agreement in Fig. 1(b)–(d) is insensitive to removing or varying this offset. The deliberately programmed +0.67 MHz test is a good consistency check of the mechanism, but it does not establish the uniqueness or necessity of the −0.2 MHz value.","section":"Sec. VI.C and Fig. 4"},{"comment":"The bond-dimension scaling in Appendix E shows that the connected correlations Cn in the strong-interaction regime are not converged at χ=512 already at t ≈ 1000 ns, and the text itself states that these simulations 'remain valuable for assessing qualitative trends.' Since the conclusion claims the framework captures 'both local observables and connected two-body correlations,' the validation of Cn in Fig. 5(b) is not established by the current classical reference. The manuscript needs either a converged reference for the strong-interaction correlation data or a documented convergence-induced error bar on the noise envelope before claiming that the QPU correlation measurements are consistent with the noise model in that regime.","section":"Sec. VI.D and Appendix E, Fig. 8(d)"},{"comment":"The central cross-device claim rests primarily on visual agreement between experimental markers and shaded uncertainty envelopes; no quantitative statistic is reported, such as the fraction of data points outside the 2.5–97.5 percentile envelope, per-device normalized residuals, or a goodness-of-fit measure for each protocol. Given that the paper's headline result is that a single calibrated model transfers across three QPUs, the manuscript should quantify the agreement per device and per protocol, and discuss any systematic per-device trends that may be hidden by the envelope width.","section":"Fig. 1(b)–(d) and Sec. III"}],"minor_comments":[{"comment":"'we use a smalldtas a time step' should read 'a small dt as a time step'; there is also a typo 'asessing' in Sec. VI.D.","section":"Appendix E"},{"comment":"The main text states that n_shots = 300 bitstrings are used for the annealing observable, while Appendix E1 states N_shots = 10^3; these numbers should be reconciled.","section":"Sec. IIIA and Appendix E1"},{"comment":"The emulator is cited via documentation URLs rather than a versioned release; please provide a stable version or DOI so that the numerical results are reproducible.","section":"References [15], [25]"},{"comment":"For the strong-interaction correlation panel, it would be helpful to show the χ=128 and χ=256 curves in the same figure, or at least to state in the caption that the envelope is not converged with respect to bond dimension.","section":"Fig. 5(b)"}],"recommendation":"major_revision","confidential_remarks":"The paper is authored entirely by Pasqal personnel and validates the authors' own hardware and software. That is not disqualifying, but it strengthens the need for the quantitative cross-device residuals and the clear restatement of the fitted post hoc parameter requested above. I would not raise this in the report itself beyond the scientific request for a precise predictive claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is worth a look: the Pasqal group has shipped an open-source, full-cycle Monte Carlo noise emulator built on Pulser/emu-mps, and they have validated it on three same-generation QPUs with a single calibration set. The integrated validation is the new part. Individual noise channels are known, but no one has shown that one parameter set predicts both annealing and post-quench observables across three devices. The paper is unusually honest about its own limits: the trajectory-statistics discussion, the MPS bond-dimension scaling study, and the explicit admission of residual discrepancies all read as genuine engineering care. The deliberately programmed +0.67 MHz detuning offset is a good control experiment, and the noise-budget decomposition in Figs. 2 and 3 is practically useful for algorithm and hardware decisions.\n\nThe main soft spot is the central claim. Section VI.C introduces a -0.2 MHz detuning offset inferred a posteriori from the FC1 comparison. That is one parameter fitted to the validation set, and at twice the calibration precision it is not an uncertainty already inside the model. So the honest claim is 'one model with one parameter fitted to FC1 transfers to SA1 and Ruby,' not 'single noise model, calibrated on one device, reproduces the behavior of all three' with no caveat. The paper does not quantify how much of the cross-device agreement depends on that fitted offset, even though the annealing benchmark and the weak-interaction quench are independent of it. The strong-interaction plateau agreement in Fig. 1(d) presumably includes the offset in the envelope; the reader needs to see what that envelope looks like without it.\n\nThe second soft spot is the correlations in the strong-interaction regime, which are not fully converged at chi=512 and dt=5 ns. The authors admit this in Sec. VI.D and Appendix E. The local population is converged, so the main validation holds, but the correlation comparison in that regime is qualitative. That is a clearly stated limitation, not a hidden flaw.\n\nThe paper deserves a serious referee. A careful revision should restate the transfer claim precisely, show sensitivity to the fitted offset, and ideally release raw data and parameter files alongside the code. The emulator and the parameter table are concrete, reproducible assets. I would cite this in work on validation methodology for analog QPUs, and I would bring it to the reading group as an example of an industrial group doing open and fairly self-critical cross-device assessment.","headline":"A genuinely useful open-source noise emulator for neutral-atom QPUs, but the headline 'single model, no refitting' claim needs a qualifier because one detuning offset was fitted to FC1.","tokens_in":27295,"tokens_out":2413,"would_cite":true,"duration_ms":23853,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single noise model calibrated on one neutral-atom quantum processor, applied without device-specific refitting, reproduces the measured annealing and post-quench observables on all three devices of the same generation.","keywords":["neutral-atom quantum computing","Rydberg arrays","noise-aware emulation","quantum annealing","post-quench dynamics","cross-device validation","Monte Carlo wavefunction sampling","matrix product states"],"falsifier":"Run the post-quench protocol at $\\Omega/U = 1.1$ with a bond dimension 1024 emulation; if the connected correlation $C_n(t)$ shifts by more than the noise-envelope width at times beyond about 1000 ns, the agreement between QPU data and the envelope does not establish the noise model. A separate device-level check would be to calibrate the parameter set on FC1, then measure the same protocols on a fourth same-generation machine with a deliberately different trap temperature: if the data fall outside the predicted envelope, the transferability claim is refuted.","tokens_in":26006,"feed_emoji":"⚛️","tokens_out":6947,"duration_ms":60556,"temperature":0.7,"pith_summary":"This paper sets out to show that analog neutral-atom quantum processors can be described by one portable noise model: a parameter set calibrated on a single device predicts what the other devices of the same generation will measure. The authors build an emulator that pushes the dominant hardware noise sources through the full computation cycle and test it on 36 atoms arranged on a $6\\times 6$ square lattice, running both an antiferromagnetic annealing protocol and a post-quench Ising dynamics protocol on three machines. The measured staggered magnetization and Rydberg population curves fall inside the emulator's uncertainty envelopes for all three devices, with the same parameters used throughout. The practical stake is that a device's output can be interpreted as the programmed Hamiltonian plus known, transferable noise, rather than as an idiosyncratic machine artifact.","feed_headline":"One calibrated noise model predicts all three neutral-atom QPUs","feed_subtitle":"Emulator calibrated on one device reproduces annealing and quench data on all three without refitting.","key_machinery":"The carrying mechanism is Monte Carlo wavefunction sampling over a stochastic distribution of Hamiltonian parameters, implemented in the open-source Pulser library with the emu-mps matrix-product-state backend. Each trajectory draws atom-specific positions, Rabi frequency and detuning corrections, preparation errors, and Lindblad jumps, then evolves the state with the time-dependent variational principle; observables are estimated from the trajectory ensemble, and the uncertainty envelope is the 2.5th--97.5th percentile spread across trajectories, interpreted as shot-to-shot hardware variability rather than the sampling error of the mean. This construction lets the same parameter table act as a predictive model of the QPU cycle end to end.","core_discovery":"The paper claims that the dominant noise mechanisms of a Rydberg QPU--thermal motion and Doppler shifts, laser intensity and phase fluctuations, state preparation and measurement errors, and effective decay and dephasing channels--can be collected into a single parameter set whose Monte Carlo emulation reproduces hardware observables quantitatively. The central assertion is cross-device transfer: parameters calibrated on the FC1 machine, without refitting, produce envelopes that contain the data from SA1 and Ruby in both protocols. For annealing, the framework captures the growth and saturation of the staggered magnetization as a function of the final ramp-down duration; for post-quench dynamics, it captures the weakly interacting Rabi oscillations and the depressed plateau in the strongly interacting regime, and it reproduces the connected nearest-neighbor correlations where the classical simulation is converged. Residual discrepancies are explained as static detuning offsets beyond calibration precision, and a control experiment with a deliberately programmed offset supports that explanation.","pith_inferences":["If the transferability claim survives recalibration and drift tests, the recurring cost of validating analog QPUs shifts from per-device noise characterization to one-time calibration plus periodic drift checks.","The residual detuning-offset result suggests that routine online measurement of the effective detuning before each protocol could tighten the envelopes; the paper does not itself propose this procedure.","Because connected correlations in the strong-interaction regime are not fully converged at bond dimension 512, the strongest test of the noise model in that regime would come from improved classical tensor-network simulations rather than from the current correlation data.","A natural stress test would be to change a single noise parameter, such as the trap temperature, and verify that the emulator envelope moves in the predicted direction."],"forward_implications":["A single calibration campaign on one device can serve as a portable description of same-generation hardware, so cross-device comparisons can separate reproducible physics from machine-specific deviations.","The per-regime noise budget identifies which hardware upgrades matter: thermal positional disorder dominates the strongly interacting quench, decoherence dominates long annealing ramps, and laser fluctuations are comparatively benign for adiabatic preparation.","Protocol parameters such as the annealing ramp-down time can be optimized in emulation before running hardware, giving a quantitative trade-off between diabatic errors at short times and noise degradation at long times.","The framework can be extended to regimes beyond classical reach by validating the model where simulations exist, then using the same parameter set to assess larger systems."],"supporting_citations":[{"why":"Supplies the annealing control schedule and the staggered magnetization observable used as the order parameter for adiabatic state preparation.","marker":"[11]"},{"why":"Pulser is the open-source pulse-sequence control library that hosts the noise model and the emulation backends.","marker":"[24]"},{"why":"emu-mps is the matrix-product-state backend used for all noisy and noiseless emulations in the main text.","marker":"[25]"},{"why":"Provides the single-atom imperfection analysis whose noise mechanisms the model propagates.","marker":"[8]"},{"why":"Prior benchmarking of a noisy analog simulator whose approach this framework generalizes into a predictive, transferable model.","marker":"[10]"},{"why":"Theoretical result that local observables remain robust under errors, justifying validation of local quantities without global fidelity.","marker":"[12]"}],"fun_headline_variants":["One noise model, calibrated once, predicts three QPUs","Calibrate on one device, predict all three QPUs","Cross-device noise emulator: one fit, three Rydberg QPUs","Single calibration transfers across neutral-atom QPUs","Noise-aware emulator reproduces three QPUs without refitting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the classical matrix-product-state emulation, which truncates the quantum state to a maximum bond dimension of 512 with 5 ns time steps and 100 Monte Carlo trajectories, is accurate enough to serve as the reference when judging whether the hardware data match the noise model.","fun_headline_variants_meta":{"raw":{"variants":["One noise model, calibrated once, predicts three QPUs","Calibrate on one device, predict all three QPUs","Cross-device noise emulator: one fit, three Rydberg QPUs","Single calibration transfers across neutral-atom QPUs","Noise-aware emulator reproduces three QPUs without refitting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000462,"raw_usage":{"total_tokens":2279,"prompt_tokens":879,"completion_tokens":1400,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":1308}},"tokens_in":495,"tokens_out":1400,"duration_ms":9720,"temperature":1.0,"reasoning_tokens":1308,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:22:16.797198+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the post-quench protocol at $\\Omega/U = 1.1$ with a bond dimension 1024 emulation; if the connected correlation $C_n(t)$ shifts by more than the noise-envelope width at times beyond about 1000 ns, the agreement between QPU data and the envelope does not establish the noise model. A separate device-level check would be to calibrate the parameter set on FC1, then measure the same protocols on a fourth same-generation machine with a deliberately different trap temperature: if the data fall outside the predicted envelope, the transferability claim is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Pulser is the open-source pulse-sequence control library that hosts the noise model and the emulation backends."},{"cited_title":"Ebadi, T","cited_arxiv_id":null,"evidence_quote":"emu-mps is the matrix-product-state backend used for all noisy and noiseless emulations in the main text."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the single-atom imperfection analysis whose noise mechanisms the model propagates."}],"review_version":2}