{"id":"2c45be1e-c93b-47f1-9568-43483c48bff3","arxiv_id":"1909.01229","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Bayesian optimization with a binomial measurement-noise model finds high-fidelity quantum control solutions with single-shot measurements, drastically cutting the number of experimental runs needed.","lead":"This paper develops a Bayesian optimization algorithm for quantum control that models measurement shot noise as a binomial process, allowing good control pulses to be found from very few experimental repetitions, even single shots. It demonstrates on simulated and real quantum hardware that this approach reduces the experimental effort compared with standard Gaussian-noise Bayesian optimization and gradient-free methods.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GHZ benchmark confounds binomial vs Gaussian likelihood with different Matérn kernels (R0 vs R2), so the claimed advantage is not cleanly attributable to binomial noise modeling.","rationale":"The reader identified the adequacy of the Gaussian-process prior as the weakest assumption; the kernel mismatch in Fig. 4 is a more specific and concrete instance of that concern, since the comparison varies the prior together with the likelihood model. The paper's central claim is not overturned: the binomial approach is still demonstrated to find good controls on several examples, including a NISQ experiment, and the SPSA comparisons are independent of the kernel-selection issue. However, the main quantitative evidence for the specific benefit of binomial noise modeling is weakened by the uncontrolled comparison, because the observed gap could come from the different Matérn kernels rather than from the noise model. The reader's conditional verdict already reflects concerns about reproducibility and prior adequacy, so no change in verdict is needed; a fixed-kernel ablation would either confirm or remove this concern. The Laplace-approximation limitation acknowledged in Sec. III B 2 is also real but secondary, since it mainly affects the regime where Gaussian modeling becomes competitive.","tokens_in":30324,"tokens_out":6680,"duration_ms":69779,"concrete_test":"Recompute the GHZ comparison in Fig. 4 with a fixed kernel order across noise models: run binomial and Gaussian BO with R2 for both, and again with R0 for both, keeping all other settings unchanged. If the binomial advantage remains roughly an order of magnitude in both fixed-kernel comparisons, the attribution to the binomial likelihood is supported; if the gap shrinks or reverses when both use R2, the headline advantage is confounded with the prior roughness. Also report the fitted length-scale values in each case to show whether the marginal-likelihood optimization is selecting comparable effective smoothness.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The key benchmark in Fig. 4 does not isolate the binomial likelihood: Sec. III B 2 states that the binomial modeling uses the Matérn kernel R0 (Eq. 20) while the Gaussian modeling uses R2 (Eq. 22). The two methods therefore differ in two respects simultaneously: the noise model (binomial vs Gaussian) and the smoothness prior for the surrogate (non-differentiable vs twice-differentiable). The order-of-magnitude advantage attributed to binomial noise modeling could in principle be caused by the rougher R0 kernel fitting the binary data better, or by an interaction between the two choices. No ablation or justification for the differing kernels is given. Since the central claim is that the binomial noise model is the source of the improvement, this is a load-bearing gap. The paper's own Laplace-approximation caveat (Sec. III B 2) is an additional source of uncertainty in the exact single-shot regime, but the kernel confound is more directly testable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a Bayesian optimization framework for quantum control in which measurement noise is modeled with the exact binomial likelihood rather than with a Gaussian approximation. A warped Gaussian process is used as a prior over each measurement probability, predictions are obtained through a Laplace approximation, and an acquisition function balances the posterior mean and variance. The method is demonstrated on a pedagogical single-qubit landscape, on GHZ-state preparation from a six-parameter circuit, on a Bose-Hubbard filling-error control problem, and on single-qubit state preparation on an IBM Q device. The central claim is that this binomial-noise Bayesian optimization finds high-fidelity controls even with single-shot measurements, and that it substantially outperforms both Gaussian-noise Bayesian optimization and SPSA for the same total number of experimental runs.","tokens_in":30485,"tokens_out":9652,"duration_ms":103977,"significance":"If the central claim holds, this is a practically valuable result: it offers a data-efficient, model-free route to quantum control in the single-shot limit, a regime where standard expectation-value estimation is prohibitively expensive. The paper has several genuine strengths: the statistical formulation is standard and clearly presented; the benchmarks report medians and interquartile ranges over random instances; comparisons are made against a well-known stochastic approximation method; and the NISQ demonstration provides real-device evidence. The main weakness is that the headline GHZ comparison changes two modeling ingredients simultaneously, so the attribution of the improvement to the binomial likelihood is not yet established.","major_comments":[{"comment":"The central comparison between binomial and Gaussian modeling is confounded. Section III B 2 states that binomial modeling uses the Matérn kernel R0 (Eq. 20) while Gaussian modeling uses R2 (Eq. 22). The two approaches therefore differ not only in the likelihood but also in the smoothness of the prior over the control landscape. The order-of-magnitude advantage attributed to binomial noise modeling in Fig. 4 could in principle be caused by the rougher R0 prior being better suited to sparse binary observations, or by an interaction between kernel choice and likelihood. Since the paper's abstract and conclusions attribute the improvement to the binomial modeling, an ablation that crosses or matches the kernels (for example, binomial with R2 and Gaussian with R0) is required before that claim is supported.","section":"III B 2, Fig. 4"},{"comment":"The run-count formula is not sufficiently justified. The text says that 'since the three observables S5, S6 and S7 commute,' assessing the fidelity for M control-parameter sets with N repetitions requires Nr = 5MN runs. Because S1, S2, S3, S4, and S5, S6, S7 each have common eigenbases, three measurement settings per control-parameter set would suffice if simultaneous measurements are used, giving Nr = 3MN rather than 5MN. If instead five separate local settings are used (XXX, ZZZ, XYY, YXY, YYX), that should be stated explicitly. The factor directly sets the horizontal axes of Figs. 4, 5, 6, 8 and 9, so the discrepancy matters for every quantitative claim about experimental effort.","section":"III B 1, Eq. (27), Nr = 5MN"}],"minor_comments":[{"comment":"The text refers to 'the Laplace approximation (Sec. II C),' but the Laplace approximation is introduced in Sec. II F 2 a, not in Sec. II C; the cross-reference should be corrected.","section":"III B 2"},{"comment":"In the description of the single-qubit example, 'with P2 given in Eq. (22)' should read 'with R2 given in Eq. (22)' to match the polynomial notation introduced in Sec. II F 1.","section":"III A"},{"comment":"The value of the acquisition weight α in Eq. (17) is specified only for the illustrative example of Fig. 2. For the GHZ and Bose-Hubbard benchmarks, α is not reported, which makes the acquisition policy and thus the numerical results hard to reproduce exactly.","section":"III B and III C"},{"comment":"The bracketed uncertainties in Table I are described as uncertainties of the medians due to the finite number of repetitions, but the method used to compute them (e.g., bootstrap or standard error of the median) is not stated.","section":"Table I"},{"comment":"The inset omits interquartile intervals for visual clarity, but in the single-shot regime a compact indication of the spread, such as min-max whiskers, would help the reader assess the reliability of the reported medians.","section":"Fig. 4 inset"}],"recommendation":"major_revision","confidential_remarks":"The kernel confound in the GHZ comparison is the key issue. It is fixable: a matched-kernel ablation would either strengthen the central claim or show that the advantage is partially a prior-smoothness effect. The run-count ambiguity is also fixable with a clear statement of how measurements were grouped. If these are addressed, the paper's contribution would be solid and well within the scope of the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something genuinely useful: it builds Bayesian optimization for quantum control on an explicit binomial likelihood for measurement shot noise, and shows this can work in the single-shot limit where a Gaussian likelihood essentially fails. The derivations are standard and clear, the numerical benchmarks include medians and quartiles, and the NISQ demonstration is a nice touch. If the central claim were clean, this would be a strong practical contribution to the quantum-control subfield.\n\nBut the central claim is not clean. Fig. 4, the key GHZ benchmark that supports the order-of-magnitude advantage, confounds the noise model with the kernel choice: the paper states that binomial modeling uses the Matérn R0 kernel while Gaussian modeling uses R2. So the comparison simultaneously changes the likelihood and the smoothness prior. The advantage attributed to binomial modeling could be partly or entirely due to the rougher R0 prior being a better fit for sparse binary observations. No ablation isolates the likelihood. That is a load-bearing gap, and the stress-test note correctly identifies it. The paper's own Laplace-approximation caveat near 0 and 1 is a secondary but real concern, since the optimal regions are exactly where probabilities saturate.\n\nWhat the paper does well beyond the benchmark: the single-qubit example uses R2 with binomial modeling and still works, which shows the method is not simply tied to the rougher kernel. The Bose-Hubbard example and the NISQ results are consistent with the claimed trend, though the kernels used there are not stated in the text I read. The comparison against SPSA is also fair and broadly favorable, but SPSA is a weak baseline in many of these landscapes, so that alone does not carry the argument.\n\nThe main soft spots are the kernel confound, the hand-picked adaptive schedules with no general principle, and the absence of released code, which makes transfer to new problems harder to assess. The adaptive strategy also switches to Gaussian modeling in later stages, which is a reasonable engineering choice but further blurs the attribution of the overall performance to binomial noise modeling.\n\nWho is this for? Anyone working on closed-loop quantum control with limited measurement budgets will want to know this approach exists. The paper deserves a serious referee, but the referee should demand an ablation or at least a same-kernel comparison before the main claim is accepted. I would not cite the performance claim in its current form, but I would cite the framework once the confound is resolved.","headline":"A useful method paper with a real kernel-choice confound in the key benchmark; the binomial likelihood idea is worth engaging, but the headline advantage is not cleanly shown.","tokens_in":31019,"tokens_out":1665,"would_cite":false,"duration_ms":18541,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Bayesian optimization with a binomial noise model finds quantum controls from single-shot data.","keywords":["quantum optimal control","Bayesian optimization","measurement shot noise","single-shot measurements","Gaussian process surrogate","binomial likelihood","model-free quantum control","adaptive control strategy"],"falsifier":"Run the same binomial Bayesian optimizer on a control landscape that is deliberately discontinuous or has features narrower than the kernel length scale, with single-shot measurements; if its advantage over Gaussian modeling and SPSA disappears or reverses in this regime, the paper's attribution of the speedup to binomial noise modeling would be refuted. A cheaper check is to compare posterior predictive calibration at extreme probabilities near 0 and 1, where the Laplace approximation is known to distort the binomial likelihood.","tokens_in":30085,"feed_emoji":"🎯","tokens_out":5478,"duration_ms":56864,"temperature":0.7,"pith_summary":"The paper tries to show that optimal quantum control does not need accurate estimates of expectation values: if measurement noise is represented by the exact binomial distribution of detector clicks, Bayesian optimization can find good control pulses even when every measurement is a single shot. This matters because model-free control of quantum devices is limited by experimental effort, and repeating measurements to suppress shot noise can erase the benefit of avoiding theory. Across the preparation of a GHZ state, a Mott-insulating phase of a Bose-Hubbard chain, and state preparation on a publicly available quantum processor, the binomial-noise version converges substantially faster than Gaussian-noise Bayesian optimization and faster than a simultaneous-perturbation stochastic approximation baseline. In the GHZ case, binomial modeling reaches infidelities about an order of magnitude lower than Gaussian modeling at equal numbers of runs for $N<1000$, and the single-shot limit still yields percent-level infidelity.","feed_headline":"Single-shot measurements still yield optimal quantum control","feed_subtitle":"Modeling detector clicks with binomial statistics lets Bayesian search outpace Gaussian and gradient baselines.","key_machinery":"The load-bearing object is the binomial likelihood of Eq. (5): for a surrogate probability $f(\\theta)$, the probability of $n$ successes in $N$ repetitions is $\\binom{N}{n} f(\\theta)^n (1-f(\\theta))^{N-n}$, replacing the phenomenological Gaussian noise model of Eq. (24). It sits inside a Gaussian-process prior with a stationary Matérn kernel (Eqs. (19)-(22)), transformed by the cumulative distribution function of a standard normal so that surrogate landscapes stay in $[0,1]$. The Laplace approximation converts the non-Gaussian posterior over the sampled landscape points into a Gaussian, so the predictive distribution and the acquisition function $a(\\theta)=\\langle f(\\theta)\\rangle + \\alpha\\,\\sigma(\\theta)$ can be evaluated analytically; this lets the algorithm select the next probe and, through the adaptive strategy of Sec. III B 3, concentrate shots on promising regions.","core_discovery":"The central claim is that the dominant obstacle to data-driven quantum control is not the sparsity of data but the wrong statistical model of noise. The paper models each measured probability as a Gaussian-process surrogate squeezed into $[0,1]$ by a normal cumulative distribution function, and feeds the actual binomial likelihood of observing $n$ clicks in $N$ repetitions into Bayes' rule; the Laplace approximation keeps the resulting predictive distribution tractable. With this likelihood, an acquisition function balancing expected fidelity and uncertainty selects the next control pulse, and an adaptive schedule starts with few repetitions per measurement and later increases $N$ while shrinking the search domain. The authors report that this procedure finds good controls in the single-shot limit, that it robustly outperforms Gaussian modeling and SPSA in every numerical and experimental example tested, and that the resulting control can then be refined to infidelities near $10^{-8}$ in the GHZ problem.","pith_inferences":["Editorial extension: a natural test is to move the binomial likelihood from state-preparation fidelities to gate-learning and variational-energy objectives, where the paper sketches applications but provides no demonstrations; the same counting statistics apply directly.","Editorial extension: the paper's rule that fewer repetitions per point is better suggests a resource-tradeoff curve, since Gaussian-process updates cost $O(M^3)$ in the number of iterations; the paper notes the scaling but does not derive the optimal batch size per pulse.","Editorial extension: in landscapes rougher than the Matérn prior, an adaptive or hierarchical kernel could be needed; comparing fixed versus learned length scales on non-smooth landscapes would separate the contribution of the noise model from that of the smoothness prior."],"forward_implications":["In the GHZ example, fixed-$N$ optimizations show that for any $N$ below 1000, binomial modeling yields infidelities roughly an order of magnitude lower than Gaussian modeling after the same number of runs.","Single-shot measurements ($N=1$) suffice to reach percent-level infidelity, so experimental effort can be shifted from repeating measurements to probing more control pulses.","An adaptive schedule that starts with five repetitions, then raises $N$ and shrinks the parameter domain, refines GHZ infidelities down to about $10^{-8}$ while remaining faster than SPSA at every run budget.","In the presence of unitary noise and readout errors, binomial Bayesian optimization still converges 2 to 5 times faster than SPSA, and a substantial fraction of SPSA runs get trapped in minima with about 50% infidelity.","On the public NISQ device, binomial Bayesian optimization outperformed both Gaussian Bayesian optimization and SPSA in all tested settings, especially with only 300 total runs."],"supporting_citations":[{"why":"Supplies the Gaussian-process framework, Matérn kernels, Laplace approximation, and marginal-likelihood hyperparameter fitting that the surrogate model and integration rely on.","marker":"[50]"},{"why":"Defines the SPSA gradient-approximation algorithm used as the main baseline throughout the numerical and experimental comparisons.","marker":"[52]"},{"why":"Provides the publicly available quantum processor used for the experimental demonstration of the method.","marker":"[51]"},{"why":"Review of Bayesian optimization that grounds the acquisition-function and hyperparameter-adaptation practices adopted in the implementation.","marker":"[27]"},{"why":"Earlier demonstration that Bayesian optimization can prepare ordered states in ultracold gases, which this work extends to poor measurement statistics.","marker":"[34]"}],"fun_headline_variants":["Bayesian optimization with binomial likelihood tames sparse quantum data","Single-shot data still yield optimal quantum control via correct noise model","Binomial statistics beat Gaussian for quantum control from few measurements","Optimal quantum control from one shot per setting? Bayesian method says yes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The algorithm assumes the true control landscape is a smooth sample from a stationary Gaussian process with a Matérn kernel, so that very sparse binary observations can be interpolated reliably; if the real landscape is much rougher or has a different length scale, the acquisition function can be systematically misled before enough data arrives to correct the prior.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian optimization with binomial likelihood tames sparse quantum data","Single-shot data still yield optimal quantum control via correct noise model","Binomial statistics beat Gaussian for quantum control from few measurements","Optimal quantum control from one shot per setting? Bayesian method says yes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1206,"prompt_tokens":884,"completion_tokens":322,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":252}},"tokens_in":500,"tokens_out":322,"duration_ms":4093,"temperature":1.0,"reasoning_tokens":252,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:23:35.090115+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same binomial Bayesian optimizer on a control landscape that is deliberately discontinuous or has features narrower than the kernel length scale, with single-shot measurements; if its advantage over Gaussian modeling and SPSA disappears or reverses in this regime, the paper's attribution of the speedup to binomial noise modeling would be refuted. A cheaper check is to compare posterior predictive calibration at extreme probabilities near 0 and 1, where the Laplace approximation is known to distort the binomial likelihood.","supporting_citations":[{"cited_title":"Optimal control of quantum-mechanical systems: Existence, numerical approximation, and applica- tions,","cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian-process framework, Matérn kernels, Laplace approximation, and marginal-likelihood hyperparameter fitting that the surrogate model and integration rely on."},{"cited_title":"Training schr¨ odinger’s cat: quantum optimal control,","cited_arxiv_id":null,"evidence_quote":"Defines the SPSA gradient-approximation algorithm used as the main baseline throughout the numerical and experimental comparisons."},{"cited_title":"Quantum op- timal control theory,","cited_arxiv_id":null,"evidence_quote":"Provides the publicly available quantum processor used for the experimental demonstration of the method."},{"cited_title":"Optimal control of complex atomic quantum sys- tems,","cited_arxiv_id":null,"evidence_quote":"Review of Bayesian optimization that grounds the acquisition-function and hyperparameter-adaptation practices adopted in the implementation."},{"cited_title":"Control of chemical reactions by feedback-optimized phase-shaped femtosecond laser pulses,","cited_arxiv_id":null,"evidence_quote":"Earlier demonstration that Bayesian optimization can prepare ordered states in ultracold gases, which this work extends to poor measurement statistics."}],"review_version":1}