{"id":"b97954d5-dd15-40cd-8b01-a150d39f5a09","arxiv_id":"2604.24475","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Physically bounded extrapolation models for zero-noise extrapolation reduce unphysical predictions and improve stability compared to unbounded fits on large synthetic benchmarks and real hardware.","lead":"Zero-noise extrapolation on quantum devices can produce results outside the physical range of observables when standard models are fitted without constraints. This paper adds explicit bounds on the zero-noise estimate during optimization for polynomial, exponential, and hybrid models.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest assumption correctly flags the practical bias question for real-device use, but the strongest claim itself is an empirical statement about the controlled synthetic benchmark and does not assert unbiased accuracy or superiority on hardware. No internal inconsistency or unsupported step in the reported benchmark results is apparent from the given description.","tokens_in":1751,"tokens_out":273,"duration_ms":47529,"concrete_test":"Re-run the unphysical-prediction and stability metrics on a random 10% held-out subset of the 180k synthetic circuits using the same bounded vs. unbounded fitting procedures; if the relative reductions remain within 5% of the full-benchmark figures, the headline claim is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is narrowly scoped to performance on the synthetic benchmark (180k circuits, 3.6M experiments) with known ground truth: bounded variants reduce unphysical predictions and improve stability for exponential and poly-exponential families, with little change for polynomials. The benchmark directly measures these quantities against the true noise-free values, so the reported improvements stand independent of whether the constraint introduces bias in real-device regimes. The paper itself notes that simulation models do not fully capture hardware variability, but this does not invalidate the synthetic results.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces physically bounded variants of polynomial, exponential, and polynomial-exponential extrapolation models for zero-noise extrapolation (ZNE) by explicitly parameterizing the zero-noise estimate and constraining it to the physical interval during optimization. It evaluates the approach on a large synthetic benchmark of 180,000 circuits and 3.6 million ZNE experiments generated from realistic IBM-derived noise models, plus preliminary hardware tests on GHZ and W-state circuits. Key results show that bounded models substantially reduce unphysical predictions and improve stability for exponential and poly-exponential families (with little difference for polynomials), and exhibit similar qualitative behavior on hardware despite device limitations.","tokens_in":1821,"tokens_out":462,"duration_ms":61766,"significance":"If the results hold, this provides a practical, low-overhead improvement to ZNE, a widely used error-mitigation technique for near-term quantum devices. The scale of the synthetic benchmark (with known ground truth) supplies robust statistical support for reduced unphysical extrapolations and enhanced stability in specific model families. The method integrates easily into existing workflows. Hardware results, though preliminary, add relevance while underscoring gaps between simulation and real devices.","major_comments":[],"minor_comments":[{"comment":"Abstract and results: the synthetic benchmark size is given as 'approximately 3.6 million' experiments; report the exact count and breakdown by model family and noise level in the methods or results section for reproducibility.","section":null},{"comment":"Results section: define the stability metric explicitly (e.g., variance of extrapolated values across amplification factors or bootstrap resampling) and report quantitative effect sizes or statistical tests for the claimed improvements in stability and reduction of unphysical predictions.","section":null},{"comment":"Hardware experiments: the validation is described as preliminary; include the exact number of circuits, shots per circuit, and quantitative metrics (e.g., fraction of unphysical predictions) to allow direct comparison with the synthetic benchmark.","section":null},{"comment":"Methods: provide the explicit parameterization of the zero-noise estimate and the form of the constrained optimization objective (e.g., via Lagrange multipliers or reparameterization) to enable straightforward implementation by readers.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive summary of our manuscript and for recommending minor revision. The assessment correctly identifies the core contribution—physically bounded extrapolation models that reduce unphysical predictions—and the value of the large-scale synthetic benchmark. We will incorporate minor editorial improvements and clarifications in the revised version.","responses":[],"tokens_in":1297,"tokens_out":78,"duration_ms":18892,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main result is that reparameterizing common ZNE fits to keep the zero-noise estimate inside the physical interval reduces bad predictions and improves stability for exponential and poly-exponential models. Polynomials barely change. They back this with 180k synthetic circuits and 3.6M experiments that have known ground truth from IBM-style noise models, which gives the comparison real weight. The bounded versions simply avoid the cases where unconstrained fits spit out values outside [0,1] or whatever the observable range is. That is a concrete, usable tweak and the scale of the test set is the part that stands out. The hardware runs on GHZ and W states show the same qualitative pattern but are described as preliminary, and the authors note that real devices have more variability than the simulations capture. The weakest part is the lack of direct checks on whether the constraint systematically shifts the estimate away from the true noise-free value when the noise model is imperfect. The synthetic numbers hold up on their own terms, but anyone adopting this would still want to see how often the bound is active and what it does to accuracy on real hardware. This is aimed at people who already run ZNE on near-term devices and want a low-effort way to make it more reliable. It is not a deep theoretical advance, but the benchmark is large enough and the change is simple enough that it deserves referee time. A review could tighten the hardware section and add a short bias diagnostic without much extra work.","headline":"Bounded ZNE models cut unphysical extrapolations on a large synthetic benchmark, but hardware results stay preliminary and bias checks are light.","tokens_in":2295,"tokens_out":369,"would_cite":false,"duration_ms":31081,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Physically bounding the zero-noise estimate during fitting reduces unphysical predictions in zero-noise extrapolation.","keywords":["zero-noise extrapolation","quantum error mitigation","physically bounded models","extrapolation fitting","near-term quantum devices","unphysical predictions"],"falsifier":"On small circuits whose exact noise-free expectation value is known classically, measure whether the mean squared error of bounded-model predictions is not materially larger than that of unbounded models while the fraction of unphysical outputs drops sharply.","tokens_in":2644,"feed_emoji":"⚛","tokens_out":636,"duration_ms":36117,"temperature":0.7,"pith_summary":"The paper introduces physically bounded variants of polynomial, exponential, and polynomial-exponential extrapolation models for zero-noise extrapolation by explicitly parameterizing the zero-noise value and constraining it to the valid physical range during optimization. This prevents the common problem of fitted models producing expectation values outside the possible range for quantum observables. On a benchmark of 180,000 synthetic circuits and millions of ZNE experiments under realistic IBM-derived noise, the bounded versions cut unphysical outputs and stabilize the exponential-family models, while polynomial models change little. Preliminary real-hardware tests on GHZ and W-state circuits show the same qualitative pattern of avoiding pathological extrapolations.","feed_headline":"Bounding the zero-noise value cuts unphysical ZNE outputs","feed_subtitle":"Constraining the extrapolated estimate to physical ranges during optimization stabilizes models on large synthetic benchmarks of realistic量子","key_machinery":"Physically bounded extrapolation models, formed by isolating the zero-noise term as a free parameter that is clamped to the valid observable range [0,1] (or equivalent) inside the fitting procedure.","core_discovery":"By reparameterizing the extrapolation function so that the zero-noise estimate appears as an explicit parameter that is then constrained to the physical interval during optimization, bounded models for zero-noise extrapolation substantially reduce unphysical predictions and improve stability for exponential and polynomial-exponential families relative to their unconstrained counterparts.","pith_inferences":["The same reparameterization trick could be applied to other extrapolation-based mitigation techniques that currently ignore physical bounds.","Device-specific calibration data could be used to set tighter, non-uniform bounds rather than the universal [0,1] interval.","The observed gap between simulation and hardware suggests that bounded models may help surface when noise models are incomplete."],"forward_implications":["Bounded extrapolation cuts unphysical predictions across the 180,000-circuit synthetic benchmark.","Exponential and polynomial-exponential models gain measurable stability; polynomial models show little change.","On real hardware the bounded variants avoid pathological extrapolations and give a more usable accuracy-coverage trade-off.","The method slots into existing ZNE pipelines with only a change to the optimizer constraints."],"fun_headline_variants":["Bounded ZNE cuts unphysical predictions","Physical constraints stabilize ZNE extrapolation","Constrained zero-noise bounds reduce invalid outputs","Bounded models limit unphysical ZNE extrapolations","ZNE with physical bounds avoids unphysical results"],"cache_read_input_tokens":64,"weakest_assumption_plain":"Constraining the zero-noise estimate to the physical interval during optimization does not systematically bias the extrapolated value away from the true noise-free result under realistic device noise.","fun_headline_variants_meta":{"raw":{"variants":["Bounded ZNE cuts unphysical predictions","Physical constraints stabilize ZNE extrapolation","Constrained zero-noise bounds reduce invalid outputs","Bounded models limit unphysical ZNE extrapolations","ZNE with physical bounds avoids unphysical results"]},"model":"grok-4.3","cost_usd":0.006108,"raw_usage":{"total_tokens":2808,"prompt_tokens":676,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":61078000,"prompt_tokens_details":{"text_tokens":676,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2067,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":676,"tokens_out":65,"duration_ms":47830,"temperature":1.0,"reasoning_tokens":2067,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-08T04:12:00.573645+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"On small circuits whose exact noise-free expectation value is known classically, measure whether the mean squared error of bounded-model predictions is not materially larger than that of unbounded models while the fraction of unphysical outputs drops sharply.","supporting_citations":[],"review_version":1}