{"id":"4150a050-f19f-4199-b582-9b926c399efa","arxiv_id":"2607.24947","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Partially fault-tolerant [[4,2,2]] Iceberg-code simulations on ibm_boston improve local Ising observables over unencoded baselines by a few percent in 1D and over 200% in 2D at late times via Observable-Ranked Postselection.","lead":"Encoded Ising simulations on IBM hardware beat bare-qubit runs on local observables by using a lightweight error-detecting code plus selective postselection. The result shows a practical near-term path where partial fault tolerance helps scientific simulations before full error correction is available.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"ORP is validated by only one half-sample split; plateau stability around an internally constructed reference may not ensure stable, unbiased late-time encoded gains.","rationale":"The reader identified the same general soft spot: the credibility of the quantitative encoded advantage depends on ORP’s data-driven detector ranking and plateau selection. My agreement is partial rather than complete because the specific worry that ORP favors detectors moving estimates toward the MPS reference is not directly supported—the selection rules described in Methods A use device data, and the holdout protects against the simplest form of reference-chasing. The remaining concern is validation insufficiency: split-to-split selection variance, sensitivity to the three plateau hyperparameters, and common-mode residual bias are not demonstrated to be negligible, especially for the late-time 2D ratio that carries the largest headline number.\n\nThis does not undermine the entire paper. The CDFs of absolute per-qubit errors, the substantial reduction in median 2D error, and the 1D result in which much deeper encoded circuits still outperform provide independent evidence that the encoding helps. Thus rejection would be too strong. But because the most striking quantitative claim relies heavily on ORP and cannot presently be re-run from released artifacts, the reader’s CONDITIONAL verdict remains appropriate. The proposed repeated-split audit would either convert that condition into stronger acceptance or show that the >200% figure needs to be replaced by a more stable absolute-error or Δα statement.","tokens_in":91515,"tokens_out":7654,"duration_ms":274375,"concrete_test":"Using the raw shot/detector bitstrings for the latest-time 2D R=1 runs, perform 100 independent half-splits. On each training half, recompute the Δj ranking, Oref, and k* with fmin=0.005, fref=0.05, z=1; evaluate ⟨Z⟩, αenc, and the percentage α improvement over the unencoded run on the complementary half. Repeat with one-at-a-time perturbations fmin∈{0.0025,0.01}, fref∈{0.025,0.10}, z∈{0.5,2}. Report the k* distribution and full improvement distribution. If the lower 95% bound remains above 200% and k* is stable, this concern largely does not land; otherwise the late-time headline should be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most exposed step is the ORP selection procedure in Methods A.2. Detectors are ranked by empirical Δj, a reference Oref is constructed from heavily filtered cut levels satisfying fmin≤fk≤fref, and k* is selected from the longest plateau within zσk. The half-sample holdout is a meaningful safeguard against naive winner’s-curse bias, and MPS results are not used in the ranking, so I do not see direct circularity. Nevertheless, one fixed split plus bootstrap uncertainty at a fixed k does not quantify model-selection variability. With finite shots, rare detector flips can make both the Δj ranking and k* split-dependent, while plateau consistency around Oref cannot exclude common-mode residual errors shared by the high-k levels. This matters most for the >200% late-time 2D claim: Appendix C says that gain is largely due to postselection, and Fig. 4 reports a percentage change in fitted α whose unencoded denominator is becoming small. Modest ORP selection variance can therefore have an outsized effect. The paper asserts robustness to the plateau parameters but does not display repeated-split or hyperparameter-stability results, and the raw data needed to audit this are not available from the manuscript.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper reports partially fault-tolerant quantum simulations of the mixed-field Ising model on IBM's ibm_boston, using 21 blocks of the [[4,2,2]] Iceberg code (42 logical qubits, up to 136 physical qubits) with fault-tolerant syndrome extraction but non-fault-tolerant logical operations. The authors introduce Observable-Ranked Postselection (ORP), which ranks error detectors by their empirical correlation with a target observable and postselects only on the highest-ranked detectors, with the cut level chosen by a plateau-finding algorithm on a held-out half of the data. Comparing encoded and unencoded circuits at equal shot count against MPS benchmarks, they report improvements in local-observable accuracy of 2–6% at intermediate times in 1D and more than 200% (in a signal-survival metric α) at late times in 2D, where the encoding also reduces circuit depth relative to direct heavy-hex embedding. Supporting material includes circuit-depth/gate tables, classical verification of gadget-level fault-tolerance conditions, memory experiments, and an analytical treatment of the ORP ranking metric.","tokens_in":91857,"tokens_out":3524,"duration_ms":117969,"significance":"If the results hold, this is a useful data point in the early-fault-tolerance landscape: it demonstrates beyond-baseline local-observable estimation with a lightweight QED code on widely accessible commercial hardware, quantifies the crossover regime in which partial FT beats both NISQ and full-FT strategies, and shows concretely that encoding can relax connectivity constraints (the 2D case) rather than only suppress errors. The explicit circuit accounting, MPS cross-checks with bootstrap uncertainties, classical verification of gadget fault tolerance, and honest reporting of failure modes (device dependence, reset tradeoffs, R-dependence reversal) are strengths that make the work reproducible in principle and falsifiable in practice. The impact is incremental rather than transformative — improvements are modest in 1D and device-specific — but it is exactly the kind of careful hardware characterization the field needs.","major_comments":[{"comment":"ORP selection variability is asserted but not quantified, and it is load-bearing for the headline numbers. The plateau algorithm ranks detectors by Δj = |⟨O⟩_{v_j=0} − ⟨O⟩_{v_j=1}|, constructs O_ref from an internally chosen band f_min ≤ f_k ≤ f_ref, and selects k* from the longest plateau within zσ_k, all on a single half-sample split. The holdout (k* chosen on one half, applied to the other) protects against the crudest winner's-curse bias, and the external MPS comparison rules out gross bias toward the target, but one fixed split plus bootstrap-at-fixed-k does not propagate model-selection variance into the reported uncertainties. With 32k shots, the Δj ranking of rare-firing detectors is itself shot-noise dominated, so k* can be split-dependent. The text states 'the quality of ⟨O⟩_k is observed to be similar for a range of parameters' without displaying any such scan. Requested: (i)","section":"Methods A.2 (Detector Selection), Fig. 5"},{"comment":"The abstract and Fig. 4 headline a '>200% improvement' in 2D at the latest times, but this is a percentage change in fitted α whose unencoded denominator is small and shrinking (Fig. 3d shows unencoded α near zero at late t). The text partially acknowledges this ('driven by rapidly decohering unencoded results'), but the abstract presents the number without that context, and Appendix C states the gain is 'largely due to postselection as opposed to reduced gate depth.' Since the α-percentage diverges as α_unenc → 0, the figure of merit is unstable precisely where the headline is taken from. Requested: report absolute α values (encoded and unencoded) alongside the percentage in the abstract/main text, and state the acceptance fraction at k* for the latest-time points (Appendix F data should be referenced inline). This is a presentation-of-claim issue, not a correctness issue, but it curren","section":"Abstract; Sec. III, Fig. 4; Appendix C"},{"comment":"There is a residual circularity risk in conditioning the postselection ranking on the same observable subsequently reported. Δj is computed from correlations of detector v_j with observable O, and the encoded ⟨O⟩_{k*} is then the reported quantity. The holdout half mitigates this within one observable, but the selection is observable-matched throughout. A direct, cheap test would resolve this: compute the Δj ranking and k* using one observable (e.g., ⟨Z_n⟩) and evaluate the resulting filter on a different observable (e.g., ⟨X_n⟩ or ⟨Z_nZ_{n+1}⟩) against MPS. If the filter transfers, the 'detectable-error removal' interpretation of ORP is supported; if not, the gains may partly reflect observable-specific selection. The Pauli-propagation/lightcone analysis in Appendix A.5 (Table IV) is suggestive but is itself observable-conditioned. Given that this test requires only re-analysis of exist","section":"Methods A.2; Sec. II.B; Appendix A.5"},{"comment":"The detector ranking used for the hardware claims is validated against Pauli propagation only in a 4-block, p=0.003 depolarizing simulation (Table IV), and the agreement is qualitative: the ∆j and γj orderings shown differ substantially even within the top 10 (e.g., v(0,1) ranks 1st by Δj and 5th by γj), and the text concedes only set-level overlap. Since the causal-lightcone argument is the main structural justification for why ORP should work (footnote 4, Appendix A.5), the manuscript should either show the ranking correspondence at a scale closer to the 21-block experiments or temper the claim that ORP-selected detectors 'coincide with the backwards lightcone' to the weaker set-overlap statement actually demonstrated.","section":"Appendix A.5, Table IV; footnote 4"}],"minor_comments":[{"comment":"Typo: 'results obtained using our scheme show improvements for both 1D and 1D simulations' — the second should be 2D.","section":"Sec. IV (Discussion), first paragraph"},{"comment":"The phrase 'resulting in biased unencoded results with small error bars' is unclear as written; presumably this means the equal-shot comparison gives the unencoded baseline tighter statistics than the postselected encoded runs. Please rephrase.","section":"Methods D, first paragraph"},{"comment":"The CDFs in Fig. 2d use t ≤ 4 while Fig. 2c shows t = 8, and the rationale (t = 4–8 excluded, shaded in Fig. 2e) is only visible in the caption of panel e. A one-sentence justification in the main text for why the excluded window is excluded — beyond 'where an encoding improvement is observed' — would help, since as written the window choice looks outcome-selected.","section":"Fig. 2d–e and caption"},{"comment":"Footnote 5 states that Δj depends on which data-qubit representative of O (modulo stabilizers) is used. This gauge dependence deserves a sentence in the main text of Methods A.2, since different representatives can in principle produce different rankings; please state which representative is used for the reported results.","section":"Methods A.2, footnote 5"},{"comment":"The two-cycle no-reset detector construction (Eq. 8) is well explained, but Eq. (4) in the main text uses identical notation with the same two-cycle skip; please note explicitly in the main text that the skip is a consequence of omitted resets, since a reader of Sec. II.B alone may assume a typo.","section":"Sec. II.B, Eq. (4) vs. Methods A.1"},{"comment":"Methods E references 'Fig. 20' and 'Fig. 21' for late-time and 5×5 simulations that do not appear in the main narrative; please verify the figure numbering and cross-references in the compiled version.","section":"Methods E"},{"comment":"Raw shot-level data and analysis code do not appear to be released with the manuscript. Given that the central validation of ORP (k* selection, plateau finding) is only auditable from shot-level data, a data/code availability statement with a public repository is strongly recommended.","section":"Data availability"},{"comment":"In Eq. (7), α is defined through a fit constrained to pass through the origin. Please comment on the sensitivity of α to allowing an intercept, since coherent rotation errors in the MFIM dynamics could shift ⟨Z⟩ by an offset that a through-origin fit would absorb into α.","section":"Sec. III, Eq. (7)"}],"recommendation":"major_revision","confidential_remarks":"ORP is closely related to existing correlation-based postselection ideas (the authors themselves cite Refs. [13, 22, 77] for the Δj interpretation), so the novelty is in the ranking-plus-plateau selection machinery and the scale of the demonstration rather than the concept. The manuscript is otherwise well within scope. I would encourage the editor to require public release of the shot-level data as a condition of publication, since the key validation analyses my report requests are only auditable with the raw data."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a careful experimental paper that actually runs something useful: 21 blocks of [[4,2,2]] on ibm_boston (42 logical qubits), FT syndrome extraction paired with nFT logical rotations, and Ising quench dynamics in 1D and 2D, compared head-to-head with matched unencoded circuits and no noise-learning mitigation. That isolation is the right design choice.\n\nWhat is new is the scale plus the connectivity win and ORP. The logical square grid lets them do 2D with lower depth than a direct heavy-hex unencoded embedding (Tables II–III make this concrete). ORP ranks detectors by empirical correlation with the target local observable and cuts at a plateau found on a hold-out half-sample; that is a practical answer to exponential shot death under full postselection. Hardware vs MPS, CDFs of absolute error, and the signal-survival α fits are cleanly reported. The 1D 2–6% intermediate-time gains look modest and believable. The >200% late-time 2D figure is mostly the unencoded baseline collapsing while encoded α decays slower; Appendix C is honest that postselection, not depth alone, drives most of it.\n\nSoft spots, in proportion. ORP is data-driven on the same class of observables later reported. Hold-out and external MPS benchmarks reduce the circularity risk, and the stress-test overstates “bias toward MPS” because MPS is not in the ranking. Still, one fixed split and fixed plateau hyperparameters (f_min, f_ref, z) leave model-selection variance unquantified; that matters most when the unencoded denominator is small. Advantage is also device- and R-dependent (worse machines and high R lose it), which the discussion already flags. None of that sinks the central claim that partial FT plus selective filtering can beat a fair unencoded baseline on local observables for this model class.\n\nMath and circuit constructions look solid; citations cover Iceberg, partial FT, and Ising quenches without obvious gaps. This is for people doing near-term quantum simulation on superconducting hardware and anyone thinking about lightweight QED before full FT. I would send it to referees. Engage with it if you care about practical encoded simulation; the 2D connectivity point and ORP are worth stealing even if you re-tune the cut.","headline":"Solid hardware demo of partial FT Iceberg Ising sims at 42 logical qubits with a useful selective postselection trick; gains are real but device- and ORP-dependent, especially the late-time 2D number.","tokens_in":93191,"tokens_out":579,"would_cite":true,"duration_ms":15165,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Partially fault-tolerant Iceberg-code simulations on IBM hardware beat matched unencoded baselines on local Ising observables, by 2–6% in 1D and over 200% in 2D at late times.","keywords":["quantum error detection","Iceberg code","partial fault tolerance","Observable-Ranked Postselection","Ising model","quantum simulation","heavy-hex connectivity","ibm_boston"],"falsifier":"Rerun the identical encoded and unencoded circuits on a device whose two-qubit error rate or coherence is modestly worse than ibm_boston; if the encoding advantage in the signal-survival factor α disappears or reverses, the claimed crossover improvement is hardware-specific rather than generic.","tokens_in":93074,"feed_emoji":"⚛️","tokens_out":1015,"duration_ms":17112,"temperature":0.7,"pith_summary":"This paper claims that lightweight quantum error detection, used only partially, already improves real Hamiltonian simulations on present superconducting hardware. The authors encode 42 logical qubits as 21 blocks of the [[4,2,2]] Iceberg code on ibm_boston, keep syndrome extraction fault-tolerant, and leave the logical rotations non-fault-tolerant so that circuit depth stays manageable. They introduce Observable-Ranked Postselection, which ranks detectors by how strongly each correlates with the target local observable and keeps only the most informative checks, avoiding the exponential shot loss of full postselection. On quench dynamics of the mixed-field Ising model the encoded runs improve local-observable accuracy over matched unencoded baselines by a few percent at intermediate times in one dimension and by more than a factor of two in two dimensions at the latest times studied. The encoding also supplies a natural square logical lattice, so the two-dimensional simulation is actually shallower than a direct unencoded embedding on heavy-hex connectivity. A sympathetic reader cares because the work places today’s devices in a usable crossover regime between NISQ heuristics and full fault tolerance, where modest encoding already helps scientific observables without waiting for a complete logical gate set.","feed_headline":"Encoded Ising sims beat bare ones by 200% on IBM hardware","feed_subtitle":"Partial fault tolerance plus ranked postselection lifts local observables without full error correction","key_machinery":"Observable-Ranked Postselection (ORP): detectors are ranked by Δ_j = |⟨O⟩_{v_j=0} − ⟨O⟩_{v_j=1}|, the absolute shift they induce in the target observable; a plateau-finding cut retains only the highest-ranked detectors, recovering most of the benefit of full postselection without exponential ensemble collapse.","core_discovery":"Partially fault-tolerant simulations that pair fault-tolerant syndrome extraction of the [[4,2,2]] Iceberg code with non-fault-tolerant logical operations, plus Observable-Ranked Postselection, improve estimates of local observables in mixed-field Ising quenches on ibm_boston relative to matched unencoded circuits: gains of 2–6% at intermediate times in 1+1D grow with depth to more than 200% in 2+1D at the latest times examined.","pith_inferences":["The same block-and-ORP pattern should transfer to other two-qubit-per-site models (Schwinger, Fermi–Hubbard) whose natural units match the [[4,2,2]] code.","Combining ORP with modest physical-noise learning on the retained shots could push the crossover regime deeper without requiring full QEC.","If ancilla resets become cheap enough not to cancel the coherence budget, more frequent syndrome rounds may reopen the early-time advantage seen only for small R."],"forward_implications":["Near-term lattice simulations can already gain from distance-2 block codes without a full fault-tolerant gate set.","Logical connectivity supplied by multi-block Iceberg encodings can make 2D (and potentially 3D) lattices cheaper in depth than direct heavy-hex embeddings.","Selective, observable-aware postselection is a practical alternative to discarding every syndrome event when many blocks are used.","Encoding advantage is largest when dynamics are slow and the observable’s backward light-cone is small, so localized or gapped regimes benefit first.","Further gains are expected once physical error rates improve only modestly past the present operating point."],"fun_headline_variants":["Partial FT Iceberg Ising sims beat bare by 200% on IBM","Ranked postselection lifts encoded Ising locals over 200% in 2+1D","[[4,2,2]] partial FT quenches top unencoded Ising on ibm_boston","Encoded 1+1D/2+1D Ising gains 2-6% to 200% vs bare circuits","FT syndrome plus ranked filter improves Ising observables on IBM"],"cache_read_input_tokens":82048,"weakest_assumption_plain":"The data-driven ranking of detectors and the plateau cut that follows must remove harmful detectable errors without systematically pushing the reported expectation value toward the classical reference beyond what residual undetectable noise allows.","fun_headline_variants_meta":{"raw":{"variants":["Partial FT Iceberg Ising sims beat bare by 200% on IBM","Ranked postselection lifts encoded Ising locals over 200% in 2+1D","[[4,2,2]] partial FT quenches top unencoded Ising on ibm_boston","Encoded 1+1D/2+1D Ising gains 2-6% to 200% vs bare circuits","FT syndrome plus ranked filter improves Ising observables on IBM"]},"model":"grok-4.5","effort":"low","cost_usd":0.005511,"raw_usage":{"total_tokens":1547,"prompt_tokens":834,"num_sources_used":0,"completion_tokens":104,"cost_in_usd_ticks":55108000,"prompt_tokens_details":{"text_tokens":834,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":609,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":834,"tokens_out":104,"duration_ms":10557,"temperature":1.0,"reasoning_tokens":609,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T05:08:52.910672+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rerun the identical encoded and unencoded circuits on a device whose two-qubit error rate or coherence is modestly worse than ibm_boston; if the encoding advantage in the signal-survival factor α disappears or reverses, the claimed crossover improvement is hardware-specific rather than generic.","supporting_citations":[],"review_version":1}