{"id":"8def0c5b-d83e-4fc1-aa75-2cd670d6ad8e","arxiv_id":"2504.12628","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Adding delay-gate-induced amplitude damping improves NDAR's best MaxCut solutions on IBM Heron, and a classical Hamming-weight-controlled variant reaches similar or better solution quality.","lead":"Researchers tested a way to make a quantum optimization method called Noise-Directed Adaptive Remapping stronger by deliberately adding a pause before reading out the qubits, which increases a specific noise effect. Longer pauses improved the best MaxCut solutions found on IBM's Heron processor, and a classical version of the same trick worked on 300-node problems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Delay-time improvement is shown on one instance per graph class; the headline claim needs multi-instance replication before it can be read as a general NDAR property.","rationale":"The reader's CONDITIONAL verdict is appropriate, and this stress-test identifies the lack of multi-instance replication as the most load-bearing gap because the central claim is empirical and stated generally. The all-zeros-attractor assumption flagged by the reader is physically plausible (idle-time amplitude damping with T1 ~ 180 microseconds gives a per-qubit 1-to-0 flip probability near 0.43 at Td=100 microseconds) and, even if the attractor is only approximately |0...0>, the gauge-remapping logic still biases samples toward the current best; so a failure of that assumption would weaken the mechanism story but would not necessarily overturn the observed trend. The parameter-transfer confound is real for QAOA, but the random-circuit results independently show a delay-time effect, so it cannot fully explain the claim. The most direct threat to 'increasing delay improves NDAR' is that the trend may not survive a change of instance. The proposed replication is the concrete check that would settle this: if the ordering holds broadly across fresh instances, the central claim is supported; if not, the claim should be narrowed to the tested instances. The reader's rationale already mentions the single-instance limitation, but their formally identified weakest assumption was different; hence 'partial' agreement. I recommend no change to the reader's verdict: CONDITIONAL, requiring either multi-instance evidence or a narrowed claim.","tokens_in":13952,"tokens_out":18143,"duration_ms":203016,"concrete_test":"Generate 20 fresh 80-node unweighted-sparse MaxCut instances (and, if budget allows, 20 weighted-dense instances) with the same generator and edge density from Eq. (12). For each instance, run the Section IV protocol: QAOA p=1 and depth-2 random circuits, M=1000, niter=8 (12 for weighted-dense), 10 runs per configuration, Td=0/50/100 microseconds, same fixed QAOA parameters and the same IBM Heron backend. Record the final mean EBest for each Td and circuit per instance. Test the hypothesis that mean EBest(Td=100) > mean EBest(Td=50) > mean EBest(Td=0) separately for QAOA and random circuits using a paired Wilcoxon test across the 20 instances, and also require the per-instance ordering to hold in at least 18 of 20 instances for whichever circuit is used for the headline claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — increasing delay time Td improves EBest — rests on Fig. 1 and Fig. 4, which report averages over 10 runs of a single 80-node instance per graph family (unweighted-sparse and weighted-dense). The quoted standard errors, e.g., EBest/ESA = 0.926(3) at Td=100, quantify run-to-run and shot noise for that one graph; they do not quantify instance-to-instance variability. MaxCut landscapes vary substantially with graph realization even at fixed edge density, and the NDAR mechanism — a low-Hamming-weight bias near the current best solution — is not guaranteed to help on every instance: if the optimal region has atypical Hamming weight or the landscape is rugged, stronger delay-induced damping can trap the search. The abstract states the result generally ('increasing delay time in the NDAR method improves the best objective value'), but the experimental design contains no evidence about how often this ordering holds across graphs. The 300-node classical results appear to share the same single-instance limitation. Without multi-instance data, the observed ordering Td=100 > Td=50 > Td=0 could be a property of the particular graphs tested rather than of delay-gate-induced NDAR.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a hardware-oriented enhancement of the Noise-Directed Adaptive Remapping (NDAR) method: inserting delay gates before measurement to strengthen amplitude damping noise and thereby steer sampling toward low-Hamming-weight states, which are assumed to be the noise attractor |0...0>. The authors report experiments on IBM's Heron processor for 80-node MaxCut problems in two settings, unweighted-sparse and fully connected weighted graphs, comparing single-layer QAOA and random circuits at delay times Td = 0, 50, and 100 microseconds. They report EBest/ESA values around 0.926(3) for the unweighted sparse case at Td = 100 us, with longer delays giving better trajectories, and a QAOA advantage over random circuits in the weighted dense case at Td = 50 us. They also introduce a classical NDAR variant that samples bits with bit-suppress probability q, showing that higher q (sharper concentration around the attractor) improves solution quality and extends the trend to 300-node instances. The central claim is that increasing delay time improves NDAR performance because stronger amplitude damping focuses sampling near the attractor state.","tokens_in":14136,"tokens_out":5234,"duration_ms":56516,"significance":"If the central claim holds, the paper makes a useful empirical contribution to the emerging line of work that exploits hardware noise rather than mitigating it. The delay-time knob is simple, hardware-relevant, and the comparison between QAOA and random circuits addresses a genuine question about whether problem information in the sampling circuit matters. The classical NDAR variant is a helpful conceptual bridge and provides a concrete, falsifiable baseline with a small number of parameters. The authors are also commendably transparent about limitations: they explicitly note that delay gates affect other noise channels, that parameter transferability may be harmed by noise, and that the classical algorithm is a baseline rather than a simulation of the quantum method. The main weakness is that the headline conclusion is supported by a very narrow empirical base: one graph instance per problem class on the quantum hardware.","major_comments":[{"comment":"The central claim stated in the abstract and in Section IV—that increasing delay time improves the best objective value—rests on a single graph instance per problem class. The 10 independent runs reported in Figs. 1 and 4 are repeated sampling runs on the same instance; they quantify shot noise and run-to-run device drift, not instance-to-instance variability. MaxCut landscapes vary substantially with graph realization even at fixed edge density, so the observed ordering Td = 100 us > Td = 50 us > Td = 0 us may be a property of the particular tested graphs. The authors should either restrict the claim to the tested instances or add multi-instance replication, ideally reporting the fraction of instances on which the delay-time ordering holds and the spread of EBest/ESA across instances.","section":"§IV-A, Fig. 1 and Fig. 4"},{"comment":"The mechanism attributed to the delay gate is amplitude damping toward the all-zeros attractor, but inserting a delay before measurement also increases dephasing, energy relaxation with a different T1/T2 ratio, and any time-dependent gate or measurement drift. The paper acknowledges in Section VI that 'the current delay-gate-induced approach inevitably affects other types of noise besides amplitude damping noise,' but this caveat sits in tension with the title and the causal claim in Section III. Without process characterization (for example, estimating the effective single-qubit channel or comparing with a calibrated amplitude-damping channel), the paper has not established that the observed improvement is specifically due to amplitude damping rather than to a generic reduction of coherence. At minimum, the authors should temper the mechanism language or provide direct evidence that the added delay acts predominantly as amplitude damping.","section":"§III, Algorithm 1"},{"comment":"The qualitative conclusions about the energy and Hamming-weight distributions—especially the claim that QAOA provides a 'more sophisticated exploration strategy' at Td = 50 us—are based on a single 'representative' run selected from ten runs, with no stated selection criterion. Figure 5 shows one run per configuration, and Figure 6 likewise shows one run per configuration. Because the underlying distributions fluctuate across runs, a hand-picked run can overstate or misstate the typical behavior. The authors should either specify a deterministic selection rule (e.g., the run closest to the mean trajectory) or show the spread across all ten runs for the key comparisons.","section":"§IV-B, Figs. 5 and 6"},{"comment":"The QAOA circuits use fixed parameters optimized once on a noiseless MPS simulator for the original Hamiltonian, then applied unchanged across all NDAR iterations and all delay times. This is an intentional parameter-transfer strategy, and the paper correctly notes that re-optimization at each step was used in the earlier NDAR work [15]. However, the conclusion that 'QAOA outperforms random circuits' in the weighted dense case is sensitive to this choice: the fixed parameters may be poorly adapted to the noisy, delay-affected circuit, making the comparison unfair to QAOA or, conversely, making QAOA look better because its noise-biased output happens to align with the attractor. The authors should either re-optimize parameters per iteration (as in [15]) or explicitly test parameter transferability across Td values before drawing conclusions about the relative merit of QAOA as an exploration strategy.","section":"§IV-B, QAOA parameter setting"}],"minor_comments":[{"comment":"The caption reads 'weighted MaxCut problem on a 300-node graph with edge density dedge ≈ 0.3', but the corresponding text in Section V-B describes the low-density unweighted MaxCut case. This is inconsistent and should be corrected.","section":"Fig. 11 caption"},{"comment":"The Require block lists nshots as the number of sampling shots, but line 3 of the pseudocode says 'Sample M bitstrings'; the variable M is not defined in Algorithm 2. Use nshots consistently.","section":"Algorithm 2"},{"comment":"The sentence 'This classical implementation samples solutions near the attractor state and iteratively refines refines it' contains a duplicated word 'refines'. Please fix the typo.","section":"§VI"},{"comment":"The classical NDAR results are presented as evidence that 'quantum NDAR would work effectively even for larger problem instances', but the classical algorithm uses independent bit-suppression sampling and does not emulate the correlations or noise structure of the quantum circuit. The paper already says it is a baseline rather than a simulation, so the summary sentence in Section VI should be softened to avoid implying that the classical results validate quantum scaling.","section":"§V-B"},{"comment":"The paper reports standard errors such as 0.926(3), but it does not report any statistical test for the ordering of Td values. A paired test across the 10 runs, or a statement that the standard errors make the ordering significant, would strengthen the claim.","section":"§IV-B"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid empirical study from an industry group, and the delay-time idea is worth publishing if the claims are appropriately scoped. The main editorial concern is that the headline claim is presented as a general property of NDAR while the quantum evidence is a single instance per graph class. I would recommend asking for multi-instance replication or a clearly scoped claim, plus a more careful treatment of the amplitude-damping mechanism given the acknowledged presence of other noise channels. The classical NDAR section is a useful contribution and could be strengthened by clarifying that it is a stylized model rather than a quantum emulation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the paper. The concrete new thing is the delay-gate insertion before measurement as a controllable way to strengthen amplitude damping in NDAR, and the classical bit-suppress analogue (Algorithm 2) that approximates random-circuit NDAR. Both are simple and usable. The quantum experiments on Heron are reported with standard errors over 10 runs, and the comparison against simulated annealing gives a reasonable sense of absolute quality. QAOA vs random circuit is a sensible probe for whether the encoded Hamiltonian matters under strong damping. I agree with the reader that there is no circularity or derivation gap: this is empirical, and the assumptions about the |0...0> attractor are inherited from Maciejewski et al., not invented here.\n\nThe soft spots are real but not disqualifying. The headline claim—longer delay improves EBest—is supported by one 80-node graph per problem class (plus a single 300-node classical case). Ten runs on the same instance quantify shot/run noise, not instance-to-instance variability. MaxCut landscapes vary a lot with graph realization, so the ordering Td=100 > 50 > 0 could easily be an artifact of the particular graphs. The paper states the result generally in the abstract, and the 'representative run' figures have no stated selection criterion. The random circuit setup (depth 2, gate set) is under-specified enough that another group would struggle to reproduce the exact circuits. There is also no head-to-head against the original NDAR protocol with per-step parameter re-optimization [15] on the same instances, so the incremental value of delay gates is not fully isolated.\n\nNone of this sinks the paper. The delay-gate idea is a useful knob and the classical NDAR is a fine baseline for people tuning NDAR-like heuristics. It just needs to be framed as a demonstration on a few instances, not a general property. A serious referee should push for multi-instance results, clearer random-circuit specification, and code/data release.\n\nWho is this for? Practitioners running NDAR-style experiments on superconducting hardware and researchers interested in noise exploitation for QAOA. It deserves peer review—it is incremental but concrete, and the field needs more honest hardware studies like this. I would engage with it as a referee.","headline":"Delay-gate control of NDAR is a neat, honestly reported knob, but the central trend rests on one instance per graph class and should be read as a proof-of-concept rather than a general result.","tokens_in":14738,"tokens_out":2085,"would_cite":true,"duration_ms":20595,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that adding delay gates before measurement strengthens amplitude damping and improves NDAR's best solutions on MaxCut, with QAOA and random circuits performing similarly except on dense weighted instances.","keywords":["NDAR","amplitude damping","delay gates","MaxCut","QAOA","exploration-exploitation","Hamming weight","local search"],"falsifier":"On the same device and MaxCut instance, run NDAR at 100 microseconds while measuring the Hamming-weight distribution of the sampled bitstrings each iteration: if EBest improves but the distribution does not shift toward lower Hamming weight, the stated amplitude-damping mechanism is wrong; conversely, if a 100-microsecond delay added to a purely dephasing channel never improves EBest/ESA over zero delay, the claim that delay-gate-induced damping is the cause fails.","tokens_in":13689,"feed_emoji":"⚛️","tokens_out":6900,"duration_ms":69135,"temperature":0.7,"pith_summary":"The paper claims that NDAR can be improved by deliberately strengthening amplitude-damping noise through delay gates inserted before measurement, and that increasing delay time improves the best objective value found in each NDAR iteration. On a 133-qubit superconducting processor, both QAOA and random circuits reach best-energy ratios around 0.9 of simulated annealing on 80-node sparse unweighted MaxCut when the delay is 50 or 100 microseconds, while zero delay shows no iteration-to-iteration improvement. On dense weighted MaxCut, longer delays also help but solution quality is lower, and at 50 microseconds QAOA edges out random circuits. The paper's classical analogue, which samples each bit as 0 with probability q to emulate the noise attractor, reaches EBest/ESA of about 0.96 on 80-node sparse MaxCut and about 0.955 on 300-node instances, supporting the picture of NDAR as local search in Hamming space. This matters because it turns a hardware noise source into a tunable search bias rather than an obstacle.","feed_headline":"Longer delays improve NDAR's best MaxCut solutions","feed_subtitle":"On sparse 80-node MaxCut, 100 microsecond delays reach about 93% of simulated annealing quality; QAOA and random circuits match.","key_machinery":"The operational core is the bit-flip gauge transformation $P_y H P_y$, which changes the signs of the Hamiltonian's fields and couplings while preserving its spectrum. Under amplitude damping with $|0\\cdots0\\rangle$ as the assumed attractor, remapping the best sampled bitstring $y$ to the attractor means that the next iteration's samples concentrate around a higher-quality solution; the delay gate of duration $T_d$ controls how tightly they concentrate. The classical NDAR isolates the same mechanism by setting each bit to 0 with probability $q$, so $q$ plays the role of the damping strength. A supporting assumption is QAOA parameter concentration, which lets the authors fix one set of variational parameters across all NDAR iterations.","core_discovery":"The central claim is that longer delay gates produce stronger amplitude damping, which concentrates measurement outcomes around the all-zeros attractor, and each NDAR gauge transformation remaps the current best solution to that attractor. Under a 100 microsecond delay, EBest/ESA reaches 0.926 on unweighted sparse MaxCut at 80 nodes; under 50 microseconds it reaches 0.908, and without a delay gate NDAR does not improve over iterations. For the fully connected weighted case, QAOA at 100 microseconds reaches 0.717 and at 50 microseconds reaches 0.619, with QAOA slightly outperforming random sampling at 50 microseconds. The similar QAOA and random-circuit trajectories on sparse graphs indicate that the damping-induced neighborhood, not the cost-encoded circuit, is doing most of the work; the classical NDAR result that larger bit-suppress probability q yields better solutions reinforces this interpretation by showing that Hamming-weight concentration around the attractor is the controlling mechanism.","pith_inferences":["An implicit consequence is that NDAR performance can be predicted from the radial distribution of samples in Hamming space, which means classical distance-biased samplers could be used to screen problem instances and delay settings before spending hardware time.","A testable extension would replace the Bernoulli bit-suppress sampler with an explicit Hamming-distance-targeted sampler or an energy-weighted sampler; comparing these would separate the contribution of proximity to the attractor from the contribution of energy guidance.","The authors hint at manually constructing attractor states; the classical results suggest that any bias concentrating samples near the current best, not necessarily amplitude damping, should drive the same local-search dynamics, which could be tested by applying NDAR to classical local-search neighborhoods with controlled radii."],"forward_implications":["If longer delay times genuinely strengthen NDAR, then a delay gate is a zero-cost exploitation dial: one can vary the exploration-exploitation balance without changing the circuit structure or variational parameters.","Because random circuits match QAOA on sparse unweighted MaxCut, NDAR's gain on these problems does not require the sampler to encode the Hamiltonian; a classically simulable distance-biased sampler can reproduce much of it.","Classical NDAR at 300 nodes reaching EBest/ESA around 0.955 suggests the same iterative remapping should remain effective beyond the roughly 100-qubit limit of current devices.","The QAOA advantage seen on dense weighted MaxCut at 50 microseconds indicates that in harder landscapes, Hamiltonian-encoding exploration may raise the ceiling once damping is not so strong that it erases the circuit's information."],"supporting_citations":[{"why":"Defines the NDAR loop, the bit-flip gauge transformation, and the |0...0> attractor assumption that this paper inherits.","marker":"[15]"},{"why":"Defines QAOA, the cost-mixer ansatz used as one of the two exploration circuits.","marker":"[9]"},{"why":"Supplies the MPS simulator used to fix QAOA parameters at iteration zero.","marker":"[25]"},{"why":"Provides the simulated annealing baseline whose best energies define the EBest/ESA ratios reported.","marker":"[26]"},{"why":"Supports the parameter-concentration assumption that justifies reusing one QAOA parameter set across NDAR iterations.","marker":"[21]"},{"why":"Supplies the T1 and T2 decoherence timescales used to choose the 0, 50, and 100 microsecond delay settings.","marker":"[24]"}],"fun_headline_variants":["100µs delays boost quantum optimiser to 93% of classical baseline","Delay gates sharpen NDAR on sparse MaxCut: QAOA ties random circuits","Amplitude damping via delay improves quantum optimisation results","Longer delays, better solutions: NDAR on 80-node MaxCut hits 93%","QAOA and random circuits match under damping, but QAOA wins dense"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument presumes that inserting delay gates strengthens amplitude damping toward the all-zeros attractor fast enough that gauge remapping still focuses sampling on better solutions, rather than merely dephasing the circuit into useless noise; if the delay mostly adds decoherence without a directed Hamming-weight bias, the method's improvement would not follow.","fun_headline_variants_meta":{"raw":{"variants":["100µs delays boost quantum optimiser to 93% of classical baseline","Delay gates sharpen NDAR on sparse MaxCut: QAOA ties random circuits","Amplitude damping via delay improves quantum optimisation results","Longer delays, better solutions: NDAR on 80-node MaxCut hits 93%","QAOA and random circuits match under damping, but QAOA wins dense"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000463,"raw_usage":{"total_tokens":2357,"prompt_tokens":1029,"completion_tokens":1328,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":1229}},"tokens_in":645,"tokens_out":1328,"duration_ms":10791,"temperature":1.0,"reasoning_tokens":1229,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:27:31.921372+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On the same device and MaxCut instance, run NDAR at 100 microseconds while measuring the Hamming-weight distribution of the sampled bitstrings each iteration: if EBest improves but the distribution does not shift toward lower Hamming weight, the stated amplitude-damping mechanism is wrong; conversely, if a 100-microsecond delay added to a purely dephasing channel never improves EBest/ESA over zero delay, the claim that delay-gate-induced damping is the cause fails.","supporting_citations":[{"cited_title":"A quantum approximate optimization algorithm,","cited_arxiv_id":null,"evidence_quote":"Defines QAOA, the cost-mixer ansatz used as one of the two exploration circuits."},{"cited_title":"Javadi-Abhari, M","cited_arxiv_id":null,"evidence_quote":"Supplies the MPS simulator used to fix QAOA parameters at iteration zero."}],"review_version":1}