{"id":"279c606c-b479-415e-9cf1-bf7838b0c3a8","arxiv_id":"2607.15076","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"SQD recovers near-exact ground-state energies from sampling circuits reduced by operator pruning (to 75% of the pool) and Clifford rounding (25% of parameters), enabling up to 2.8× hardware depth reduction at unchanged accuracy.","lead":"This paper shows that a quantum-chemistry pipeline called SQD still finds accurate ground-state energies when the quantum circuit feeding it is aggressively simplified — operators pruned, parameters snapped to Clifford angles. Across 21 molecules the median error stays below the 1.6 mHa chemical-accuracy threshold for moderate compression, with up to 33× faster classical simulation and up to 2.8× shallower hardware circuits.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SQD compression claim is untested in the sparse-sampling regime: hardware and simulation use tiny CS spaces / exact-statevector sampling, where near-exhaustive sampling makes identical SQD energies a small-space artifact rather than evidence of tolerance to imprecise circuits.","rationale":"The reader's conditional verdict is appropriate. I am elevating a concern the reader listed as issue 4: the experiments that are supposed to demonstrate SQD's tolerance to circuit degradation do not isolate that tolerance from small-space saturation. Hardware uses 3–6 CS qubits (8–64-dimensional spaces) with 100k shots, so the selected-CI subspace can be essentially the full space for every configuration; simulation uses exact statevector sampling, bypassing shot noise. This affects both legs of the evidence, whereas the reader's named weakest assumption (single-term gradient proxy) only affects attribution of the pruning benefit, not the core compression phenomenon. A single sparse-sampling simulation on a larger CS space would settle whether SQD really absorbs sampler error in the regime where the method is relevant. If that test fails, the paper's headline should be weakened to a small-space proof of principle; if it passes, the core insight is substantially confirmed. Hence the verdict should remain conditional, not accept or reject.","tokens_in":12897,"tokens_out":11124,"duration_ms":118906,"concrete_test":"Run the existing pipeline on a molecule with ≥16 CS qubits (or a larger active space) with a fixed sparse sample budget—e.g., 2,000 shots per SQD iteration instead of 100k—and compare combined compression (fg=0.5, fc=0.5) against baseline (fg=1, no rounding). Include an exhaustive-sample control (10^6 shots or all basis states) as the floor. If the compressed-circuit SQD error separates from baseline by more than chemical accuracy (1.6 mHa) as the budget shrinks, the central claim fails in the sparse-sampling regime that motivates SQD; if it remains within threshold, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that SQD only needs bitstrings with sufficient ground-state overlap, so circuits can be pruned/rounded aggressively. The evidence for this claim does not actually exercise the sparse-overlap regime. Hardware validation (Table III, Figure 5) uses 3–6 CS qubits; with 100k shots the 2^3–2^6 configuration space is essentially sampled to exhaustion. SQD then diagonalizes nearly the full CS Hamiltonian, and any circuit with nonzero support on the important determinants yields the same near-FCI energy. The paper effectively concedes this in §IV-D: 'the CS qubit counts here are small enough that even a significantly perturbed circuit distributes probability over the correct region.' Simulation (§IV-A.d) samples 'from the exact statevector distribution' with no shot count or sparsity control; with 3–14 CS qubits, near-exhaustive coverage is again plausible. Thus the 'zero loss' hardware result and the 21-molecule ablation may reflect SQD's behavior when sampling is not a limiting resource, not the claimed tolerance to an imprecise sampler in the Hilbert-space regime where SQD is meant to matter. The paper's own limitations section restricts to ≤20 CS qubits and classically simulable systems, so the practical conclusion is unsupported by the data as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that Sample-based Quantum Diagonalization (SQD) is robust to circuit imprecision, and exploits this by compressing the sampling circuit through gradient-based operator pruning and Clifford rounding. The pipeline first projects the molecular Hamiltonian into a contextual subspace, runs a compressed VQE, samples bitstrings, and then classically diagonalizes the resulting selected-CI subspace. The authors report a 21-molecule ablation showing that median SQD error remains within chemical accuracy at moderate compression, a 33× simulation speedup, and hardware experiments on 6 molecules with up to 2.8× transpiled-depth reduction and identical SQD errors across all compression settings.","tokens_in":13123,"tokens_out":5411,"duration_ms":57062,"significance":"If the central claim holds, the paper reframes SQD from a post-processing accuracy booster into an enabler of aggressive circuit compression, which would be a valuable practical insight for near-term quantum chemistry. The paper contains a systematic two-axis ablation, hardware validation on IBM quantum hardware, and a public implementation. The main strength is the clarity of the core observation: SQD only requires ground-state overlap, not accurate energy expectation values. However, the evidence as presented does not exercise the sparse-sampling regime where this robustness would matter, and the headline claims overstate what the data support. The core idea is plausible and worth publishing after substantial revision.","major_comments":[{"comment":"The abstract claims 'median SQD error stays within chemical accuracy even at 50% compression on both axes,' but Table I shows at (fg=0.50, fc=0.50) a median error of 3.79 mHa and at (0.50, 0.25) 3.82 mHa; all entries in the fg=0.50 row exceed the 1.6 mHa threshold. The supported operating point is fg≥0.75 with fc≥0.25. Similarly, the abstract's '33× simulation speedup' corresponds to (fg=0.25, fc=0.00), where the median error is 37.16 mHa. These figures should be reported with their accuracy context, and the wording revised to match the data.","section":"Abstract & §IV-B(b), Table I"},{"comment":"The central robustness claim is tested only in a regime where sampling is nearly exhaustive. Hardware runs use 3–6 CS qubits with 100k shots, which essentially covers the full 2^3–2^6 configuration space; simulation samples from the exact statevector distribution. The identical SQD errors across all compression configurations is therefore consistent with a small-space artifact: SQD diagonalizes nearly the full CS Hamiltonian regardless of circuit quality. The paper itself concedes this in §IV-D ('the CS qubit counts here are small enough...'). To support the claim that compressed circuits tolerate SQD, the authors should add shot-limited experiments (e.g., 1k–10k shots) or larger CS spaces, and report the number of unique sampled determinants per batch.","section":"§IV-A(d), §IV-D, Table III"},{"comment":"Gradient pruning uses a single-term gradient proxy evaluated at θ=0 with all other parameters fixed at zero, and the text explicitly states that inter-operator interactions are not accounted for. Since pruning is the dominant source of error increase in Table I, this proxy is load-bearing. The ablation at fg≥0.75 could be explained by redundancy in the UCCSD pool or by the small CS spaces, not by the proxy's validity. The authors should validate the proxy against a coupled-gradient/ADAPT-style ranking, or at least include a random-pruning control at the same fg values to demonstrate that the gradient ranking is informative.","section":"§III-C, Table I"},{"comment":"The contextual-subspace size is selected as the smallest number of CS qubits whose subspace ground energy is within 1.6 mHa of FCI—the same threshold used as the pass/fail criterion for SQD accuracy. For molecules satisfying this condition, any near-exhaustive sample set will return the CS ground state, whose error is below threshold by construction, independent of circuit compression. To isolate the circuit-compression contribution, the authors should report SQD errors relative to the CS ground-state energy rather than FCI, and disclose the CS truncation error per molecule.","section":"§IV-A(b)"},{"comment":"The hardware 'zero loss' result is reported as identical SQD errors to 0.01 mHa across all configurations, but no statistical uncertainty, shot counts of unique postselected determinants, or repeated runs over device noise are provided. With 100k shots on 3–6 qubits, sampling variance is negligible and the invariance is expected. The authors should report per-batch energy distributions and the number of unique sampled determinants to assess whether compression changes the variance or bias of SQD estimates.","section":"§IV-D, Table III"}],"minor_comments":[{"comment":"The caption describes a 'stabilizer-tableau component' and an 'MPS component,' but the method text (§III-C) never defines or uses an MPS representation. Please reconcile the caption with the STN-based implementation.","section":"Figure 1 caption"},{"comment":"The bold markers for within-chemical-accuracy entries are not visible in the manuscript text provided; please ensure they render in the final version.","section":"Table I"},{"comment":"The statement that 'error stays flat across columns until fc=0.00' is not strictly true at fg=0.50 (3.62→3.79 mHa from fc=0.75 to fc=0.50). The qualitative point stands, but the wording should be softened.","section":"§IV-B(b)"},{"comment":"The definition of 'Clifford angle' as kπ/2 is clear, but the set of k values should be stated (integer k) and the convention for wrapping angles should be specified, since parameter values outside [0,2π) may occur after rounding.","section":"§III-C"}],"recommendation":"major_revision","confidential_remarks":"The core idea is attractive and the experimental structure is appropriate, but the current presentation overstates the results: the abstract's 50%-compression claim is contradicted by Table I, and the hardware and simulation evidence do not yet exercise the sparse-sampling regime that would distinguish the proposed mechanism from a small-Hilbert-space artifact. The paper should not be rejected, but it needs substantive additions beyond copy-editing: shot-limited experiments, random-pruning controls, and a more careful separation of contextual-subspace truncation error from circuit-compression error."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear X,\n\nThe paper's real new contribution is the two-axis characterization: prune a UCCSD pool by single-term gradient magnitude, round the least-sensitive parameters to Clifford angles, and show that SQD's recovered energy stays flat over a wide compression range. That is a genuinely useful measurement, and it is not in QSCI, SQD, or ADAPT-QSCI. The 21-molecule ablation is wide, outliers included, and the hardware depth reductions (2.4–2.8× transpiled-depth) are concrete and mechanical. They also make a fair point that most converged VQE parameters sit near Clifford angles, which explains why the rounding is nearly free.\n\nThe soft spots are real but localized. First, the abstract says \"chemical accuracy even at 50% compression on both axes,\" but Table I gives 3.79 mHa at (0.50, 0.50) against a 1.6 mHa threshold. The body's own operating recommendation is f_g ≥ 0.75. That discrepancy should be corrected. Second, the speedup table shows a strange f_c dependence at f_c=0.00 — 9–33× at fixed f_g — without a timing methodology; it looks like the STN simulator's Clifford collapse, but they do not report how runtime was measured. Third, there is no random-pruning control. Given the tiny CS spaces, random pruning might do almost as well, and that would change the interpretation of the gradient proxy. They should add it.\n\nThe bigger caveat, which the stress-test note gets right, is that the evidence never leaves the regime where sampling is not a limiting resource. Hardware uses 3–6 CS qubits with 100k shots; 2^3 to 2^6 is essentially exhausted, so SQD diagonalizes the full CS Hamiltonian and any circuit with nonzero overlap on the important determinants gives the same energy. Simulation samples from the exact statevector with no shot-count control, and the CS spaces are 3–14 qubits. The paper concedes this in §IV-D. The central qualitative claim — SQD tolerates imprecise circuits — is supported in the regime tested, but the regime where SQD matters (sparse sampling in large CS spaces) is exactly where the claim is untested. That does not kill the paper, but it does mean the conclusion should be scoped to small contextual spaces.\n\nOverall, this is a careful, honest empirical study with a correctable overclaim and one important generalization gap. I would send it to review, and I'd ask for the abstract fix, a random-pruning baseline, and some discussion — or better, a small test — of the sparse-sampling regime.\n\nMy two cents: worth a serious referee.","headline":"A genuinely useful ablation showing SQD absorbs heavy circuit compression, but the abstract overclaims and the evidence never reaches the sparse-sampling regime where the claim would matter.","tokens_in":13767,"tokens_out":2363,"would_cite":true,"duration_ms":24779,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sample-based diagonalization lets quantum chemistry circuits be compressed aggressively without losing accuracy.","keywords":["subspace quantum diagonalization","circuit compression","gradient pruning","Clifford rounding","variational quantum eigensolver","quantum chemistry","stabilizer tensor network","contextual subspace"],"falsifier":"Run the same 21-molecule ablation with a random-pruning control (same f_g fractions, operators chosen at random) and with a full ADAPT-VQE-style coupled-gradient ranking; if random pruning at f_g=0.75 also stays within 1.6 mHa, the paper's gradient proxy is not load-bearing, while if coupled gradients keep accuracy at f_g=0.25 where the single-term proxy fails, the proxy is the limiting factor. A second decisive test is to re-run the hardware experiment with far fewer shots (e.g., 1k–10k) on a molecule with at least 10 CS qubits to see whether compressed circuits still yield identical SQD ener","tokens_in":12652,"feed_emoji":"⚛️","tokens_out":4642,"duration_ms":43432,"temperature":0.7,"pith_summary":"The paper argues that Subspace Quantum Diagonalization (SQD), which recovers ground-state energies by classically diagonalizing the Hamiltonian in the space spanned by sampled bitstrings, only needs the sampling circuit to overlap the ground-state subspace—not to produce an accurate energy estimate. Because of this, the circuit can be aggressively simplified. The authors show that two compressions—gradient-based pruning of excitation operators and rounding of parameters to Clifford angles—can be applied to a variational ansatz while keeping median SQD error within chemical accuracy (1.6 mHa) at 50% compression on both axes, and on hardware the transpiled depth drops up to 2.8x with identical SQD energies. The practical upshot: SQD pipelines should be deliberately built from shallow, imprecise circuits, since the classical eigensolve absorbs sampler error. These results are demonstrated on 21 molecules in simulation and 6 on hardware, all in a minimal basis with 3–14 contextual-subspace qubits.","feed_headline":"SQD lets quantum circuits shrink 2.8x with no accuracy loss","feed_subtitle":"Pruning operators and rounding parameters to Clifford angles cuts hardware depth up to 2.8x while SQD keeps energies exact.","key_machinery":"The load-bearing object is the SQD sampler requirement, formalized as overlap on the ground-state subspace rather than energy accuracy. The paper's two compression tools operate on a UCCSD-style ansatz in a contextually reduced Hilbert space: (1) gradient pruning ranks each excitation operator by the magnitude of its single-term gradient ∂E/∂θ_j evaluated at θ_j=0 with all other parameters fixed at zero, keeping the top fraction f_g; (2) Clifford rounding, applied post-convergence, computes each parameter's distance to the nearest Clifford angle kπ/2 and snaps the fraction f_c of parameters with the smallest distances, converting non-Clifford gates into classically trackable Clifford operati","core_discovery":"The central claim is that SQD changes what the quantum circuit must achieve: standalone VQE demands an accurate expectation value, while SQD only requires bitstrings with sufficient overlap on the ground-state subspace. The authors exploit this by pruning low-gradient excitation operators and snapping the least significant parameters to the nearest Clifford angle (kπ/2). In a 4x5 ablation over 21 molecules, median SQD error remains below 1.6 mHa for gradient fractions f_g ≥ 0.75 combined with any Clifford fraction f_c ≥ 0.25, while median simulation speedup reaches 2.47x at (f_g=0.75, f_c=0.25); combining aggressive pruning (f_g=0.25) with full Clifford rounding (f_c=0.00) collapses accuracy","pith_inferences":["If the overlap-only requirement holds at larger scale, SQD pipelines could trade away variational accuracy deliberately, e.g., using hardware-efficient entanglers or even random Clifford circuits as samplers, so long as the sampled subspace retains ground-state overlap.","The single-term gradient proxy at θ=0 is untested against coupled-gradient or random-pruning baselines; a control experiment comparing pruning rankings would reveal whether the robustness is due to the ranking quality or to SQD's tolerance.","The identical hardware energies across all configurations suggest a saturation effect: at 100k shots on 3–6 CS qubits the sample set nearly exhausts or well-covers the relevant subspace, so the compressed circuit's distribution still lands in the recovery region; the sparse-sampling regime where SQD matters most remains unexplored.","The f_c=0.00 phase transition (accuracy intact at f_g=1.00, collapses at f_g≤0.75) implies a threshold in combined expressivity loss; locating that threshold for a given molecule could be done by monitoring ground-state overlap of the sampled distribution rather than final energy."],"forward_implications":["SQD's robustness property means circuit quality can be substantially degraded without changing final energy accuracy, reframing SQD from an accuracy refinement tool into an enabler of circuit compression.","Clifford rounding is nearly free in accuracy cost for f_c ≥ 0.25 across all tested molecules, with median error flat across that range.","Gradient pruning is the dominant source of error growth; f_g ≥ 0.75 is the recommended operating point for molecules with larger contextual subspaces.","On hardware, combined compression reduces transpiled circuit depth by 2.4–2.8x for molecules with more than a few parameters, directly reducing decoherence and gate-error exposure.","The two compression axes contribute roughly multiplicatively to simulation speedup (2.8x from pruning, 2.1x from rounding, 3.8x combined at f_g=0.25, f_c=0.25), and the practical operating point is f_c in [0.25, 0.50]."],"fun_headline_variants":["SQD lets quantum circuits shrink 2.8x with no accuracy loss","Prune operators and round Clifford: SQD cuts depth 2.8x","SQD compression: 33x simulation speedup, chemical accuracy kept","Compress quantum circuits 50% and still hit chemical accuracy","SQD: drop half the circuit, keep chemical accuracy"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The pruning ranking assumes each operator's importance is captured by its single-term gradient at θ=0 with all other parameters frozen, ignoring inter-operator interactions; if joint effects dominate, pruning silently removes operators needed to keep ground-state overlap, and SQD cannot recover configurations the sampler never produces.","fun_headline_variants_meta":{"raw":{"variants":["SQD lets quantum circuits shrink 2.8x with no accuracy loss","Prune operators and round Clifford: SQD cuts depth 2.8x","SQD compression: 33x simulation speedup, chemical accuracy kept","Compress quantum circuits 50% and still hit chemical accuracy","SQD: drop half the circuit, keep chemical accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000424,"raw_usage":{"total_tokens":2020,"prompt_tokens":759,"completion_tokens":1261,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":1167}},"tokens_in":503,"tokens_out":1261,"duration_ms":13813,"temperature":1.0,"reasoning_tokens":1167,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:15:46.613337+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 21-molecule ablation with a random-pruning control (same f_g fractions, operators chosen at random) and with a full ADAPT-VQE-style coupled-gradient ranking; if random pruning at f_g=0.75 also stays within 1.6 mHa, the paper's gradient proxy is not load-bearing, while if coupled gradients keep accuracy at f_g=0.25 where the single-term proxy fails, the proxy is the limiting factor. A second decisive test is to re-run the hardware experiment with far fewer shots (e.g., 1k–10k) on a molecule with at least 10 CS qubits to see whether compressed circuits still yield identical SQD ener","supporting_citations":[],"review_version":1}