{"id":"6fe8d539-db17-4673-a6fb-e5b14017b8e2","arxiv_id":"2511.02555","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Grouping qubits by mutual information and building locally optimal dual frames from reconstructed local states yields unbiased estimators with dramatically lower variance than standard classical shadows.","lead":"This paper introduces 'k-locally optimal' classical shadows: after single-qubit Pauli measurements, it groups correlated qubits and reconstructs each group's state to build measurement-adapted estimators. In simulations on molecules up to 40 qubits, these estimators cut the number of shots needed for accurate observable estimation by orders of magnitude compared to standard classical shadows.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unbiasedness in the finite-shot same-dataset setting is unverified; App. C shows the empirical-frequency variant is biased under exactly this data reuse, and the SDP case is left explicitly open.","rationale":"The reader's weakest_assumption centered on the heuristic MI-grouping premise, which is indeed a real limitation and is candidly acknowledged in Appendix A. However, I judge the more load-bearing issue to be the finite-sample unbiasedness of the actual estimation workflow. The abstract promises 'unbiased estimators,' and the method's practical appeal is that the same single-qubit measurement data can be post-processed into better estimators. But in the reported protocol, the duals are trained on the same shots used for evaluation. Standard frame-theoretic unbiasedness applies to a fixed dual frame; it does not automatically survive the dependence between the empirical frequencies and the data-dependent duals. The paper's own Appendix C explicitly shows that the closely related empirical-frequency construction is biased under exactly this reuse, and the authors leave the SDP-based case as an open question. That is a self-identified gap in the support for the central claim.\n\nI considered alternative concerns: the truncated-MPS simulations may favor local measurements, and the choice of MI grouping is heuristic and can substantially affect performance. These are legitimate but they weaken the generality of the numerical demonstrations, not the core logical validity of the estimator construction. The same-dataset reuse issue, by contrast, directly threatens the claim that the estimators are unbiased and that the reported order-of-magnitude improvements are achievable in the way the protocol is specified. The paper deserves credit for being transparent enough to include this caveat, but a conditional verdict is exactly right: the central claim needs one additional validation step before it can be accepted as stated.","tokens_in":24239,"tokens_out":5845,"duration_ms":65397,"concrete_test":"Split the S=10^6 shots into disjoint training and evaluation sets (e.g., 5×10^5 shots each). Construct k-LO duals using only the training half; compute estimates on the evaluation half for the TLD1433 40-qubit ground-state energy and for a small molecule such as H2O (14 qubits). Repeat over, say, 100 independent splits and compare the mean and root-mean-square error of the k-LO estimates to CS-Pauli evaluated on the same evaluation shots. If the k-LO estimates are centered on the exact ansatz energy and the order-of-magnitude advantage persists out-of-sample, the concern is resolved. If the estimates are biased or the advantage shrinks substantially, the reported in-sample results are not evidence for the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that k-LO shadows give unbiased estimators with large variance reductions using only single-qubit measurements and pure post-processing. The estimator is provably unbiased only when the dual frame D_m is fixed independently of the evaluation shots. In the reported protocol, the same S=10^6 shots are used both to construct the k-LO duals (via mutual-information grouping and SDP local tomography, Eq. (15)) and to compute the estimate via the frequency-weighted average in Eq. (5). This creates an in-sample conditioning problem: the empirical frequencies f_m used as weights are correlated with the duals D_m constructed from the same data, so E_f[Σ_m f_m Tr[D_m O]] need not equal Tr[ρO], even though each D_m is formally a valid dual frame.\n\nAppendix C demonstrates exactly this failure for the empirical-frequency variant: for intermediate shot counts the estimates miss the true ground-state energy by several standard errors, and the authors attribute this to a type of overfitting. They then state that they 'cannot provide a reason as to why SDP and closest PSD seem to produce unbiased estimates for the finite dataset or whether they also produce biased estimators but at a much smaller and unnoticeable scale.' Thus the strongest claim — unbiasedness with order-of-magnitude error reduction on the same dataset — rests on an unverified finite-sample property. The variance numbers in Fig. 2 use the exact probabilities (10), not the finite training data, and the TLD1433 results in Fig. 3 use the same dataset for building duals and estimating energies, so they can be optimistically biased.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a state-aware, k-locally optimal (k-LO) shadow estimation protocol. After measuring an n-qubit state with single-qubit Pauli POVMs, the same S measurement outcomes are used (i) to estimate pairwise classical mutual informations, (ii) to partition qubits into disjoint groups of size at most k, (iii) to perform local SDP-based state tomography in each group, and (iv) to build locally optimal dual frames via Eq. (15). These correlated dual frames are then applied in Eq. (5) to estimate any observable. The authors benchmark energy estimation for molecules up to 16 qubits, for TLD1433 up to 40 qubits using truncated MPS states, and for a 50-qubit TFIM, reporting large variance reductions relative to standard classical shadows and several other methods. Appendices provide toy counterexamples, a comparison of grouping algorithms, and a study of finite-shot tomography variants.","tokens_in":24554,"tokens_out":3607,"duration_ms":47077,"significance":"If the protocol performs as claimed, it is practically valuable: it requires only single-qubit measurements, remains observable-agnostic, and moves the optimization to classical post-processing. The paper is also commendably transparent: Appendix A explicitly shows that k-LO duals can be worse than canonical duals, Appendix B shows that the grouping heuristic strongly affects performance, and Appendix C acknowledges an open finite-sample bias question. The core theoretical ingredients—dual-frame reconstruction and locally optimal duals—are standard and correctly used. The significance is conditional, however, on resolving whether the same-dataset estimator is unbiased and on whether the large-system benchmarks are representative, rather than artifacts of truncated states and the chosen grouping heuristic.","major_comments":[{"comment":"The central claim of unbiased estimators is not established for the protocol as implemented. Eq. (5) is unbiased when the duals D_m are fixed independently of the evaluation data, but in the proposed workflow the same S shots determine the mutual-information grouping, the SDP tomography (14), and the frame operator (15). App. C explicitly shows that the empirical-frequency variant is biased under exactly this data reuse, and the authors state they 'cannot provide a reason as to why SDP and closest PSD seem to produce unbiased estimates.' This is a load-bearing gap: the abstract promises 'unbiased estimators,' and the numerical evidence in Fig. 2 reports exact single-shot variances, not finite-sample biases of the same-dataset estimator. Please either prove unbiasedness for the SDP/closest-PSD construction under stated conditions, or weaken the claims and report finite-shot bias diagnosti","section":"§III, Eq. (5) and App. C"},{"comment":"The largest numerical demonstrations—TLD1433 at 28 and 40 qubits—use truncated MPS representations with bond dimension at most 50. The authors acknowledge that 'MPS truncation ... may favor local measurements and k-LO duals.' Since the 40-qubit order-of-magnitude error reduction is the main evidence for scalability, this is not a minor caveat. The claim that k-LO duals reduce errors by orders of magnitude for large molecules would be substantially strengthened by tests on higher-bond-dimension or untruncated ansatz states, or by an explicit demonstration that the observed advantage is robust to the truncation parameter.","section":"§V B, Fig. 3"},{"comment":"The paper correctly frames the method as heuristic, but the abstract and conclusions use unqualified language ('we obtain unbiased estimators that outperform state-of-the-art methods'). Appendix A constructs states for which 1-LO duals have larger observable variance than canonical duals, and Appendix B shows that naive grouping nearly eliminates the advantage. Thus the central performance claim is not a theorem but an empirical observation dependent on the mutual-information grouping heuristic. The manuscript would be more accurate if the abstract and concluding claims were qualified accordingly, e.g., 'in the benchmarked systems, k-LO shadows reduce estimation errors,' rather than implying a general superiority.","section":"App. A and App. B"}],"minor_comments":[{"comment":"Notation is inconsistent: Eq. (15) writes |D^{G_i}_{m_{G_i}}> = F^{-1}_{G_i} |Pi^{G_i}_{m_{G_i}}>, but the frame operator is defined as F_{\\rho_{G_i}}. Please clarify that local optimal duals are constructed from the reconstructed RDM, not from the exact state.","section":"Eq. (15) and Sec. IV C"},{"comment":"The caption for Fig. 2 in Sec. V A appears to be a duplicate of the Fig. 1 caption inserted before the results discussion. The caption should describe the molecular variance comparison properly.","section":"Figure captions"},{"comment":"The text says 'MSE error' where 'MSE' already includes 'error'; also the phrase 'the 2-LO duals ... in fact proves to be the best' should be 'prove.' Minor language issues throughout the appendices.","section":"App. A.4"},{"comment":"The paper would benefit from a short statement on code/data availability. The benchmarks use public data sets, but the implementation details for grouping, SDP tomography, and variance calculations are not specified sufficiently for exact reproduction without access to the authors' code.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid heuristic proposal with careful honest appendices, but the same-dataset unbiasedness gap is central to its advertised contribution. I do not see this as a reject: the error can in principle be fixed by proving a concentration/independence condition, by splitting the data into construction and evaluation shots (at the cost of reduced effective sample size), or by reformulating the claims. However, without such changes the abstract's 'unbiased estimators' claim is too strong. The truncated-MPS benchmark issue is also important for the 'orders of magnitude' claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent, well-written extension of the dual-frame shadow idea, with real engineering value, but the paper claims more than it proves. The specific contribution — greedy mutual-information grouping plus SDP-based local tomography to build tensor-product locally-optimal duals — is new and sensible. The authors are unusually honest: they call the method heuristic, give explicit counterexamples where it underperforms canonical shadows, and compare against standard baselines. The variance reductions in Fig. 2 are real as far as they go, and the RMSE numbers in Appendix D are a fairer comparison because the duals are fixed from a separate large dataset.\n\nThe soft spot is the one the stress-test flags. In the protocol as stated, the same S = 10^6 shots are used both to construct the dual frame and to compute the estimate from the frequency-weighted average in Eq. (5). Unbiasedness is proven only when the duals are fixed independently of the evaluation shots. With data-dependent duals, the law-of-total-expectation argument fails because the shots are not fresh samples conditional on the constructed duals. Appendix C demonstrates exactly this failure for the empirical-frequency variant, and the authors explicitly say they cannot currently explain why SDP and closest-PSD appear unbiased. So the abstract's \"unbiased estimators\" is, at best, an unverified finite-sample guess. That matters: the main selling point is unbiasedness with an order-of-magnitude shot reduction, and a biased estimator changes the comparison.\n\nA couple of smaller issues are worth noting. The variance numbers in Fig. 2 use exact probabilities, not the fluctuations of the training data, so they won't reflect the true end-to-end mean squared error. The TLD1433 results rely on truncated MPS states with bond dimension <= 50, which the authors acknowledge may favor local measurements. And there's no code or data release, which makes the numerical claims hard to reproduce.\n\nStill, this deserves a serious referee. The method is plausible, the exposition is clear, and the open problem of finite-sample bias is explicitly posed in the paper. A referee should push for a proof or a numerical bias characterization on the same-data protocol, and for a comparison that accounts for the training overhead. If the bias turns out small, this is a useful contribution to near-term measurement estimation.\n\nRecommendation: send to peer review.","headline":"Useful practical recipe for state-aware shadows, but the headline claim of unbiasedness is not established for the same-data protocol; the paper's own appendix leaves the finite-sample case open.","tokens_in":25102,"tokens_out":4079,"would_cite":false,"duration_ms":42190,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A classical-shadows protocol that groups qubits by measured correlations and builds state-aware dual frames estimates any observable with far fewer shots than standard shadows, using only single-qubit measurements.","keywords":["classical shadows","dual frames","quantum observable estimation","measurement grouping","mutual information","state tomography","semidefinite programming","variance reduction"],"falsifier":"Run the k-LO protocol on a state whose correlations are spread across many qubits, such as a Greenberger-Horne-Zeilinger-type state, and compare single-shot variance on a global observable against canonical classical shadows; the paper's Appendix A toy examples show there are parameter regions where canonical duals win, and any such region found for k≥2 on a scalable state would falsify the practical claim of universal improvement.","tokens_in":24086,"feed_emoji":"⚛️","tokens_out":5207,"duration_ms":52494,"temperature":0.7,"pith_summary":"The paper claims that the statistical error of classical shadow estimation can be reduced without changing the measurement circuit: keep single-qubit random Pauli measurements, but replace the state-agnostic inversion rule with a state-aware one. After measuring, it uses the observed frequencies to compute pairwise mutual information between qubits, partitions the qubits into correlated blocks of at most k qubits, tomographs each block's reduced state, and constructs the variance-optimal dual frame for each block. Multiplying these local optimal duals gives an unbiased estimator for any observable in pure post-processing. In numerical tests on molecular Hamiltonians up to 40 qubits and a 50-qubit spin chain, the estimator consistently beats canonical classical shadows and most competing methods, often by orders of magnitude in variance. The advantage is heuristic: no general optimality proof is given, and Appendix A exhibits states where k-LO shadows perform worse than canonical ones.","feed_headline":"Correlated shadows from single-qubit measurements cut estimation error","feed_subtitle":"Same simple Pauli measurements, but the post-processing learns correlations — fewer shots for molecular energies and spin observables.","key_machinery":"The machinery is the locally optimal dual frame. In frame theory, an overcomplete POVM admits many dual frames; the optimal one for a given state uses the state probabilities in the frame operator and minimizes estimator variance for every observable. The paper approximates that ideal with k-local dual frames: partition qubits into groups of size at most k using pairwise classical mutual information of outcome frequencies, estimate each group's reduced state by semidefinite-program tomography, and use the resulting probabilities to build the optimal dual on each group. The tensor product of these group duals is the k-LO shadow. This machinery injects state information into the post-processin","core_discovery":"The central claim is that for informationally overcomplete local measurements, the best post-processing is not the canonical dual frame but a correlated, state-dependent one: for each group of qubits that are strongly correlated according to the measured data, one can construct the locally optimal dual frame from the group's tomographed reduced state, and the tensor product of these frames is a valid shadow with lower variance for any observable. This turns classical shadows from state-agnostic to state-aware while retaining single-qubit measurements and observable-agnostic, informationally complete acquisition. The paper validates the claim numerically, showing variance reductions by orders","pith_inferences":["Editorial inference: if the correlation-graph heuristic generalizes, the same mutual-information grouping could guide a second measurement round, making acquisition state-aware as well as post-processing.","Editorial inference: the practical value depends on whether the MI-based partition stays near-optimal for shallow-circuit states preparable on hardware; the truncated-matrix-product-state tests may favor local methods, so hardware-native states are the natural next testbed.","Editorial inference: the paper's own Appendix A suggests a research program—compare k-LO duals against globally optimized k-local duals on many states; any systematic gap would indicate room for further gains from optimization.","Editorial inference: because the duals are state-aware but observable-agnostic, they could serve as a drop-in post-processing layer for noise-mitigated energy estimation pipelines without changing the measurement hardware."],"forward_implications":["Because acquisition is identical to standard classical shadows, any existing single-qubit Pauli dataset can be re-post-processed with k-LO duals to obtain lower-variance estimates without new measurements.","The method estimates many observables simultaneously from one informationally complete dataset, including molecular energies, excitation gaps, and spin correlation functions.","In the reported benchmarks, k-LO duals achieve comparable or better precision than methods that require entangling measurement circuits or observable-tailored adaptive loops, while remaining observable-agnostic.","Even local observables benefit from correlated global post-processing, as demonstrated by the Ising-model correlation-function results.","The protocol is compatible with any local informationally complete POVM, so it can be combined with other measurement strategies rather than replacing them."],"fun_headline_variants":["Local measurements, smart post-processing: lower error shadows","State-aware shadows: same measurements, fewer shots","Correlated post-processing turns local measurements into precise shadows","Single-qubit measurements, optimal dual frames, orders of magnitude less error","Shadow estimation upgrade: learn correlations from single-qubit data"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The promise rests on an unproven heuristic: the qubit blocks chosen by pairwise correlation in single-qubit measurement outcomes are the right blocks for building low-variance shadows, and the paper itself shows two-qubit states where this loses to standard shadows.","fun_headline_variants_meta":{"raw":{"variants":["Local measurements, smart post-processing: lower error shadows","State-aware shadows: same measurements, fewer shots","Correlated post-processing turns local measurements into precise shadows","Single-qubit measurements, optimal dual frames, orders of magnitude less error","Shadow estimation upgrade: learn correlations from single-qubit data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000663,"raw_usage":{"total_tokens":2809,"prompt_tokens":634,"completion_tokens":2175,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":378,"completion_tokens_details":{"reasoning_tokens":2095}},"tokens_in":378,"tokens_out":2175,"duration_ms":14725,"temperature":1.0,"reasoning_tokens":2095,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T00:09:19.566950+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the k-LO protocol on a state whose correlations are spread across many qubits, such as a Greenberger-Horne-Zeilinger-type state, and compare single-shot variance on a global observable against canonical classical shadows; the paper's Appendix A toy examples show there are parameter regions where canonical duals win, and any such region found for k≥2 on a scalable state would falsify the practical claim of universal improvement.","supporting_citations":[],"review_version":1}