Pith. sign in

REVIEW 2 major objections 4 minor 30 references

Quantum tomography for non-iid sources

T0 review · 2 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Projected least-squares quantum tomography stays sample-optimal even when the source is fully adaptive and non-iid.

desk verdict Theorem 1 is a clean and correct extension of PLS tomography to adaptive sources; Theorem 2's O(d^6/epsilon^2) claim is not proven because the projection step is missing a spectral-norm lemma. read the letter →

arxiv 2602.22057 v2 pith:3WDBJK2N submitted 2026-02-25 quant-ph

classification quant-ph
keywords quantumstatetomographyprocessnon-iidsourcesadaptivepreparationprojectedleast-squaressamplecomplexitymatrixmartingaletime-averaged
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proves that projected least-squares (PLS) tomography, a standard estimation protocol, remains statistically optimal when the i.i.d. assumption is dropped: even if each emitted state is chosen adversarially based on all previous outcomes, the sample complexity for reconstructing the time-averaged state is the same as for i.i.d. copies, O(d r^2/epsilon^2) for rank-r states in dimension d. The same holds for quantum process tomography, with O(d^6/epsilon^2) in diamond distance. The reason is that each single-shot estimate is conditionally unbiased given the past, so errors accumulate as a martingale rather than systematic drift. Thus realistic drifting or feedback-controlled devices can be characterized with standard protocols and no loss in scaling, provided the reconstructed object is interpreted as the trajectory average.

What carries the argument

The central mechanism is conditional unbiasedness of the single-shot least-squares estimator. For a measurement from a complex projective 2-design, the estimator for outcome k is (d+1)P_k - 1; for local 2-designs it is a tensor product of (3P-1) factors. Because the prepared state rho_t is fixed at round t, the expectation of the estimator over the measurement outcome equals rho_t. Conditioning on the full past history, the errors form a centered martingale difference sequence, so the accumulated error is a matrix martingale; its spectral-norm deviations are controlled by a matrix martingale concentration inequality using the predictable quadratic variation. A rank-aware trace-norm projectio

What would settle it

A concrete test: run PLS tomography on a controllable quantum source where the preparation is adaptive and the measurement setting is revealed to the source before it chooses each state. If the empirical reconstruction error does not decrease as O(1/N) as predicted, the assumption about hidden independent settings is in fact load-bearing. Conversely, one could attempt to engineer a source whose outcome distribution deviates from the standard quantum probability rule for the nominal prepared state (e.g., through classical interference with the measurement apparatus); any such deviation that bia

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that the error of the projected least-squares estimator with respect to the time-averaged state, ||rho_hat - rho_bar||_1, is bounded by epsilon (plus a rank-tail term) using N = O(d r^2/epsilon^2) samples, regardless of how the sequence of states is generated, as long as each round is a valid quantum measurement on the prepared state. The unbiasedness of the linear estimator at each round is the load-bearing property: E[rho_hat_t | F_{t-1}] = rho_t, so the difference sequence is a martingale difference. A matrix concentration inequality for martingales then gives the same rate as i.i.d. arguments. The result extends to channels by mapping them t

Load-bearing premise

The proof assumes that at each round, once the source has chosen the state, the measurement outcome is random only through the standard quantum probability rule applied to that state, with the measurement setting independent of the state and hidden from the source; if the source could bias the outcome beyond this probabilistic rule, or know the measurement in advance, the conditional unbiasedness fails and the sample-complexity guarantee collapses.

Editorial extensions

If this is right

  • Standard PLS tomography can be applied to data from drifting, noisy, or feedback-controlled hardware without increasing the number of samples; only the target of estimation becomes the time-averaged state or channel.
  • The sample complexity for state tomography remains O(d r^2/epsilon^2) and for process tomography O(d^6/epsilon^2) up to log factors, matching known i.i.d. lower bounds, so adaptivity does not change the fundamental cost of tomography.
  • Existing experimental pipelines that aggregate outcomes into empirical frequencies remain valid even when temporal correlations are present, as long as each round's state is fixed at measurement time.
  • Process tomography inherits the same robustness through channel-state duality, assuming the input state at each round is chosen uniformly at random and independently of the implemented channel.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The martingale argument likely extends to adaptive measurement strategies chosen by the experimenter from private randomness, as long as the measurement setting is independent of the current state and hidden from the source; the paper only states the non-adaptive case, but the conditional-unbiasedness logic does not require the measurement to be fixed in advance.
  • The result suggests that other linear-in-state estimation methods (e.g., classical shadows) may also retain their i.i.d. sample complexity under adaptive preparation, since their single-shot estimators are also conditionally unbiased; this is not shown in the paper.
  • One practical caveat follows from the assumptions: to reap these guarantees, experimentalists must ensure that the source cannot learn the measurement setting before preparing the state; otherwise the estimator may become biased. This may be a real constraint in closed-loop quantum devices that share a classical control system.
  • The time-averaged object is well-defined even if individual states are far from each other; error against the average does not imply the reconstruction captures any individual snapshot, so the protocol is only appropriate for effective device characterization, not for tracking the trajectory.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper extends projected least-squares (PLS) quantum state and process tomography to the non-i.i.d. regime. The source is allowed to emit an arbitrarily adaptive sequence of states or channels; the target is the time-averaged state ρ̄_N or channel Ē_N. The key idea is that each single-shot estimator remains conditionally unbiased given the history, so the accumulated error is a matrix martingale, and Matrix Freedman provides concentration even without independence. Theorem 1 claims O(d r²/ε²) sample complexity for state tomography in trace distance for fully adaptive sources under global or local complex projective 2-design measurements. Theorem 2 claims O(d⁶/ε²) sample complexity for process tomography in diamond distance via Choi-state estimation. The state-tomography proof is largely complete and plausible, aside from a terse trace-norm conversion lemma. The process-tomography proof, however, contains a load-bearing gap: the projection step is asserted without proof, and the claimed O(d⁶/ε²) scaling is not established as written.

Significance. If Theorem 2 can be completed, the paper is a valuable and conceptual contribution: it shows that dropping the i.i.d. assumption does not degrade the sample complexity of a standard, practical tomography protocol, and it identifies the time-averaged object as the natural estimable quantity under adaptive sources. The state-tomography result (Theorem 1) is convincing, has no free parameters, and gives explicit constants; the martingale formulation is elegant and likely correct. The process-tomography extension is plausible and probably follows from known i.i.d. process-tomography analyses, but the manuscript as written does not supply the necessary projection lemma. The paper would be suitable for publication after the missing argument is supplied and the optimality claims are qualified.

major comments (2)
  1. [Section III.B, Eq. (16)-(17)] The proof of Theorem 2 hinges on the assertion that the Frobenius projection of ĥL_N onto the Choi-state set K satisfies ‖ρ̂_PLS − ρ̄_N‖ ≤ ε with the same probability as the spectral bound ‖ĥL_N − ρ̄_N‖ ≤ ε/2. This is not proven. Frobenius non-expansiveness alone gives only ‖ρ̂_PLS − ρ̄_N‖_F ≤ ‖ĥL_N − ρ̄_N‖_F, and since the space has dimension D=d², this yields ‖ρ̂_PLS − ρ̄_N‖_∞ ≤ d·(ε/2). Combining this with the quoted diamond-norm conversion ‖·‖_⋄ ≤ d²‖·‖ would force τ=ε/(2d³) and a sample complexity O(d⁸/ε²), not the claimed O(d⁶/ε²). The manuscript needs an explicit lemma—analogous to Lemma 1 but for Choi states—showing that the projection step does not worsen the spectral error by more than a constant, or a precise citation to the corresponding result in Ref. [8] with the proof reproduced or sketched. As written, this is an internal gap in the second headline result.
  2. [Section III.B, after Eq. (18)] The chain of norm conversions used to pass from the spectral-norm bound on the Choi-state estimator to the diamond-norm bound on channels is not self-contained. The paper quotes ‖·‖_⋄ ≤ d²‖·‖ from Ref. [19] without stating the normalization of the Choi matrix or proving the inequality. This constant directly affects the claimed O(d⁶/ε²) scaling. Please state the exact relation for the normalization used here and provide a proof or a precise theorem reference. If the correct conversion is actually d³ (or requires a trace-norm factor), the sample complexity in Theorem 2 would need to be revised accordingly.
minor comments (4)
  1. [Abstract and Introduction] The abstract says the sample complexity 'matches the optimal i.i.d. scaling', while the Introduction calls PLS 'near-optimal'. The theorem proves O(d r²/ε²), which is the known PLS rate but not necessarily the minimax-optimal rate for rank-r states (typically O(d r/ε²)). Please qualify the optimality claim and say explicitly that the result matches the i.i.d. PLS scaling.
  2. [Appendix B, Lemma 1] Lemma 1 is load-bearing for Theorem 1 but is stated without proof. It is attributed informally to [7]. Please provide a proof in the appendix or give the exact theorem number and a clear statement of its hypotheses. The constant 4r and the dependence on both Σ_r(ρ̄_N) and Σ_r(ρ̂_PLS) should be justified.
  3. [Appendix B, Eq. (B13)] The variance bound drops the −ρ_t² term. This is harmless for the stated O(d) bound, but the text should say '≤' with the positive term removed explicitly, so the reader does not infer an equality.
  4. [Section III] There is a typo in 'Choi–Jamio lkowski' (broken hyphenation). Also, the notation ρ̄_N is used both for the average state and the average Choi state; this is acceptable but should be flagged or distinguished.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the non-iid tomography result is derived from the Born rule, 2-design identities, and external concentration/variance results; the sole self-citation is not load-bearing.

full rationale

The paper's derivation chain is self-contained in the relevant sense. The central mechanism is the conditional unbiasedness of the single-shot estimator, Eq. (7), which follows from the Born-rule assumption Eq. (2) and the 2-design properties of the measurements; this is a mathematical derivation from stated assumptions, not a fitted parameter or a renamed input. The martingale argument uses Matrix Freedman (Theorem 3, cited to Tropp [18]) on the centered differences X_t = rho_hat_t - rho_t, and the variance and range bounds are either derived in Appendix C or cited to external prior work [7,8]. The projection step Lemma 1 is standard PLS error-conversion and is applied, not assumed into the theorem. The process-tomography extension uses the Choi isomorphism and cites external variance/range bounds from Ref. [8] and the diamond-norm conversion from Ref. [19]; these are independent published results, not self-citations. The only self-citation is Ref. [9] (Zambrano et al.) cited in the Introduction as one of the PLS tomography references; it is not used in any proof and is not load-bearing. The skeptic's concern about Theorem 2's projected-operator spectral bound is a proof-gap/correctness-risk issue, not circularity: it concerns an omitted estimate, not an input that is equivalent to the conclusion by construction. Similarly, the extra assumption that the measurement setting is hidden from the adversary is a limitation, not a circular step. There are no fitted inputs called predictions, no uniqueness theorem imported from the authors' own prior work, and no ansatz smuggled in via self-citation.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

No free parameters or invented entities. The sample complexity bounds have explicit constants derived in the proof. The central claim rests on the Born-rule model, 2-design measurement assumptions, and standard concentration/conversion lemmas.

assumptions (6)
  • domain assumption At each round t the source fixes a state ρ_t based solely on the past history F_{t-1}; the measurement outcome Y_t is then distributed by the Born rule with respect to a fixed IC POVM (Eq. (2)).
    Section II.A. This is the foundational statistical model. If ρ_t could depend on the current measurement outcome or on future information, the martingale difference X_t would not be centered.
  • domain assumption The measurement ensemble is a complex projective 2-design (global) or a tensor product of local 2-designs (Pauli).
    Section II.B. The unbiasedness of the single-shot estimator and the variance identities (Appendix C) rely on the 2-design property.
  • standard math Matrix Freedman inequality for matrix martingales.
    Appendix A, Theorem 3. Used to control the accumulated error M_N.
  • standard math Variance identity E[ρhat^2] = (d-1)ρ + d I for global 2-designs, and the local product formula.
    Appendix B, Eq. (B12), derived in Appendix C following Ref [7].
  • standard math Trace-norm conversion lemma (Lemma 1) for the projection step.
    Appendix B, Lemma 1, cited from Ref [7].
  • domain assumption Choi-Jamiolkowski isomorphism and the bound ||·||_⋄ ≤ d^2 ||·||_2 for differences of channels.
    Section III. The diamond norm bound is cited from Ref [19]; the zero-partial-trace structure of channel differences is implicit.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Quantum tomography for non-iid sources." pith.science (2026). https://pith.science/paper/3WDBJK2N

@misc{pith2026260222057,
  author       = {Pith},
  title        = {Pith review of: Quantum tomography for non-iid sources},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3WDBJK2N}},
  note         = {Machine review of arXiv:2602.22057}
}
abstract

Quantum state and process tomography are typically analyzed under the assumption that devices emit independent and identically distributed (i.i.d.) states or channels. In realistic experiments, however, noise, drift, feedback, or adversarial behavior violate this assumption. We show that projected least-squares tomography remains statistically optimal even under fully adaptive state and channel preparation. Specifically, we prove that the sample complexity for reconstructing the time-averaged state or channel matches the optimal i.i.d. scaling for non-adaptive, single-copy measurements. For rank-$r$ states, the sample complexity is $\mathcal{O}(d r^2/\epsilon^2)$ to achieve accuracy $\epsilon$ in trace distance, while for process tomography it is $\mathcal{O}(d^6/\epsilon^2)$ to achieve accuracy $\epsilon$ in diamond distance. Thus, dropping the i.i.d. assumption does not increase the fundamental sample complexity of quantum tomography, but only changes the interpretation of the reconstructed object.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

30 extracted references · 3 linked inside Pith

  1. [8]

    Surawy-Stepney, J

    T. Surawy-Stepney, J. Kahn, R. Kueng, and M. Guta, Projected least-squares quantum process tomography, Quantum6, 844 (2022)

  2. [19]

    Oufkir, Sample-optimal quantum process tomog- raphy with non-adaptive incoherent measurements, arXiv:2301.12925 (2023)

    A. Oufkir, Sample-optimal quantum process tomog- raphy with non-adaptive incoherent measurements, arXiv:2301.12925 (2023)

  3. [1]

    Hradil, Quantum-state estimation, Phys

    Z. Hradil, Quantum-state estimation, Phys. Rev. A55, R1561 (1997)

  4. [2]

    Gross, Y.-K

    D. Gross, Y.-K. Liu, S. T. Flammia, S. Becker, and J. Eis- ert, Quantum state tomography via compressed sensing, Phys. Rev. Lett.105, 150401 (2010)

  5. [3]

    O’Donnell and J

    R. O’Donnell and J. Wright, Efficient quantum tomogra- phy, arXiv:1508.01907 (2015)

  6. [4]

    J. Haah, A. W. Harrow, Z. Ji, X. Wu, and N. Yu, Sample- optimal tomography of quantum states, IEEE Trans. Inf. Theory63, 5628 (2017)

  7. [5]

    Anshu and S

    A. Anshu and S. Arunachalam, A survey on the com- plexity of learning quantum states, Nat. Rev. Phys.6, 59 (2024)

  8. [6]

    J. A. Tropp, User-friendly tail bounds for sums of random matrices, Found. Comput. Math.12, 389 (2012)

Show all 30 references
  1. [7]

    Gut ¸˘ a, J

    M. Gut ¸˘ a, J. Kahn, R. Kueng, and J. A. Tropp, Fast state tomography with optimal error bounds, J. Phys. A: Math. Theor53, 204001 (2020)

  2. [9]

    Zambrano, S

    L. Zambrano, S. Ramos-Calderer, and R. Kueng, Fast quantum measurement tomography with dimension- optimal error bounds, arXiv:2507.04500 (2025)

  3. [10]

    P. V. Klimov, J. Kelly, Z. Chen, M. Neeley, A. Megrant, B. Burkett, R. Barends, K. Arya, B. Chiaro, Y. Chen, A. Dunsworth, A. Fowler, B. Foxen, C. Gidney, M. Giustina, R. Graff, T. Huang, E. Jeffrey, E. Lucero, J. Y. Mutus, O. Naaman, C. Neill, C. Quintana, P. Roushan, D. Sank...

  4. [11]

    Proctor, M

    T. Proctor, M. Revelle, E. Nielsen, K. Rudinger, D. Lob- ser, P. Maunz, R. Blume-Kohout, and K. Young, Detect- ing and tracking drift in quantum information processors, Nat. Commun.11, 5396 (2020)

  5. [12]

    G. A. White, C. D. Hill, F. A. Pollock, L. C. Hollenberg, and K. Modi, Demonstration of non-Markovian process characterisation and control on a quantum processor, Nat. Commun.11, 6301 (2020)

  6. [13]

    McEwen, D

    M. McEwen, D. Kafri, Z. Chen, J. Atalaya, K. Satzinger, C. Quintana, P. V. Klimov, D. Sank, C. Gidney, A. Fowler,et al., Removing leakage-induced correlated errors in superconducting quantum error correction, Nat. Commun.12, 1761 (2021)

  7. [14]

    Pirandola, U

    S. Pirandola, U. L. Andersen, L. Banchi, M. Berta, D. Bunandar, R. Colbeck, D. Englund, T. Gehring, C. Lupo, C. Ottaviani, J. L. Pereira, M. Razavi, J. S. Shaari, M. Tomamichel, V. C. Usenko, G. Vallone, P. Vil- loresi, and P. Wallden, Advances in quantum cryptogra- phy, Adv. ...

  8. [15]

    Fawzi, R

    O. Fawzi, R. Kueng, D. Markham, and A. Oufkir, Learn- ing properties of quantum states without the iid assump- tion, Nat. Commun.15, 9677 (2024)

  9. [16]

    S. J. van Enk and R. Blume-Kohout, When quantum tomography goes wrong: drift of quantum sources and other errors, New J. Phys.15, 025024 (2013)

  10. [17]

    Vershynin,High-Dimensional Probability: An Intro- duction with Applications in Data Science(Cambridge University Press, 2018)

    R. Vershynin,High-Dimensional Probability: An Intro- duction with Applications in Data Science(Cambridge University Press, 2018)

  11. [18]

    J. A. Tropp, Freedman’s inequality for matrix martin- gales, Electron. Commun. Probab.16, 262 (2011)

  12. [20]

    J. M. Renes, R. Blume-Kohout, A. J. Scott, and C. M. Caves, Symmetric informationally complete quantum measurements, J. Math. Phys.45, 2171 (2004)

  13. [21]

    Klappenecker and M

    A. Klappenecker and M. Roetteler, Mutually unbiased bases are complex projective 2-designs, arXiv:0502031 (2005)

  14. [22]

    Dankert, R

    C. Dankert, R. Cleve, J. Emerson, and E. Livine, Exact and approximate unitary 2-designs and their application to fidelity estimation, Phys. Rev. A80, 012304 (2009)

  15. [23]

    Williams,Probability with martingales(Cambridge university press, 1991)

    D. Williams,Probability with martingales(Cambridge university press, 1991). 6 Appendix A: Statistical tools In the context of quantum state tomography via least- squares estimation, we model the sequence of estimators as a matrix-valued stochastic process. The appropriate math...

  16. [24]

    LetF t−1 denote theσ-algebra generated by the full experimental history up to stept−1

    Martingale structure and unbiasedness We now show that the estimation error admits a nat- ural martingale structure. LetF t−1 denote theσ-algebra generated by the full experimental history up to stept−1. This includes the sequence of measurement settings and outcomes, as well ...

  17. [25]

    Concentration bounds for global 2-design measurements Consider a POVM{Π k}M k=1 defined by Π k = d M Pk, where the set of rank-1 projectors{P k}M k=1 forms a com- plex projective 2-design. As shown in Appendix C, if the outcomeY t =kcorresponding to projectorP k is obtained at...

  18. [26]

    This means that for each qubit j= 1,

    Concentration bounds for local 2-design measurements Let the measurement consist of a POVM constructed from local 2-designs. This means that for each qubit j= 1, . . . , n, we have a local 2-design defined by rank- 1 projectors{P k}m k=1. Then, the POVM elements are Πk = d M P...

  19. [27]

    and P tr Pkj ρj = m 2 , we obtain 2 m mX kj =1 tr Pkj ρj (3Pkj +1) =ρ j + 21.(B19) The global second moment is the tensor product of these local terms: E k∼p(ρt) [ˆρ2 t ] = nO j=1 (ρj + 21) = X α∈P([n]) 2|α|trα(ρt)⊗1 ⊗α,(B20) where tr α(ρt) denotes the partial trace of the ele...

  20. [28]

    To convert these bounds into trace-norm guarantees, we use an operator-to-trace norm conversion lemma based on the effective rank of the state [7]

    T race-norm conversion and final bounds In the previous subsections, we derived concentration bounds in the spectral norm. To convert these bounds into trace-norm guarantees, we use an operator-to-trace norm conversion lemma based on the effective rank of the state [7]. For a ...

  21. [29]

    Consider a measurement de- scribed by a set of POVM elements{Π k}M k=1, where Πk = d M Pk and{P k}M k=1 is a set of rank-1 projectors forming a complex projective 2-design

    Estimator for global 2-designs Theorem 4.Letρbe a quantum state on ad- dimensional Hilbert space. Consider a measurement de- scribed by a set of POVM elements{Π k}M k=1, where Πk = d M Pk and{P k}M k=1 is a set of rank-1 projectors forming a complex projective 2-design. The le...

  22. [30]

    Suppose the measurement POVM elements are tensor products of single-qubit 2-designs

    Estimator for local 2-designs Theorem 5.Letρbe a state on ann-qubit system (d= 2 n). Suppose the measurement POVM elements are tensor products of single-qubit 2-designs. Specifically, let the outcomes be indexed byk= (k 1, . . . , kn), correspond- ing to POVM elements Πk = d M...

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.