REVIEW 3 major objections 4 minor 1 cited by
Statistical inference for quantum singular models
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper proves QWAIC is an asymptotically unbiased estimator of quantum generalization loss for singular Bayesian state models.
desk verdict Serious quantum extension of singular learning theory with real proofs, but the singular-case unbiasedness of QWAIC is conditional on an unproven algebraic-geometric assumption (S2), and the numerical support is thin. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by the simultaneous log resolution of singularities for the two average loss functions: the classical Kullback-Leibler divergence $K(\theta) = \mathrm{KL}(p(x|\theta_0)\|p(x|\theta))$ and the quantum relative entropy $K_Q(\theta) = D(\sigma(\theta_0^Q)\|\sigma(\theta))$. Hironaka-type resolution gives local normal crossing forms $K(g(u)) = u^{2k}$ and $K_Q(g(u)) = r(u) u^{2k_Q}$, and the proof needs the vanishing orders to coincide, $k = k_Q$ (Assumption S2). The real log canonical threshold $\lambda$ (the learning coefficient) and singular fluctuation $\nu$ of the classical theory enter through the factors $r_{CQ}\lambda$ and $r_{CQ}\nu$, while two new quantum quantities, $\nu_Q$ and $\chi_Q$, absorb the shadow-based training loss. Classical shadows are the second load-bearing element: the snapshot $\hat{\rho}_x$ is an unbiased estimator of $\rho$, which lets the training loss $T^Q_n$ be defined directly from measurement data and lets the quantum log-likelihood ratio $f_Q(\hat{\rho}_x,\theta)$ be analytic in $\theta$. QWAIC's correction term $C^Q_n$ is the posterior covariance of the classical log-likelihood and its shadow quantum analog, and the proof shows this covariance converges to the $\chi_Q$ term in the loss expansion.
What would settle it
Take a tomographically complete measurement and a parametric state family that satisfies the fundamental conditions but whose quantum relative entropy vanishes to higher order than the classical KL divergence on the optimal set; compute $\mathbb{E}_{X^n}[G^Q_n]-\mathbb{E}_{X^n}[\mathrm{QWAIC}]$ at increasing $n$. If the difference does not decay as $o(1/n)$, then Theorem III.4's singular-case claim fails.
Extended reading notes
Core claim
On its own terms, the central claim is that the gap between quantum training and generalization in Bayesian state estimation can be estimated from data, even for singular models. Defining the Bayesian mean $\sigma_B = \int \sigma(\theta) p(\theta|x^n)\,d\theta$ and using classical shadows $\hat{\rho}_x$, the paper sets $G^Q_n = -\operatorname{Tr}(\rho \log \sigma_B)$ and $T^Q_n = -(1/n)\sum_i \operatorname{Tr}(\hat{\rho}_{x_i} \log \sigma_B)$. It then derives expansions: in the regular case $\mathbb{E}_{X^n}[G^Q_n] = -\operatorname{Tr}(\rho \log \sigma(\theta_0)) + (\lambda_Q+\nu'_Q-\nu_Q)/n + o(1/n)$, and in the singular case $\mathbb{E}_{X^n}[G^Q_n] = -\operatorname{Tr}(\rho \log \sigma(\theta_0)) + (r_{CQ}\lambda + r_{CQ}\nu - \nu_Q)/n + o(1/n)$. The key result, Theorem III.4, states that $\mathrm{QWAIC} := T^Q_n + (1/n)\sum_i \operatorname{Cov}_\theta[\log p(x_i|\theta), \operatorname{Tr}(\hat{\rho}_{x_i} \log \sigma(\theta))]$ satisfies $\mathbb{E}_{X^n}[G^Q_n] = \mathbb{E}_{X^n}[\mathrm{QWAIC}] + o(1/n)$ under the singular-case assumptions. This makes QWAIC the quantum analog of WAIC: a quantity computed from measurement outcomes that predicts the expected out-of-sample quantum relative entropy loss.
Load-bearing premise
The main proof for singular models assumes that, after resolving singularities, the classical Kullback-Leibler loss and the quantum relative entropy vanish at the same order ($k = k_Q$); the paper states this as Assumption S2 and leaves it as a conjecture whether the fundamental conditions imply it.
Editorial extensions
If this is right
- QWAIC gives a computable model-selection score for Bayesian quantum state estimation that remains asymptotically unbiased for singular models such as neural-network quantum states and quantum Boltzmann machines.
- The learning coefficient for quantum state estimation is $\lambda_Q+\nu'_Q-\nu_Q$ in regular cases and $r_{CQ}\lambda+r_{CQ}\nu-\nu_Q$ in singular cases, so the effective model complexity depends on the ratio of quantum to classical Fisher information.
- In the realizable regular case $\lambda_Q = (1/2)\operatorname{Tr}(J_Q J^{-1})$, so the second term of QWAIC reflects the measurement penalty as well as the parameter count.
- The difference between expected quantum generalization and training loss is $2\chi_Q/n$ to leading order, so QWAIC's covariance term estimates exactly the overfitting gap.
- If the theory is right, singular quantum models, like their classical counterparts, learn faster in the sense that a smaller learning coefficient gives a smaller expected generalization loss.
Reading between the lines
- A natural extension is to drop Assumption S2 and define a two-coefficient criterion using separate real log canonical thresholds for $K$ and $K_Q$; the simple multiplicative factor $r_{CQ}$ would fail, but an additive correction might still be built.
- The covariance term $C^Q_n$ can be read as an empirical 'effective number of parameters' for a singular quantum model, and the numerics in the paper already hint that it tracks $r_{CQ}\lambda$ rather than the raw parameter dimension.
- Because the proof only needs the classical shadow's unbiasedness and analyticity of the shadow log-likelihood, the same construction should transfer to loss functions other than the quantum relative entropy, such as Petz-Rényi divergences, with a modified covariance term.
- The exponential cost of matrix logarithms in QWAIC is a practical barrier; shadow-tailored or variational approximations to the covariance term could make the criterion scalable to larger systems.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends Watanabe's singular learning theory to Bayesian quantum state estimation. It defines quantum generalization and training losses through quantum relative entropy and classical shadows, derives asymptotic expansions for these losses in both regular and singular cases using a simultaneous resolution of singularities of the KL divergence and the quantum relative entropy, and introduces QWAIC as a data-computable asymptotically unbiased estimator of the quantum generalization loss. The formal proofs are given in Appendix D under Fundamental Conditions I and II and Assumptions R1/R2 (regular) or S1/S2 (singular), and the paper includes analytic and numerical examples for one-qubit models.
Significance. If the main results hold, this is a substantial step: it provides a quantum analog of WAIC for model selection among singular quantum state models, introduces a likelihood-like object via classical shadows, and identifies an algebraic-geometric coefficient r_CQ lambda that plays the role of a quantum learning coefficient. The paper is unusually careful in stating assumptions, and the appendices contain detailed proof arguments rather than heuristic sketches. The main reservation is that the singular-case unbiasedness of QWAIC is formally conditional on Assumption S2, which is not proved and is explicitly left as Conjecture B.6; the numerical examples all satisfy KQ proportional to K, so they do not exercise the case where S2 could fail.
major comments (3)
- [Assumption S2 / Eq. (B.1); Theorems D.14-D.16 and D.21] The singular-case proof of QWAIC unbiasedness is conditional on Assumption S2, the equality k = kQ in the simultaneous log resolution, and this assumption is not proved. Lemma D.14 uses kQ = k to write fQ(...) = u^k aQ(...) and to bound the posterior moments of u^{kQ}; Theorem D.15 uses the same equality to replace E_theta[r(u)u^{2kQ}] by r_CQ E_theta[u^{2k}] before applying the classical singular expansion with lambda and nu. If kQ differs from k, the coefficient r_CQ multiplying lambda in Theorem D.16 is not justified, and the o(1/n) expansion of GQ_n and TQ_n in the singular case is not established. The paper itself lists Conjectures A.8 and B.6 as future work in Section V. The informal statement of Theorem III.4, saying QWAIC is unbiased 'even when' neither regularity condition holds, therefore overstates the formally proven result. Please either prove Assumption S2 under the stated fundamental conditions (or under an explicitly stated and verifiable additional condition), or reformulate the abstract, Theorem III.4, and the conclusions so that the singular-case claim is explicitly conditional on S2. The numerical examples in Section IV do not resolve this, since in every example KQ is proportional to K (e.g., r(u)=3), so they do not test the possibility kQ != k.
- [Lemma D.4, Eq. (D.9); Lemmas D.5 and D.14] The bounds for the training-loss cumulants require the empirical average (1/n) sum_i \hat\rho_{x_i} to be positive semidefinite, because the proof applies Tr(A|B|) <= Tr(A|B|) and H\"older-type inequalities that require a positive semidefinite weight A. Classical shadow snapshots are not physical states in general, so this empirical average need not be positive semidefinite for fixed n. The proof asserts that 'in the asymptotic limit' the average is positive semidefinite, but no uniform or high-probability argument is supplied. Since Theorem D.3's o_p(1/n) residuals for TQ_n depend on these bounds in both the regular case (Lemma D.5) and the singular case (Lemma D.14), the training-loss expansion is not fully rigorous as written. A concentration argument for the minimum eigenvalue of the empirical shadow average, or a modified inequality that avoids the PSD requirement, would close this gap.
- [Section III.C / Definition D.9 / Theorem D.21] The regular and singular cases are treated with different assumptions: Theorem D.11 uses R1 and R2, while Theorem D.21 uses S1 and S2. The main text and abstract present QWAIC as an unbiased estimator of GQ_n without prominently displaying this dependence. In particular, the singular claim depends on S2, which is conjectural. The relation between the ordinary WAIC proof (Theorem C.15, which requires only Watanabe's relatively finite variance) and the quantum proof should be clarified: the quantum singular statement is not a direct analogue unless S2 is either proved or explicitly assumed as a hypothesis of the main theorem. I recommend stating the singular-case theorem with 'under Assumptions S1 and S2, and assuming Conjecture B.6' in a displayed form, and adjusting the informal theorem statements accordingly.
minor comments (4)
- [Throughout] There are numerous typos and formatting slips that should be corrected: 'Fundamendal condition' in Appendix A, 'Wayanabe' after Proposition A.5, 'Theorms' before Section IV, 'asympto unbiased' in Theorem C.15, and a duplicated phrase ('the analysis of the properties of the analysis of the properties') after Lemma D.5.
- [Section IV.B / Figures 4-5] The captions and legends would be clearer if they distinguished the empirical values (CQ_n, QWAIC) from the analytic comparison curves (e.g., 8.08/n and 3*2/n) more explicitly; currently the labels such as '#params/#shots' are easy to misread.
- [Section III.B after Theorem III.3] The sentence 'The quantity r_CQ lambda generalizes lambda_Q' would benefit from a precise statement of the reduction: in the regular realizable case, lambda_Q = (1/2)Tr(JQ J^{-1}), while r_CQ lambda is constructed from the normal crossing data; explaining in what sense the latter reduces to the former would help the reader.
- [Glossary / Section III.A] The notation \hat\rho_x for a classical snapshot is used in Definition D.9 before it is fully specified in Section III.A; please ensure that the definition of a snapshot as a function of a measurement outcome x is stated once in a prominent place and then used consistently.
Circularity Check
No significant circularity: the QWAIC unbiasedness proof is derived from algebraic-geometric expansions and empirical-process arguments, not from the definition of QWAIC or from fitted parameters.
full rationale
The paper's central claim is that QWAIC (Definition D.9) estimates the quantum generalization loss GQ_n up to o(1/n) in both regular and singular cases. This is not circular: QWAIC is defined as a data-computable quantity, while GQ_n is defined from the true state and the posterior predictive state; the unbiasedness is established through the asymptotic expansions in Theorems D.8 and D.16, rather than by identifying the two quantities by construction. The constants λ, ν, rCQ, λQ, νQ, and related quantities are derived from normal-crossing representations and Fisher-information matrices, not fitted from the data being predicted. The singular-case proof is conditional on Assumptions S1 and S2; in particular, Assumption S2 (k = kQ) is unproved and is posed as Conjecture B.6, with resolving it explicitly listed as future work. That is a genuine correctness and scope limitation, but it is not circular: the theorem states its hypotheses, and the conclusion does not presuppose S2. The numerical comparisons use analytically computed constants (e.g., 8.08 from Tr(JQ J^{-1}) with JQ ≈ 10.565 and J ≈ 1.308, and rCQ = 3), so the simulations do not fit the predicted quantity. Self-citation to the authors' earlier QAIC paper [65] is used only as motivation and comparison, not as the load-bearing premise of Theorem III.4. Overall, no equation in the derivation reduces to its own input by construction.
Assumptions & free parameters
assumptions (4)
- standard math Hironaka's resolution of singularities produces a simultaneous log resolution for K and KQ with normal crossing forms (Eq. (B.1)).
- domain assumption Fundamental conditions I and II: ρ and σ(θ) full rank, fQ(ρ̂,θ) analytic in θ taking values in L²(q), and Θ compact semi-analytic.
- domain assumption Assumption S1: relatively finite variances for f and fQ.
- ad hoc to paper Assumption S2: k = kQ in Eq. (B.1).
Cite this review
Pith. "Pith review of Statistical inference for quantum singular models." pith.science (2026). https://pith.science/paper/VGQ2MT52
@misc{pith2026241116396,
author = {Pith},
title = {Pith review of: Statistical inference for quantum singular models},
year = {2026},
howpublished = {\url{https://pith.science/paper/VGQ2MT52}},
note = {Machine review of arXiv:2411.16396}
}
read the original abstract
Deep learning has seen substantial achievements, with numerical and theoretical evidence suggesting that singularities of statistical models are considered a contributing factor to its performance. From this remarkable success of classical statistical models, it is naturally expected that quantum singular models will play a vital role in many quantum statistical tasks. However, while the theory of quantum statistical models in regular cases has been established, theoretical understanding of quantum singular models is still limited. To investigate the statistical properties of quantum singular models, we focus on two prominent tasks in quantum statistical inference: quantum state estimation and model selection. In particular, we base our study on classical singular learning theory and seek to extend it within the framework of Bayesian quantum state estimation. To this end, we define quantum generalization and training loss functions and give their asymptotic expansions through algebraic geometrical methods. The key idea of the proof is the introduction of a quantum analog of the likelihood function using classical shadows. Consequently, we construct an asymptotically unbiased estimator of the quantum generalization loss, the quantum widely applicable information criterion (QWAIC), as a computable model selection metric from given measurement outcomes.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 1 Pith paper
-
On Quantum and Quantum-Inspired Maximum Likelihood Estimation and Filtering of Stochastic Volatility Models
The authors propose quantum and quantum-inspired hidden Markov model estimators for stochastic volatility and claim a quadratic reduction in hidden states, but the theoretical proof and likelihood formulas contain ser...
Reference graph
Works this paper leans on
-
[1]
Assumption S1 implies that σ(θQ 0 ) =σ(θQ 1 ), p (x|θ0) =p(x|θ1)
-
[2]
Assumptions S1 and R2 imply that ΘQ 0 = Θ0̸=ϕ
-
[3]
Conversely, assuming ΘQ 0 = Θ0̸=ϕ, the homogeneous condition for quantum models, that is σ(θQ 0 ) =σ(θQ 1 ) is equivalent to the one for the associated pair (q(x),p (x|θ)), that is p(x|θ0) =p(x|θ1). 28 Proof. The latter statement of (1) follows from [38, Lemma 3]. By our assumption on the supports of ρ andσ(θ), the former part also follows in the same way...
-
[4]
Assumptions R1 and R2 imply Assumption S2
-
[5]
In particular, under Assumption R1, Assumption R2 is equivalent to Assumption S2
Assumption S2 implies Assumption R2. In particular, under Assumption R1, Assumption R2 is equivalent to Assumption S2. Proof. Item (1) follows from Proposition A.5 and Lemma A.10. (2) It follows from Assumption R2 that Θ0 and ΘQ 0 consist of only one points. Combined with Assumption R2, two optimal sets Θ0 and ΘQ 0 must coincide. Furthermore, the Hessian ...
-
[6]
Algebraic geometry: log resolution of singularities In this subsection, we recall the basic notions of algebraic geometry, particularly the resolution of singularities and real log canonical thresholds, which are fundamental concepts in singular learning theory. Before introducing the real log canonical thresholds, let us briefly recall the results on the...
-
[8]
The morphism g is isomorphic over Xsm
-
[9]
The inverse image g−1(Xsing) of the singular locus is a simple normal crossing divisor. 31 For our purposes, we refer to an alternative version of this theorem that is more directly applicable to our context, as Atiyah wrote [86]. Theorem B.2 ([37, Theorem 2.3]). Letf(x) be a non-constant real analytic function from a neighborhood of the origin in Rd to R...
Show all 33 references
-
[10]
The map g induces a real analytic isomorphism between U\U0 and W\W0 where W0 :={x∈W|f(x) = 0}, U 0 :=g−1(U0)
-
[11]
Moreover, if f(x)≥ 0 for any x, then the integers κ1,··· ,κd are even
Around any p∈U0, there is a local coordinate u = (u1,··· ,ud) of U so that f(g(u)) =Suκ1 1 ··· uκd d , g′(u) =b(u)uh1 1 ··· uhd d whereS is a constant, b(u) is a real analytic function with b(p)̸= 0 and κ1,··· ,κd,h 1,··· ,hd are nonnegative integers. Moreover, if f(x)≥ 0 for ...
-
[12]
As a more general situation, even when Θ Q 0 ̸= Θ0, we can also obtain Eq. (B.1). This is a direct consequence of the simultaneous resolution of singularities
-
[13]
It is known that the smallerλ is, the worse the singularity in algebraic geometry
As we introduce in Definition C.12, we define the learning coefficient λ of a model as the RLCT for Θ 0. It is known that the smallerλ is, the worse the singularity in algebraic geometry. Theorem III.3 tells us this situation corresponds to a smaller generalization loss. This ...
-
[14]
Empirical process with log-likelihood Empirical processes provide a framework for analyzing the random behavior of empirical functions based on observed data, enabling investigations into the asymptotic behavior of estimators of parameters (or posterior distribution) and train...
-
[15]
The log-likelihood ratio function is defined by f(x,θ ) := logp(x|θ0) p(x|θ)
-
[16]
We denote by I := I(θ0) and J := J(θ0)
Matrices I(θ) and J(θ) are defined by I(θ) := EX [(∂ logp(X|θ) ∂θ )(∂ logp(X|θ) ∂θ )T] , J (θ) :=−EX [∂2 logp(X|θ) ∂θ2 ] , where J(θ) is the Hessian matrix of the KL divergence. We denote by I := I(θ0) and J := J(θ0). In the realizable case, where p(x|θ0) = q(x),∀x holds, I(θ)...
-
[17]
Definition B.9
Empirical process with classical shadow First, we introduce the following as an analog of Appendix B 2. Definition B.9. Let us define some quantities related to quantum information theory
-
[18]
The quantum log-likelihood ratio function is defined by fQ(ˆρx,θ ) := Tr ( ˆρx{logσ(θQ 0 )− logσ(θ)} ) , with a classical snapshot ˆρx
-
[19]
all models are wrong
Matrices IQ(θ) and JQ(θ) are defined by IQ(θ) := Tr ( ρ (∂ logσ(θ) ∂θ )(∂ logσ(θ) ∂θ )T) , J Q(θ) :=− Tr ( ρ∂2 logσ(θ) ∂θ2 ) , whereJQ(θ) is the Hessian matrix of the quantum relative entropy. We denote byIQ :=I(θQ 0 ) andJQ :=J(θQ 0 ). In the realizable case, where σ(θQ 0 ) =...
-
[20]
Introductory concepts and regular learning theory In this subsection, let us review the basic theorem of Bayesian statistics and the learning theory for regular models. Specifically, we focus on the asymptotic behavior of the generalization and training losses (1), which shed ...
-
[21]
The generalization loss Gn and training loss Tn can be expanded as follows: Gn = EX[− logp(X|θ0)] + d 2n + 1 2∆T nJ∆n− 1 2n Tr ( IJ−1) +op (1 n ) , Tn =− 1 n n∑ i=1 logp(xi|θ0) + d 2n− 1 2∆T nJ∆n− 1 2n Tr ( IJ−1) +op (1 n )
-
[22]
Remark C.7
The expected values of Gn and Tn with respect to given samples are EXn[Gn] = EX[− logp(X|θ0)] +λ n +o (1 n ) , EXn[Tn] = EX[− logp(X|θ0)] +λ− 2ν n +o (1 n ) , where λ = d 2, ν = 1 2 Tr ( IJ−1) . Remark C.7. If the regularity (Definition A.1) and realizability (Definition A.11)...
-
[23]
In such cases, the Hessian of the KL divergence degenerates, which hinders the formulation of a learning theory in the same way as in regular cases
Singular learning theory For singular models, the optimal parameter set Θ 0 contains multiple (singular) points in general. In such cases, the Hessian of the KL divergence degenerates, which hinders the formulation of a learning theory in the same way as in regular cases. One ...
1954
-
[24]
In other words, λ := min i {2ki + 1 hi } , where ki and hi are defined through a log resolution of singularities Eq
The learning coefficient λ is defined as the real log canonical threshold associated to Θ 0. In other words, λ := min i {2ki + 1 hi } , where ki and hi are defined through a log resolution of singularities Eq. (B.1)
-
[25]
The renormalized log likelihood function a(x,u ) is defined from the standard form f(x,g (u)) = uk·a(x,u ) in [38, Definition 14]
-
[26]
The renormalized empirical processξn(u) is introduced in Eq. (B.6)
-
[27]
The origin of this is the state density formula and the Mellin transformation
Let t :=n·u2k∈ R. The origin of this is the state density formula and the Mellin transformation. In this paper, we treat it simply as a variable
-
[28]
Note that V (ξn) is a functional of ξn because Vθ[·] depends on ξn
Let us define the functional variance and singular fluctuation as V (ξn) := EX[Vθ[ √ ta(X,u )]], ν := 1 2 Eξ[V (ξ)]. Note that V (ξn) is a functional of ξn because Vθ[·] depends on ξn. It is worth reminding that, contrary to regular cases, the posterior distribution cannot be ...
-
[29]
The generalization and training losses are expanded as Gn = EX [− logp(X|θ0)] + 1 n ( λ + 1 2 Eθ[ √ tξn(u)]− 1 2V (ξn) ) +op (1 n ) , Tn = 1 n n∑ i=1 {− logp(xi|θ0)} + 1 n ( λ− 1 2 Eθ[ √ tξn(u)]− 1 2V (ξn) ) +op (1 n )
-
[30]
Now, though we omit the details of the proof here, let us state that WAIC is a generalization of AIC in the following sense
Their expectations are given by EXn[Gn] = EX[− logp(X|θ0)] +λ n +o (1 n ) , EXn[Tn] = EX[− logp(X|θ0)] +λ− 2ν n +o (1 n ) . Now, though we omit the details of the proof here, let us state that WAIC is a generalization of AIC in the following sense. Theorem C.15 ([38, Theorem 1...
-
[31]
Definition D.1
Basic Bayes theorem To study the asymptotic behaviors of GQ n and TQ n , we introduce the cumulant generating function of the quantum log-likelihood. Definition D.1. Let α∈ R. The cumulant generating function of the quantum log-likelihood is defined by sQ(ˆρ,α ) := Tr(ˆρ log Φ...
-
[32]
This assumption allows us to utilize the benefits coming from the classical and quantum statistics in view of the Fisher information matrix
Regular cases In this subsection, we consider regular cases, where quantum models satisfy the regularity condition (Assumption R1 and R2). This assumption allows us to utilize the benefits coming from the classical and quantum statistics in view of the Fisher information matri...
-
[33]
To define WAIC, one had to assume an L2 and finiteness property of the classical log-likelihood ratio function
Singular cases Finally, we develop a theory to deal with quantum singular models. To define WAIC, one had to assume an L2 and finiteness property of the classical log-likelihood ratio function. To analyze the singularities arising from a quantum models (ρ,σ (θ)), we introduce ...
-
[88]
We present the theorem relevant to our discussion: Theorem B.1
for real spaces and [89] for complex spaces. We present the theorem relevant to our discussion: Theorem B.1. LetX be a real analytic space and Xsm (resp. Xsing) be its smooth (resp. singular) part. Then, a log resolution of singularities exists. In other words, there exists a ...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.