Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Statistical inference for quantum singular models

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper proves QWAIC is an asymptotically unbiased estimator of quantum generalization loss for singular Bayesian state models.

desk verdict Serious quantum extension of singular learning theory with real proofs, but the singular-case unbiasedness of QWAIC is conditional on an unproven algebraic-geometric assumption (S2), and the numerical support is thin. read the letter →

arxiv 2411.16396 v1 pith:VGQ2MT52 submitted 2024-11-25 quant-ph cs.LGmath.AGstat.ML

classification quant-phcs.LGmath.AGstat.ML MSC 62F1562B1014E1581P50
keywords quantumstateestimationsingularlearningtheoryBayesianinferencemodelselectionWAICclassicalshadowsrelativeentropyalgebraicgeometry
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to bring Watanabe's singular learning theory into quantum state estimation. It defines a quantum generalization loss $G^Q_n = -\operatorname{Tr}(\rho \log \sigma_B)$ and a training loss $T^Q_n$ built from classical shadows, and proves asymptotic expansions for both in regular and singular regimes. In the singular regime, where neither the classical nor the quantum regularity conditions hold, the expansions are governed by the real log canonical threshold of a simultaneous desingularization of the KL divergence and the quantum relative entropy. On this foundation the paper defines QWAIC and proves $\mathbb{E}_{X^n}[G^Q_n] = \mathbb{E}_{X^n}[\mathrm{QWAIC}] + o(1/n)$, making QWAIC a computable model-selection score for singular quantum state models. A sympathetic reader should care because model selection for quantum tomographic models has so far relied on regular-model criteria such as AIC, which break down at singularities.

What carries the argument

The argument is carried by the simultaneous log resolution of singularities for the two average loss functions: the classical Kullback-Leibler divergence $K(\theta) = \mathrm{KL}(p(x|\theta_0)\|p(x|\theta))$ and the quantum relative entropy $K_Q(\theta) = D(\sigma(\theta_0^Q)\|\sigma(\theta))$. Hironaka-type resolution gives local normal crossing forms $K(g(u)) = u^{2k}$ and $K_Q(g(u)) = r(u) u^{2k_Q}$, and the proof needs the vanishing orders to coincide, $k = k_Q$ (Assumption S2). The real log canonical threshold $\lambda$ (the learning coefficient) and singular fluctuation $\nu$ of the classical theory enter through the factors $r_{CQ}\lambda$ and $r_{CQ}\nu$, while two new quantum quantities, $\nu_Q$ and $\chi_Q$, absorb the shadow-based training loss. Classical shadows are the second load-bearing element: the snapshot $\hat{\rho}_x$ is an unbiased estimator of $\rho$, which lets the training loss $T^Q_n$ be defined directly from measurement data and lets the quantum log-likelihood ratio $f_Q(\hat{\rho}_x,\theta)$ be analytic in $\theta$. QWAIC's correction term $C^Q_n$ is the posterior covariance of the classical log-likelihood and its shadow quantum analog, and the proof shows this covariance converges to the $\chi_Q$ term in the loss expansion.

What would settle it

Take a tomographically complete measurement and a parametric state family that satisfies the fundamental conditions but whose quantum relative entropy vanishes to higher order than the classical KL divergence on the optimal set; compute $\mathbb{E}_{X^n}[G^Q_n]-\mathbb{E}_{X^n}[\mathrm{QWAIC}]$ at increasing $n$. If the difference does not decay as $o(1/n)$, then Theorem III.4's singular-case claim fails.

Watch

Extended reading notes

Core claim

On its own terms, the central claim is that the gap between quantum training and generalization in Bayesian state estimation can be estimated from data, even for singular models. Defining the Bayesian mean $\sigma_B = \int \sigma(\theta) p(\theta|x^n)\,d\theta$ and using classical shadows $\hat{\rho}_x$, the paper sets $G^Q_n = -\operatorname{Tr}(\rho \log \sigma_B)$ and $T^Q_n = -(1/n)\sum_i \operatorname{Tr}(\hat{\rho}_{x_i} \log \sigma_B)$. It then derives expansions: in the regular case $\mathbb{E}_{X^n}[G^Q_n] = -\operatorname{Tr}(\rho \log \sigma(\theta_0)) + (\lambda_Q+\nu'_Q-\nu_Q)/n + o(1/n)$, and in the singular case $\mathbb{E}_{X^n}[G^Q_n] = -\operatorname{Tr}(\rho \log \sigma(\theta_0)) + (r_{CQ}\lambda + r_{CQ}\nu - \nu_Q)/n + o(1/n)$. The key result, Theorem III.4, states that $\mathrm{QWAIC} := T^Q_n + (1/n)\sum_i \operatorname{Cov}_\theta[\log p(x_i|\theta), \operatorname{Tr}(\hat{\rho}_{x_i} \log \sigma(\theta))]$ satisfies $\mathbb{E}_{X^n}[G^Q_n] = \mathbb{E}_{X^n}[\mathrm{QWAIC}] + o(1/n)$ under the singular-case assumptions. This makes QWAIC the quantum analog of WAIC: a quantity computed from measurement outcomes that predicts the expected out-of-sample quantum relative entropy loss.

Load-bearing premise

The main proof for singular models assumes that, after resolving singularities, the classical Kullback-Leibler loss and the quantum relative entropy vanish at the same order ($k = k_Q$); the paper states this as Assumption S2 and leaves it as a conjecture whether the fundamental conditions imply it.

Editorial extensions

If this is right

  • QWAIC gives a computable model-selection score for Bayesian quantum state estimation that remains asymptotically unbiased for singular models such as neural-network quantum states and quantum Boltzmann machines.
  • The learning coefficient for quantum state estimation is $\lambda_Q+\nu'_Q-\nu_Q$ in regular cases and $r_{CQ}\lambda+r_{CQ}\nu-\nu_Q$ in singular cases, so the effective model complexity depends on the ratio of quantum to classical Fisher information.
  • In the realizable regular case $\lambda_Q = (1/2)\operatorname{Tr}(J_Q J^{-1})$, so the second term of QWAIC reflects the measurement penalty as well as the parameter count.
  • The difference between expected quantum generalization and training loss is $2\chi_Q/n$ to leading order, so QWAIC's covariance term estimates exactly the overfitting gap.
  • If the theory is right, singular quantum models, like their classical counterparts, learn faster in the sense that a smaller learning coefficient gives a smaller expected generalization loss.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to drop Assumption S2 and define a two-coefficient criterion using separate real log canonical thresholds for $K$ and $K_Q$; the simple multiplicative factor $r_{CQ}$ would fail, but an additive correction might still be built.
  • The covariance term $C^Q_n$ can be read as an empirical 'effective number of parameters' for a singular quantum model, and the numerics in the paper already hint that it tracks $r_{CQ}\lambda$ rather than the raw parameter dimension.
  • Because the proof only needs the classical shadow's unbiasedness and analyticity of the shadow log-likelihood, the same construction should transfer to loss functions other than the quantum relative entropy, such as Petz-Rényi divergences, with a modified covariance term.
  • The exponential cost of matrix logarithms in QWAIC is a practical barrier; shadow-tailored or variational approximations to the covariance term could make the criterion scalable to larger systems.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper extends Watanabe's singular learning theory to Bayesian quantum state estimation. It defines quantum generalization and training losses through quantum relative entropy and classical shadows, derives asymptotic expansions for these losses in both regular and singular cases using a simultaneous resolution of singularities of the KL divergence and the quantum relative entropy, and introduces QWAIC as a data-computable asymptotically unbiased estimator of the quantum generalization loss. The formal proofs are given in Appendix D under Fundamental Conditions I and II and Assumptions R1/R2 (regular) or S1/S2 (singular), and the paper includes analytic and numerical examples for one-qubit models.

Significance. If the main results hold, this is a substantial step: it provides a quantum analog of WAIC for model selection among singular quantum state models, introduces a likelihood-like object via classical shadows, and identifies an algebraic-geometric coefficient r_CQ lambda that plays the role of a quantum learning coefficient. The paper is unusually careful in stating assumptions, and the appendices contain detailed proof arguments rather than heuristic sketches. The main reservation is that the singular-case unbiasedness of QWAIC is formally conditional on Assumption S2, which is not proved and is explicitly left as Conjecture B.6; the numerical examples all satisfy KQ proportional to K, so they do not exercise the case where S2 could fail.

major comments (3)
  1. [Assumption S2 / Eq. (B.1); Theorems D.14-D.16 and D.21] The singular-case proof of QWAIC unbiasedness is conditional on Assumption S2, the equality k = kQ in the simultaneous log resolution, and this assumption is not proved. Lemma D.14 uses kQ = k to write fQ(...) = u^k aQ(...) and to bound the posterior moments of u^{kQ}; Theorem D.15 uses the same equality to replace E_theta[r(u)u^{2kQ}] by r_CQ E_theta[u^{2k}] before applying the classical singular expansion with lambda and nu. If kQ differs from k, the coefficient r_CQ multiplying lambda in Theorem D.16 is not justified, and the o(1/n) expansion of GQ_n and TQ_n in the singular case is not established. The paper itself lists Conjectures A.8 and B.6 as future work in Section V. The informal statement of Theorem III.4, saying QWAIC is unbiased 'even when' neither regularity condition holds, therefore overstates the formally proven result. Please either prove Assumption S2 under the stated fundamental conditions (or under an explicitly stated and verifiable additional condition), or reformulate the abstract, Theorem III.4, and the conclusions so that the singular-case claim is explicitly conditional on S2. The numerical examples in Section IV do not resolve this, since in every example KQ is proportional to K (e.g., r(u)=3), so they do not test the possibility kQ != k.
  2. [Lemma D.4, Eq. (D.9); Lemmas D.5 and D.14] The bounds for the training-loss cumulants require the empirical average (1/n) sum_i \hat\rho_{x_i} to be positive semidefinite, because the proof applies Tr(A|B|) <= Tr(A|B|) and H\"older-type inequalities that require a positive semidefinite weight A. Classical shadow snapshots are not physical states in general, so this empirical average need not be positive semidefinite for fixed n. The proof asserts that 'in the asymptotic limit' the average is positive semidefinite, but no uniform or high-probability argument is supplied. Since Theorem D.3's o_p(1/n) residuals for TQ_n depend on these bounds in both the regular case (Lemma D.5) and the singular case (Lemma D.14), the training-loss expansion is not fully rigorous as written. A concentration argument for the minimum eigenvalue of the empirical shadow average, or a modified inequality that avoids the PSD requirement, would close this gap.
  3. [Section III.C / Definition D.9 / Theorem D.21] The regular and singular cases are treated with different assumptions: Theorem D.11 uses R1 and R2, while Theorem D.21 uses S1 and S2. The main text and abstract present QWAIC as an unbiased estimator of GQ_n without prominently displaying this dependence. In particular, the singular claim depends on S2, which is conjectural. The relation between the ordinary WAIC proof (Theorem C.15, which requires only Watanabe's relatively finite variance) and the quantum proof should be clarified: the quantum singular statement is not a direct analogue unless S2 is either proved or explicitly assumed as a hypothesis of the main theorem. I recommend stating the singular-case theorem with 'under Assumptions S1 and S2, and assuming Conjecture B.6' in a displayed form, and adjusting the informal theorem statements accordingly.
minor comments (4)
  1. [Throughout] There are numerous typos and formatting slips that should be corrected: 'Fundamendal condition' in Appendix A, 'Wayanabe' after Proposition A.5, 'Theorms' before Section IV, 'asympto unbiased' in Theorem C.15, and a duplicated phrase ('the analysis of the properties of the analysis of the properties') after Lemma D.5.
  2. [Section IV.B / Figures 4-5] The captions and legends would be clearer if they distinguished the empirical values (CQ_n, QWAIC) from the analytic comparison curves (e.g., 8.08/n and 3*2/n) more explicitly; currently the labels such as '#params/#shots' are easy to misread.
  3. [Section III.B after Theorem III.3] The sentence 'The quantity r_CQ lambda generalizes lambda_Q' would benefit from a precise statement of the reduction: in the regular realizable case, lambda_Q = (1/2)Tr(JQ J^{-1}), while r_CQ lambda is constructed from the normal crossing data; explaining in what sense the latter reduces to the former would help the reader.
  4. [Glossary / Section III.A] The notation \hat\rho_x for a classical snapshot is used in Definition D.9 before it is fully specified in Section III.A; please ensure that the definition of a snapshot as a function of a measurement outcome x is stated once in a prominent place and then used consistently.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the QWAIC unbiasedness proof is derived from algebraic-geometric expansions and empirical-process arguments, not from the definition of QWAIC or from fitted parameters.

full rationale

The paper's central claim is that QWAIC (Definition D.9) estimates the quantum generalization loss GQ_n up to o(1/n) in both regular and singular cases. This is not circular: QWAIC is defined as a data-computable quantity, while GQ_n is defined from the true state and the posterior predictive state; the unbiasedness is established through the asymptotic expansions in Theorems D.8 and D.16, rather than by identifying the two quantities by construction. The constants λ, ν, rCQ, λQ, νQ, and related quantities are derived from normal-crossing representations and Fisher-information matrices, not fitted from the data being predicted. The singular-case proof is conditional on Assumptions S1 and S2; in particular, Assumption S2 (k = kQ) is unproved and is posed as Conjecture B.6, with resolving it explicitly listed as future work. That is a genuine correctness and scope limitation, but it is not circular: the theorem states its hypotheses, and the conclusion does not presuppose S2. The numerical comparisons use analytically computed constants (e.g., 8.08 from Tr(JQ J^{-1}) with JQ ≈ 10.565 and J ≈ 1.308, and rCQ = 3), so the simulations do not fit the predicted quantity. Self-citation to the authors' earlier QAIC paper [65] is used only as motivation and comparison, not as the load-bearing premise of Theorem III.4. Overall, no equation in the derivation reduces to its own input by construction.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper's central theorems are conditional on standard algebraic geometry (resolution of singularities), Watanabe's singular learning theory, and three domain assumptions (Fundamental conditions I/II, S1, S2). No parameters are fitted to data; all constants are derived from the model geometry or from Fisher information.

assumptions (4)
  • standard math Hironaka's resolution of singularities produces a simultaneous log resolution for K and KQ with normal crossing forms (Eq. (B.1)).
    Used in Appendix B.1 and Theorem B.2 to obtain u^{2k} and r(u)u^{2kQ} representations; foundational for defining RLCT and the asymptotic expansions.
  • domain assumption Fundamental conditions I and II: ρ and σ(θ) full rank, fQ(ρ̂,θ) analytic in θ taking values in L²(q), and Θ compact semi-analytic.
    Stated in Appendix A; required for empirical process convergence (Propositions B.8, B.10) and for the desingularization argument.
  • domain assumption Assumption S1: relatively finite variances for f and fQ.
    Appendix A; quantum analog of Watanabe's condition, used in Lemmas D.13, D.14, D.18, D.20 to control log-likelihood ratio functions and standard forms.
  • ad hoc to paper Assumption S2: k = kQ in Eq. (B.1).
    Appendix A; needed for the singular-case expansions and QWAIC unbiasedness. Unproven; Conjecture B.6 is an open problem.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Statistical inference for quantum singular models." pith.science (2026). https://pith.science/paper/VGQ2MT52

@misc{pith2026241116396,
  author       = {Pith},
  title        = {Pith review of: Statistical inference for quantum singular models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VGQ2MT52}},
  note         = {Machine review of arXiv:2411.16396}
}
read the original abstract

Deep learning has seen substantial achievements, with numerical and theoretical evidence suggesting that singularities of statistical models are considered a contributing factor to its performance. From this remarkable success of classical statistical models, it is naturally expected that quantum singular models will play a vital role in many quantum statistical tasks. However, while the theory of quantum statistical models in regular cases has been established, theoretical understanding of quantum singular models is still limited. To investigate the statistical properties of quantum singular models, we focus on two prominent tasks in quantum statistical inference: quantum state estimation and model selection. In particular, we base our study on classical singular learning theory and seek to extend it within the framework of Bayesian quantum state estimation. To this end, we define quantum generalization and training loss functions and give their asymptotic expansions through algebraic geometrical methods. The key idea of the proof is the introduction of a quantum analog of the likelihood function using classical shadows. Consequently, we construct an asymptotically unbiased estimator of the quantum generalization loss, the quantum widely applicable information criterion (QWAIC), as a computable model selection metric from given measurement outcomes.

Figures

Figures reproduced from arXiv: 2411.16396 by the authors.

Figure 1
Figure 1. FIG. 1: Our setting in quantum state estimation. [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2: Resolution of singularities of a parameter space. Technically, obtaining the normal crossing representation [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3: Resolution of the singularity of [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: FIG. 4: Numerical results for a regular model. [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5 [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: FIG. 6: Relationship between the conditions of the model [106, Figure 1.6]. [PITH_FULL_IMAGE:figures/full_fig_p026_6.png]
Figure 7
Figure 7. Figure 7: FIG. 7: Relationship between the properties of quantum and classical models surrounding the realizability condition. [PITH_FULL_IMAGE:figures/full_fig_p029_7.png]
Figure 8
Figure 8. Figure 8: FIG. 8: Relationship between the conditions of the quantum and classical models. [PITH_FULL_IMAGE:figures/full_fig_p030_8.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Quantum and Quantum-Inspired Maximum Likelihood Estimation and Filtering of Stochastic Volatility Models

    quant-ph 2025-07 reject novelty 5.0 of 10

    The authors propose quantum and quantum-inspired hidden Markov model estimators for stochastic volatility and claim a quadratic reduction in hidden states, but the theoretical proof and likelihood formulas contain ser...

Reference graph

Works this paper leans on

33 extracted references · 33 canonical work pages · cited by 1 Pith paper

  1. [1]

    Assumption S1 implies that σ(θQ 0 ) =σ(θQ 1 ), p (x|θ0) =p(x|θ1)

  2. [2]

    Assumptions S1 and R2 imply that ΘQ 0 = Θ0̸=ϕ

  3. [3]

    homogeneous

    Conversely, assuming ΘQ 0 = Θ0̸=ϕ, the homogeneous condition for quantum models, that is σ(θQ 0 ) =σ(θQ 1 ) is equivalent to the one for the associated pair (q(x),p (x|θ)), that is p(x|θ0) =p(x|θ1). 28 Proof. The latter statement of (1) follows from [38, Lemma 3]. By our assumption on the supports of ρ andσ(θ), the former part also follows in the same way...

  4. [4]

    Assumptions R1 and R2 imply Assumption S2

  5. [5]

    In particular, under Assumption R1, Assumption R2 is equivalent to Assumption S2

    Assumption S2 implies Assumption R2. In particular, under Assumption R1, Assumption R2 is equivalent to Assumption S2. Proof. Item (1) follows from Proposition A.5 and Lemma A.10. (2) It follows from Assumption R2 that Θ0 and ΘQ 0 consist of only one points. Combined with Assumption R2, two optimal sets Θ0 and ΘQ 0 must coincide. Furthermore, the Hessian ...

  6. [6]

    Before introducing the real log canonical thresholds, let us briefly recall the results on the log resolution of singularities, originally proved by Hironaka [88, 89]

    Algebraic geometry: log resolution of singularities In this subsection, we recall the basic notions of algebraic geometry, particularly the resolution of singularities and real log canonical thresholds, which are fundamental concepts in singular learning theory. Before introducing the real log canonical thresholds, let us briefly recall the results on the...

  7. [8]

    The morphism g is isomorphic over Xsm

  8. [9]

    31 For our purposes, we refer to an alternative version of this theorem that is more directly applicable to our context, as Atiyah wrote [86]

    The inverse image g−1(Xsing) of the singular locus is a simple normal crossing divisor. 31 For our purposes, we refer to an alternative version of this theorem that is more directly applicable to our context, as Atiyah wrote [86]. Theorem B.2 ([37, Theorem 2.3]). Letf(x) be a non-constant real analytic function from a neighborhood of the origin in Rd to R...

Show all 33 references
  1. [10]

    The map g induces a real analytic isomorphism between U\U0 and W\W0 where W0 :={x∈W|f(x) = 0}, U 0 :=g−1(U0)

  2. [11]

    Moreover, if f(x)≥ 0 for any x, then the integers κ1,··· ,κd are even

    Around any p∈U0, there is a local coordinate u = (u1,··· ,ud) of U so that f(g(u)) =Suκ1 1 ··· uκd d , g′(u) =b(u)uh1 1 ··· uhd d whereS is a constant, b(u) is a real analytic function with b(p)̸= 0 and κ1,··· ,κd,h 1,··· ,hd are nonnegative integers. Moreover, if f(x)≥ 0 for ...

  3. [12]

    As a more general situation, even when Θ Q 0 ̸= Θ0, we can also obtain Eq. (B.1). This is a direct consequence of the simultaneous resolution of singularities

  4. [13]

    It is known that the smallerλ is, the worse the singularity in algebraic geometry

    As we introduce in Definition C.12, we define the learning coefficient λ of a model as the RLCT for Θ 0. It is known that the smallerλ is, the worse the singularity in algebraic geometry. Theorem III.3 tells us this situation corresponds to a smaller generalization loss. This ...

  5. [14]

    Empirical process with log-likelihood Empirical processes provide a framework for analyzing the random behavior of empirical functions based on observed data, enabling investigations into the asymptotic behavior of estimators of parameters (or posterior distribution) and train...

  6. [15]

    The log-likelihood ratio function is defined by f(x,θ ) := logp(x|θ0) p(x|θ)

  7. [16]

    We denote by I := I(θ0) and J := J(θ0)

    Matrices I(θ) and J(θ) are defined by I(θ) := EX [(∂ logp(X|θ) ∂θ )(∂ logp(X|θ) ∂θ )T] , J (θ) :=−EX [∂2 logp(X|θ) ∂θ2 ] , where J(θ) is the Hessian matrix of the KL divergence. We denote by I := I(θ0) and J := J(θ0). In the realizable case, where p(x|θ0) = q(x),∀x holds, I(θ)...

  8. [17]

    Definition B.9

    Empirical process with classical shadow First, we introduce the following as an analog of Appendix B 2. Definition B.9. Let us define some quantities related to quantum information theory

  9. [18]

    The quantum log-likelihood ratio function is defined by fQ(ˆρx,θ ) := Tr ( ˆρx{logσ(θQ 0 )− logσ(θ)} ) , with a classical snapshot ˆρx

  10. [19]

    all models are wrong

    Matrices IQ(θ) and JQ(θ) are defined by IQ(θ) := Tr ( ρ (∂ logσ(θ) ∂θ )(∂ logσ(θ) ∂θ )T) , J Q(θ) :=− Tr ( ρ∂2 logσ(θ) ∂θ2 ) , whereJQ(θ) is the Hessian matrix of the quantum relative entropy. We denote byIQ :=I(θQ 0 ) andJQ :=J(θQ 0 ). In the realizable case, where σ(θQ 0 ) =...

  11. [20]

    Introductory concepts and regular learning theory In this subsection, let us review the basic theorem of Bayesian statistics and the learning theory for regular models. Specifically, we focus on the asymptotic behavior of the generalization and training losses (1), which shed ...

  12. [21]

    The generalization loss Gn and training loss Tn can be expanded as follows: Gn = EX[− logp(X|θ0)] + d 2n + 1 2∆T nJ∆n− 1 2n Tr ( IJ−1) +op (1 n ) , Tn =− 1 n n∑ i=1 logp(xi|θ0) + d 2n− 1 2∆T nJ∆n− 1 2n Tr ( IJ−1) +op (1 n )

  13. [22]

    Remark C.7

    The expected values of Gn and Tn with respect to given samples are EXn[Gn] = EX[− logp(X|θ0)] +λ n +o (1 n ) , EXn[Tn] = EX[− logp(X|θ0)] +λ− 2ν n +o (1 n ) , where λ = d 2, ν = 1 2 Tr ( IJ−1) . Remark C.7. If the regularity (Definition A.1) and realizability (Definition A.11)...

  14. [23]

    In such cases, the Hessian of the KL divergence degenerates, which hinders the formulation of a learning theory in the same way as in regular cases

    Singular learning theory For singular models, the optimal parameter set Θ 0 contains multiple (singular) points in general. In such cases, the Hessian of the KL divergence degenerates, which hinders the formulation of a learning theory in the same way as in regular cases. One ...

  15. [24]

    In other words, λ := min i {2ki + 1 hi } , where ki and hi are defined through a log resolution of singularities Eq

    The learning coefficient λ is defined as the real log canonical threshold associated to Θ 0. In other words, λ := min i {2ki + 1 hi } , where ki and hi are defined through a log resolution of singularities Eq. (B.1)

  16. [25]

    The renormalized log likelihood function a(x,u ) is defined from the standard form f(x,g (u)) = uk·a(x,u ) in [38, Definition 14]

  17. [26]

    The renormalized empirical processξn(u) is introduced in Eq. (B.6)

  18. [27]

    The origin of this is the state density formula and the Mellin transformation

    Let t :=n·u2k∈ R. The origin of this is the state density formula and the Mellin transformation. In this paper, we treat it simply as a variable

  19. [28]

    Note that V (ξn) is a functional of ξn because Vθ[·] depends on ξn

    Let us define the functional variance and singular fluctuation as V (ξn) := EX[Vθ[ √ ta(X,u )]], ν := 1 2 Eξ[V (ξ)]. Note that V (ξn) is a functional of ξn because Vθ[·] depends on ξn. It is worth reminding that, contrary to regular cases, the posterior distribution cannot be ...

  20. [29]

    The generalization and training losses are expanded as Gn = EX [− logp(X|θ0)] + 1 n ( λ + 1 2 Eθ[ √ tξn(u)]− 1 2V (ξn) ) +op (1 n ) , Tn = 1 n n∑ i=1 {− logp(xi|θ0)} + 1 n ( λ− 1 2 Eθ[ √ tξn(u)]− 1 2V (ξn) ) +op (1 n )

  21. [30]

    Now, though we omit the details of the proof here, let us state that WAIC is a generalization of AIC in the following sense

    Their expectations are given by EXn[Gn] = EX[− logp(X|θ0)] +λ n +o (1 n ) , EXn[Tn] = EX[− logp(X|θ0)] +λ− 2ν n +o (1 n ) . Now, though we omit the details of the proof here, let us state that WAIC is a generalization of AIC in the following sense. Theorem C.15 ([38, Theorem 1...

  22. [31]

    Definition D.1

    Basic Bayes theorem To study the asymptotic behaviors of GQ n and TQ n , we introduce the cumulant generating function of the quantum log-likelihood. Definition D.1. Let α∈ R. The cumulant generating function of the quantum log-likelihood is defined by sQ(ˆρ,α ) := Tr(ˆρ log Φ...

  23. [32]

    This assumption allows us to utilize the benefits coming from the classical and quantum statistics in view of the Fisher information matrix

    Regular cases In this subsection, we consider regular cases, where quantum models satisfy the regularity condition (Assumption R1 and R2). This assumption allows us to utilize the benefits coming from the classical and quantum statistics in view of the Fisher information matri...

  24. [33]

    To define WAIC, one had to assume an L2 and finiteness property of the classical log-likelihood ratio function

    Singular cases Finally, we develop a theory to deal with quantum singular models. To define WAIC, one had to assume an L2 and finiteness property of the classical log-likelihood ratio function. To analyze the singularities arising from a quantum models (ρ,σ (θ)), we introduce ...

  25. [88]

    We present the theorem relevant to our discussion: Theorem B.1

    for real spaces and [89] for complex spaces. We present the theorem relevant to our discussion: Theorem B.1. LetX be a real analytic space and Xsm (resp. Xsing) be its smooth (resp. singular) part. Then, a log resolution of singularities exists. In other words, there exists a ...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.