{"id":"7e5cae67-9014-4e4d-b551-44096aefea7c","arxiv_id":"2411.16396","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors define quantum generalization and training losses via classical shadows and prove that QWAIC is an asymptotically unbiased estimator of the quantum generalization loss for singular quantum state models.","lead":"This paper extends classical singular learning theory to quantum state estimation, using classical shadows to define quantum generalization and training losses. It introduces QWAIC, an information criterion that estimates the quantum generalization loss even for singular quantum models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unproven Assumption S2 (k=kQ) is the load-bearing hinge for the singular-case proof of QWAIC unbiasedness; if it fails, Theorem III.4 is not established for singular quantum models.","rationale":"The paper is a serious extension of Watanabe's singular learning theory to quantum state estimation, with detailed proofs of the basic Bayesian expansion, the classical-shadow construction, and the QWAIC correction term. The reader's weakest-assumption analysis correctly identifies Assumption S2 as the principal gap: the singular-case proofs of the asymptotic expansions (Theorem D.15/D.16) and hence of QWAIC unbiasedness (Theorem D.21) rely on k=kQ, which is left as Conjecture B.6 rather than proved. I considered whether the definition of r_CQ as a constant extracted from a posterior expectation was a second load-bearing gap, since r(u) can vary along the zero set and the weighted average may depend on the limiting posterior distribution; however, the r_CQ terms cancel in the difference GQ_n - TQ_n, so the unbiasedness claim is less sensitive to that issue than to S2. I also considered whether tomographic completeness might actually force S2 through the injectivity of the measurement map on tangent directions; if so, S2 is true but the paper has not shown it, and the theorem remains conditional as stated. The formal verification is absent and the numerical tests are limited to models with KQ proportional to K, so they provide no evidence about S2. Therefore the honest verdict is conditional: the central claim is plausible and likely correct under a proof of S2, but as written the singular-case unbiasedness theorem is not established unconditionally. This does not change the reader's CONDITIONAL verdict, so I mark the verdict UNCHANGED.","tokens_in":55296,"tokens_out":16230,"duration_ms":165427,"concrete_test":"Settle Conjecture B.6 by proving or disproving that tomographic completeness plus full-rank analyticity force equal vanishing orders. Concretely, use the linear measurement map M: Δ ↦ (Tr(Π_x Δ))_x, which is injective on traceless Hermitians for a T.C. POVM. Near a full-rank optimal state, D(ρ||σ) is locally equivalent to ||Δ||^2 and KL(q||p) is locally equivalent to ||MΔ||^2, which should imply KQ and K have the same first nonzero vanishing order. Make this rigorous; if it holds, upgrade S2 from conjecture to theorem. As a numerical cross-check, search over analytic 1-qubit families, e.g., σ(θ)=I/2 + f_1(θ) X + f_2(θ) Y + f_3(θ) Z with f_j having different lowest-order homogeneous terms, and compute the vanishing orders of K(θ) in Eq. (6) and KQ(θ) in Eq. (7).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, Theorem III.4 / Theorem D.21, is formally proved only under Assumptions S1 and S2. Assumption S2 is the equality k=kQ of the vanishing orders in the simultaneous log resolution (Eq. B.1). It is used twice in essential places. First, Lemma D.14 uses exactly the step 'fQ(...)=u^{kQ}aQ=u^k aQ' to control the higher-order cumulants in the basic expansion Theorem D.3; without k=kQ, the posterior moments of u^{kQ} need not scale as O(n^{-ℓ/2}) and the o(1/n) basic expansion can fail. Second, Theorem D.15 evaluates Eθ[r(u)u^{2kQ}] as r_CQ Eθ[u^{2k}] and then applies the classical singular expansion with λ and ν; if kQ≠k, the coefficient multiplying the RLCT changes and the simple r_CQλ term in Theorem D.16 is not justified. The paper does not prove S2; it is proposed as Conjecture B.6, and the conclusion explicitly lists resolving Conjectures A.8 and B.6 as future work. The informal statement of Theorem III.4 says unbiasedness holds 'even when' the model is singular, but the singular-case proof is conditional on this unverified algebraic-geometric equality. A counterexample with kQ<k would not contradict the classical data-processing inequality K≤KQ, so the assumption is not forced by that inequality alone. The numerical examples all have KQ proportional to K (e.g., r(u)=3), so they do not test S2.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends Watanabe's singular learning theory to Bayesian quantum state estimation. It defines quantum generalization and training losses through quantum relative entropy and classical shadows, derives asymptotic expansions for these losses in both regular and singular cases using a simultaneous resolution of singularities of the KL divergence and the quantum relative entropy, and introduces QWAIC as a data-computable asymptotically unbiased estimator of the quantum generalization loss. The formal proofs are given in Appendix D under Fundamental Conditions I and II and Assumptions R1/R2 (regular) or S1/S2 (singular), and the paper includes analytic and numerical examples for one-qubit models.","tokens_in":55620,"tokens_out":6375,"duration_ms":69949,"significance":"If the main results hold, this is a substantial step: it provides a quantum analog of WAIC for model selection among singular quantum state models, introduces a likelihood-like object via classical shadows, and identifies an algebraic-geometric coefficient r_CQ lambda that plays the role of a quantum learning coefficient. The paper is unusually careful in stating assumptions, and the appendices contain detailed proof arguments rather than heuristic sketches. The main reservation is that the singular-case unbiasedness of QWAIC is formally conditional on Assumption S2, which is not proved and is explicitly left as Conjecture B.6; the numerical examples all satisfy KQ proportional to K, so they do not exercise the case where S2 could fail.","major_comments":[{"comment":"The singular-case proof of QWAIC unbiasedness is conditional on Assumption S2, the equality k = kQ in the simultaneous log resolution, and this assumption is not proved. Lemma D.14 uses kQ = k to write fQ(...) = u^k aQ(...) and to bound the posterior moments of u^{kQ}; Theorem D.15 uses the same equality to replace E_theta[r(u)u^{2kQ}] by r_CQ E_theta[u^{2k}] before applying the classical singular expansion with lambda and nu. If kQ differs from k, the coefficient r_CQ multiplying lambda in Theorem D.16 is not justified, and the o(1/n) expansion of GQ_n and TQ_n in the singular case is not established. The paper itself lists Conjectures A.8 and B.6 as future work in Section V. The informal statement of Theorem III.4, saying QWAIC is unbiased 'even when' neither regularity condition holds, therefore overstates the formally proven result. Please either prove Assumption S2 under the stated fundamental conditions (or under an explicitly stated and verifiable additional condition), or reformulate the abstract, Theorem III.4, and the conclusions so that the singular-case claim is explicitly conditional on S2. The numerical examples in Section IV do not resolve this, since in every example KQ is proportional to K (e.g., r(u)=3), so they do not test the possibility kQ != k.","section":"Assumption S2 / Eq. (B.1); Theorems D.14-D.16 and D.21"},{"comment":"The bounds for the training-loss cumulants require the empirical average (1/n) sum_i \\hat\\rho_{x_i} to be positive semidefinite, because the proof applies Tr(A|B|) <= Tr(A|B|) and H\\\"older-type inequalities that require a positive semidefinite weight A. Classical shadow snapshots are not physical states in general, so this empirical average need not be positive semidefinite for fixed n. The proof asserts that 'in the asymptotic limit' the average is positive semidefinite, but no uniform or high-probability argument is supplied. Since Theorem D.3's o_p(1/n) residuals for TQ_n depend on these bounds in both the regular case (Lemma D.5) and the singular case (Lemma D.14), the training-loss expansion is not fully rigorous as written. A concentration argument for the minimum eigenvalue of the empirical shadow average, or a modified inequality that avoids the PSD requirement, would close this gap.","section":"Lemma D.4, Eq. (D.9); Lemmas D.5 and D.14"},{"comment":"The regular and singular cases are treated with different assumptions: Theorem D.11 uses R1 and R2, while Theorem D.21 uses S1 and S2. The main text and abstract present QWAIC as an unbiased estimator of GQ_n without prominently displaying this dependence. In particular, the singular claim depends on S2, which is conjectural. The relation between the ordinary WAIC proof (Theorem C.15, which requires only Watanabe's relatively finite variance) and the quantum proof should be clarified: the quantum singular statement is not a direct analogue unless S2 is either proved or explicitly assumed as a hypothesis of the main theorem. I recommend stating the singular-case theorem with 'under Assumptions S1 and S2, and assuming Conjecture B.6' in a displayed form, and adjusting the informal theorem statements accordingly.","section":"Section III.C / Definition D.9 / Theorem D.21"}],"minor_comments":[{"comment":"There are numerous typos and formatting slips that should be corrected: 'Fundamendal condition' in Appendix A, 'Wayanabe' after Proposition A.5, 'Theorms' before Section IV, 'asympto unbiased' in Theorem C.15, and a duplicated phrase ('the analysis of the properties of the analysis of the properties') after Lemma D.5.","section":"Throughout"},{"comment":"The captions and legends would be clearer if they distinguished the empirical values (CQ_n, QWAIC) from the analytic comparison curves (e.g., 8.08/n and 3*2/n) more explicitly; currently the labels such as '#params/#shots' are easy to misread.","section":"Section IV.B / Figures 4-5"},{"comment":"The sentence 'The quantity r_CQ lambda generalizes lambda_Q' would benefit from a precise statement of the reduction: in the regular realizable case, lambda_Q = (1/2)Tr(JQ J^{-1}), while r_CQ lambda is constructed from the normal crossing data; explaining in what sense the latter reduces to the former would help the reader.","section":"Section III.B after Theorem III.3"},{"comment":"The notation \\hat\\rho_x for a classical snapshot is used in Definition D.9 before it is fully specified in Section III.A; please ensure that the definition of a snapshot as a function of a measurement outcome x is stated once in a prominent place and then used consistently.","section":"Glossary / Section III.A"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serious contribution with a promising construction and careful formal scaffolding, but the singular-case headline theorem is conditional on a conjectural algebraic-geometric equality (Assumption S2). I do not think this requires rejection: the authors could either prove S2 or transparently present QWAIC's singular-case unbiasedness as conditional, and the PSD gap in Lemma D.4 is likely fixable with a concentration argument. Given the journal context, I would recommend major revision rather than acceptance at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, know this: the paper is a legitimate extension of Watanabe's singular learning theory to Bayesian quantum state estimation, not a repackaging. It defines quantum generalization and training losses using classical shadows, introduces a quantum log-likelihood ratio, proves asymptotic expansions for both regular and singular models, and constructs QWAIC. The proofs in Appendix D are detailed and the assumptions are stated clearly. The regular-case results (Theorems D.6, D.8, D.11) look solid, and the idea of a posterior covariance between the classical log-likelihood and the quantum log-likelihood with snapshots is a sensible generalization of WAIC. The paper is honest: it flags Conjectures A.8 and B.6 as open.\n\nThe soft spot is exactly what the stress-test note says. Assumption S2 (k = kQ in the simultaneous log resolution) is used in two load-bearing places: Lemma D.14 relies on fQ(...) = u^k aQ to get the higher-order cumulant bounds, and Theorem D.15 uses E_theta[r(u)u^{2kQ}] = r_CQ E_theta[u^{2k}] to pull out the r_CQ lambda term in the singular expansion. If kQ differs from k, that simple factor fails and the singular-case proof of QWAIC unbiasedness (Theorem D.21) collapses. The paper does not prove S2; it is Conjecture B.6, and the conclusion lists it as future work. The informal statement of Theorem III.4 says 'even when' the model is singular, but the proof is conditional. Since the data-processing inequality K <= K_Q does not force k = kQ, this is a real gap, not a technicality.\n\nThe numerical validation is also limited: one qubit, mostly classical state models, no error bars, no shipped code or data. The examples do verify S2 for the families tested, but they do not test the general claim.\n\nIs it worth your time? Yes, if you work on quantum state estimation or singular learning theory. The framework is new and the regular-case results are usable. The singular-case claims need an S2 proof or a rephrasing as conditional results. I would send this to a serious referee — the math is substantive — but the referee should push on S2 and on more meaningful numerics. If the conjecture is resolved, this becomes a much stronger paper.","headline":"Serious quantum extension of singular learning theory with real proofs, but the singular-case unbiasedness of QWAIC is conditional on an unproven algebraic-geometric assumption (S2), and the numerical support is thin.","tokens_in":56132,"tokens_out":2445,"would_cite":true,"duration_ms":23582,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62B10","14E15","81P50"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves QWAIC is an asymptotically unbiased estimator of quantum generalization loss for singular Bayesian state models.","keywords":["quantum state estimation","singular learning theory","Bayesian inference","model selection","WAIC","classical shadows","quantum relative entropy","algebraic geometry"],"falsifier":"Take a tomographically complete measurement and a parametric state family that satisfies the fundamental conditions but whose quantum relative entropy vanishes to higher order than the classical KL divergence on the optimal set; compute $\\mathbb{E}_{X^n}[G^Q_n]-\\mathbb{E}_{X^n}[\\mathrm{QWAIC}]$ at increasing $n$. If the difference does not decay as $o(1/n)$, then Theorem III.4's singular-case claim fails.","tokens_in":55078,"feed_emoji":"⚛️","tokens_out":6234,"duration_ms":52824,"temperature":0.7,"pith_summary":"This paper tries to bring Watanabe's singular learning theory into quantum state estimation. It defines a quantum generalization loss $G^Q_n = -\\operatorname{Tr}(\\rho \\log \\sigma_B)$ and a training loss $T^Q_n$ built from classical shadows, and proves asymptotic expansions for both in regular and singular regimes. In the singular regime, where neither the classical nor the quantum regularity conditions hold, the expansions are governed by the real log canonical threshold of a simultaneous desingularization of the KL divergence and the quantum relative entropy. On this foundation the paper defines QWAIC and proves $\\mathbb{E}_{X^n}[G^Q_n] = \\mathbb{E}_{X^n}[\\mathrm{QWAIC}] + o(1/n)$, making QWAIC a computable model-selection score for singular quantum state models. A sympathetic reader should care because model selection for quantum tomographic models has so far relied on regular-model criteria such as AIC, which break down at singularities.","feed_headline":"Singular quantum models get a WAIC-style criterion","feed_subtitle":"A shadow-based Bayesian score estimates the expected quantum relative entropy loss even when regularity conditions fail.","key_machinery":"The argument is carried by the simultaneous log resolution of singularities for the two average loss functions: the classical Kullback-Leibler divergence $K(\\theta) = \\mathrm{KL}(p(x|\\theta_0)\\|p(x|\\theta))$ and the quantum relative entropy $K_Q(\\theta) = D(\\sigma(\\theta_0^Q)\\|\\sigma(\\theta))$. Hironaka-type resolution gives local normal crossing forms $K(g(u)) = u^{2k}$ and $K_Q(g(u)) = r(u) u^{2k_Q}$, and the proof needs the vanishing orders to coincide, $k = k_Q$ (Assumption S2). The real log canonical threshold $\\lambda$ (the learning coefficient) and singular fluctuation $\\nu$ of the classical theory enter through the factors $r_{CQ}\\lambda$ and $r_{CQ}\\nu$, while two new quantum quantities, $\\nu_Q$ and $\\chi_Q$, absorb the shadow-based training loss. Classical shadows are the second load-bearing element: the snapshot $\\hat{\\rho}_x$ is an unbiased estimator of $\\rho$, which lets the training loss $T^Q_n$ be defined directly from measurement data and lets the quantum log-likelihood ratio $f_Q(\\hat{\\rho}_x,\\theta)$ be analytic in $\\theta$. QWAIC's correction term $C^Q_n$ is the posterior covariance of the classical log-likelihood and its shadow quantum analog, and the proof shows this covariance converges to the $\\chi_Q$ term in the loss expansion.","core_discovery":"On its own terms, the central claim is that the gap between quantum training and generalization in Bayesian state estimation can be estimated from data, even for singular models. Defining the Bayesian mean $\\sigma_B = \\int \\sigma(\\theta) p(\\theta|x^n)\\,d\\theta$ and using classical shadows $\\hat{\\rho}_x$, the paper sets $G^Q_n = -\\operatorname{Tr}(\\rho \\log \\sigma_B)$ and $T^Q_n = -(1/n)\\sum_i \\operatorname{Tr}(\\hat{\\rho}_{x_i} \\log \\sigma_B)$. It then derives expansions: in the regular case $\\mathbb{E}_{X^n}[G^Q_n] = -\\operatorname{Tr}(\\rho \\log \\sigma(\\theta_0)) + (\\lambda_Q+\\nu'_Q-\\nu_Q)/n + o(1/n)$, and in the singular case $\\mathbb{E}_{X^n}[G^Q_n] = -\\operatorname{Tr}(\\rho \\log \\sigma(\\theta_0)) + (r_{CQ}\\lambda + r_{CQ}\\nu - \\nu_Q)/n + o(1/n)$. The key result, Theorem III.4, states that $\\mathrm{QWAIC} := T^Q_n + (1/n)\\sum_i \\operatorname{Cov}_\\theta[\\log p(x_i|\\theta), \\operatorname{Tr}(\\hat{\\rho}_{x_i} \\log \\sigma(\\theta))]$ satisfies $\\mathbb{E}_{X^n}[G^Q_n] = \\mathbb{E}_{X^n}[\\mathrm{QWAIC}] + o(1/n)$ under the singular-case assumptions. This makes QWAIC the quantum analog of WAIC: a quantity computed from measurement outcomes that predicts the expected out-of-sample quantum relative entropy loss.","pith_inferences":["A natural extension is to drop Assumption S2 and define a two-coefficient criterion using separate real log canonical thresholds for $K$ and $K_Q$; the simple multiplicative factor $r_{CQ}$ would fail, but an additive correction might still be built.","The covariance term $C^Q_n$ can be read as an empirical 'effective number of parameters' for a singular quantum model, and the numerics in the paper already hint that it tracks $r_{CQ}\\lambda$ rather than the raw parameter dimension.","Because the proof only needs the classical shadow's unbiasedness and analyticity of the shadow log-likelihood, the same construction should transfer to loss functions other than the quantum relative entropy, such as Petz-Rényi divergences, with a modified covariance term.","The exponential cost of matrix logarithms in QWAIC is a practical barrier; shadow-tailored or variational approximations to the covariance term could make the criterion scalable to larger systems."],"forward_implications":["QWAIC gives a computable model-selection score for Bayesian quantum state estimation that remains asymptotically unbiased for singular models such as neural-network quantum states and quantum Boltzmann machines.","The learning coefficient for quantum state estimation is $\\lambda_Q+\\nu'_Q-\\nu_Q$ in regular cases and $r_{CQ}\\lambda+r_{CQ}\\nu-\\nu_Q$ in singular cases, so the effective model complexity depends on the ratio of quantum to classical Fisher information.","In the realizable regular case $\\lambda_Q = (1/2)\\operatorname{Tr}(J_Q J^{-1})$, so the second term of QWAIC reflects the measurement penalty as well as the parameter count.","The difference between expected quantum generalization and training loss is $2\\chi_Q/n$ to leading order, so QWAIC's covariance term estimates exactly the overfitting gap.","If the theory is right, singular quantum models, like their classical counterparts, learn faster in the sense that a smaller learning coefficient gives a smaller expected generalization loss."],"supporting_citations":[{"why":"Supplies singular learning theory: RLCT, singular fluctuation, and WAIC as the classical blueprint QWAIC extends.","marker":"[37]"},{"why":"Provides the standard forms, renormalized empirical processes, and Bayesian asymptotic theorems used throughout the proofs.","marker":"[38]"},{"why":"Introduces classical shadows, the unbiased snapshot representation used to define the quantum training loss and the quantum log-likelihood ratio.","marker":"[67]"},{"why":"Earlier QAIC for regular quantum models, the direct predecessor whose assumptions and formulas QWAIC generalizes.","marker":"[65]"},{"why":"Hironaka's resolution of singularities, invoked to obtain the simultaneous normal crossing representations of the two loss functions.","marker":"[88]"}],"fun_headline_variants":["Quantum WAIC: a computable score for singular model selection","QWAIC computes quantum relative entropy loss even when regularity fails","Classical shadows make WAIC work for singular quantum models","Bayesian quantum state estimation gets a singularity-friendly criterion","QWAIC: a Bayesian shadow criterion for singular quantum models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main proof for singular models assumes that, after resolving singularities, the classical Kullback-Leibler loss and the quantum relative entropy vanish at the same order ($k = k_Q$); the paper states this as Assumption S2 and leaves it as a conjecture whether the fundamental conditions imply it.","fun_headline_variants_meta":{"raw":{"variants":["Quantum WAIC: a computable score for singular model selection","QWAIC computes quantum relative entropy loss even when regularity fails","Classical shadows make WAIC work for singular quantum models","Bayesian quantum state estimation gets a singularity-friendly criterion","QWAIC: a Bayesian shadow criterion for singular quantum models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000971,"raw_usage":{"total_tokens":4199,"prompt_tokens":1083,"completion_tokens":3116,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":699,"completion_tokens_details":{"reasoning_tokens":3033}},"tokens_in":699,"tokens_out":3116,"duration_ms":17673,"temperature":1.0,"reasoning_tokens":3033,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:09:40.546011+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a tomographically complete measurement and a parametric state family that satisfies the fundamental conditions but whose quantum relative entropy vanishes to higher order than the classical KL divergence on the optimal set; compute $\\mathbb{E}_{X^n}[G^Q_n]-\\mathbb{E}_{X^n}[\\mathrm{QWAIC}]$ at increasing $n$. If the difference does not decay as $o(1/n)$, then Theorem III.4's singular-case claim fails.","supporting_citations":[{"cited_title":"We present the theorem relevant to our discussion: Theorem B.1","cited_arxiv_id":null,"evidence_quote":"Hironaka's resolution of singularities, invoked to obtain the simultaneous normal crossing representations of the two loss functions."}],"review_version":1}