{"id":"9daed99f-a058-4a25-acd5-a8c0fa7963bb","arxiv_id":"2608.12496","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"The paper proposes a polynomial shrinkage estimator for the precision matrix in total least squares fingerprinting, jointly estimating variability inflation factors and providing uncertainty quantification for climate change scaling factors.","lead":"A new statistical method for climate fingerprinting estimates how much observed warming is caused by human and natural factors while accounting for errors in the observations and in the climate model signals. It uses nonlinear shrinkage of the covariance matrix and estimates model variability inflation, with the goal of narrower, more reliable confidence intervals for attribution statements.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Asymptotic optimality and valid inference are claimed for the data-driven shrinkage selector f-hat, but Theorems 3.3/3.8/3.11/3.13 cover only fixed f; no uniform or selection-consistency argument is supplied.","rationale":"The reader's weakest_assumption is the common-proportional-covariance model (1)--(3), which the paper's own residual test rejects for large continental domains. That is a serious limitation of the application, and the paper acknowledges it. However, the more load-bearing concern for the central claim is the gap between the fixed-f theorems and the data-driven f-hat used in the actual procedure. The sentence immediately after (19) asserts asymptotic optimality within the polynomial class, but no theorem states or proves the required uniform convergence or selection consistency. The reader's rationale does note that 'the central claim of asymptotic optimality and valid inference for the data-driven selector f-hat is not proven,' so there is partial overlap. I agree with the CONDITIONAL verdict: the fixed-f theory appears plausible and is supported by substantial random-matrix-technical development, but the missing f-hat argument is a real correctness risk that is fixable by adding uniform convergence and selection consistency results. The numerical check I propose would settle whether the gap actually manifests in finite samples: if f-hat behaves like a fixed f in simulations, the concern would be mitigated; if it does not, the inference claims would need to be weakened. No change to the reader's verdict is needed, because CONDITIONAL already captures this unresolved dependence on a missing proof.","tokens_in":38201,"tokens_out":4227,"duration_ms":45581,"concrete_test":"Add a theorem establishing uniform convergence of trace(\\hat{\\Xi}(f)) to a deterministic limit over the compact admissible class F, plus consistency of f-hat to the oracle minimizer f*, or provide a counterexample. As a numerical settlement check, run the Section 4 Setting 3 simulation with m = 624, N = 624, n1 = n2 = 40, and 500 replicates: compute f-hat by the paper's grid search, then compare empirical coverage of the 90% confidence interval using f-hat against the coverage obtained with the oracle f* (the minimizer of trace(\\Xi(f)) computed with true \\sigma_X^2 and \\sigma_Z^2). If coverage with f-hat falls below 88% or deviates from the oracle coverage by more than 1.5 percentage points, the data-driven selection step invalidates the claimed inference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central methodological claim is that the data-driven selector f-hat, defined in (19) as the minimizer of trace(\\hat{\\Xi}(f)) over the admissible polynomial family F, achieves asymptotic optimality in total variance and yields asymptotically valid confidence intervals. However, every asymptotic normality theorem (Theorems 3.3, 3.8, 3.11) and the residual test (Theorem 3.13) is stated for a fixed admissible f. The proofs in the supplement, including Lemmas S.2.1--S.5.1 and the CLT arguments in S.4, treat f as fixed: for example, Lemma 3.2 establishes pointwise convergence of \\Delta(f), \\Pi(f), \\Omega(f), K(f), and the trace of \\hat{\\Xi}(f), but nothing is shown about uniform convergence over F or stochastic equicontinuity of the map f -> trace(\\hat{\\Xi}(f)). Consequently, the minimizer f-hat may not converge to the oracle minimizer of the limiting trace, and even if it does, the asymptotic distribution of \\hat{\\beta}(f-hat) is not automatically the same as that of \\hat{\\beta}(f) for any fixed f. The same gap affects the residual test, whose denominator involves K(f-hat). This is not a contradiction within the fixed-f theory, but it is an unproven bridge between the theory and the actual procedure used in simulations and the application: the reported confidence intervals and p-values use f-hat, so their claimed validity depends on an argument that the paper does not provide. The assumption that model (1)--(3) holds is also challenged by the paper's own residual test for large continental domains, but that is an acknowledged limitation of the application; the missing uniform selection theory is more directly load-bearing for the central claim of asymptotic optimality.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a regularized optimal fingerprinting framework for climate change detection and attribution. It extends the total least squares (TLS) estimator to high-dimensional settings by using a polynomial (more generally, analytic) shrinkage estimator of the error precision matrix, jointly estimating the scaling factors β and the variability inflation factors σ²_X and σ²_Z. The authors provide asymptotic normality theorems for the TLS estimator for any fixed admissible shrinkage function f, consistent estimators for the inflation factors, Wald-type tests and confidence intervals, and a residual consistency test for model adequacy. The method is evaluated in simulations and applied to annual mean temperature over 1951–2020 for several continental and subcontinental regions. The central methodological claim is that the data-driven shrinkage selector f̂ defined in Eq. (19) achieves asymptotic optimality and yields valid inference, but the formal theorems are stated for fixed f only.","tokens_in":38549,"tokens_out":4230,"duration_ms":40570,"significance":"If the claimed results hold, the paper would make a useful contribution to high-dimensional errors-in-variables regression and to climate fingerprinting, where covariance estimation uncertainty and model–observation variance heterogeneity are practically important. The manuscript contains a substantial theoretical supplement with detailed martingale and contour-integral arguments for fixed-f asymptotics, consistent estimation of inflation factors, and a residual diagnostic that is rarely addressed in this literature. The simulation evidence for efficiency gains under nonlinear covariance structures is encouraging, and the application illustrates the practical relevance of estimating the inflation factors rather than fixing them to one. However, the strength of the central claim depends on the behavior of the data-driven selector f̂, and that behavior is not covered by the theorems as stated.","major_comments":[{"comment":"Theorems 3.3, 3.8, and 3.11 establish asymptotic normality for a fixed admissible f, and Theorem 3.13 does the same for the residual test with a fixed f. The actual procedure selects f̂ as the minimizer of trace(Ξ̂(f)) in Eq. (19), and f̂ is used in all simulations and in the application. No theorem supplies uniform convergence of trace(Ξ̂(f)) over the class F, stochastic equicontinuity of the map f ↦ trace(Ξ̂(f)), or selection consistency of f̂ toward the oracle minimizer. Consequently, the claimed asymptotic optimality and the validity of the confidence intervals and p-values based on f̂ are not established by the supplied proofs; the proofs in the supplement, including Lemmas S.2.1–S.5.1 and the CLT argument in Section S.4, treat f as fixed. This is a load-bearing gap between the theory and the implemented method.","section":"Section 3.1, Eq. (19); Theorems 3.3, 3.8, 3.11, 3.13"},{"comment":"The residual consistency test of Theorem 3.13 is applied in Section 5 and rejects the null model at the 0.05 level for the large continental domains GL, NH, NHM, and EA, while the application still reports attribution intervals for these domains under model (1)–(3). Since the paper's own diagnostic indicates that the assumed proportional covariance and constant inflation structure is misspecified for these regions, the claimed empirical validity of the resulting confidence intervals and attribution statements is called into question. The authors should either restrict the inference claims to the regions where the test is not rejected or provide a robustness analysis showing that the conclusions are insensitive to the suspected misspecification.","section":"Section 3.3.3 and Section 5"},{"comment":"The residual test is evaluated in the simulation study under the true values of σ²_X and σ²_Z, and the authors note that with estimated factors the test over-rejects the null at around 15% for a nominal 5% level. In the application, however, the inflation factors are estimated, not known. This means the calibration of the diagnostic in the setting where it is actually used is not established, and the statement in Section 6 that the residual test is 'rigorously established' and 'consistent' overstates what is shown for the estimated-factor case. The finite-sample calibration under estimated factors should be addressed or the claims should be weakened accordingly.","section":"Section 4, Table 4.3 and Remark preceding it"},{"comment":"The sentence 'The proposed strategy achieves asymptotic optimality with respect to the total variance of the TLS estimator, within the class of polynomial shrinkage estimators' is presented without a formal theorem or proof. The lemmas preceding it establish pointwise limits for fixed f, but no result shows that the minimizer of trace(Ξ̂(f)) converges to the minimizer of the corresponding limiting trace, nor that the asymptotic variance at f̂ equals the oracle variance. A precise optimality statement with conditions would be needed to support this claim.","section":"Paragraph after Eq. (19)"}],"minor_comments":[{"comment":"The caption lists Setting 2 as Σ_PL and Setting 3 as Σ_LS, which is inconsistent with the order of settings defined in Section 4 and used in Tables 4.1–4.3; the caption should be corrected.","section":"Figure 4.1 caption"},{"comment":"The caption says 'the 5 regions analyzed in the study,' but the table lists eight regions (including global and subcontinental domains); the number should be updated.","section":"Table 5.1 caption"},{"comment":"Step 5 says to repeat Steps 2–4 with W = f̂(S), but no stopping rule or convergence criterion is specified for this iterative refinement; a brief remark on the number of iterations or on why one pass is sufficient would improve reproducibility.","section":"Algorithm 2, Step 5"},{"comment":"The discussion of multiple minimizers states that any of them may be selected and that they are 'equally efficient with respect to the total variance criterion,' but the asymptotic equivalence of the resulting estimators is not proven; at minimum this should be flagged as a heuristic rather than a theorem.","section":"End of Section 3.1"},{"comment":"The notation M_2 in Theorem 3.6 is defined as M_2^2 = 2N^{-1} tr(Σ²), which is slightly awkward because M_2 itself never appears; using a single symbol such as H would be clearer.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid fixed-f asymptotic core and a well-organized supplement, but the data-driven selector f̂ is the method actually used, and the missing uniform-convergence or selection-consistency argument is the key obstacle. This seems fixable either by adding a theorem that closes the gap under additional compactness or Lipschitz conditions on F, or by reframing the theoretical claims to apply to fixed f and treating the selection step as an empirically validated heuristic. I would also ask the authors to confront the tension between the residual test's rejection of the model for large domains and the reported confidence intervals for those domains."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nI've read the Li & Li paper. The genuinely new pieces are the polynomial shrinkage family for the precision matrix in TLS fingerprinting, the data-driven choice of shrinkage function, and joint estimation of the variance inflation factors. The fixed-f asymptotic results (Theorems 3.3, 3.8, 3.11) look plausible and are worked out in impressive detail in the supplement. Simulations show real efficiency gains when the true covariance is nonlinear, and the CMIP6 application is thoughtful. The residual consistency test is a useful addition to a literature that often skips model checking.\n\nThe soft spot is exactly what the stress-test flags: the procedure actually used selects f-hat by minimizing the estimated trace in (19), but no theorem covers the data-driven selector. Theorems 3.3, 3.8, 3.11, and 3.13 all require a fixed admissible f. Without uniform convergence or a selection-consistency argument, the claimed asymptotic optimality and interval coverage for the actual estimator are unproven. I don't see a contradiction in the fixed-f theory—it looks sound—but the bridge to f-hat is missing. This is load-bearing and should be fixable with stochastic equicontinuity or similar arguments, but the paper as written overclaims.\n\nA second concern is the proportional covariance model (1)–(3). It's standard, but the paper's own residual test rejects it for the larger continental domains, and the test itself over-rejects (around 15%) when the inflation factors are estimated. In the application they use estimated factors, so the reported p-values for those domains are likely optimistic. The authors acknowledge this, but it tempers the empirical conclusions.\n\nOverall, this is a promising paper for anyone working on regularized fingerprinting or high-dimensional errors-in-variables regression. The gap between theory and practice is real but likely repairable. I'd send it to a serious referee—ideally someone fluent in both random matrix theory and climate attribution—with the request to either prove the selection claims or substantially temper them. It deserves peer review, not desk rejection.","headline":"Solid fixed-f theory for polynomial shrinkage fingerprinting, but the data-driven selector actually used in simulations and the application is unproven.","tokens_in":39137,"tokens_out":3170,"would_cite":true,"duration_ms":31157,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H12","62F12","62J05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Nonlinear shrinkage of the error covariance restores the asymptotic optimality of total least squares scaling-factor estimates in high-dimensional fingerprinting.","keywords":["Measurement error","Fingerprinting","Statistical Regularization","Covariance matrix estimation","Total least squares","High-dimensional inference","Climate change detection and attribution","Variance inflation factor"],"falsifier":"Simulate the model (1)–(3) with a block-diagonal $\\Sigma$ whose two spatial blocks carry different inflation factors, using the same dimensions as the application ($N\\approx 600$, $m=151$, $n_i=40$), then compute the empirical coverage of the proposed 90% confidence intervals for the ANT scaling factor; if coverage falls substantially below 90%, the proportional-covariance assumption is the part of the framework that cannot be relaxed without changing the claims.","tokens_in":37944,"feed_emoji":"🌡️","tokens_out":7054,"duration_ms":56748,"temperature":0.7,"pith_summary":"Climate detection and attribution regresses observed temperature on model-simulated fingerprints of external forcings, but both the response and the fingerprints carry internal-variability noise whose covariance matrix must be estimated from a limited number of control runs. The paper proposes to estimate that covariance by a nonlinear, rotation-invariant shrinkage: a polynomial function of the sample covariance matrix whose coefficients are chosen to minimize the estimated asymptotic variance of the total least squares estimator of the scaling factors. It further proposes joint estimation of the two variability inflation factors that scale the model noise relative to observations, and derives a residual consistency test for the fitted model. The upshot is a claim that scaling-factor estimates are asymptotically unbiased and their confidence intervals achieve nominal coverage even when the dimension is comparable to the number of control runs, correcting a known source of bias and undercoverage.","feed_headline":"Polynomial shrinkage tightens climate attribution uncertainty","feed_subtitle":"A data-chosen polynomial shrinkage restores valid confidence intervals for climate fingerprint scaling factors.","key_machinery":"The central object is the polynomial shrinkage function $f(S)=\\ell_k S^k+\\cdots+\\ell_1 S+I_N$ applied to the sample covariance matrix of the control runs, with degree $k$ fixed (recommended $k=3$) and coefficients chosen from an admissible set that keeps $f$ monotone and bounded on the spectral support. Because admissible polynomials have separated roots, the asymptotic variance of the TLS estimator can be written as a partial-fraction formula in terms of resolvent-based quantities $\\Delta(f)$, $\\Pi(f)$, $\\Omega(f)$, and $K(f)$, computed from $S$, the fingerprints, and the estimated inflation factors. Minimizing the estimated trace of $\\Xi(f)$ over this set performs the data-driven shrinkage selection, and the same resolvent calculus delivers the estimators of $\\sigma_X^2$ and $\\sigma_Z^2$ and the residual test statistic.","core_discovery":"Under the errors-in-variables model (1)–(3), where observation error, fingerprint noise, and control runs are normal with covariances $\\Sigma$, $\\sigma_X^2\\Sigma$, and $\\sigma_Z^2\\Sigma$, the paper shows that the TLS estimator using the admissible polynomial shrinkage weight matrix $f(S)$ is asymptotically normal with an explicit covariance $\\Xi(f)$, and that selecting $f$ by minimizing the estimated trace of $\\Xi(f)$ is asymptotically optimal within the class of polynomial shrinkage estimators. When $\\sigma_X^2$ and $\\sigma_Z^2$ are unknown, the proposed estimators are consistent, and the TLS estimator retains asymptotic normality with a covariance augmented by a term $\\Gamma(f)$ that reflects the cost of estimating $\\sigma_X^2$; the paper also constructs Wald tests, confidence regions, and a residual consistency test based on these results.","pith_inferences":["A direct testable extension is to allow spatially varying inflation factors within a region; the paper's own residual diagnostic suggests this would resolve the rejections at continental scales.","Because the machinery is expressed purely through the spectrum of $S$, the approach should carry over to other high-dimensional errors-in-variables settings, such as genomics or econometrics, where a separate sample estimates the noise covariance.","The normality assumption is used for the martingale arguments; whether sub-Gaussian or heavier-tailed errors preserve the coverage claims is an open question that simulations could answer before policy use."],"forward_implications":["Scaling-factor confidence intervals keep nominal coverage even when the number of spatial-temporal dimensions $N$ is comparable to the number of control runs $m$, a regime where the raw sample covariance matrix is singular or unstable.","When the true covariance has nonlinear spectral structure, polynomial shrinkage produces shorter intervals than optimal linear shrinkage while maintaining coverage, improving efficiency in attribution statements.","Estimating the inflation factors instead of fixing them at one removes the systematic bias from model–observation variance mismatch and changes detection conclusions in several continental regions.","The residual consistency test provides a formal check on the proportional-covariance assumption, and it rejects that assumption for large continental domains while passing subcontinental regions.","The same framework extends to any analytic shrinkage function through contour integrals, so domain knowledge about $\\Sigma$ can be incorporated without losing the theoretical guarantees."],"supporting_citations":[{"why":"Establishes the weighted TLS formulation for optimal fingerprinting with proportional covariance and an inflation factor.","marker":"Allen and Stott (2003)"},{"why":"Introduced regularized optimal fingerprinting with a well-conditioned linear shrinkage covariance estimate.","marker":"Ribes et al. (2009)"},{"why":"Refined the linear shrinkage weight matrix to achieve optimality within the linear class for TLS estimation.","marker":"Li and Li (2025)"},{"why":"Provides the well-conditioned linear shrinkage estimator used as an initial weight and benchmark.","marker":"Ledoit and Wolf (2004)"},{"why":"Shows the optimal shrinkage function is nonlinear, motivating polynomial shrinkage.","marker":"Ledoit and Wolf (2017)"},{"why":"Supplies the random-matrix theory, including limiting spectral distributions and resolvent asymptotics, used in the proofs.","marker":"Bai and Silverstein (2010)"},{"why":"Gives the standard errors-in-variables and TLS framework from which the estimator expression is adapted.","marker":"Fuller (2009)"},{"why":"Provides the HadCRUT5 observations used in the empirical application.","marker":"Morice et al. (2021)"}],"fun_headline_variants":["Polynomial shrinkage sharpens climate attribution confidence","Nonlinear shrinkage refines climate fingerprint intervals","Adaptive shrinkage yields narrower climate attribution intervals","Precision matrix shrinkage improves climate detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that observation error, fingerprint noise, and control-run variability all have the same covariance shape $\\Sigma$ up to constant multipliers $\\sigma_X^2$ and $\\sigma_Z^2$, and are normally distributed; if the model–observation variability difference is not a common scale multiple of a single matrix, the prewhitening and the resulting confidence intervals are misspecified, a condition the paper's own residual test rejects for large continental regions.","fun_headline_variants_meta":{"raw":{"variants":["Polynomial shrinkage sharpens climate attribution confidence","Nonlinear shrinkage refines climate fingerprint intervals","Adaptive shrinkage yields narrower climate attribution intervals","Precision matrix shrinkage improves climate detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000348,"raw_usage":{"total_tokens":1890,"prompt_tokens":920,"completion_tokens":970,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":916}},"tokens_in":536,"tokens_out":970,"duration_ms":8286,"temperature":1.0,"reasoning_tokens":916,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:07:26.208476+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the model (1)–(3) with a block-diagonal $\\Sigma$ whose two spatial blocks carry different inflation factors, using the same dimensions as the application ($N\\approx 600$, $m=151$, $n_i=40$), then compute the empirical coverage of the proposed 90% confidence intervals for the ANT scaling factor; if coverage falls substantially below 90%, the proportional-covariance assumption is the part of the framework that cannot be relaxed without changing the claims.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the weighted TLS formulation for optimal fingerprinting with proportional covariance and an inflation factor."}],"review_version":1}