{"id":"7c6966d4-0ad8-4559-b416-712450b70d17","arxiv_id":"2505.04070","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A data-driven linear shrinkage choice and consistent variance estimator give scaling-factor confidence intervals with near-nominal coverage and shorter length in climate detection and attribution.","lead":"This paper improves how climate scientists estimate confidence ranges when matching observed temperature changes to model-simulated fingerprints of human and natural causes. The new method chooses the regularization strength automatically and produces shorter, more reliable error bars for these estimates.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Data-driven lambda selection and post-selection inference are claimed but not proven; fixed-lambda asymptotics do not cover the reported intervals.","rationale":"I focused on the paper's advertised move from fixed-lambda asymptotics to a data-driven lambda_hat. The fixed-lambda Theorem 1 is imported from Li et al. (2023) and is plausible; Propositions 2-4 give the ingredients for consistent estimation of Xi(lambda) at fixed lambda. But the paper's central contribution is the 'linearly optimal weight matrix', obtained by selecting lambda_hat = argmin tr(Xi_hat(lambda)). The text states that this selection is asymptotically optimal, yet no theorem shows that the estimated trace converges uniformly on the search interval, that lambda_hat converges to the minimizer of the limiting trace, or that the normal approximation at lambda_hat remains valid. Lemma 1 is fixed-lambda; it cannot by itself justify inference after selection. The simulations are the only evidence at lambda_hat, and at the practically important m=50 they show 90.6-93.2% coverage, i.e., undercoverage. Thus the central claim is not fully supported. I would therefore keep the reader's conditional verdict: the fixed-lambda machinery is likely sound, but the paper should either prove the required uniform/selection results or explicitly state the optimality and coverage claims as empirical rather than asymptotic. The reader's chosen weakest assumption (covariance equality/Gaussianity) is a legitimate external limitation, but the selection gap is internal to the argument and more directly tied to the headline claim, which is why I partially agree rather than fully agreeing with the reader's choice.","tokens_in":21663,"tokens_out":7972,"duration_ms":80582,"concrete_test":"Re-run the Table S1 simulation with m=50, N=275 and the Sigma_UN, gamma=0.5 setting, but replace lambda_hat by the oracle lambda* that minimizes the true tr(Xi(lambda)) (computable since X and Sigma are known in simulation). Construct the proposed normal-based CIs at lambda* and compare their empirical coverage and average length to the reported lambda_hat-based Optim row. If coverage at lambda* is near 95% while lambda_hat coverage is 93.2%, the residual undercoverage is caused by the selection step, confirming the missing post-selection theory; if coverage at lambda* is also around 93%, the variance estimator itself is biased at this sample size.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Propositions 2-4 and Lemma 1 establish the plug-in variance estimator only for a fixed lambda. The paper then defines lambda_hat = argmin_{lambda in [lambda_l, lambda_u]} tr(Xi_hat(lambda)) and asserts in the Abstract and Introduction that this choice is asymptotically optimal within the linear shrinkage class and yields near-nominal coverage. No theorem proves (i) uniform consistency of tr(Xi_hat(lambda)) over [lambda_l, lambda_u], (ii) lambda_hat -> lambda_opt (or even boundedness away from the boundary), or (iii) the key distributional result sqrt(N)(beta_hat(lambda_hat)-beta) converging to N(0, Xi(lambda_opt)) that would justify the normal-based confidence intervals at the selected lambda. Without (iii), plugging lambda_hat into the fixed-lambda normal approximation has no asymptotic justification. The simulation evidence does not fill this gap, since coverage at m=50 is 90.6-93.2%, below the nominal 95%, and the interval-length comparisons in Figure 1 and Table S1 are made at the selected lambda. Because the advertised 'linearly optimal weight matrix' is precisely the lambda_hat-based estimator, this missing proof is load-bearing: the central claim of the paper is not established by the provided arguments.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies regularized fingerprinting for detection and attribution of climate change, in which the regression error covariance Σ is estimated by a linear shrinkage matrix Σ̂(λ) = S + λI and the total-least-squares estimator β̂(λ) is used for the scaling factors. The main theoretical result is a fixed-λ asymptotic normality theorem for β̂(λ), followed by a plug-in estimator Ξ̂(λ) of the asymptotic covariance matrix. The paper then selects λ by minimizing tr(Ξ̂(λ)) over an interval, calls the resulting weight matrix \"linearly optimal\", and constructs marginal and joint confidence intervals at the selected λ. Simulation studies compare the method with calibrated and uncalibrated competitors, and an application to annual mean temperature data for several regions is presented. The central advertised contribution is the data-driven optimal choice of λ together with valid confidence intervals under that choice.","tokens_in":21827,"tokens_out":4843,"duration_ms":54018,"significance":"If the fixed-λ variance estimator and the λ-selection procedure were fully justified, the paper would make a practical contribution to a widely used climate attribution method: it would replace bootstrap calibration by a computationally cheap plug-in variance estimate and would offer a principled criterion for choosing the shrinkage parameter. The paper has several strengths: the high-dimensional asymptotic framework is clearly laid out, the fixed-λ variance formulas are supported by random-matrix theory, and the simulation design is realistic and reasonably extensive, with results reported for multiple covariance structures, signal strengths, and sample sizes. The real-data application is also useful in showing how the method behaves in practice. However, the paper's main claim, asymptotic optimality of the data-selected λ, is not supported by a theorem, and the plug-in confidence intervals at the selected λ lack a post-selection distributional result. These gaps are load-bearing for the advertised contribution.","major_comments":[{"comment":"The consistency statement in Lemma 1 is proved only for a fixed λ, while the paper defines λ̂_opt = argmin_{λ∈[λ_l,λ_u]} tr(Ξ̂(λ)) and then uses this selected value in the marginal and joint confidence intervals. No theorem establishes uniform consistency of Ξ̂(λ) over the search interval, consistency of λ̂_opt to the population minimizer λ_opt, or the post-selection central limit theorem √N(β̂(λ̂_opt)−β) → N(0, Ξ(λ_opt)). The simulation results in Table S1 and Figure 1 are computed at the selected λ, so they do not fill this gap. Because the paper's title and abstract advertise a \"linearly optimal weight matrix\" obtained precisely through λ̂_opt, this missing proof is the central unresolved issue.","section":"§2.3, Lemma 1 and the definition of λ̂_opt"},{"comment":"Assumptions 3–6 are stated \"for any fixed λ\", and the proofs of Theorem 1 and Propositions 2–3 invoke concentration results that hold for fixed λ. To justify minimization of tr(Ξ̂(λ)) over an interval, the authors need uniform analogues of these assumptions and concentration bounds on [λ_l, λ_u], or else a proof that λ̂_opt converges to an interior minimizer and that the fixed-λ approximation is valid uniformly near that minimizer. Without such a result, λ̂_opt may select a boundary point or a region where the asymptotic approximation is poor, and the optimized interval lengths reported in Figure 1 are not statements about the population minimizer λ_opt.","section":"Assumptions 3–6 and the fixed-λ arguments in Supplementary A.4–A.5"},{"comment":"Propositions 2 and 3 are proved using the Gaussian representation X̃ = X + Σ^{1/2} W D^{1/2} with W standard normal, and the model in Section 2 assumes η_ik ~ N(0, Σ) with the same covariance Σ as the observational error. The Discussion states that the method \"imposes no additional assumptions\" compared with approaches such as Hannart (2016). This overstates the paper's assumptions: the bias corrections in Propositions 2 and 3 and the plug-in variance estimator in Lemma 1 rely on normality and on exact covariance equality. The authors should either prove robustness to these assumptions, relax them, or explicitly list them as limitations of the theoretical results.","section":"Supplementary A.5 and Discussion, Section 5"},{"comment":"Lemma 1, which asserts consistency of Ξ̂(λ) for fixed λ, is stated without a proof in either the main text or the supplied Supplementary Materials; the proof section A.5 covers only Propositions 2 and 3, and the citation given after Proposition 4 refers to Lemma 2 of Chen et al. (2011), not to Lemma 1. Since Lemma 1 is the basis for both the variance estimator and the λ-selection criterion, a complete proof or a precise citation to a source that proves this exact statement should be provided.","section":"Lemma 1, main text"}],"minor_comments":[{"comment":"The text says that coverage rates remain \"around 91%\" even at m = 50, and the abstract and introduction refer to \"near-nominal\" coverage; Table S1 reports values between 90.6% and 93.2% for m = 50 across settings. This should be stated more precisely as finite-sample undercoverage that improves as m grows, rather than as close-to-nominal performance in the most challenging setting.","section":"Table S1, rows with m = 50"},{"comment":"The table caption says \"5 regions\", but the table lists eight rows (GL, NH, NHM, EA, NA, WNA, CNA, ENA); the caption should be corrected.","section":"Section 4.1, Table 1"},{"comment":"The sentence \"A publicly available software implementation further facilitates this task\" suggests that code is available, but no repository or URL is given; please include the relevant link or remove the sentence.","section":"Section 5, Discussion"},{"comment":"The notation uses λ and λ for the lower and upper bounds of the search interval, which is easily confused with the true parameter λ; using λ_l and λ_u would improve readability.","section":"Section 2.3, definition of the search interval"},{"comment":"The statement \"There exist constants α and α\" appears to be a typographical duplication of the lower and upper bound constants; it should read α and ̄α (or α and α_max).","section":"Assumption 1"}],"recommendation":"major_revision","confidential_remarks":"The missing post-selection theory is not a cosmetic issue: the paper's central advertised contribution is the optimally selected weight matrix and the confidence intervals at that selected λ. The fixed-λ variance estimator may be publishable on its own, but as written the manuscript oversells the data-driven selection. I would recommend asking for a rigorous treatment of uniform consistency and post-selection inference, or a substantial reframing of the claims to fixed-λ results only. The simulation evidence, while suggestive, cannot replace the missing theorem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has a real, useful core: a consistent plug-in estimator for the asymptotic covariance of scaling factors under a linear shrinkage weight matrix, plus a data-driven choice of lambda by minimizing the estimated trace. The fixed-lambda theory is plausible, the derivations are careful, and the simulations show genuine improvement over existing methods, especially once m is not tiny. The application to HadCRUT5/CMIP6 data is competent, and the authors are honest about the covariance-equality assumption that underpins the whole framework.\n\nThe soft spots are real and roughly match the stress-test note. Theorem 1 is explicitly imported from Li et al. (2023), so the asymptotic distribution is inherited, not new. Propositions 2–4 and Lemma 1 establish consistency of the variance estimator only for fixed lambda. Then the paper defines lambda_hat by minimizing tr(Xi_hat(lambda)) and claims in the abstract and introduction that this choice is asymptotically optimal and yields near-nominal coverage. No theorem proves uniform consistency of tr(Xi_hat(lambda)) over the search interval, nor lambda_hat -> lambda_opt, nor the post-selection normality that would justify normal-based intervals at the selected lambda. That is a load-bearing gap, because the central selling point is precisely the lambda_hat-based 'optimal' interval. The simulation evidence does not close it: at m=50, coverage sits at 90.6–93.2%, below nominal, and interval lengths are reported at the selected lambda.\n\nThere are also two smaller issues worth naming. The paper promises public software but no code appears in the preprint or supplement, which makes the empirical results hard to check. And the discussion claims the method 'imposes no additional assumptions' beyond the model, which is too strong: the proofs in Supplementary A.5 rely explicitly on Gaussian noise (W standard normal) and equal covariance between simulations and observations. Those are standard in fingerprinting, but they are assumptions all the same.\n\nNone of this makes the paper unserious. The fixed-lambda variance estimator is a clean, useful contribution, and the lambda-selection heuristic is reasonable in practice. What is missing is either a theorem covering the selected lambda or a modest rewording that presents lambda_hat as an empirically motivated choice rather than a proven optimum. The authors can fix this with a heavy revision, or by softening the claims. I would send this to peer review, but the referee should demand one of those two outcomes.","headline":"Solid fixed-lambda variance estimator for regularized fingerprinting, but the advertised 'linearly optimal' lambda selection and post-selection confidence intervals go beyond what is proven.","tokens_in":22394,"tokens_out":1512,"would_cite":false,"duration_ms":17493,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62H12","62J05","62P12"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that within the class of linear shrinkage weight matrices, choosing the regularization parameter to minimize the estimated asymptotic covariance of the scaling-factor estimates yields near-nominal confidence coverage in…","keywords":["measurement error","linear shrinkage estimator","optimal fingerprinting","detection and attribution","errors-in-variables regression","confidence intervals","total least squares","high-dimensional covariance estimation"],"falsifier":"Simulate observations with error covariance $\\Sigma_{\\mathrm{obs}}=2\\,\\Sigma_{\\mathrm{model}}$ while drawing control runs from $\\Sigma_{\\mathrm{model}}$, estimate $\\Sigma$ with the sample covariance $S$, build the proposed intervals on many replicates, and check whether the empirical coverage falls below the nominal 95%; a clear drop would show the covariance-equality premise is load-bearing.","tokens_in":21395,"feed_emoji":"🌡️","tokens_out":9304,"duration_ms":83039,"temperature":0.7,"pith_summary":"Climate detection and attribution usually works through optimal fingerprinting, a regression of observed temperatures onto model-simulated fingerprints in which noise in both sides shares a covariance matrix that must be estimated from control runs. When that covariance is estimated from a limited number of runs, existing regularized methods produce confidence intervals that are too narrow and undercover. This paper claims that within the class of linear shrinkage weight matrices $\\widehat{\\Sigma}(\\lambda)=S+\\lambda I$, one can consistently estimate the asymptotic covariance of the scaling-factor estimates and choose $\\lambda$ to minimize its trace. The resulting intervals reach near-nominal coverage without bootstrap calibration and are shorter on average, both in simulations and in an application to observed near-surface temperature. If correct, this restores reliable inference in the low-control-run regime typical of real attribution studies.","feed_headline":"Fingerprinting intervals reach 95% coverage without bootstrap","feed_subtitle":"Choosing the shrinkage parameter by estimated variance shortens intervals and restores nominal coverage.","key_machinery":"The central object is the linear shrinkage weight matrix $\\widehat{\\Sigma}(\\lambda)=S+\\lambda I_N$, where $S$ is the sample covariance of $m$ control runs and $\\lambda>0$ is a ridge parameter. The argument is carried by the asymptotic covariance matrix $\\Xi(\\lambda)=(1+\\beta^T D\\beta)\\,\\Delta_1^{-1}(\\lambda)\\{\\Delta_2(\\lambda)+K(\\lambda)(D^{-1}+\\beta\\beta^T)^{-1}\\}\\Delta_1^{-1}(\\lambda)$ of the total-least-squares scaling-factor estimator, together with its plug-in estimate $\\widehat{\\Xi}(\\lambda)$, whose bias corrections $\\Theta_1(\\lambda)$ and $\\Theta_2(\\lambda)$ account for measurement error in the fingerprints and for estimating $S$. The optimal $\\lambda$ is found by a grid search minimizing $\\operatorname{tr}(\\widehat{\\Xi}(\\lambda))$, which is asymptotically optimal within the linear shrinkage class.","core_discovery":"The paper's central claim is that regularized fingerprinting can recover valid inference without calibration by turning the shrinkage parameter into a variance-minimizing choice. Theorem 1 states that $\\sqrt{N}(\\widehat{\\beta}(\\lambda)-\\beta)$ converges in distribution to a centered normal with covariance $\\Xi(\\lambda)$ as $N,m\\to\\infty$ with $N/m\\to c$, and that the plug-in estimator $\\widehat{\\Xi}(\\lambda)$ is consistent. Propositions 2 and 3 give explicit bias corrections $\\Theta_1(\\lambda)$ and $\\Theta_2(\\lambda)$ that absorb the contamination of the ensemble-averaged fingerprints and the randomness of $S$, so that $\\widehat{\\Delta}_1(\\lambda)$, $\\widehat{\\Delta}_2(\\lambda)$, and $\\widehat{K}(\\lambda)$ can be computed from $\\widetilde{X}$ and $S$ alone. Selecting $\\lambda$ to minimize $\\operatorname{tr}(\\widehat{\\Xi}(\\lambda))$ then yields the linearly optimal weight matrix within the shrinkage class, and the resulting confidence intervals attain near-nominal coverage with shorter length than existing calibrated methods in simulation and in the 1951–2020 temperature application.","pith_inferences":["The covariance-equality premise is the linchpin; a natural and testable extension would be a sensitivity analysis or an estimator that allows a known or estimated mismatch between observational and control-run variability, which the paper does not address.","Since the aspect ratio $c=N/m$ enters through the Marčenko–Pastur corrections, spatial aggregation, which sets $N$, changes the inferred optimal $\\lambda$ and interval width; rerunning the same analysis at several grid resolutions would reveal how much of the regional differences is resolution-driven.","The linear shrinkage class is convenient but restrictive; a similar variance-minimizing argument might carry over to nonlinear shrinkage or structured shrinkage targets, but the bias corrections would have to be re-derived for those families."],"forward_implications":["Within the linear shrinkage class, the $\\lambda$ that minimizes $\\operatorname{tr}(\\widehat{\\Xi}(\\lambda))$ is asymptotically optimal, so the weight matrix can be chosen by a one-dimensional grid search rather than by cross-validation or bootstrap.","Confidence intervals built from the normal approximation with $\\widehat{\\Xi}(\\widehat{\\lambda}_{\\mathrm{opt}})$ keep coverage near the nominal level even with only about fifty control runs, where calibrated bootstrap competitors still undercover.","Because the variance estimator is consistent for every weight matrix in the linear shrinkage class, the standard regularized fingerprinting estimator also receives valid uncertainty quantification without extra computation.","In the real-data temperature analysis, the proposed intervals are shorter than those from the linear-shrinkage baseline and comparable to or shorter than those from the minimum-variance calibrated baseline, and they change several regional detection and attribution conclusions."],"supporting_citations":[{"why":"Establishes the total-least-squares errors-in-variables framework that the paper's weight-matrix construction builds on.","marker":"(Allen & Stott 2003)"},{"why":"Supplies the linear shrinkage covariance estimator that defines the class of weight matrices studied here.","marker":"(Ledoit & Wolf 2004)"},{"why":"Introduced linear shrinkage into optimal fingerprinting, the method whose properties this paper refines.","marker":"(Ribes et al. 2009)"},{"why":"Defines regularized optimal fingerprinting and provides the simulation design and comparison baseline used in the numerical studies.","marker":"(Ribes et al. 2013)"},{"why":"Documents the undercoverage problem and proposes the calibration bootstrap that the new direct variance estimator is designed to replace.","marker":"(Li et al. 2021)"},{"why":"Shows existing regularized weights are not MSE-optimal and supplies the theorem used in the proof of this paper's Theorem 1.","marker":"(Li et al. 2023)"},{"why":"Provides the concentration result used as Proposition 4 and the technical lemmas behind the bias corrections.","marker":"(Chen et al. 2011)"},{"why":"Gives the quadratic-form concentration lemmas that yield the bias corrections in Propositions 2 and 3.","marker":"(El Karoui & Kösters 2011)"}],"fun_headline_variants":["Linearly optimal weights fix fingerprinting coverage","Shrinkage that minimizes variance shortens climate intervals","Variance-minimizing fingerprint weights yield reliable intervals","Optimal fingerprinting via shrinkage: shorter, reliable intervals","Tuning shrinkage by variance gives climate detection gains"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The construction assumes that internal variability in climate-model simulations has exactly the same covariance matrix as the observational error, and the proofs further require Gaussian noise; if either fails, the bias corrections and the estimated variance no longer target the quantity the intervals claim to cover.","fun_headline_variants_meta":{"raw":{"variants":["Linearly optimal weights fix fingerprinting coverage","Shrinkage that minimizes variance shortens climate intervals","Variance-minimizing fingerprint weights yield reliable intervals","Optimal fingerprinting via shrinkage: shorter, reliable intervals","Tuning shrinkage by variance gives climate detection gains"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1392,"prompt_tokens":973,"completion_tokens":419,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":345}},"tokens_in":589,"tokens_out":419,"duration_ms":4643,"temperature":1.0,"reasoning_tokens":345,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:39:34.536746+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate observations with error covariance $\\Sigma_{\\mathrm{obs}}=2\\,\\Sigma_{\\mathrm{model}}$ while drawing control runs from $\\Sigma_{\\mathrm{model}}$, estimate $\\Sigma$ with the sample covariance $S$, build the proposed intervals on many replicates, and check whether the empirical coverage falls below the nominal 95%; a clear drop would show the covariance-equality premise is load-bearing.","supporting_citations":[{"cited_title":"Geometric sensitivity of random matrix results: consequences for shrinkage estimators of covariance and related statistical methods","cited_arxiv_id":"1105.1404","evidence_quote":"Establishes the total-least-squares errors-in-variables framework that the paper's weight-matrix construction builds on."}],"review_version":1}