REVIEW 2 major objections 4 minor 19 references
Functional worst risk minimization
T0 review · 2 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper proves that for functional structural equation models with a bounded resolvent operator, the worst out-of-sample risk over a covariance-dominated shift set equals an exact linear combination of the two observed risks, and…
desk verdict A promising functional worst-risk framework, but the central decomposition theorem is false as stated: closure of the shift set does not give the needed approximability inside the constraint set. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the solution operator $\mathcal{S}=(I-\mathcal{T})^{-1}$ acting on $L^2([T_1,T_2])^{p+1}$; requiring only that $\mathcal{S}$ be bounded lets $\mathcal{T}$ itself be unbounded, for example a derivative operator. The shift set $\mathcal{C}_\gamma^\mathcal{A}(A)$ is defined by the requirement that $\int g(s)K_{A'}(s,t)g(t)^\top\,ds\,dt \le \gamma\int g(s)K_A(s,t)g(t)^\top\,ds\,dt$ for all test functions $g$, a Mercer-type kernel domination condition. The proof expands the risk in an orthonormal basis, uses the assumption that the noise $\varepsilon^A$ is an independent copy of a fixed zero-mean process to cancel shift-noise cross terms, and then reduces the supremum over the entire shift set to the single shift $\sqrt{\gamma}A$.
What would settle it
Simulate a functional SEM with bounded $\mathcal{S}$, and generate an observational environment with noise variance $\sigma_O^2$ and a shifted environment where either the noise variance is different ($\sigma_A^2\neq\sigma_O^2$) or the shift is correlated with the noise ($A=c\varepsilon$). Compute both sides of Theorem 3.7 over a one-dimensional shift set; a nontrivial discrepancy between the empirical supremum and $\frac{1}{2}R_+(\beta)+(\gamma-\frac{1}{2})R_\Delta(\beta)$ would refute the decomposition.
Extended reading notes
Core claim
For an environment generated as $(Y^A,X^A)=\mathcal{S}(A+\varepsilon^A)$, where $\mathcal{S}=(I-\mathcal{T})^{-1}$ is bounded and linear, and for a shift set $\mathcal{C}_\gamma^\mathcal{A}(A)$ consisting of shifts whose covariance kernel is dominated by $\gamma$ times the observed shift's kernel, the worst future out-of-sample risk is exactly $$\sup_{A'\in\mathcal{C}_\gamma^\mathcal{A}(A)} R_{A'}(\$\beta$)=\frac{1}{2}R_+(\$\beta$)+\left(\gamma-\frac{1}{2}\right)R_\$\Delta$(\$\beta$),$$ for every regression kernel $\beta\in(L^2([T_1,T_2]))^p$. The paper establishes this decomposition under the sole structural condition that $\sqrt{\gamma}A$ lies in the closure of the shift space, and it uses the same decomposition to give necessary and sufficient conditions for a unique worst-risk minimizer in square-integrable kernels. The minimizer is expressed in an arbitrary orthonormal basis for the target and an eigenbasis of a regularized covariate operator, with the consequence that estimation no longer requires first estimating unknown eigenfunctions of the data operators.
Load-bearing premise
In every environment, the noise process $\varepsilon^A$ is an independent copy of the same zero-mean process $\varepsilon$ used in the observational environment; if shifts are correlated with the noise or the noise distribution changes across environments, the worst-risk decomposition can fail.
Editorial extensions
If this is right
- Out-of-sample robustness guarantees for functional regression can be derived from just two observed environments, with the tuning parameter $\gamma$ interpolating between pooled-risk prediction at $\gamma=1/2$ and increasingly conservative, regularized solutions.
- Worst-risk minimizers can be estimated consistently without estimating the eigenfunctions of the target and covariate covariance operators, because the population solution is available in any arbitrary orthonormal basis.
- The framework covers structural operators that are unbounded, such as differentiation, as long as the resolvent $(I-\mathcal{T})^{-1}$ is bounded, extending functional SEMs beyond compact or Hilbert-Schmidt settings.
- Theorem 4.1 provides a robust, multivariate generalization of the basic theorem for functional linear models, with existence and uniqueness tied to an explicit Hilbert-Schmidt summability condition and injectivity of a regularized covariance operator.
- The decomposition gives a practical target: if an analyst can estimate the pooled risk and the risk difference from two environments, the worst risk over a whole covariance-dominated shift family is known without observing any future shifted environment.
Reading between the lines
- The proof structure suggests a diagnostic that the paper leaves implicit: hold out a genuinely shifted third environment, compute both sides of the decomposition from the two training environments, and compare; a mismatch would localize exactly where the noise-invariance assumption fails.
- Because the cancellation of shift-noise cross terms uses only linearity of the score functionals, the decomposition may extend to certain nonlinear functionals of the processes, though the paper does not pursue that direction.
- For wide-sense stationary shifts, Proposition 3.6 turns the shift-set condition into a spectral one, offering a practical pre-check: estimate the spectral density difference $\gamma\hat{K}_A-\hat{K}_{A'}$ and reject candidate shifts where it ceases to be positive semidefinite on a positive-frequency set.
- The consistency theorems require sample splitting so that numerator and denominator terms are estimated independently; in finite samples this is an implicit cost that practitioners should budget for.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a functional analogue of worst-risk (worst average loss) minimization for functional structural equation models. It models environments through a linear, possibly unbounded operator T with bounded inverse S=(I-T)^{-1}, defines a covariance-based out-of-sample shift set C^gamma_A(A), and claims in Theorem 3.7 that the worst risk over this shift set decomposes exactly as 1/2 R_+(beta)+(gamma-1/2)R_Delta(beta). On that basis it characterizes worst-risk minimizers in the space of square-integrable kernels, gives a solution in an arbitrary orthonormal basis, and constructs consistent estimators, with a simulation example.
Significance. If the main theorem were correct, this would be a meaningful extension of anchor-regression/causal-regularization ideas to functional data, and the claim that the minimizer can be expressed in an arbitrary ON-basis without estimating eigenfunctions would be a useful practical advance. The paper contains substantial proof detail and clearly connects to the non-functional analogue. However, the central theorem is false under the stated closure assumption, and the appendix contains coefficient errors in the proof of the basis-space results; these issues propagate to the minimizer and estimation theorems. The framework remains promising, but the central claim needs repair before the downstream results can be trusted.
major comments (2)
- [Theorem 3.7 and Appendix A.4, Step 7] The theorem is false as stated. The step 'choose tilde A_Delta in C^gamma_A(A) with ||tilde A_Delta - sqrt(gamma)A||_V small' does not follow from sqrt(gamma)A in closure(A) and Lemma 3.1: convergence in V only gives convergence of the covariance forms, while the defining inequality of C^gamma_A(A) in Definition 3.2 is one-sided, so a sequence in A can converge to sqrt(gamma)A from the outside without ever entering C^gamma_A(A). Concretely, take p=1, S=I_2, deterministic shifts of the form (a f,0) with a in {1/2} union (1,infinity), observed shift A=(f,0), gamma=1, and beta=0. Let A be the set of such shifts; then sqrt(gamma)A lies in the closure of A, but C^1_A(A) contains only (f/2,0). The left-hand side of Theorem 3.7 equals 1/4||f||^2 + E[epsilon_1^2], while the right-hand side equals ||f||^2 + E[epsilon_1^2]. The proof can be repaired by strengthening the hypothesis to sqrt(gamma)A in A, or by requiring A to be closed, or by an explicit condition that C^gamma_A(A) contains a sequence converging to sqrt(gamma)A; the theorem statement and Corollaries 3.8-3.9 must be adjusted accordingly. Because Theorem 4.1, Theorem 4.3, and the consistency results all invoke Theorem 3.7, this is a load-bearing issue.
- [Appendix A.7 and A.9, proofs of Theorems 4.3 and 5.4] The second-moment matrix of X^{sqrt(gamma)A} is repeatedly written with coefficients sqrt(gamma) and 1-sqrt(gamma), for example in the line 'E_{sqrt(gamma)A}[F_{1:n}(X^{sqrt(gamma)A})^T F_{1:n}(X^{sqrt(gamma)A})] = sqrt(gamma) E_A[...] + (1-sqrt(gamma))E_O[...]' and in the definitions of hat G_{n,1}(M) and hat G_{n,2}(M). Since X^{sqrt(gamma)A}=S(sqrt(gamma)A+epsilon), the correct coefficients are gamma and 1-gamma, exactly as derived in Theorem 4.1 and as used in Section 5.3's definition of hat G_n(M). Taken literally, the appendix proofs target a different Gram matrix from the estimator in the theorem statement and do not establish the claimed consistency. Please correct these coefficients throughout A.7 and A.9 and re-verify the subsequent bounds.
minor comments (4)
- [Title/Abstract] The abstract contains the typo 'plays the the part of B'; please correct.
- [Appendix A.9] The heading 'A.9 Proof of Theorem 4.3' appears to contain the proof of the estimator result stated as Theorem 5.4, not another proof of the population Theorem 4.3; please relabel.
- [Definition 3.2 and Proposition 3.3] The notation for the shift set is inconsistent: C^gamma_A(A) alternates with C^gamma(A) in the surrounding text and proofs; please unify the notation.
- [Section 2.2] The noise-invariance assumption (epsilon^A is an independent copy of a fixed zero-mean process epsilon with identical distribution across environments) is essential to Claims A.1 and A.2 and hence to the decomposition; the paper should state this assumption prominently as a limitation, since correlated or environment-dependent noise would break the worst-risk decomposition.
Circularity Check
No significant circularity: the worst-risk decomposition is derived from the SEM and noise-invariance assumptions rather than from a fitted input or a self-citation chain; the only self-citation (Kania and Wit 2022) is non-load-bearing.
full rationale
The central result, Theorem 3.7, is a population-level mathematical derivation, not a fitted prediction. The proof expands the risk in terms of score coordinates, uses the noise invariance assumptions (Claims A.1 and A.2) to cancel shift-noise cross terms and to identify the noise covariance across environments, and then optimizes over the shift set C^gamma_A(A). The conclusion sup R_{A'} = 1/2 R_+ + (gamma - 1/2) R_Delta is not written into Definition 3.2; it is derived from the structural equation (2.3), the bounded/invertible operator S, and the covariance-order condition (3.9). No parameter is fitted to a subset of data and then renamed a prediction. The paper does cite Kania and Wit (2022), co-authored by one of the present authors, but only as an analogue: 'the above decomposition is an exact analogue of the non-functional case (Kania and Wit, 2022).' This citation is not used to prove Theorem 3.7; the functional proof is self-contained. There is a notable mathematical gap in Step 7 of the proof of Theorem 3.7, where the authors assert that sqrt(gamma)A in closure(A) lets them choose tilde A_Delta in C^gamma_A(A) arbitrarily close to sqrt(gamma)A; closure in V does not by itself ensure approximability from within the constrained set C^gamma_A(A). This is a correctness issue, not a circularity: it is an invalid inference rather than a reduction of the conclusion to its premises. Similarly, the abstract's claim that the solution is expressed in an arbitrary ON-basis and 'completely removes any necessity of estimating eigenfunctions' is contradicted by Theorem 4.1, which uses the eigenfunctions {psi_n} of K; this is an internal overclaim, not a circular derivation. Because the derivation does not reduce by construction to its inputs and the self-citation is not load-bearing, the circularity score is minimal.
Assumptions & free parameters
free parameters (2)
- gamma =
user-specified
- e(n) =
user-specified sequence
assumptions (5)
- domain assumption The noise epsilon has zero mean, and for each shift A, the noise epsilon_A is a copy of epsilon and is independent of A.
- domain assumption The operator S = (I-T)^{-1} is bounded and linear on L2^{p+1}.
- domain assumption The set A and the shift A satisfy sqrt(gamma)A in closure(A) in V.
- domain assumption For the estimation theory, the sample paths are assumed cadlag.
- ad hoc to paper Existence of consistent estimators hat(psi)_{l,n} of the eigenfunctions psi_l of the operator K.
Cite this review
Pith. "Pith review of Functional worst risk minimization." pith.science (2026). https://pith.science/paper/FLZ37CW7
@misc{pith2026241200412,
author = {Pith},
title = {Pith review of: Functional worst risk minimization},
year = {2026},
howpublished = {\url{https://pith.science/paper/FLZ37CW7}},
note = {Machine review of arXiv:2412.00412}
}
abstract
The aim of this paper is to extend worst risk minimization, also called worst average loss minimization, to the functional realm. This means finding a functional regression representation that will be robust to future distribution shifts on the basis of data from two environments. In the classical non-functional realm, structural equations are based on a transfer matrix $B$. In section~\ref{sec:sfr}, we generalize this to consider a linear operator $\mathcal{T}$ on square integrable processes that plays the the part of $B$. By requiring that $(I-\mathcal{T})^{-1}$ is bounded -- as opposed to $\mathcal{T}$ -- this will allow for a large class of unbounded operators to be considered. Section~\ref{sec:worstrisk} considers two separate cases that both lead to the same worst-risk decomposition. Remarkably, this decomposition has the same structure as in the non-functional case. We consider any operator $\mathcal{T}$ that makes $(I-\mathcal{T})^{-1}$ bounded and define the future shift set in terms of the covariance functions of the shifts. In section~\ref{sec:minimizer}, we prove a necessary and sufficient condition for existence of a minimizer to this worst risk in the space of square integrable kernels. Previously, such minimizers were expressed in terms of the unknown eigenfunctions of the target and covariate integral operators (see for instance \cite{HeMullerWang} and \cite{YaoAOS}). This means that in order to estimate the minimizer, one must first estimate these unknown eigenfunctions. In contrast, the solution provided here will be expressed in any arbitrary ON-basis. This completely removes any necessity of estimating eigenfunctions. This pays dividends in section~\ref{sec:estimation}, where we provide a family of estimators, that are consistent with a large sample bound. Proofs of all the results are provided in the appendix.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
A new look at the statistical model identification
Hirotugu Akaike. A new look at the statistical model identification. IEEE transactions on automatic control, 19 0 (6): 0 716--723, 1974
work page 1974
-
[2]
Martin Arjovsky, L \'e on Bottou, and Ishaan Gulrajani. Invariant risk minimization. arXiv preprint arXiv:1907.02893, 2019
arXiv 1907
-
[3]
Linear Processes in Function Spaces: Theory and Applications
Denis Bosq. Linear Processes in Function Spaces: Theory and Applications. Springer, 2000. ISBN 978-1-4612-1154-9
work page 2000
-
[4]
Functional additive regression
Yingying Fan, Gareth M James, and Peter Radchenko. Functional additive regression. The Annals of Statistics, 43 0 (5): 0 2296--2325, 2015
work page 2015
-
[5]
The interpretation of mallows’s cp-statistic
Steven G Gilmour. The interpretation of mallows’s cp-statistic. Journal of the Royal Statistical Society Series D: The Statistician, 45 0 (1): 0 49--56, 1996
work page 1996
-
[6]
Functional linear regression via canonical analysis
Guozhong He, Hans-Georg M \"u ller, Jane-Ling Wang, and Wenjing Yang. Functional linear regression via canonical analysis . Bernoulli, 16 0 (3): 0 705 -- 729, 2010. doi:10.3150/09-BEJ228. URL https://doi.org/10.3150/09-BEJ228
-
[7]
Pritish Kamath, Akilesh Tangella, Danica Sutherland, and Nathan Srebro. Does invariant risk minimization capture invariance? In International Conference on Artificial Intelligence and Statistics, pages 4069--4077. PMLR, 2021
work page 2021
-
[8]
Causal regularization: On the trade-off between in-sample risk and out-of-sample risk guarantees
Lucas Kania and Ernst Wit. Causal regularization: On the trade-off between in-sample risk and out-of-sample risk guarantees. arXiv preprint arXiv:2205.01593, 2022
Show all 19 references
-
[9]
Risk minimization in multi-factor portfolios: What is the best strategy? Annals of Operations Research, 266: 0 255--291, 2018
Philipp J Kremer, Andreea Talmaciu, and Sandra Paterlini. Risk minimization in multi-factor portfolios: What is the best strategy? Annals of Operations Research, 266: 0 255--291, 2018
2018
-
[10]
Functional structural equation model
Kuang-Yao Lee and Lexin Li. Functional structural equation model. Journal of the Royal Statistical Society Series B: Statistical Methodology, 84 0 (2): 0 600--629, 2022
2022
-
[11]
Functional additive models
Hans-Georg M \"u ller and Fang Yao. Functional additive models. Journal of the American Statistical Association, 103 0 (484): 0 1534--1544, 2008
2008
-
[12]
Causal inference by using invariant prediction: identification and confidence intervals
Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen. Causal inference by using invariant prediction: identification and confidence intervals. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 78 0 (5): 0 947--1012, 2016. doi:10.1111/rssb.12167
2016 doi
-
[13]
The risks of invariant risk minimization
Elan Rosenfeld, Pradeep Ravikumar, and Andrej Risteski. The risks of invariant risk minimization. arXiv preprint arXiv:2010.05761, 2020
2010 arXiv
-
[14]
a usler, Nicolai Meinshausen, Peter B \
Dominik Rothenh \"a usler, Nicolai Meinshausen, Peter B \"u hlmann, and Jonas Peters. Anchor regression: Heterogeneous data meet causality. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 83 0 (2): 0 215--246, 2021
2021
-
[15]
Causal dantzig: Fast inference in linear structural equation models with hidden variables under additive interventions
Dominik Rothenhäusler, Peter Bühlmann, and Nicolai Meinshausen. Causal dantzig: Fast inference in linear structural equation models with hidden variables under additive interventions. Ann. Statist., 47 0 (3): 0 1688--1722, 06 2019. doi:10.1214/18-AOS1732
2019 doi
-
[16]
Schwartz
L. Schwartz. Radon Measures on Arbitrary Topological Spaces and Cylindrical Measures. Studies in mathematics. Tata Institute of Fundamental Research, 1973. ISBN 9780195605167. URL https://books.google.se/books?id=bKKBjgEACAAJ
1973
-
[17]
The cross-validated adaptive epsilon-net estimator
Mark J Van der Laan, Sandrine Dudoit, and Aad W van der Vaart. The cross-validated adaptive epsilon-net estimator. Statistics & Decisions, 24 0 (3): 0 373--395, 2006
2006
-
[18]
Super learner
Mark J Van der Laan, Eric C Polley, and Alan E Hubbard. Super learner. Statistical applications in genetics and molecular biology, 6 0 (1), 2007
2007
-
[19]
Generalized additive models: an introduction with R
Simon N Wood. Generalized additive models: an introduction with R. chapman and hall/CRC, 2017
2017
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.