Pith. sign in

REVIEW 2 major objections 4 minor 19 references

Functional worst risk minimization

T0 review · 2 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper proves that for functional structural equation models with a bounded resolvent operator, the worst out-of-sample risk over a covariance-dominated shift set equals an exact linear combination of the two observed risks, and…

desk verdict A promising functional worst-risk framework, but the central decomposition theorem is false as stated: closure of the shift set does not give the needed approximability inside the constraint set. read the letter →

arxiv 2412.00412 v3 pith:FLZ37CW7 submitted 2024-11-30 math.ST math.PRstat.TH

classification math.STmath.PRstat.TH MSC 62R1062G0562F35
keywords functionaldataanalysisworstriskminimizationdistributionshiftstructuralequationmodelunboundedlinearoperatorrobustregressionorthonormalbasisestimationout-of-sample
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to make worst risk minimization—choosing a predictor that is robust to future distribution shifts—work directly on functional data, where the observations are curves rather than vectors. It models a structural system by a linear, possibly unbounded operator whose resolvent $\mathcal{S}=(I-\mathcal{T})^{-1}$ is bounded, and it considers two observed environments: an observational one and one shifted by a random process $A$. The central result is that the worst out-of-sample risk over a shift set defined by a covariance-kernel domination condition equals $\frac{1}{2}R_+(\beta)+(\gamma-\frac{1}{2})R_\Delta(\beta)$, a linear combination of the pooled risk and the risk difference of the two observed environments. If this is right, robustness guarantees for functional regression can be obtained from two environments without estimating the unknown shift distribution, and worst-risk minimizers can be computed in any orthonormal basis, avoiding eigenfunction estimation.

What carries the argument

The carrying object is the solution operator $\mathcal{S}=(I-\mathcal{T})^{-1}$ acting on $L^2([T_1,T_2])^{p+1}$; requiring only that $\mathcal{S}$ be bounded lets $\mathcal{T}$ itself be unbounded, for example a derivative operator. The shift set $\mathcal{C}_\gamma^\mathcal{A}(A)$ is defined by the requirement that $\int g(s)K_{A'}(s,t)g(t)^\top\,ds\,dt \le \gamma\int g(s)K_A(s,t)g(t)^\top\,ds\,dt$ for all test functions $g$, a Mercer-type kernel domination condition. The proof expands the risk in an orthonormal basis, uses the assumption that the noise $\varepsilon^A$ is an independent copy of a fixed zero-mean process to cancel shift-noise cross terms, and then reduces the supremum over the entire shift set to the single shift $\sqrt{\gamma}A$.

What would settle it

Simulate a functional SEM with bounded $\mathcal{S}$, and generate an observational environment with noise variance $\sigma_O^2$ and a shifted environment where either the noise variance is different ($\sigma_A^2\neq\sigma_O^2$) or the shift is correlated with the noise ($A=c\varepsilon$). Compute both sides of Theorem 3.7 over a one-dimensional shift set; a nontrivial discrepancy between the empirical supremum and $\frac{1}{2}R_+(\beta)+(\gamma-\frac{1}{2})R_\Delta(\beta)$ would refute the decomposition.

Watch

Extended reading notes

Core claim

For an environment generated as $(Y^A,X^A)=\mathcal{S}(A+\varepsilon^A)$, where $\mathcal{S}=(I-\mathcal{T})^{-1}$ is bounded and linear, and for a shift set $\mathcal{C}_\gamma^\mathcal{A}(A)$ consisting of shifts whose covariance kernel is dominated by $\gamma$ times the observed shift's kernel, the worst future out-of-sample risk is exactly $$\sup_{A'\in\mathcal{C}_\gamma^\mathcal{A}(A)} R_{A'}(\$\beta$)=\frac{1}{2}R_+(\$\beta$)+\left(\gamma-\frac{1}{2}\right)R_\$\Delta$(\$\beta$),$$ for every regression kernel $\beta\in(L^2([T_1,T_2]))^p$. The paper establishes this decomposition under the sole structural condition that $\sqrt{\gamma}A$ lies in the closure of the shift space, and it uses the same decomposition to give necessary and sufficient conditions for a unique worst-risk minimizer in square-integrable kernels. The minimizer is expressed in an arbitrary orthonormal basis for the target and an eigenbasis of a regularized covariate operator, with the consequence that estimation no longer requires first estimating unknown eigenfunctions of the data operators.

Load-bearing premise

In every environment, the noise process $\varepsilon^A$ is an independent copy of the same zero-mean process $\varepsilon$ used in the observational environment; if shifts are correlated with the noise or the noise distribution changes across environments, the worst-risk decomposition can fail.

Editorial extensions

If this is right

  • Out-of-sample robustness guarantees for functional regression can be derived from just two observed environments, with the tuning parameter $\gamma$ interpolating between pooled-risk prediction at $\gamma=1/2$ and increasingly conservative, regularized solutions.
  • Worst-risk minimizers can be estimated consistently without estimating the eigenfunctions of the target and covariate covariance operators, because the population solution is available in any arbitrary orthonormal basis.
  • The framework covers structural operators that are unbounded, such as differentiation, as long as the resolvent $(I-\mathcal{T})^{-1}$ is bounded, extending functional SEMs beyond compact or Hilbert-Schmidt settings.
  • Theorem 4.1 provides a robust, multivariate generalization of the basic theorem for functional linear models, with existence and uniqueness tied to an explicit Hilbert-Schmidt summability condition and injectivity of a regularized covariance operator.
  • The decomposition gives a practical target: if an analyst can estimate the pooled risk and the risk difference from two environments, the worst risk over a whole covariance-dominated shift family is known without observing any future shifted environment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The proof structure suggests a diagnostic that the paper leaves implicit: hold out a genuinely shifted third environment, compute both sides of the decomposition from the two training environments, and compare; a mismatch would localize exactly where the noise-invariance assumption fails.
  • Because the cancellation of shift-noise cross terms uses only linearity of the score functionals, the decomposition may extend to certain nonlinear functionals of the processes, though the paper does not pursue that direction.
  • For wide-sense stationary shifts, Proposition 3.6 turns the shift-set condition into a spectral one, offering a practical pre-check: estimate the spectral density difference $\gamma\hat{K}_A-\hat{K}_{A'}$ and reject candidate shifts where it ceases to be positive semidefinite on a positive-frequency set.
  • The consistency theorems require sample splitting so that numerator and denominator terms are estimated independently; in finite samples this is an implicit cost that practitioners should budget for.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper develops a functional analogue of worst-risk (worst average loss) minimization for functional structural equation models. It models environments through a linear, possibly unbounded operator T with bounded inverse S=(I-T)^{-1}, defines a covariance-based out-of-sample shift set C^gamma_A(A), and claims in Theorem 3.7 that the worst risk over this shift set decomposes exactly as 1/2 R_+(beta)+(gamma-1/2)R_Delta(beta). On that basis it characterizes worst-risk minimizers in the space of square-integrable kernels, gives a solution in an arbitrary orthonormal basis, and constructs consistent estimators, with a simulation example.

Significance. If the main theorem were correct, this would be a meaningful extension of anchor-regression/causal-regularization ideas to functional data, and the claim that the minimizer can be expressed in an arbitrary ON-basis without estimating eigenfunctions would be a useful practical advance. The paper contains substantial proof detail and clearly connects to the non-functional analogue. However, the central theorem is false under the stated closure assumption, and the appendix contains coefficient errors in the proof of the basis-space results; these issues propagate to the minimizer and estimation theorems. The framework remains promising, but the central claim needs repair before the downstream results can be trusted.

major comments (2)
  1. [Theorem 3.7 and Appendix A.4, Step 7] The theorem is false as stated. The step 'choose tilde A_Delta in C^gamma_A(A) with ||tilde A_Delta - sqrt(gamma)A||_V small' does not follow from sqrt(gamma)A in closure(A) and Lemma 3.1: convergence in V only gives convergence of the covariance forms, while the defining inequality of C^gamma_A(A) in Definition 3.2 is one-sided, so a sequence in A can converge to sqrt(gamma)A from the outside without ever entering C^gamma_A(A). Concretely, take p=1, S=I_2, deterministic shifts of the form (a f,0) with a in {1/2} union (1,infinity), observed shift A=(f,0), gamma=1, and beta=0. Let A be the set of such shifts; then sqrt(gamma)A lies in the closure of A, but C^1_A(A) contains only (f/2,0). The left-hand side of Theorem 3.7 equals 1/4||f||^2 + E[epsilon_1^2], while the right-hand side equals ||f||^2 + E[epsilon_1^2]. The proof can be repaired by strengthening the hypothesis to sqrt(gamma)A in A, or by requiring A to be closed, or by an explicit condition that C^gamma_A(A) contains a sequence converging to sqrt(gamma)A; the theorem statement and Corollaries 3.8-3.9 must be adjusted accordingly. Because Theorem 4.1, Theorem 4.3, and the consistency results all invoke Theorem 3.7, this is a load-bearing issue.
  2. [Appendix A.7 and A.9, proofs of Theorems 4.3 and 5.4] The second-moment matrix of X^{sqrt(gamma)A} is repeatedly written with coefficients sqrt(gamma) and 1-sqrt(gamma), for example in the line 'E_{sqrt(gamma)A}[F_{1:n}(X^{sqrt(gamma)A})^T F_{1:n}(X^{sqrt(gamma)A})] = sqrt(gamma) E_A[...] + (1-sqrt(gamma))E_O[...]' and in the definitions of hat G_{n,1}(M) and hat G_{n,2}(M). Since X^{sqrt(gamma)A}=S(sqrt(gamma)A+epsilon), the correct coefficients are gamma and 1-gamma, exactly as derived in Theorem 4.1 and as used in Section 5.3's definition of hat G_n(M). Taken literally, the appendix proofs target a different Gram matrix from the estimator in the theorem statement and do not establish the claimed consistency. Please correct these coefficients throughout A.7 and A.9 and re-verify the subsequent bounds.
minor comments (4)
  1. [Title/Abstract] The abstract contains the typo 'plays the the part of B'; please correct.
  2. [Appendix A.9] The heading 'A.9 Proof of Theorem 4.3' appears to contain the proof of the estimator result stated as Theorem 5.4, not another proof of the population Theorem 4.3; please relabel.
  3. [Definition 3.2 and Proposition 3.3] The notation for the shift set is inconsistent: C^gamma_A(A) alternates with C^gamma(A) in the surrounding text and proofs; please unify the notation.
  4. [Section 2.2] The noise-invariance assumption (epsilon^A is an independent copy of a fixed zero-mean process epsilon with identical distribution across environments) is essential to Claims A.1 and A.2 and hence to the decomposition; the paper should state this assumption prominently as a limitation, since correlated or environment-dependent noise would break the worst-risk decomposition.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the worst-risk decomposition is derived from the SEM and noise-invariance assumptions rather than from a fitted input or a self-citation chain; the only self-citation (Kania and Wit 2022) is non-load-bearing.

full rationale

The central result, Theorem 3.7, is a population-level mathematical derivation, not a fitted prediction. The proof expands the risk in terms of score coordinates, uses the noise invariance assumptions (Claims A.1 and A.2) to cancel shift-noise cross terms and to identify the noise covariance across environments, and then optimizes over the shift set C^gamma_A(A). The conclusion sup R_{A'} = 1/2 R_+ + (gamma - 1/2) R_Delta is not written into Definition 3.2; it is derived from the structural equation (2.3), the bounded/invertible operator S, and the covariance-order condition (3.9). No parameter is fitted to a subset of data and then renamed a prediction. The paper does cite Kania and Wit (2022), co-authored by one of the present authors, but only as an analogue: 'the above decomposition is an exact analogue of the non-functional case (Kania and Wit, 2022).' This citation is not used to prove Theorem 3.7; the functional proof is self-contained. There is a notable mathematical gap in Step 7 of the proof of Theorem 3.7, where the authors assert that sqrt(gamma)A in closure(A) lets them choose tilde A_Delta in C^gamma_A(A) arbitrarily close to sqrt(gamma)A; closure in V does not by itself ensure approximability from within the constrained set C^gamma_A(A). This is a correctness issue, not a circularity: it is an invalid inference rather than a reduction of the conclusion to its premises. Similarly, the abstract's claim that the solution is expressed in an arbitrary ON-basis and 'completely removes any necessity of estimating eigenfunctions' is contradicted by Theorem 4.1, which uses the eigenfunctions {psi_n} of K; this is an internal overclaim, not a circular derivation. Because the derivation does not reduce by construction to its inputs and the self-citation is not load-bearing, the circularity score is minimal.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The central theory rests on the structural equation model and the assumptions of noise invariance and boundedness of (I-T)^{-1}. No new entities are introduced. The main free parameter is gamma, the shift-set radius, which is user-chosen. The paper does not fit any constants to data.

free parameters (2)
  • gamma = user-specified
    Regularization level defining the shift set C_gamma^A(A); controls how extreme future shifts are considered. Not estimated from data.
  • e(n) = user-specified sequence
    Truncation level for the number of basis functions in the estimators; chosen by the user, subject to e(n) <= E(n) for the consistency theorems.
assumptions (5)
  • domain assumption The noise epsilon has zero mean, and for each shift A, the noise epsilon_A is a copy of epsilon and is independent of A.
    Used in Claims A.1 and A.2 to cancel cross terms between shifts and noise and to identify the noise covariance across environments. Without this, the worst-risk decomposition does not hold.
  • domain assumption The operator S = (I-T)^{-1} is bounded and linear on L2^{p+1}.
    Ensures that the solutions (Y^A,X^A)=S(A+epsilon_A) are well-defined square-integrable processes and that the risk functions are continuous in the shift.
  • domain assumption The set A and the shift A satisfy sqrt(gamma)A in closure(A) in V.
    Required in Theorem 3.7 to approximate the worst-case shift sqrt(gamma)A by elements of A.
  • domain assumption For the estimation theory, the sample paths are assumed cadlag.
    Used in Section 5 to ensure the random partitions Pi_n have finitely many points, enabling discretization of the integrals.
  • ad hoc to paper Existence of consistent estimators hat(psi)_{l,n} of the eigenfunctions psi_l of the operator K.
    Theorem 5.1 assumes such estimators exist, but the paper does not provide a construction or conditions. This is a significant unaddressed practical requirement.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Functional worst risk minimization." pith.science (2026). https://pith.science/paper/FLZ37CW7

@misc{pith2026241200412,
  author       = {Pith},
  title        = {Pith review of: Functional worst risk minimization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FLZ37CW7}},
  note         = {Machine review of arXiv:2412.00412}
}
abstract

The aim of this paper is to extend worst risk minimization, also called worst average loss minimization, to the functional realm. This means finding a functional regression representation that will be robust to future distribution shifts on the basis of data from two environments. In the classical non-functional realm, structural equations are based on a transfer matrix $B$. In section~\ref{sec:sfr}, we generalize this to consider a linear operator $\mathcal{T}$ on square integrable processes that plays the the part of $B$. By requiring that $(I-\mathcal{T})^{-1}$ is bounded -- as opposed to $\mathcal{T}$ -- this will allow for a large class of unbounded operators to be considered. Section~\ref{sec:worstrisk} considers two separate cases that both lead to the same worst-risk decomposition. Remarkably, this decomposition has the same structure as in the non-functional case. We consider any operator $\mathcal{T}$ that makes $(I-\mathcal{T})^{-1}$ bounded and define the future shift set in terms of the covariance functions of the shifts. In section~\ref{sec:minimizer}, we prove a necessary and sufficient condition for existence of a minimizer to this worst risk in the space of square integrable kernels. Previously, such minimizers were expressed in terms of the unknown eigenfunctions of the target and covariate integral operators (see for instance \cite{HeMullerWang} and \cite{YaoAOS}). This means that in order to estimate the minimizer, one must first estimate these unknown eigenfunctions. In contrast, the solution provided here will be expressed in any arbitrary ON-basis. This completely removes any necessity of estimating eigenfunctions. This pays dividends in section~\ref{sec:estimation}, where we provide a family of estimators, that are consistent with a large sample bound. Proofs of all the results are provided in the appendix.

Figures

Figures reproduced from arXiv: 2412.00412 by the authors.

Figure 1
Figure 1. Observational environment: a functional system that serves as an illustration of a structural system [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Interventional environment: the structural functional system is also observed under a slightly [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Sample (Y A, XA(1), XA(2)) from the shifted environment. Note that both X(1) and X(2) seem quite predictive for Y , but only X(1) is causal — and therefore X(1) has the most robust out-of-sample risk behaviour, if Y is not intervened, as in this example. where ϵ O ∼ N(0, Σ) with Σ = I30×30. In our case, we assume homogeneous effects across all the basis functions and choose, B =     0 bx1y 0 0 0 0 byx2 0 0   … view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Functional regression coefficients βx1y and βx2y shown on the same x-y-z scale: (a) True causal parameters; (b) population values, pooling the two data-environments, “mistakenly” finds that X(2) affects Y ; (c) population values minimizing the out-of-sample risk in Cγ=…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

19 extracted references · 15 canonical work pages

  1. [1]

    A new look at the statistical model identification

    Hirotugu Akaike. A new look at the statistical model identification. IEEE transactions on automatic control, 19 0 (6): 0 716--723, 1974

  2. [2]

    Invariant risk minimization

    Martin Arjovsky, L \'e on Bottou, and Ishaan Gulrajani. Invariant risk minimization. arXiv preprint arXiv:1907.02893, 2019

  3. [3]

    Linear Processes in Function Spaces: Theory and Applications

    Denis Bosq. Linear Processes in Function Spaces: Theory and Applications. Springer, 2000. ISBN 978-1-4612-1154-9

  4. [4]

    Functional additive regression

    Yingying Fan, Gareth M James, and Peter Radchenko. Functional additive regression. The Annals of Statistics, 43 0 (5): 0 2296--2325, 2015

  5. [5]

    The interpretation of mallows’s cp-statistic

    Steven G Gilmour. The interpretation of mallows’s cp-statistic. Journal of the Royal Statistical Society Series D: The Statistician, 45 0 (1): 0 49--56, 1996

  6. [6]

    Functional linear regression via canonical analysis

    Guozhong He, Hans-Georg M \"u ller, Jane-Ling Wang, and Wenjing Yang. Functional linear regression via canonical analysis . Bernoulli, 16 0 (3): 0 705 -- 729, 2010. doi:10.3150/09-BEJ228. URL https://doi.org/10.3150/09-BEJ228

  7. [7]

    Does invariant risk minimization capture invariance? In International Conference on Artificial Intelligence and Statistics, pages 4069--4077

    Pritish Kamath, Akilesh Tangella, Danica Sutherland, and Nathan Srebro. Does invariant risk minimization capture invariance? In International Conference on Artificial Intelligence and Statistics, pages 4069--4077. PMLR, 2021

  8. [8]

    Causal regularization: On the trade-off between in-sample risk and out-of-sample risk guarantees

    Lucas Kania and Ernst Wit. Causal regularization: On the trade-off between in-sample risk and out-of-sample risk guarantees. arXiv preprint arXiv:2205.01593, 2022

Show all 19 references
  1. [9]

    Risk minimization in multi-factor portfolios: What is the best strategy? Annals of Operations Research, 266: 0 255--291, 2018

    Philipp J Kremer, Andreea Talmaciu, and Sandra Paterlini. Risk minimization in multi-factor portfolios: What is the best strategy? Annals of Operations Research, 266: 0 255--291, 2018

  2. [10]

    Functional structural equation model

    Kuang-Yao Lee and Lexin Li. Functional structural equation model. Journal of the Royal Statistical Society Series B: Statistical Methodology, 84 0 (2): 0 600--629, 2022

  3. [11]

    Functional additive models

    Hans-Georg M \"u ller and Fang Yao. Functional additive models. Journal of the American Statistical Association, 103 0 (484): 0 1534--1544, 2008

  4. [12]

    Causal inference by using invariant prediction: identification and confidence intervals

    Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen. Causal inference by using invariant prediction: identification and confidence intervals. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 78 0 (5): 0 947--1012, 2016. doi:10.1111/rssb.12167

  5. [13]

    The risks of invariant risk minimization

    Elan Rosenfeld, Pradeep Ravikumar, and Andrej Risteski. The risks of invariant risk minimization. arXiv preprint arXiv:2010.05761, 2020

  6. [14]

    a usler, Nicolai Meinshausen, Peter B \

    Dominik Rothenh \"a usler, Nicolai Meinshausen, Peter B \"u hlmann, and Jonas Peters. Anchor regression: Heterogeneous data meet causality. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 83 0 (2): 0 215--246, 2021

  7. [15]

    Causal dantzig: Fast inference in linear structural equation models with hidden variables under additive interventions

    Dominik Rothenhäusler, Peter Bühlmann, and Nicolai Meinshausen. Causal dantzig: Fast inference in linear structural equation models with hidden variables under additive interventions. Ann. Statist., 47 0 (3): 0 1688--1722, 06 2019. doi:10.1214/18-AOS1732

  8. [16]

    Schwartz

    L. Schwartz. Radon Measures on Arbitrary Topological Spaces and Cylindrical Measures. Studies in mathematics. Tata Institute of Fundamental Research, 1973. ISBN 9780195605167. URL https://books.google.se/books?id=bKKBjgEACAAJ

  9. [17]

    The cross-validated adaptive epsilon-net estimator

    Mark J Van der Laan, Sandrine Dudoit, and Aad W van der Vaart. The cross-validated adaptive epsilon-net estimator. Statistics & Decisions, 24 0 (3): 0 373--395, 2006

  10. [18]

    Super learner

    Mark J Van der Laan, Eric C Polley, and Alan E Hubbard. Super learner. Statistical applications in genetics and molecular biology, 6 0 (1), 2007

  11. [19]

    Generalized additive models: an introduction with R

    Simon N Wood. Generalized additive models: an introduction with R. chapman and hall/CRC, 2017

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.