{"id":"ba5d222a-eaf4-4546-b55b-b1c57659085e","arxiv_id":"2411.18398","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A functional data method is extended to estimate the derivatives of multivariate functional data and reconstruct them through a multivariate Karhunen-Loève expansion.","lead":"This paper develops two methods for estimating derivatives of multivariate functional data, where each observation consists of several curves measured together. It tests both methods in simulations and uses them on coronary angiogram data to recover how vessel diameter and blood flow change along the artery.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DMKL is handicapped in the simulations by truncating the initial MFPCA at K=3; the claimed DMFPCA superiority over DMKL may be an artifact of this asymmetric truncation.","rationale":"The reader's weakest_assumption focused on Assumption 1 and the smoothness of raw curves; I agree this is a limitation, especially for the coronary application where derivatives are properties of the smoothed curves rather than the raw measurement process. However, the most load-bearing issue for the paper's central, simulation-based claim of superiority is the asymmetric truncation in the DMKL comparison. The paper's own acknowledgement that increasing the truncation number would help DMKL signals that the reported comparison is not a fair assessment of the method. A simple rerun would settle this. Because the concern does not invalidate the methodological framework but does undercut a headline claim, the appropriate verdict remains CONDITIONAL: the authors should either run the fair comparison or qualify the abstract. The reader already reached a conditional verdict, partly on the related point of abstract overstatement, so my stress test does not change the overall verdict.","tokens_in":16264,"tokens_out":12035,"duration_ms":113087,"concrete_test":"Rerun the four Section 5 simulation settings with DMKL modified so the initial MFPCA on X uses K'=10 (or a data-driven K' capturing at least 99% of variance), while keeping the final MFPCA on derivative estimates at K=3 as reported. Compare ISE for eigenfunctions, RE for eigenvalues, and RMISE for derivatives against DMFPCA and P-splines+MFPCA. If DMKL's third-eigencomponent errors drop to levels comparable to or better than DMFPCA, the paper's superiority claim is an artifact of truncation; if DMKL remains worse, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The simulation study (Section 5) compares DMFPCA, DMKL, and P-splines+MFPCA with the same truncation K=3 for all methods. For DMKL (Algorithm 2), the same K is used twice: first to compute an MFPCA of the original curves X, then to compute the final MFPCA of the derivative estimates. Consequently, the DMKL derivative estimates are constrained to the span of the derivatives of the first 3 MFPCs of X. This 3-dimensional subspace is generically not the optimal 3-dimensional subspace for representing the derivatives, so the large ISE/RE reported for DMKL's third eigencomponent is a direct result of this truncation choice, not of the method's derivative estimation per se. The paper itself states that 'Increasing the truncation number could enable DMKL to yield eigencomponents closer to the true result', but this variant is not run. Therefore, the abstract's claim that DMFPCA outperforms DMKL is not established by the reported simulations; it may reverse if DMKL is allowed a larger initial truncation (e.g., K'=10) before final reduction to K=3. The medium-sparsity result already shows P-splines+MFPCA beating DMFPCA, so the reported hierarchy is sensitive to implementation choices. The smoothness assumption is a real limitation, but in both the simulations (smooth latent curves) and the application (pre-smoothed curves) the derivative target is well-defined, so the unfair truncation is the more load-bearing defect for the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two methods for estimating the principal components, scores, and reconstructed derivatives of multivariate functional data: DMFPCA, which differentiates the univariate covariance functions and then applies MFPCA, and DMKL, which differentiates the multivariate Karhunen-Loève expansion of the original data. The methods are compared with a direct approach (P-splines followed by MFPCA) in simulations covering dense, medium-sparse, and high-sparse settings, and applied to coronary angiogram data. The central claim is that DMFPCA outperforms the other methods, particularly for densely observed data.","tokens_in":16572,"tokens_out":4266,"duration_ms":36638,"significance":"If the claims hold, the paper fills a genuine gap by extending derivative-based FPCA to the multivariate setting. The two algorithms are clearly described and rely on standard theory (Kadota's differentiation theorem and Happ and Greven's relationship between univariate and multivariate FPCA). The simulation study is reasonably comprehensive and the application to coronary angiogram data is relevant, showing potential for clinical classification of coronary artery disease patterns. However, the central claim of DMFPCA's superiority is not fully established because of an asymmetric truncation in the DMKL comparison and because the medium-sparsity results contradict the abstract's general statement.","major_comments":[{"comment":"The comparison between DMFPCA and DMKL is asymmetric with respect to truncation. In Algorithm 2, the initial MFPCA of the original curves is performed with K=3, so the intermediate derivative estimates are restricted to the span of derivatives of the first three MFPCs of X. This subspace is generically not the optimal three-dimensional subspace for the derivatives, and the large ISE/RE for DMKL's third component is a direct consequence of this truncation choice. The paper itself acknowledges this on page 8: 'Increasing the truncation number could enable DMKL to yield eigencomponents closer to the true result,' yet no such variant is run. The abstract's claim that 'DMFPCA outperforms DMKL' is therefore not supported by the reported simulations; it may reverse if DMKL is given a larger initial truncation (e.g., K'=10) before the final reduction to K=3. The authors should provide this sensitivity analysis or otherwise justify the truncation.","section":"Section 5, Simulation (Figures 2–4)"},{"comment":"The central claim of DMFPCA's superiority is contradicted by the medium-sparsity results. The text on page 9 states: 'For medium sparsity case, DMFPCA yields slightly higher ISE than the P-splines + MFPCA method for all three eigenfunctions,' and Figure 4 shows that P-splines+MFPCA has the lowest RMISE (around 8%) compared with DMFPCA (around 10%). The abstract and the concluding discussion claim that 'DMFPCA outperforms DMKL and the direct approach, particularly for densely observed data,' but the only setting in which DMFPCA is consistently best is the dense case. The authors should either qualify the claim to 'for dense and high-sparsity settings' or provide an explanation for why medium sparsity is an exception, rather than making a blanket statement.","section":"Section 5, medium sparsity setting"},{"comment":"The estimators used in the simulation are not fully specified, which compromises reproducibility and the comparability of the methods. In Algorithm 1 (DMFPCA), the estimation of ∂^d∂^d C^{[p,p]} is left to the reader, but the simulation text only mentions that 'we use P-splines with same smoothing parameter to estimate ∂^d∂^d C^{[p,p]} for both medium and high sparsity cases.' The actual smoothing parameter is not given. Similarly, in Algorithm 2 (DMKL), the derivative of the estimated eigenfunctions is stated as 'left to the reader,' but the simulation must have used a specific method (e.g., finite differences), which is not reported. Without these details, the numerical results cannot be reproduced, and differences between methods may be driven by implementation choices rather than the algorithms themselves.","section":"Sections 4.1 and 4.2, Algorithm details"},{"comment":"The application pre-smooths the raw coronary angiogram curves with P-splines before applying DMFPCA and DMKL. The theoretical framework in Assumption 1 requires the d-th derivative of the observed process to exist almost surely, but raw curves are irregular and noisy, so the derivatives are well-defined only for the smoothed curves. The paper does not discuss that the DMFPCs and scores estimated in Section 6 describe the smoothed process, not the raw measurement process. While pre-smoothing is a common practical device, the authors should explicitly state that the interpretation of the estimated components is conditional on the smoothing step and clarify how this relates to the theoretical assumptions.","section":"Section 6, Application and Assumption 1"}],"minor_comments":[{"comment":"In the definition of ISE, the estimated eigenfunction is written as bψ^{(p)}_{d,k}(t_p), but elsewhere the notation is bψ^{[p]}_{d,k}; the parentheses appear to be a typo and the brackets are used consistently throughout the paper.","section":"Section 5, ISE definition"},{"comment":"The ground truth is obtained by applying MFPCA to the true derivatives and selecting K=3. It would be helpful to state the total variation explained by the first three components of the true derivative process (the text says '93%' but does not report the individual percentages), since this is the benchmark for the comparison.","section":"Section 5, Ground truth"},{"comment":"The captions say 'plus (blue) and minus (purple) the estimated eigenfunctions,' which is ambiguous; it should read 'the estimated mean derivative functions (red), the mean derivative plus the eigenfunction (blue), and the mean derivative minus the eigenfunction (purple)' to make the interpretation clear.","section":"Section 6, Figure 6 and 7 captions"},{"comment":"The discussion repeats the claim that DMFPCA 'should be the preferred choice over the other two methods, particularly for practical applications,' which is too strong given the medium-sparsity results. The wording should be consistent with the qualified simulation findings.","section":"Section 7, Discussion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of stat.ME and addresses a relevant problem. The main issue is the unfair DMKL comparison, which is a load-bearing point for the central claim. The authors should be required to run a sensitivity analysis on the truncation number for DMKL and to temper the abstract and discussion accordingly. The application is illustrative but would benefit from a clearer statement about the pre-smoothing step."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nHere's the short version: this is a competent, clearly written extension of univariate derivative PCA to multivariate functional data, but the abstract's headline claim is not fully supported by the simulations. The two methods—DMFPCA and DMKL—are essentially the two approaches from Grith et al. (2018) transplanted into the Happ–Greven MFPCA framework. The translation is legitimate and useful: the algorithms are detailed, the simulation study is well executed (500 replications, dense/sparse/noisy settings), and the coronary angiogram application demonstrates interpretable DMFPCs and scores.\n\nThe main soft spot is the simulation comparison. DMKL is handicapped by using the same truncation K=3 for both the initial MFPCA of the raw curves and the final MFPCA of the derivative estimates. That restricts DMKL to the span of derivatives of the first three MFPCs of X, which is generically not the best three-dimensional subspace for the derivatives. The paper even acknowledges that increasing the truncation number would help, but never runs that variant. So the claim that DMFPCA outperforms DMKL is not established; the hierarchy could easily shift if DMKL were given a larger initial truncation. The medium-sparsity results already show P-splines+MFPCA beating DMFPCA, making the reported ordering sensitive to implementation choices.\n\nA smaller concern is Assumption 1: the method requires the d-th derivative of the process to exist almost surely and the covariance to be sufficiently smooth. In the real data, the raw curves are irregular, so the derivatives are really properties of the pre-smoothed curves. That is a natural limitation but it deserves a clearer statement about the target of inference. Also, no code or theoretical convergence results are provided; that is not fatal for an applied method paper, but it limits reproducibility.\n\nI would send this to peer review. The methods are likely correct, the extension is genuinely useful for practitioners, and the simulation study is rigorous enough to justify referee time. The authors should fix the abstract, re-run DMKL with a larger initial truncation, and clarify the smoothness assumption. This is a solid tool paper, not a breakthrough, but it deserves a proper review and, with revisions, publication.\n\nBest regards","headline":"Solid multivariate extension of known univariate derivative PCA, but the simulations handicap DMKL and the abstract overclaims the DMFPCA advantage.","tokens_in":17094,"tokens_out":3013,"would_cite":true,"duration_ms":24301,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62R10","62H25"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper extends functional principal component analysis to derivatives of multivariate functional data, proposing DMFPCA and DMKL, and reports that DMFPCA is the more accurate method for densely observed data.","keywords":["derivative estimation","multivariate functional data","functional principal component analysis","Karhunen-Loève expansion","derivative covariance","coronary angiogram","DMFPCA","DMKL"],"falsifier":"Run a simulation in which the true curves violate Assumption 1, for example estimating second derivatives from curves that are only once differentiable or that have jumps, and compare both methods against analytic derivatives; if DMFPCA's eigencomponents remain accurate despite the violation, the identity is not the load-bearing step, and if they degrade sharply the assumption is confirmed as the gatekeeper. In a setting with known true derivatives, a more direct check is to compare $\\partial^d\\partial^d$ of the estimated featurewise covariance with the empirical covariance of the true derivatives, since a systematic discrepancy would falsify equation (4) as the estimation route.","tokens_in":16079,"feed_emoji":"📈","tokens_out":9030,"duration_ms":72181,"temperature":0.7,"pith_summary":"The paper's aim is to move derivative-based functional principal component analysis from the univariate setting to multivariate functional data, where each observation consists of several curves recorded together. It proposes two ways to estimate the eigencomponents and scores of the derivatives: DMFPCA, which estimates derivative covariances featurewise and then assembles multivariate principal components, and DMKL, which differentiates the multivariate Karhunen-Loève expansion of the original data. A direct approach, which smooths each curve, differentiates it, and then runs MFPCA, serves as the benchmark. The paper claims that DMFPCA gives the most accurate eigenfunctions, eigenvalues, and reconstructed derivatives for densely observed data and remains competitive under sparsity, and that the resulting derivative scores can be used to classify coronary artery disease patterns from angiograms.","feed_headline":"Derivatives of multivariate curves get their own principal components","feed_subtitle":"Two estimators, DMFPCA and DMKL, reconstruct derivative curves across features and beat the direct approach on dense data.","key_machinery":"The load-bearing object is the derivative covariance operator $\\Gamma_d$ with kernel $C_d(s,t)=E[\\,\\partial^d X(s)\\,\\partial^d X(t)^\\top\\,]$, together with identity (4), which equates its featurewise diagonal blocks to differentiated covariance surfaces, $\\partial^d\\partial^d C^{[p,p]}(s_p,t_p)=E[\\,\\partial^d X^{[p]}(s_p)\\,\\partial^d X^{[p]}(t_p)\\,]$. This identity lets DMFPCA estimate the derivative covariance from smoothed versions of the observed curves without unstable numerical differentiation of individual trajectories. From the spectral decomposition of $\\Gamma_d$ come the derivative multivariate functional principal components $\\psi_{d,k}$ and scores $\\rho_{d,k}$ that reconstruct derivatives through the multivariate Karhunen-Loève expansion; Happ and Greven's Proposition 5 supplies the bridge from per-feature eigencomponents to the joint multivariate ones, and the BLUP construction of Dai et al. predicts univariate scores from noisy discrete observations.","core_discovery":"On the paper's own terms, the central discovery is that the derivative process $\\partial^d X$ has its own multivariate Karhunen-Loève expansion, $\\partial^d X(t)=\\sum_{k=1}^\\infty \\rho_{d,k}\\psi_{d,k}(t)$, and that the covariance operator of this derivative process can be reached without differentiating the raw curves: under Assumption 1, the $d$-th derivative of each featurewise covariance equals the covariance of the $d$-th derivatives, $\\partial^d\\partial^d C^{[p,p]}(s_p,t_p)=C_d^{[p,p]}(s_p,t_p)$. DMFPCA exploits this identity to estimate derivative eigencomponents from smoothed covariance surfaces and then links the univariate pieces into multivariate principal components, while DMKL instead differentiates the original eigenfunctions and reuses the original scores; because derivatives of eigenfunctions are not orthogonal, DMKL's expansion is suboptimal for a fixed truncation. The simulation study supports the claim that DMFPCA outperforms both DMKL and the direct P-splines-plus-MFPCA approach for dense data, and the angiogram application shows the derivative components and scores carrying information about where diameter narrows and QFR drops fastest.","pith_inferences":["The same covariance-differentiation template should extend to second and higher derivatives as long as Assumption 1 holds; a natural next experiment is to compare DMFPCA's second-derivative estimates against analytic derivatives on the paper's simulation family.","DMFPCA's featurewise smoothing parameter is fixed in the simulation; adaptively choosing it per feature and per sampling density could close the gap to the direct approach in the medium-sparsity setting.","Because derivative scores summarize dynamics rather than levels, they may separate CAD patterns better than scores from the original diameter and QFR curves; the paper shows score boxplots by pattern but does not run a formal classifier, so that is a testable extension.","Identity (4) also opens a route to derivative estimation when curves from different features are observed on different grids, since the featurewise covariance derivatives can be estimated separately before the multivariate assembly step."],"forward_implications":["In dense, smooth settings, practitioners should expect DMFPCA to give the lowest errors in eigenfunctions, eigenvalues, and reconstructed derivatives, with average RMISE below 5% in the simulated no-noise and sigma=0.5 noise cases.","DMKL remains a serviceable route for reconstructing derivatives even under high sparsity, with average RMISE below 15%, but with a fixed truncation it can badly miss later eigencomponents because the differentiated original eigenfunctions are not orthogonal.","The direct approach of smoothing each curve, differentiating, and then running MFPCA is the safest fallback and the best of the three at medium sparsity, so the choice of method depends on sampling density.","Derivative scores from DMFPCA and DMKL are usable as predictors in follow-up analyses, such as classifying coronary artery disease patterns into focal, serial, diffuse, and mixed types.","In the coronary application, the methods locate the vessel position where diameter is narrowest and QFR is dropping fastest, information relevant to stent placement."],"supporting_citations":[{"why":"Supplies the term-by-term differentiability of the Karhunen-Loève expansion and the interchange of expectation and differentiation that underlies identity (4).","marker":"Kadota (1967)"},{"why":"Provides the multivariate functional PCA framework and Proposition 5 linking univariate to multivariate eigencomponents, used in DMFPCA.","marker":"Happ and Greven (2018)"},{"why":"Contributes the BLUP construction for predicting univariate derivative scores from noisy discrete observations.","marker":"Dai et al. (2018)"},{"why":"Earlier univariate derivative estimation via differentiated FPCA; DMKL extends this idea to the multivariate setting.","marker":"Liu and Müller (2009)"},{"why":"Prior extension of derivative FPCA to multidimensional domains and the two-approach template of differentiating the expansion versus expanding the derivatives.","marker":"Grith et al. (2018)"},{"why":"P-spline derivative estimation for longitudinal data, used to smooth curves and compute derivatives in the direct approach and the application.","marker":"Simpkin et al. (2018)"},{"why":"Alternative Gram-matrix MFPCA implementation offered alongside Happ and Greven for the direct estimation path.","marker":"Golovkine et al. (2023)"}],"fun_headline_variants":["Derivative-specific principal components for multivariate curves","DMFPCA outperforms other methods for dense multivariate derivatives","Multivariate Karhunen-Loève expansion for derivative curves","Estimating derivative principal components for multivariate data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument stands on Assumption 1: the $d$-th derivative of the underlying process exists almost surely and lies in $H$, and the mixed partial derivative of each feature's covariance $\\partial^d\\partial^d C^{[p,p]}$ exists and is continuous, so that expectation and differentiation can be interchanged; if the observed curves are not smooth enough for these derivatives to exist, the estimated derivative components and scores are not well defined.","fun_headline_variants_meta":{"raw":{"variants":["Derivative-specific principal components for multivariate curves","DMFPCA outperforms other methods for dense multivariate derivatives","Multivariate Karhunen-Loève expansion for derivative curves","Estimating derivative principal components for multivariate data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000782,"raw_usage":{"total_tokens":3467,"prompt_tokens":971,"completion_tokens":2496,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":2434}},"tokens_in":587,"tokens_out":2496,"duration_ms":18331,"temperature":1.0,"reasoning_tokens":2434,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:14:53.179010+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a simulation in which the true curves violate Assumption 1, for example estimating second derivatives from curves that are only once differentiable or that have jumps, and compare both methods against analytic derivatives; if DMFPCA's eigencomponents remain accurate despite the violation, the identity is not the load-bearing step, and if they degrade sharply the assumption is confirmed as the gatekeeper. In a setting with known true derivatives, a more direct check is to compare $\\partial^d\\partial^d$ of the estimated featurewise covariance with the empirical covariance of the true derivatives, since a systematic discrepancy would falsify equation (4) as the estimation route.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the term-by-term differentiability of the Karhunen-Loève expansion and the interchange of expectation and differentiation that underlies identity (4)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the BLUP construction for predicting univariate derivative scores from noisy discrete observations."},{"cited_title":"and M \\\"u ller, H.-G","cited_arxiv_id":null,"evidence_quote":"Earlier univariate derivative estimation via differentiated FPCA; DMKL extends this idea to the multivariate setting."},{"cited_title":"K., and Kneip, A","cited_arxiv_id":null,"evidence_quote":"Prior extension of derivative FPCA to multidimensional domains and the two-approach template of differentiating the expansion versus expanding the derivatives."},{"cited_title":"J., Durban, M., Lawlor, D","cited_arxiv_id":null,"evidence_quote":"P-spline derivative estimation for longitudinal data, used to smooth curves and compute derivatives in the direct approach and the application."},{"cited_title":"J., and Bargary, N","cited_arxiv_id":null,"evidence_quote":"Alternative Gram-matrix MFPCA implementation offered alongside Happ and Greven for the direct estimation path."}],"review_version":1}