REVIEW 3 major objections 7 minor 1 cited by
Regularized Generalized Covariance (RGCov) Estimator
T0 review · 3 major / 7 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper proposes a ridge-regularized Generalized Covariance estimator that stays consistent and asymptotically normal in high dimensions, and recovers the GCov efficiency bound when the shrinkage decays to zero.
desk verdict Useful ridge regularization for GCov with good simulations, but the asymptotic theory is missing a rate condition on δ_T and the identifiability claim is unverified for the paper's own designs. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the regularized variance matrix $\hat\Gamma_T(0;\theta,\delta)=\delta_T I_K+\hat\Gamma_T(0;\theta)$, whose inverse appears in each term of the objective $L_T(\theta,\delta_T)=\sum_{h=1}^H \operatorname{Tr}[\hat\Gamma_T(h;\theta)\hat\Gamma_T(0;\theta,\delta_T)^{-1}\hat\Gamma_T(h;\theta)'\hat\Gamma_T(0;\theta,\delta_T)^{-1}]$. The ridge term inflates the smallest sample eigenvalues and keeps the inverse numerically stable when $K=Jn$ is large, which is what fails for the unregularized GCov. The schedule $\delta_T\to\delta\ge 0$ controls the limiting variance: fixed $\delta$ yields a sandwich-type covariance matrix, while $\delta_T\to 0$ collapses it to the efficient GCov information matrix. A Sherman–Morrison recursion updates the inverse observation by observation, avoiding repeated high-dimensional inversions in the numerical optimization.
What would settle it
Fit the RGCov estimator to a mixed causal–noncausal VAR using only transformations that are functionally dependent, such as $u$ and $u^2$ with a known relation, and check whether the minimizer is unique across starting values and whether the RNLSD test keeps its nominal size; non-uniqueness or size distortion would show that identifiability, not the ridge term, is carrying the asymptotic claims.
Extended reading notes
Core claim
The central claim is that the regularized estimator $\hat\theta_T(\delta_T)$ is consistent for $\theta_0$ under standard regularity conditions, and that $\sqrt{T}(\hat\theta_T-\theta_0)$ converges in distribution to $N(0,J(\theta_0,\delta)^{-1}I(\theta_0,\delta)J(\theta_0,\delta)^{-1})$ when $\delta_T\to\delta\ge 0$. When $\delta_T\to 0$, this distribution simplifies to $N(0,J(\theta_0)^{-1})$, the same limit as the unregularized GCov estimator, so the regularized version is asymptotically semiparametrically efficient. The paper also proves that the residual-based RGCov specification test and the regularized NLSD test have asymptotic chi-square distributions with $K^2H-\dim(\theta)$ degrees of freedom when $\delta_T\to 0$, and weighted sums of chi-squares when $\delta$ is fixed. The regularization is confined to the weighting matrix in the objective; the VAR coefficients themselves are not shrunk.
Load-bearing premise
The load-bearing premise is that $\theta_0$ is uniquely identified by the finite restrictions $\Gamma(h;\theta)=0$ for $h=1,\ldots,H$, an assumption the paper states but does not prove for the transformation sets it actually uses; if those transforms fail to pin down $\theta_0$, neither consistency nor either chi-square limit follows no matter how $\delta$ is chosen.
Editorial extensions
If this is right
- High-dimensional GCov estimation becomes feasible without imposing sparsity on the VAR coefficients, because the only matrix that needs to be inverted is made regular by the ridge term.
- If $\delta_T$ is chosen to vanish with $T$, inference from RGCov is asymptotically equivalent to inference from GCov, so the regularization can be used as a computational stabilizer without sacrificing semiparametric efficiency.
- The RGCov residual-based specification test and the RNLSD test extend portmanteau testing to cases with many variables or many nonlinear transformations, with known chi-square degrees of freedom when the shrinkage decays.
- For fixed $\delta>0$, the null distribution is a weighted sum of chi-squares whose weights are products of eigenvalues of $\Gamma(0)^{-1/2}\Gamma(0,\delta)\Gamma(0)^{-1/2}$, giving a principled testing procedure when the shrinkage is not allowed to vanish.
- In the empirical application, RGCov identifies a mixed causal–noncausal VAR for twelve green-energy stock series where GCov yields near-zero eigenvalues of $\hat\Gamma(0)$, and the estimated causal and noncausal components are used to build portfolios that outperform the index in cumulative returns.
Reading between the lines
- The paper documents a bias–variance trade-off in $\delta$ but stops short of a data-driven selector; cross-validating $\delta$ on the objective or on test size is a natural next step that the simulations already support.
- The statement that $J\ge pn$ transformations are needed for identifiability sits awkwardly with the paper's own $n=15$, $J=2$ simulation, where identification succeeds; a fair reading is that identification depends on the informational content of the transformations, not just their count, and a formal condition would strengthen the result.
- Because the regularization targets the weighting matrix rather than the coefficients, the method does not produce a sparse VAR; combining RGCov with coefficient shrinkage is an extension the paper does not explore.
- For moderate samples with fixed $\delta$, a bootstrap calibrated to the weighted chi-square mixture could give better size control than the asymptotic approximation; this is a testable refinement not covered in the paper.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a regularized version of the Generalized Covariance (GCov) estimator, called RGCov, which replaces the inverse of the sample variance matrix Γ̂_T(0;θ) in the GCov objective by (δI + Γ̂_T(0;θ))^{-1}. The authors claim that, for a deterministic sequence δ_T → δ ≥ 0, the RGCov estimator is consistent and asymptotically normal with a covariance matrix they provide, and that when δ_T → 0 it achieves the same limiting distribution as the GCov estimator and thus semiparametric efficiency. They also extend the GCov specification test and the NLSD test to the regularized setting, obtaining a weighted chi-square limit for δ_T → δ>0 and a chi-square limit for δ_T → 0. Finite-sample performance is examined in two simulation designs (high n and high J), and the method is applied to a mixed causal-noncausal VAR for 12 green energy stock prices.
Significance. If the theoretical claims are correct, the paper offers a practical, computationally feasible way to estimate mixed causal-noncausal VARs in high dimensions, with the Sherman-Morrison updates and the diagonal estimator providing useful computational alternatives. The simulation evidence of strong variance reduction relative to the unregularized GCov is striking. However, the identification condition underlying all consistency results is not verified, and the proofs are presented as sketches that rely heavily on prior work. The paper is therefore a potentially useful contribution, but its central theorems need to be placed on firmer ground.
major comments (3)
- [Section 3 / Assumption A.1(iii)] The identifiability condition A.1(iii), which requires that Γ(h;θ)=0 for h=1,...,H has θ0 as its unique solution, is never verified for the transformation sets used in Section 4.1 (J=2), Section 4.2 (J=10) or Section 5 (J=4). The only supporting statement, 'it is necessary to use nonlinear transformations J≥pn to ensure the identifiability of the model' (Section 3), is contradicted by the paper's own simulation in Section 4.1, where n=15, p=1 and J=2 yet identification is reported as successful. Because Propositions 1, 2 and 4 all invoke A.1(iii), this is a load-bearing gap.
- [Appendix A, step d)] The expansion of the first-order conditions asserts that the term √T ∂²L_T(θ0,δ)/∂θ∂δ′ (δ_T−δ) is negligible whenever δ_T→δ. This requires √T ∂²L_T/∂θ∂δ′ to be stochastically bounded; the manuscript neither proves this nor states primitive conditions. If this derivative is not O_p(1) after scaling, a rate condition on δ_T is needed. The proof of Proposition 1(ii) and the δ_T→0 efficiency claim are therefore not fully established as written.
- [Section 3.2.2, Proposition 3] The proof of Proposition 3 is given only for the RNLSD case (dim θ=0). The general case with dim θ>0, which justifies the degrees of freedom K²H−dimθ used in Proposition 4, is dismissed with 'the general case is similar' and no details are provided. Since the test statistics are a central claimed contribution, a rigorous proof or a precise citation for the estimation effect is needed.
minor comments (7)
- [Section 2.1, Eq. (2.2)] The displayed definition of v_t(θ) repeats the same blocks twice; it should list a_j[g_i(ỹ_t;θ)] for i=1,...,n and j=1,...,J once.
- [Section 3.1, Definition 1] In (3.2), the shrinkage coefficient in R̂_T^2 is written as δ, but the definition uses δ_T; the argument should be made consistent.
- [Section 4.2, Table 4] The caption of Table 4 says 'fourteen inside and one outside' but the DGP in Section 4.2 has two eigenvalues inside and one outside the unit circle.
- [Section 4.1, text after Table 1] The text refers to 'Table??' when discussing identification frequencies; this should be Table 2.
- [Section 4.1, simulation design] In the δ_T=η/T setting, the largest η values (e.g., η=800 with T=800) imply δ_T=1, so the simulations do not actually explore the δ_T→0 regime of Proposition 2; the finite-sample evidence is therefore primarily about fixed (or slowly varying) δ.
- [Section 5] The selection of δ=0.3 via 'a higher distance' of eigenvalues from unity is ad hoc and not grounded in the theoretical results; a data-driven selection rule would be preferable.
- [Global] The index name is spelled both 'Rennix' and 'Renixx' at different places; please make it consistent.
Circularity Check
No circularity: the δ→0 case is an explicitly acknowledged algebraic special case of GCov, while the fixed-δ limit theorems are derived directly from the regularized objective.
full rationale
The RGCov estimator is defined by Eqs. (3.1)-(3.3) as the minimizer of a δ-regularized modification of the GCov objective. Proposition 1's consistency proof uses the standard Jennrich argument: the objective converges uniformly to L(θ,δ)=Σ_h Tr[Γ(h;θ)Γ(0;θ,δ)^{-1}Γ(h;θ)'Γ(0;θ,δ)^{-1}], which is zero at θ0 and, by Assumption A.1(iii), has no other zero, so the minimizer converges to θ0. The asymptotic normality result is obtained from a first-order expansion of the first-order conditions and the classical CLT for sample autocovariances of i.i.d. vectors (Chitturi 1976; Hannan 1976), giving the matrices J(θ0,δ) and I(θ0,δ) explicitly. Proposition 3 derives the mixture-of-chi-square limit by writing √T vec Γ̂T(h) as X(h)~N(0,Γ(0)⊗Γ(0)) and diagonalizing Γ(0;δ)^{-1}⊗Γ(0;δ)^{-1}; this is a direct quadratic-form calculation rather than an imported distributional conclusion. When δ=0, Γ(0;θ,0)=Γ(0;θ), so the RGCov objective, estimator, and test statistics coincide algebraically with the GCov versions; the paper states this as a special case ('This includes as a special case δT=δ>0' and 'in the special case δ=0... simplified') rather than relabeling an input as a prediction. The efficiency statement at δ→0 cites Gourieroux and Jasiak (2023), but that citation is to published, parameter-free GCov theory whose assumptions do not include the RGCov result; it is prior evidence, not a circular input. Similarly, Jasiak and Neyazi (2023) is cited for the NLSD test, but the RNLSD limit is derived in Proposition 3 and only reduces to the NLSD chi-square when δ=0, again an algebraic limit. The unverified identifiability condition A.1(iii), and the apparent tension between Section 3's 'J≥pn' statement and Section 4.1's J=2 design, are correctness risks rather than circular reductions, since every stated theorem is conditional on A.1(iii). No equation in the paper reduces a claimed output to a fitted parameter or to a self-citation by construction.
Assumptions & free parameters
free parameters (2)
- Shrinkage coefficient delta =
delta = 0.3 in application; grid 0.5, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0 in simulations
- Eta in delta_T = eta/T =
eta from 100 to 800 in simulations
assumptions (5)
- domain assumption Asymptotic identifiability: Gamma(h;theta)=0 for h=1,...,H implies theta=theta_0 (Assumption A.1(iii))
- domain assumption Regularity conditions: compact parameter space, geometric ergodicity, continuous invariant distribution, twice differentiability, stochastic Lipschitz equicontinuity (Assumptions A.1-A.2)
- domain assumption Known semiparametric efficiency of the GCov estimator (Gourieroux and Jasiak 2023)
- ad hoc to paper Unstated rate condition: sqrt(T)(delta_T - delta) -> 0 for asymptotic normality
- standard math Limiting Gaussianity of sqrt(T) vec Gamma_hat_T(h) under serial independence (Chitturi 1976, Hannan 1976)
Cite this review
Pith. "Pith review of Regularized Generalized Covariance (RGCov) Estimator." pith.science (2026). https://pith.science/paper/N6YE4NI3
@misc{pith2026250418678,
author = {Pith},
title = {Pith review of: Regularized Generalized Covariance (RGCov) Estimator},
year = {2026},
howpublished = {\url{https://pith.science/paper/N6YE4NI3}},
note = {Machine review of arXiv:2504.18678}
}
read the original abstract
We introduce a regularized Generalized Covariance (RGCov) estimator as an extension of the GCov estimator to high dimensional setting that results either from high-dimensional data or a large number of nonlinear transformations used in the objective function. The approach relies on a ridge-type regularization for high-dimensional matrix inversion in the objective function of the GCov. The RGCov estimator is consistent and asymptotically normally distributed. We provide the conditions under which it can reach semiparametric efficiency and discuss the selection of the optimal regularization parameter. We also examine the diagonal GCov estimator, which simplifies the computation of the objective function. The GCov-based specification test, and the test for nonlinear serial dependence (NLSD) are extended to the regularized RGCov specification and RNLSD tests with asymptotic Chi-square distributions. Simulation studies show that the RGCov estimator and the regularized tests perform well in the high dimensional setting. We apply the RGCov to estimate the mixed causal and noncausal VAR model of stock prices of green energy companies.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Nonfundamentalness or missing information ? Evidence from causal-noncausal VARs in macro-finance
In the Stock–Watson monetary VAR, noncausal roots mostly disappear after factor filtering, and filtered IRFs no longer show the price puzzle.
Reference graph
Works this paper leans on
-
[1]
Bierens, H. J. (1990). A consistent conditional moment test of functional form. Econometrica: Journal of the Econometric Society\/ , 1443--1458
work page 1990
-
[2]
Breidt, F. J., R. A. Davis, K.-S. Lii, and M. Rosenblatt (1991). Maximum likelihood estimation for noncausal autoregressive processes. Journal of Multivariate Analysis\/ 36\/ (2), 175--198
work page 1991
- [3]
-
[4]
Chitturi, R. V. (1976). Distribution of multivariate white noise autocorrelations. Journal of the American Statistical Association\/ 71\/ (353), 223--226
work page 1976
-
[5]
Sequential Monte Carlo for Noncausal Processes
Cubadda, G., F. Giancaterini, and S. Grassi (2025). Sequential monte carlo for noncausal processes. arXiv preprint arXiv:2501.03945\/
work page Pith review arXiv 2025
-
[6]
Cubadda, G., F. Giancaterini, A. Hecq, and J. Jasiak (2024). Optimization of the generalized covariance estimator in noncausal processes. Statistics and Computing\/ 34\/ (4), 127
work page 2024
-
[7]
Cubadda, G. and A. Hecq (2011). Testing for common autocorrelation in data-rich environments. Journal of Forecasting\/ 30\/ (3), 325--335
work page 2011
-
[8]
Cubadda, G., A. Hecq, and S. Telg (2019). Detecting co-movements in non-causal time series. Oxford Bulletin of Economics and Statistics\/ 81\/ (3), 697--715
work page 2019
Show all 32 references
-
[9]
Hecq, and E
Cubadda, G., A. Hecq, and E. Voisin (2023). Detecting common bubbles in multivariate mixed causal--noncausal models. Econometrics\/ 11\/ (1), 9
2023
-
[10]
Davis, R. A. and L. Song (2020). Noncausal vector ar processes with application to economic time series. Journal of Econometrics\/ 216\/ (1), 246--267
2020
-
[11]
Engle, R. F. and C. W. Granger (1987). Co-integration and error correction: representation, estimation, and testing. Econometrica: journal of the Econometric Society\/ , 251--276
1987
-
[12]
Engle, R. F. and S. Kozicki (1993). Testing for common features. Journal of Business & Economic Statistics\/ 11\/ (4), 369--380
1993
-
[13]
and J.-M
Fries, S. and J.-M. Zakoian (2019). Mixed causal-noncausal ar processes and the modelling of explosive bubbles. Econometric Theory\/ 35\/ (6), 1234--1270
2019
-
[14]
Hencic, and J
Gourieroux, C., A. Hencic, and J. Jasiak (2021). Forecast performance and bubble analysis in noncausal mar (1, 1) processes. Journal of Forecasting\/ 40\/ (2), 301--326
2021
-
[15]
Gourieroux, C. and J. Jasiak (2016). Filtering, prediction and simulation methods for noncausal processes. Journal of Time Series Analysis\/ 37\/ (3), 405--430
2016
-
[16]
Gourieroux, C. and J. Jasiak (2017). Noncausal vector autoregressive process: Representation, identification and semi-parametric estimation. Journal of Econometrics\/ 200\/ (1), 118--134
2017
-
[17]
Gourieroux, C. and J. Jasiak (2022). Nonlinear forecasts and impulse responses for causal-noncausal (s) var models. arXiv preprint arXiv:2205.09922\/
2022 arXiv
-
[18]
Gourieroux, C. and J. Jasiak (2023). Generalized covariance estimator. Journal of Business & Economic Statistics\/ 41\/ (4), 1315--1327
2023
-
[19]
Jasiak, and M
Gourieroux, C., J. Jasiak, and M. Tong (2021). Convolution-based filtering and forecasting: An application to wti crude oil prices. Journal of Forecasting\/ 40\/ (7), 1230--1244
2021
-
[20]
and J.-M
Gouri \'e roux, C. and J.-M. Zako \"i an (2017). Local explosion modelling by non-causal process. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 79\/ (3), 737--756
2017
-
[21]
Hall, M. K. and J. Jasiak (2024). Modelling common bubbles in cryptocurrency prices. Economic Modelling\/ 139 , 106782
2024
-
[22]
Hannan, E. J. (1976). The asymptotic distribution of serial covariances. The Annals of Statistics\/ 4\/ (2), 396--399
1976
-
[23]
Lieb, and S
Hecq, A., L. Lieb, and S. Telg (2016). Identification of mixed causal-noncausal models in finite samples. Annals of Economics and Statistics/Annales d' \'E conomie et de Statistique\/ (123/124), 307--331
2016
-
[24]
Hecq, A. and E. Voisin (2021). Forecasting bubbles with mixed causal-noncausal autoregressive models. Econometrics and Statistics\/ 20 , 29--45
2021
-
[25]
Hencic, A. and C. Gouri \'e roux (2015). Noncausal autoregressive model in application to bitcoin/usd exchange rates. Econometrics of risk\/ 583 , 17--40
2015
-
[26]
Jasiak, J. and A. M. Neyazi (2023). Gcov-based portmanteau test. arXiv preprint arXiv:2312.05373\/
2023
-
[27]
Lanne, M. and P. Saikkonen (2011). Noncausal autoregressions for economic time series. Journal of Time Series Econometrics\/ 3\/ (3), Article 2
2011
-
[28]
Lanne, M. and P. Saikkonen (2013). Noncausal vector autoregression. Econometric Theory\/ 29\/ (3), 447--481
2013
-
[29]
Lof, M. and H. Nyberg (2017). Noncausality and the commodity currency hypothesis. Energy Economics\/ 65 , 424--433
2017
-
[30]
Sherman, J. and W. J. Morrison (1949). Adjustment of an inverse matrix corresponding to a change in one element of a given matrix. In Annals of Mathematical Statistics , Volume 20, pp.\ 317--317
1949
-
[31]
Sherman, J. and W. J. Morrison (1950). Adjustment of an inverse matrix corresponding to a change in one element of a given matrix. The Annals of Mathematical Statistics\/ 21\/ (1), 124--127
1950
-
[32]
Swensen, A. R. (2022). On causal and non-causal cointegrated vector autoregressive time series. Journal of Time Series Analysis\/ 43\/ (2), 178--196
2022
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.