{"id":"97a5f533-90e9-415b-b257-a4f87224bbb6","arxiv_id":"2504.18678","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A ridge-regularized Generalized Covariance estimator is proposed for high-dimensional mixed causal-noncausal VAR models, with asymptotic normality and chi-square tests when the shrinkage goes to zero.","lead":"This paper introduces a ridge-regularized version of the GCov estimator so it can be used when many variables or many nonlinear transformations make the covariance matrix hard to invert. If the shrinkage parameter shrinks to zero with the sample size, the regularized estimator inherits the efficiency of the original GCov estimator, and its tests keep chi-square limits.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption A.1(iii) (identifiability from H autocovariances) is unverified for the paper's own transform sets, and Section 3's J≥pn claim conflicts with Section 4.1's n=15, J=2 design; all limit theorems depend on it.","rationale":"The reader's weakest-assumption analysis and mine coincide: the unverified identification condition A.1(iii) is the load-bearing premise, because consistency, asymptotic normality, and the chi-square limits all fail if the limiting objective has multiple zeros. I also checked the other candidate concern, the omitted rate condition on δ_T in Appendix A. At θ0 all population autocovariances vanish, so the √T(δ_T-δ) term in the FOC expansion is multiplied by a second derivative that is O_p(T^{-1/2}); the rate condition is not the main gap. The identifiability issue is more serious because it is not merely a missing regularity detail but an unproven primitive for the exact transform sets in the simulations and application. The paper's Section 3 dimension statement is internally inconsistent with Section 4.1, which weakens the authors' own support for the assumption. The recommended verdict remains CONDITIONAL (UNCHANGED): the RGCov idea is plausible and the simulations are suggestive, but the theorems should not be used for a full-VAR high-dimensional claim until identification is verified or explicitly assumed with a worked example. No ad hominem is intended; the concern is about the argument's missing support.","tokens_in":23987,"tokens_out":14150,"duration_ms":153101,"concrete_test":"Compute the population RGCov objective L(θ,δ)=Σ_{h=1}^H Tr[Γ(h;θ)Γδ(0;θ)^{-1}Γ(h;θ)'Γδ(0;θ)^{-1}] for the Section 4.1 DGP and check whether it has minimizers other than θ0. First report whether the simulation's Φ is diagonal or full. Then, for a full n×n Φ with n=15, p=1, use the linear/quadratic transforms (J=2) and evaluate L on a grid around θ0, including the time-reversed/noncausal parameterization; any θ≠θ0 with L=0 (or with L equal to the value at θ0) rejects A.1(iii) for that design. Complement this with a multi-start optimization at T=800: if different starting values converge to distinct θ with near-zero objective, identifiability is not supported and Propositions 1-4 cannot be applied to the claimed high-dimensional setting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The consistency and distribution results in Propositions 1, 2, and 4 all rest on Assumption A.1(iii): the population autocovariances Γ(h;θ)=0 for h=1,...,H must have θ0 as their unique solution. The paper states this as a high-level condition but never verifies it for the transformation sets actually used: linear/quadratic transforms in Section 4.1, the ten transforms in Section 4.2, or the powers used in Section 5. The only supporting dimension argument in Section 3, 'it is necessary to use nonlinear transformations J≥pn to ensure the identifiability of the model,' is inconsistent with Section 4.1, where n=15, p=1, J=2 and identification is reported as successful. If the DGP there has a diagonal Φ, the parameter space is only 15-dimensional and the J≥pn statement is not informative for a full VAR; if Φ is full, the J≥pn claim is false. In either case, the identification condition is not established. If A.1(iii) fails, a θ≠θ0 minimizes the limiting RGCov objective and the claimed consistency, asymptotic normality, and chi-square test limits fail simultaneously.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a regularized version of the Generalized Covariance (GCov) estimator, called RGCov, which replaces the inverse of the sample variance matrix Γ̂_T(0;θ) in the GCov objective by (δI + Γ̂_T(0;θ))^{-1}. The authors claim that, for a deterministic sequence δ_T → δ ≥ 0, the RGCov estimator is consistent and asymptotically normal with a covariance matrix they provide, and that when δ_T → 0 it achieves the same limiting distribution as the GCov estimator and thus semiparametric efficiency. They also extend the GCov specification test and the NLSD test to the regularized setting, obtaining a weighted chi-square limit for δ_T → δ>0 and a chi-square limit for δ_T → 0. Finite-sample performance is examined in two simulation designs (high n and high J), and the method is applied to a mixed causal-noncausal VAR for 12 green energy stock prices.","tokens_in":24233,"tokens_out":17120,"duration_ms":148856,"significance":"If the theoretical claims are correct, the paper offers a practical, computationally feasible way to estimate mixed causal-noncausal VARs in high dimensions, with the Sherman-Morrison updates and the diagonal estimator providing useful computational alternatives. The simulation evidence of strong variance reduction relative to the unregularized GCov is striking. However, the identification condition underlying all consistency results is not verified, and the proofs are presented as sketches that rely heavily on prior work. The paper is therefore a potentially useful contribution, but its central theorems need to be placed on firmer ground.","major_comments":[{"comment":"The identifiability condition A.1(iii), which requires that Γ(h;θ)=0 for h=1,...,H has θ0 as its unique solution, is never verified for the transformation sets used in Section 4.1 (J=2), Section 4.2 (J=10) or Section 5 (J=4). The only supporting statement, 'it is necessary to use nonlinear transformations J≥pn to ensure the identifiability of the model' (Section 3), is contradicted by the paper's own simulation in Section 4.1, where n=15, p=1 and J=2 yet identification is reported as successful. Because Propositions 1, 2 and 4 all invoke A.1(iii), this is a load-bearing gap.","section":"Section 3 / Assumption A.1(iii)"},{"comment":"The expansion of the first-order conditions asserts that the term √T ∂²L_T(θ0,δ)/∂θ∂δ′ (δ_T−δ) is negligible whenever δ_T→δ. This requires √T ∂²L_T/∂θ∂δ′ to be stochastically bounded; the manuscript neither proves this nor states primitive conditions. If this derivative is not O_p(1) after scaling, a rate condition on δ_T is needed. The proof of Proposition 1(ii) and the δ_T→0 efficiency claim are therefore not fully established as written.","section":"Appendix A, step d)"},{"comment":"The proof of Proposition 3 is given only for the RNLSD case (dim θ=0). The general case with dim θ>0, which justifies the degrees of freedom K²H−dimθ used in Proposition 4, is dismissed with 'the general case is similar' and no details are provided. Since the test statistics are a central claimed contribution, a rigorous proof or a precise citation for the estimation effect is needed.","section":"Section 3.2.2, Proposition 3"}],"minor_comments":[{"comment":"The displayed definition of v_t(θ) repeats the same blocks twice; it should list a_j[g_i(ỹ_t;θ)] for i=1,...,n and j=1,...,J once.","section":"Section 2.1, Eq. (2.2)"},{"comment":"In (3.2), the shrinkage coefficient in R̂_T^2 is written as δ, but the definition uses δ_T; the argument should be made consistent.","section":"Section 3.1, Definition 1"},{"comment":"The caption of Table 4 says 'fourteen inside and one outside' but the DGP in Section 4.2 has two eigenvalues inside and one outside the unit circle.","section":"Section 4.2, Table 4"},{"comment":"The text refers to 'Table??' when discussing identification frequencies; this should be Table 2.","section":"Section 4.1, text after Table 1"},{"comment":"In the δ_T=η/T setting, the largest η values (e.g., η=800 with T=800) imply δ_T=1, so the simulations do not actually explore the δ_T→0 regime of Proposition 2; the finite-sample evidence is therefore primarily about fixed (or slowly varying) δ.","section":"Section 4.1, simulation design"},{"comment":"The selection of δ=0.3 via 'a higher distance' of eigenvalues from unity is ad hoc and not grounded in the theoretical results; a data-driven selection rule would be preferable.","section":"Section 5"},{"comment":"The index name is spelled both 'Rennix' and 'Renixx' at different places; please make it consistent.","section":"Global"}],"recommendation":"major_revision","confidential_remarks":"The manuscript relies heavily on prior work by the same group (Gourieroux and Jasiak 2023; Jasiak and Neyazi 2023), and the extension to regularization is incremental. The paper may be suitable for a journal that values methodological extensions in time series econometrics, but the editor should ensure that the proofs are verifiable before acceptance. The paper also contains several typos and cross-reference errors that suggest a need for careful copy-editing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you spend time on this. First, the regularized GCov estimator is a sensible answer to a real numerical problem: inverting Γ(0) when K is large is fragile, and adding a ridge term stabilizes it. The simulations do show the fix working, and the application to green-energy stock prices is a legitimate stress test. Second, the asymptotic theory as written is not fully established. The proof of Proposition 1 expands the FOC and then drops √T(δ_T−δ) with only the assumption δ_T→δ. That needs a rate condition such as δ_T−δ = o_p(1/√T). As stated, Propositions 1 and 2 hold for δ_T = η/T, which is the case the simulations actually use, but not for the more general claim. That is a fixable gap, but it is load-bearing.\n\nWhat is genuinely new: the ridge-regularized objective, the fixed-δ mixture-of-chi-square limit for the regularized specification test and the RNLSD test, and the delta→0 chi-square special case. The Sherman-Morrison updating is a practical contribution. The paper is also honest about the bias-variance tradeoff in choosing δ, and the Monte Carlo work is thorough. The self-citation to Gourieroux and Jasiak for the δ=0 efficiency is legitimate; that case is the published GCov estimator, and importing its efficiency is not circular.\n\nThe soft spots are in proportion. The identifiability assumption A.1(iii) is stated as a high-level condition and never verified for the transformation sets actually used. Section 3's claim that J≥pn is necessary contradicts the Section 4.1 design (n=15, p=1, J=2), where identification is reported as successful only because the DGP is essentially diagonal; a full VAR with that design would not be identified from the linear and quadratic moments. That tension needs to be resolved, either by proving identifiability for the specific transforms or by softening the dimension claim. The application also selects δ after inspecting the estimated eigenvalues, which makes the causal-noncausal conclusion somewhat fragile.\n\nMy bottom line: this is a useful paper with a plausible estimator, but the theorems need a rate condition and the identifiability discussion needs care before the results are used as stated. It deserves a serious referee, not a desk reject. I would send it out, with a request for the authors to fix the δ_T rate and either prove or qualify the identifiability claim.","headline":"Useful ridge regularization for GCov with good simulations, but the asymptotic theory is missing a rate condition on δ_T and the identifiability claim is unverified for the paper's own designs.","tokens_in":24784,"tokens_out":2848,"would_cite":false,"duration_ms":30160,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62M10","62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a ridge-regularized Generalized Covariance estimator that stays consistent and asymptotically normal in high dimensions, and recovers the GCov efficiency bound when the shrinkage decays to zero.","keywords":["high-dimensional time series","Generalized Covariance estimator","ridge regularization","mixed causal-noncausal VAR","semiparametric efficiency","nonlinear serial dependence test","portmanteau test","RNLSD test"],"falsifier":"Fit the RGCov estimator to a mixed causal–noncausal VAR using only transformations that are functionally dependent, such as $u$ and $u^2$ with a known relation, and check whether the minimizer is unique across starting values and whether the RNLSD test keeps its nominal size; non-uniqueness or size distortion would show that identifiability, not the ridge term, is carrying the asymptotic claims.","tokens_in":23757,"feed_emoji":"📈","tokens_out":11249,"duration_ms":101034,"temperature":0.7,"pith_summary":"This paper claims that the Generalized Covariance (GCov) estimator, a semiparametric method for mixed causal–noncausal vector autoregressions, can be made to work when the covariance matrix it has to invert is high-dimensional. The proposed fix replaces the sample variance matrix $\\hat\\Gamma_T(0;\\theta)$ by the ridge-regularized matrix $\\delta_T I + \\hat\\Gamma_T(0;\\theta)$ inside the objective function. The resulting RGCov estimator is consistent and asymptotically normal, and when the shrinkage coefficient decays to zero with the sample size it recovers the GCov limiting distribution and its semiparametric efficiency. The same regularization extends the GCov specification test and the nonlinear serial dependence (NLSD) test, with chi-square limits when $\\delta_T\\to 0$. Simulations show gains over GCov and diagonal GCov, and an application to green-energy stock prices illustrates the method on a 12-variable system.","feed_headline":"A ridge term makes high-dimensional covariance estimation work","feed_subtitle":"Regularizing the matrix before inversion keeps the GCov estimator stable; as δ shrinks, efficiency returns.","key_machinery":"The load-bearing object is the regularized variance matrix $\\hat\\Gamma_T(0;\\theta,\\delta)=\\delta_T I_K+\\hat\\Gamma_T(0;\\theta)$, whose inverse appears in each term of the objective $L_T(\\theta,\\delta_T)=\\sum_{h=1}^H \\operatorname{Tr}[\\hat\\Gamma_T(h;\\theta)\\hat\\Gamma_T(0;\\theta,\\delta_T)^{-1}\\hat\\Gamma_T(h;\\theta)'\\hat\\Gamma_T(0;\\theta,\\delta_T)^{-1}]$. The ridge term inflates the smallest sample eigenvalues and keeps the inverse numerically stable when $K=Jn$ is large, which is what fails for the unregularized GCov. The schedule $\\delta_T\\to\\delta\\ge 0$ controls the limiting variance: fixed $\\delta$ yields a sandwich-type covariance matrix, while $\\delta_T\\to 0$ collapses it to the efficient GCov information matrix. A Sherman–Morrison recursion updates the inverse observation by observation, avoiding repeated high-dimensional inversions in the numerical optimization.","core_discovery":"The central claim is that the regularized estimator $\\hat\\theta_T(\\delta_T)$ is consistent for $\\theta_0$ under standard regularity conditions, and that $\\sqrt{T}(\\hat\\theta_T-\\theta_0)$ converges in distribution to $N(0,J(\\theta_0,\\delta)^{-1}I(\\theta_0,\\delta)J(\\theta_0,\\delta)^{-1})$ when $\\delta_T\\to\\delta\\ge 0$. When $\\delta_T\\to 0$, this distribution simplifies to $N(0,J(\\theta_0)^{-1})$, the same limit as the unregularized GCov estimator, so the regularized version is asymptotically semiparametrically efficient. The paper also proves that the residual-based RGCov specification test and the regularized NLSD test have asymptotic chi-square distributions with $K^2H-\\dim(\\theta)$ degrees of freedom when $\\delta_T\\to 0$, and weighted sums of chi-squares when $\\delta$ is fixed. The regularization is confined to the weighting matrix in the objective; the VAR coefficients themselves are not shrunk.","pith_inferences":["The paper documents a bias–variance trade-off in $\\delta$ but stops short of a data-driven selector; cross-validating $\\delta$ on the objective or on test size is a natural next step that the simulations already support.","The statement that $J\\ge pn$ transformations are needed for identifiability sits awkwardly with the paper's own $n=15$, $J=2$ simulation, where identification succeeds; a fair reading is that identification depends on the informational content of the transformations, not just their count, and a formal condition would strengthen the result.","Because the regularization targets the weighting matrix rather than the coefficients, the method does not produce a sparse VAR; combining RGCov with coefficient shrinkage is an extension the paper does not explore.","For moderate samples with fixed $\\delta$, a bootstrap calibrated to the weighted chi-square mixture could give better size control than the asymptotic approximation; this is a testable refinement not covered in the paper."],"forward_implications":["High-dimensional GCov estimation becomes feasible without imposing sparsity on the VAR coefficients, because the only matrix that needs to be inverted is made regular by the ridge term.","If $\\delta_T$ is chosen to vanish with $T$, inference from RGCov is asymptotically equivalent to inference from GCov, so the regularization can be used as a computational stabilizer without sacrificing semiparametric efficiency.","The RGCov residual-based specification test and the RNLSD test extend portmanteau testing to cases with many variables or many nonlinear transformations, with known chi-square degrees of freedom when the shrinkage decays.","For fixed $\\delta>0$, the null distribution is a weighted sum of chi-squares whose weights are products of eigenvalues of $\\Gamma(0)^{-1/2}\\Gamma(0,\\delta)\\Gamma(0)^{-1/2}$, giving a principled testing procedure when the shrinkage is not allowed to vanish.","In the empirical application, RGCov identifies a mixed causal–noncausal VAR for twelve green-energy stock series where GCov yields near-zero eigenvalues of $\\hat\\Gamma(0)$, and the estimated causal and noncausal components are used to build portfolios that outperform the index in cumulative returns."],"supporting_citations":[{"why":"Introduces the GCov estimator for noncausal VAR models, whose objective and regularity framework the paper extends.","marker":"Gourieroux and Jasiak (2017)"},{"why":"Establishes the semiparametric efficiency and chi-square tests of GCov that RGCov recovers when $\\delta_T\\to 0$.","marker":"Gourieroux and Jasiak (2023)"},{"why":"Defines the NLSD portmanteau test that the paper regularizes into the RNLSD test.","marker":"Jasiak and Neyazi (2023)"},{"why":"Provides the asymptotic normal distribution of sample autocovariances used in the derivation of the score and test limits.","marker":"Chitturi (1976)"},{"why":"Supplies the same distributional result for serial covariances, cited alongside Chitturi in the proof of Proposition 1.","marker":"Hannan (1976)"},{"why":"Gives the rank-one inverse-update formula used to compute the regularized inverse recursively.","marker":"Sherman and Morrison (1949)"},{"why":"Restates the inverse-update formula applied in Corollaries 1–2 for efficient numerical implementation.","marker":"Sherman and Morrison (1950)"},{"why":"Justifies using autocovariances of nonlinear transforms to reveal nonlinear serial dependence, the basis for the GCov objective.","marker":"Chan et al. (2006)"}],"fun_headline_variants":["Ridge regularization tames high-dimensional covariance estimation","RGCov: Ridge-stabilized estimator for big covariance problems","High-dimensional GCov? Add a ridge to keep it efficient","Ridge-inverted GCov stays consistent and efficient","Regularizing the inverse, not coefficients, boosts GCov"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that $\\theta_0$ is uniquely identified by the finite restrictions $\\Gamma(h;\\theta)=0$ for $h=1,\\ldots,H$, an assumption the paper states but does not prove for the transformation sets it actually uses; if those transforms fail to pin down $\\theta_0$, neither consistency nor either chi-square limit follows no matter how $\\delta$ is chosen.","fun_headline_variants_meta":{"raw":{"variants":["Ridge regularization tames high-dimensional covariance estimation","RGCov: Ridge-stabilized estimator for big covariance problems","High-dimensional GCov? Add a ridge to keep it efficient","Ridge-inverted GCov stays consistent and efficient","Regularizing the inverse, not coefficients, boosts GCov"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00081,"raw_usage":{"total_tokens":3552,"prompt_tokens":940,"completion_tokens":2612,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":2531}},"tokens_in":556,"tokens_out":2612,"duration_ms":16582,"temperature":1.0,"reasoning_tokens":2531,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:12:47.519630+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit the RGCov estimator to a mixed causal–noncausal VAR using only transformations that are functionally dependent, such as $u$ and $u^2$ with a known relation, and check whether the minimizer is unique across starting values and whether the RNLSD test keeps its nominal size; non-uniqueness or size distortion would show that identifiability, not the ridge term, is carrying the asymptotic claims.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the GCov estimator for noncausal VAR models, whose objective and regularity framework the paper extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the semiparametric efficiency and chi-square tests of GCov that RGCov recovers when $\\delta_T\\to 0$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the NLSD portmanteau test that the paper regularizes into the RNLSD test."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the asymptotic normal distribution of sample autocovariances used in the derivation of the score and test limits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the same distributional result for serial covariances, cited alongside Chitturi in the proof of Proposition 1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the rank-one inverse-update formula used to compute the regularized inverse recursively."},{"cited_title":"Ho, and H","cited_arxiv_id":null,"evidence_quote":"Justifies using autocovariances of nonlinear transforms to reveal nonlinear serial dependence, the basis for the GCov objective."}],"review_version":1}