REVIEW 4 major objections 4 minor 28 references
Analysis of Multiple Long-Run Relations in Panel Data Models
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A pooled covariance matrix built from sub-sample time averages recovers the number and coefficients of multiple long-run relations in panels where the number of units far exceeds the time dimension, a regime no existing panel method covers.
desk verdict A genuinely new panel cointegration estimator for n >> T, with a strong common-cointegration assumption that needs testing; send to review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central device is the sub-sample time average. Each unit's $T$ observations are cut into $q\ge 2$ non-overlapping blocks of length $T/q$; the deviation of a block mean from the full-sample mean removes the fixed effect $a_i$, damps the stationary component to $O_p(T^{-1/2})$, and lets the $I(1)$ partial-sum component dominate at $O_p(T^{1/2})$. The exact moment $E(Q_{\bar s\bar s})=\frac{q-1}{6}(\frac{1}{q}+\frac{1}{T^2})\Sigma_i$ turns that dominance into the limit $\frac{q-1}{6q}\Psi_n$, handing the pooled matrix the rank defect of $\Psi_n$. The $r_0$ smallest eigenvectors of $Q_{\bar w\bar w}$ are the PME estimator of the relations $B_0$, the threshold rule $\tilde r=\sum_j \mathbb{1}(\tilde\lambda_j<T^{-\delta})$ on the correlation matrix estimates how many there are, and exact identifying restrictions rotate the eigenvectors into named economic coefficients with standard errors from a plug-in covariance estimator.
What would settle it
Simulate a panel with $n=500$ units and $T=50$ where half the units are generated to cointegrate along $w_1-w_2$ and half along $w_1-2w_2$, holding short-run dynamics identical across groups. This violates the common-space assumption, so the paper's theory predicts the pooled eigenvectors will mix the two relations and the threshold estimator will become unreliable; a clean recovery of either group's vector would show the assumption is unnecessary. A complementary check adds an $I(1)$ latent common factor, which Assumption 4 excludes: the relations that hold after removing the factor will be invisible to PME in levels, and the method should report no stationary relations.
Extended reading notes
Core claim
For the model $w_{it}=a_i+G_i f_t+C_i s_{it}+v_{it}$, with $s_{it}$ a partial sum of innovations and $v_{it}$ a stationary linear process, the pooled sub-sample covariance matrix satisfies $Q_{\bar w\bar w}=\frac{q-1}{6q}\Psi_n+O_p(n^{-1/2})+O_p(T^{-1})$, where $\Psi_n=n^{-1}\sum_i C_i\Sigma_i C_i'$. Since $\Psi_n$ has rank $m-r_0$, its null space is the space of common long-run relations $B_0$ with $B_0'C_i=0$, and that null space is consistently estimated by the $r_0$ smallest eigenvectors of $Q_{\bar w\bar w}$ as $n,T\to\infty$ jointly with $T\approx n^d$, $d>0$. The number $r_0$ is estimated by eigenvalue thresholding, $\tilde r=\sum_j \mathbb{1}(\tilde\lambda_j<T^{-\delta})$ applied to the correlation matrix, with $\delta=1/4$ recommended. Under $r_0^2$ exact identifying restrictions $R B_0=A$, the PME estimator is asymptotically normal at rate $\sqrt{nT}$ when $d>1/2$, with a covariance matrix estimator built from the same sub-sample quantities, so restrictions such as unit coefficients in financial ratios can be tested without estimating short-run dynamics.
Load-bearing premise
All units' random-walk components must span the same $(m-r_0)$-dimensional space, so the pooled matrix has exactly $r_0$ zero eigenvalues; if different units cointegrate along different vectors, the pooled eigenvectors blend relations that no individual unit actually obeys.
Editorial extensions
If this is right
- In panels with $n\gg T$, multiple long-run relations can be estimated without modelling short-run dynamics, choosing lag orders, or knowing the direction of long-run causality.
- Economic restrictions can be tested directly: the paper rejects unit long-run elasticities for most financial ratios it examines, while supporting a stationary short-term-debt-to-assets ratio.
- The number of relations is estimated rather than assumed, turning the procedure into a discovery tool; the macro application finds three relations among four variables, including a previously overlooked exports-productivity relation.
- The estimated relations are super-consistent and can feed second-stage error-correction models for adjustment speeds, forecasting, and counterfactual analysis.
Reading between the lines
- Remark 5's local-to-zero eigenvalue discussion suggests the threshold rule could be calibrated as a formal panel cointegration test: at eigenvalue rates $n^{-b}$ with $b\ge 1/2$ the estimator should lose power, giving a boundary that a designed test could exploit.
- PME's exact invariance to normalization lets it serve as a specification check for single-equation panel estimators: large disagreements, like the exports-productivity gap in the application, point to misspecified short-run dynamics or causality assumptions in the single-equation approach.
- The common-space assumption suggests a heterogeneity diagnostic within the method: applying PME to subgroups and comparing eigenvector estimates, as the authors do for advanced versus emerging economies, can reveal when pooling is invalid, and a formal test of eigenvector alignment would extend the procedure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a pooled minimum eigenvalue (PME) estimator for the number r0 and the coefficients B0 of common long-run relations in panels where the cross-section dimension n is large relative to the time dimension T. The method splits each unit's sample into q non-overlapping sub-samples, forms deviations of sub-sample time averages from the full-sample average, and uses the eigen-structure of the pooled covariance matrix Q_wbar_wbar of these deviations. Under the model w_it = a_i + G_i f_t + C_i s_it + v_it with stationary latent factors f_t, the paper shows that Q_wbar_wbar converges to ((q-1)/(6q))Ψ, so that the eigenvectors associated with the r0 smallest eigenvalues consistently estimate the common long-run relations B0 (Theorem 1), and that exactly identified coefficients are √(nT)-asymptotically normal when T ≈ n^d with d > 1/2 (Theorem 2). A thresholding rule based on eigenvalues of the sample correlation matrix is proposed to estimate r0. The paper reports extensive Monte Carlo evidence and two empirical applications, one to Compustat firm data and one to Penn World Table macro data.
Significance. If the results hold, the PME estimator fills a genuine gap: it addresses multiple long-run relations in panels with n >> T without estimating short-run dynamics or specifying the direction of long-run causality, under the relatively weak relative-rate condition d > 1/2. The asymptotic expansion in Proposition 1 is clean, the exact moment in Lemma 2 is analytically verified, and the Monte Carlo design is unusually broad, covering VAR and VARMA processes, interactive effects, non-Gaussian errors, and (in the supplement) GARCH and threshold autoregressions. The proposed estimator is also computationally trivial. However, the scope of the central claim is narrower than the title suggests: all identification rests on Assumption 3(a), which requires a common cointegrating space across all units, and the claimed robustness to interactive time effects holds only for stationary factors. These are not internal inconsistencies, but they are substantive scope restrictions that need to be stated prominently and, ideally, addressed with a diagnostic.
major comments (4)
- [Section 3, Assumption 3(a) (eq. (3)) and Section 6 (eq. (37))] The entire identification of r0 and B0 rests on the assumption that there exists a common (m-r0)-dimensional null space of Ψ_n, equivalently a common m×r0 matrix B0 with B0'C_i = 0 for all i. If units have unit-specific cointegrating vectors, rank(Ψ_n) = m generically, all population eigenvalues of Q_wbar_wbar are positive, and the thresholding estimator (37) selects r0 = 0 even though every unit cointegrates. This is not merely a technical condition: the paper's own macro application (Tables 11-13) finds β11 = -0.882 for advanced economies versus -1.005 for emerging economies and analyzes the two groups separately, which is precisely the heterogeneity that makes Assumption 3(a) fail under pooling. The paper should state the common-cointegrating-space requirement as a substantive scope condition in the introduction, provide a formal diagnostic or pre-test for Assumption 3(a), and discuss what PME estimates when the assumption fails, for example whether the estimator can be interpreted as recovering the intersection of unit-specific cointegrating spaces.
- [Section 5.3, eq. (36)] The variance estimator for vec(Θ_hat) appears to be missing a factor of T. From Corollary 1 (eqs. (33)-(34)), √(nT)vec(Θ_hat - Θ0) converges to N(0, Ω_θq) with Ω_θq = (6q/(q-1))^2 (I⊗Ψ22^{-1})Ω_q,22(I⊗Ψ22^{-1}). Since Q_wbar_wbar,22 → ((q-1)/(6q))Ψ22, the correct estimator is (1/(nT))Q22^{-1}Ω_hat_q,22Q22^{-1}, not (1/(nT^2))Q22^{-1}Ω_hat_q,22Q22^{-1}. The printed formula would deflate standard errors by a factor of √T and would make t-tests severely oversized; the near-nominal Monte Carlo sizes in Tables 3-6 indicate that the code used for the simulations must employ the 1/(nT) normalization. Please correct (36) and also check the corresponding formula in Supplement Section S2 for unbalanced panels.
- [Section 7, Assumption 4] The robustness claim for 'interactive time effects' holds only for covariance-stationary factors f_t. The paper explicitly excludes unit-root factors: 'we do not allow for the possibility of unit roots in latent factors' (lines following eq. (48)). In many macro applications, common stochastic trends are the source of cointegration; if f_t is I(1), the term Q_fi_fi in (49) is O_p(1) rather than O_p(T^{-2}), and the expansion (51) breaks down. The abstract and introduction should qualify the claim to 'stationary interactive time effects', and the paper should either extend the results to I(1) factors or explain clearly why such cases are outside the intended scope.
- [Section 6 vs Section 8.2] The thresholding consistency proof in Section 6 is developed for eigenvalues of Q_wbar_wbar (eq. (37)), while the implemented and recommended estimator (eq. (55)) thresholds eigenvalues of the correlation matrix R_wbar_wbar. The rank and null-space properties are preserved under the diagonal normalization in (54), but no formal argument is given for the ordered eigenvalues of R_wbar_wbar or for the threshold choice under this normalization. Since the Monte Carlo results in Tables 1-2 are for (55), the paper should add a formal proposition covering the correlation-matrix version, or explicitly state that (55) is a scale-normalized implementation justified by the same convergence applied to a transformed population matrix.
minor comments (4)
- [Section 10.1] The first sentence reads 'The fist application'; it should be 'The first application'.
- [Section 5.2, eq. (24)] The remainder term is printed as O_p(T^1); from the derivation and the proof in (A.11) it should be O_p(T^{-1}).
- [Tables 3-6 notes] The phrase 'Simulated power are computed' should be 'Simulated powers are computed'; the same grammatical issue appears in the notes to Tables 4-6.
- [Section 8.1] The conjecture q_T ≈ T^{1/3} is stated without formal support. Because the theorems treat q as fixed, the paper should clarify that the T^{1/3} rule is a practical recommendation rather than a result covered by the proofs.
Circularity Check
No significant circularity: PME consistency is proved against an external population object Ψ_n whose null space is the common cointegrating space by a rank assumption, not by construction.
full rationale
The paper's central derivation is self-contained and not circular. The target object B0 is defined economically by B0' C_i = 0 for all units, while the pooled matrix is Ψ_n = n^{-1} Σ C_i Σ_i C_i'. Because each Σ_i is positive definite, the null space of Ψ_n is exactly the set of vectors annihilating every C_i, i.e. span(B0) under Assumption 3(a). Thus the claim that the r0 smallest eigenvectors of Q_{w̄w̄} recover B0 is a consistency theorem against an external population limit, not a definitional identity: Proposition 1 and Theorem 1 prove Q_{w̄w̄} → ((q−1)/(6q))Ψ and then use the eigen-structure of Ψ, which is derived from C_i and Σ_i. The rank condition in Assumption 3 is an identification assumption, not a fitted input; the eigenvalue thresholding estimator for r0 relies on the genuine eigen-gap λ_{r0+1} > 0, which is implied by that rank condition. Turning to implementation, the choices C = 1 and δ = 1/4 are calibrated using the authors' own Monte Carlo experiments, but they are tuning parameters for the threshold rule and are not presented as out-of-sample predictions; nothing in the consistency or normality proofs depends on these specific values. The self-citations to Chudik, Pesaran, and Smith (2023a, 2023b) enter only as comparison estimators and as a Monte Carlo design, not as the load-bearing justification for the PME estimator. The principal fragility of the paper, Assumption 3(a)'s requirement of a common cointegrating space, is a substantive modeling assumption about the data generating process; the authors' decision to analyze advanced and emerging economies separately in the application is a response to that heterogeneity, but it does not make the estimator's target equal to its input. No equation or fitted parameter is renamed as a prediction, and no self-citation chain is invoked to force the central result. I therefore find no circular step that can be exhibited with a quote and a specific reduction.
Assumptions & free parameters
free parameters (4)
- q (number of sub-samples) =
2 recommended; 4 tested
- delta (threshold exponent) =
1/4 preferred; 1/2 alternative
- C (threshold constant) =
1 (with correlation-matrix scaling)
- T_ave for unbalanced threshold =
T_ave = n^{-1} sum T_i
assumptions (5)
- domain assumption Assumption 1: u_it independently distributed across i and t, zero mean, covariance Sigma_i with eigenvalues bounded away from zero and infinity, and finite 4+epsilon moments.
- domain assumption Assumption 2: C_i has rank m - r0 for all i, C*_i,ell decay exponentially in ell, and the innovation covariance components are uniformly bounded.
- domain assumption Assumption 3: Psi_n = n^{-1} sum C_i Sigma_i C_i' has rank m - r0 for all large n, with a strictly positive eigenvalue gap lambda_{r0+1} > 0.
- domain assumption Assumption 4: latent factors f_t are covariance stationary with no unit roots, and loadings G_i are uniformly bounded.
- standard math Model (1): Beveridge-Nelson-type decomposition w_it = a_i + G_i f_t + C_i s_it + v_it with block-diagonal C and C*(L) across units.
Cite this review
Pith. "Pith review of Analysis of Multiple Long-Run Relations in Panel Data Models." pith.science (2026). https://pith.science/paper/5CTL4ZRN
@misc{pith2026250602135,
author = {Pith},
title = {Pith review of: Analysis of Multiple Long-Run Relations in Panel Data Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/5CTL4ZRN}},
note = {Machine review of arXiv:2506.02135}
}
abstract
The literature on panel cointegration is extensive but does not cover data sets where the cross section dimension, $n$, is larger than the time series dimension $T$. This paper proposes a novel methodology that filters out the short run dynamics using sub-sample time averages as deviations from their full-sample counterpart, and estimates the number of long-run relations and their coefficients using eigenvalues and eigenvectors of the pooled covariance matrix of these sub-sample deviations. We refer to this procedure as pooled minimum eigenvalue (PME). We show that PME estimator is consistent and asymptotically normal as $n$ and $T \rightarrow \infty$ jointly, such that $T\approx n^{d}$, with $d>0$ for consistency and $d>1/2$ for asymptotic normality. Extensive Monte Carlo studies show that the number of long-run relations can be estimated with high precision, and the PME estimators have good size and power properties. The utility of our approach is illustrated by micro and macro applications using Compustat and Penn World Tables.
Figures
Reference graph
Works this paper leans on
-
[1]
Breitung, J. (2005). A parametric approach to the estimation of cointegration vectors in panel data https://doi.org/10.1081/etc-200067895. Econometric Reviews\/ 24 , 151--173
-
[2]
Breitung, J. and M. H. Pesaran (2008). Unit roots and cointegration in panels. In L. Matyas and P. Sevestre (Eds.), The Econometrics of Panel Data , Volume 46, pp.\ 279--322. Berlin, Heidelberg: Elsevier BV
work page 2008
-
[3]
Bykhovskaya, A. and V. Gorin (2022). Cointegration in large VARs . Annals of Statistics\/ 50 , 1593--1617
work page 2022
-
[4]
Choi, I. (2015, 01). Panel cointegration. In The Oxford Handbook of Panel Data . Oxford University Press
work page 2015
-
[5]
Chudik, A., M. H. Pesaran, and R. P. Smith (2023a). Pooled Bewley estimator of long-run relationships in dynamic heterogenous panels https://doi.org/10.1016/j.ecosta.2023.11.001. Econometrics and Statistics\/ . In Press, Corrected Proof. Available online 3 November 2023
-
[6]
Chudik, A., M. H. Pesaran, and R. P. Smith (2023b). Revisiting the Great Ratios Hypothesis https://doi.org/10.24149/gwp415. Oxford Bulletin of Economics and Statistics\/ 85 , 1023--1047
-
[7]
Coles, J. L. and Z. F. Li (2023). An Empirical Assessment of Empirical Corporate Finance https://doi.org/10.1017/S0022109022000448. Journal of financial and quantitative analysis\/ 58\/ (4), 1391--1430
-
[8]
Davidson, J. (1994). Stochastic Limit Theory . Oxford University Press
work page 1994
Show all 28 references
-
[9]
Hajda, E
Geelen, T., J. Hajda, E. Morellec, and A. Winegar (2024). Asset life, leverage, and debt maturity matching https://doi.org/10.1016/j.jfineco.2024.103796. Journal of Financial Economics\/ 154 , 103796
2024
-
[10]
Groen, J. J. J. and F. Kleibergen (2003). Likelihood-Based Cointegration Analysis in Panels of Vector Error-Correction Models https://doi.org/10.1198/073500103288618972. Journal of Business and Economic Statistics\/ 21 , 295--318
2003 doi
-
[11]
Hamilton, J. D. (1994). Time Series Analysis . Princeton University Press
1994
-
[12]
Im, K. S., M. Pesaran, and Y. Shin (2003). Testing for unit roots in heterogeneous panels https://doi.org/10.1016/S0304-4076(03)00092-7. Journal of Econometrics\/ 115 , 53--74
2003 doi
-
[13]
Johansen, S. (1988). Statistical analysis of cointegration vectors https://doi.org/10.1016/0165-1889(88)90041-3. Journal of Economic Dynamics and Control\/ 12\/ (2), 231--254
1988 doi
-
[14]
Johansen, S. (1991). Estimation and hypothesis testing of cointegration vectors in gaussian vector autoregressive models. Econometrica\/ 59 , 1551
1991
-
[15]
Johansen, S. (1995). Likelihood Based Inference on Cointegration in the Vector Autoregressive Model . Oxford: Oxford University Press
1995
-
[16]
Larsson, R. and J. Lyhagen (2007). Inference in panel cointegration models with long panels https://doi.org/10.1198/073500106000000549. Journal of Business and Economic Statistics\/ 25 , 473--483
2007 doi
-
[17]
Lev, B. and S. Sunder (1979). Methodological issues in the use of financial ratios https://doi.org/10.1016/0165-4101(79)90007-7. Journal of Accounting and Economics\/ 1 , 187--210
1979 doi
-
[18]
Mark, N. C. and D. Sul (2003). Cointegration vector estimation by panel DOLS and long-run money demand https://doi.org/10.1111/j.1468-0084.2003.00066.x. Oxford Bulletin of Economics and Statistics\/ 65\/ (5), 655--680
2003
-
[19]
M \"u ller, U. K. and M. W. Watson (2018). Long-run covariability https://doi.org/10.3982/ECTA15047. Econometrica\/ 86 , 775--804
2018 doi
-
[20]
Onatski, A. and C. Wang (2018). Alternative asymptotics for cointegration tests in large VARs . Econometrica\/ 86 , 1465--1478
2018
-
[21]
Onatski, A. and C. Wang (2019). Extreme canonical correlations and high-dimensional cointegration analysis. Journal of Econometrics\/ 212 , 307--322
2019
-
[22]
Pedroni, P. (1996). Fully Modified OLS for Heterogeneous Cointegrated Panels and the Case of Purchasing Power Parity https://web.williams.edu/Economics/pedroni/WP-96-20.pdf. Indiana University working papers in economics no. 96-020 (June 1996)
1996
-
[23]
Pedroni, P. (2001a). Fully modified OLS for heterogeneous cointegrated panels https://doi.org/10.1016/S0731-9053(00)15004-2. In B. Baltagi, T. Fomby, and R. C. Hill (Eds.), Nonstationary Panels, Panel Cointegration, and Dynamic Panels, (Advances in Econometrics, Vol. 15) , pp....
2001 doi
-
[24]
Pedroni, P. (2001b). Purchasing Power Parity Tests in Cointegrated Panels https://doi.org/10.1162/003465301753237803. The Review of Economics and Statistics\/ 83 , 727--731
2001 doi
-
[25]
Pesaran, M. H., Y. Shin, and R. P. Smith (1999). Pooled mean group estimation of dynamic heterogeneous panels https://doi.org/10.1080/01621459.1999.10474156. Journal of the American Statistical Association\/ 94 , 621--634
1999
-
[26]
Petrov, V. V. (1992). Moments of Sums of Independent Random Variables https://doi.org/10.1007/BF01362802. Journal of Soviet Mathematics\/ 61 , 1905--1906
1992 doi
-
[27]
Phillips, P. C. B. and S. Ouliaris (1988). Testing for cointegration using principal component methods https://doi.org/10.1016/0165-1889(88)90040-1. Journal of Economic Dynamics and Control\/ 12 , 205--230
1988 doi
-
[28]
Phillips, P. C. B. and S. Ouliaris (1990). Asymptotic Properties of Residual Based Tests for Cointegration http://www.jstor.org/stable/2938339. Econometrica\/ 58 , 165--193
1990
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.