Pith. sign in

REVIEW 5 major objections 5 minor 36 references

High-Dimensional Binary Variates: Maximum Likelihood Estimation with Nonstationary Covariates and Factors

T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper establishes the first maximum-likelihood asymptotics for high-dimensional binary factor models with integrated covariates and factors, showing that coefficient estimators converge at two different rates depending on direction.

desk verdict Novel extension of nonstationary binary choice to factor models, but Assumption 4 conflicts with the model's common-factor structure and the proofs are all deferred; the factor CLT is likely mixed normal. read the letter →

arxiv 2505.22417 v1 pith:HGGJKY6R submitted 2025-05-28 math.ST stat.TH

classification math.STstat.TH MSC 62F1262H2562M1062P05
keywords BinaryresponsemodelsNonstationaryfactorsIntegratedcovariatesMaximumlikelihoodestimationFactorCointegrationJumparrivalsHigh-dimensionalpaneldata
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper supplies asymptotic theory for maximum likelihood estimation in a high-dimensional binary response factor model in which both the covariates $x_{it}$ and the common factors $f_{0t}$ are integrated of order one. In the nonstationary case, where the single index $z_{0it}=\beta'_{0i}x_{it}+\lambda'_{0i}f_{0t}$ is itself $I(1)$, the MLE of the combined coefficient vector $\alpha_{0i}=(\beta'_{0i},\lambda'_{0i})'$ converges at rate $T^{1/4}$ along the direction $\alpha_{0i}/\|\alpha_{0i}\|$ and at the faster rate $T^{3/4}$ in orthogonal directions, while the factor estimators converge at rate $t^{-\delta/2}N^{1/2}$ with $\delta\in(1/4,3/4)$ determined by the link function. If the single index is cointegrated, those rates improve to $T^{1/2}$ and $T$ for the coefficients and to $N^{1/2}$ for the factors. A sympathetic reader would care because existing high-dimensional binary factor asymptotics assume stationarity, and this is the first result of its kind for nonstationary binary factor models, with an application to jump-arrival factors in asset pricing.

What carries the argument

The carrying object is the coordinate rotation $Q_i=(Q_i^{(1)},Q_i^{(2)})$ with $Q_i^{(1)}=\alpha_{0i}/\|\alpha_{0i}\|$, which splits every estimation problem into a scale component along the true direction and a directional component orthogonal to it. Working in this rotated system, the score and Hessian separate into diagonal blocks in $i$ and in $t$, and the squared score kernel $K=M^{2}\Psi(1-\Psi)$ (for logit, $e^{x}/(1+e^{x})^{2}$; for probit, $\varphi^{2}/(\Phi(1-\Phi))$) controls everything: its integrability makes most of the Hessian vanish, and its expected decay $\mathbb{E}[K(z)]\asymp\sigma^{-2\delta}$ sets the factor rate exponent $\delta$. The local time of the Brownian motion driving the index enters the limiting covariance matrices, which is why the nonstationary limits are mixture normals.

What would settle it

A direct simulation can settle the claim: generate logit or probit panels with $I(1)$ covariates and factors, estimate the MLE, and measure the error of $\hat{\alpha}_i$ in the direction of $\alpha_{0i}$ and in orthogonal directions as $T$ grows with $N$ fixed. If the orthogonal error does not shrink like $T^{-3/4}$ while the parallel error shrinks like $T^{-1/4}$, or if the factor error does not grow like $t^{\delta/2}/\sqrt{N}$ for the simulated $\delta$, the dual-rate theory fails; the consistency condition $T^{1/2}/N\to 0$ can be checked in the same design by fixing $N$ and increasing $T$.

Watch

Extended reading notes

Core claim

The central claim is that the maximum likelihood estimator in the single-index general factor model $y_{it}=\Psi(\beta'_{0i}x_{it}+\lambda'_{0i}f_{0t})+u_{it}$ has a directional rate structure: after an orthogonal rotation placing $\alpha_{0i}/\|\alpha_{0i}\|$ on the first axis, the estimation error of $\hat{\alpha}_i$ decays at $T^{1/4}$ along that axis and at $T^{3/4}$ along the remaining axes in the nonstationary case, and the normalized direction estimator $\hat{\alpha}_i/\|\hat{\alpha}_i\|$ inherits the faster $T^{3/4}$ rate. The factor estimator $\hat{f}_t$ is consistent when $T^{\delta}/N\to 0$, with $\hat{f}_t-f_{0t}$ of order $t^{\delta/2}/\sqrt{N}$; the time dependence arises because the Hessian block for the factors drifts toward zero as $t$ grows. In the cointegrated single-index case the coefficient rates increase to $(T^{1/2},T)$, factors recover the conventional $N^{1/2}$ rate, and the limiting distributions stop depending on $t$. The paper also derives mixture-normal limits driven by the local time of a Brownian motion, plug-in covariance estimators, a local-time estimator, and a rank-based rule for selecting the number of factors.

Load-bearing premise

The whole rate picture rests on one decay assumption: the average curvature contributed by the link function must fall like the inverse square of the spread of the index, with an exponent strictly between 1/4 and 3/4, and that exponent must be the same for every unit and every time; if the decay is faster, slower, or heterogeneous, the stated rates and the required sample sizes no longer follow.

Editorial extensions

If this is right

  • Collective consistency of the coefficients in the nonstationary case requires $T^{1/2}/N\to 0$, so the cross-section must grow faster than the square root of the time horizon.
  • Factor estimators are consistent only when $T^{\delta}/N\to 0$ with $\delta\in(1/4,3/4)$; because the Hessian for later time points approaches zero, accurate factor estimation at large $t$ demands a larger cross-section.
  • In the cointegrated single-index case both coefficients and factors converge at $\min(\sqrt{N},\sqrt{T})$, with coefficient direction estimates improving to rate $T$, and the factor limiting distributions no longer depend on $t$.
  • Normalizing the coefficient estimator to the unit sphere, $\hat{\alpha}_i/\|\hat{\alpha}_i\|$, yields a faster rate than the unnormalized estimator, so estimating the direction of $\alpha_{0i}$ is more accurate than estimating its magnitude.
  • In the empirical application, the estimated jump-arrival factors are nonstationary, and adding them to the Fama–French–Carhart five-factor model raises explained variation by roughly 30 percent.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the time-dependent factor rate implies a simple diagnostic — split the panel into early and late time blocks and compare factor estimate dispersions; the late block should show variance inflated by a factor of order $t^{\delta}$, providing a direct empirical check of the Hessian-to-zero mechanism.
  • Beyond the paper: because $\delta$ is simulated near $0.5$ for logit and probit, one could estimate $\delta$ from the plug-in Hessian and test whether Assumption 4(i)'s homogeneous-exponent requirement holds in a given dataset.
  • Beyond the paper: the sharp contrast between the nonstationary and cointegrated rates suggests practitioners should pre-test the single index for cointegration, since the cointegrated specification, when warranted, gives strictly faster estimation.
  • Beyond the paper: the rank-threshold rule for selecting the number of factors presumes strong factors; a natural extension would be to ask how the rule distorts under weak or locally vanishing loadings, which the current proofs do not address.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper studies maximum likelihood estimation in a high-dimensional binary panel model with integrated (I(1)) covariates and factors, allowing for either nonstationary or cointegrated single indices. The main theoretical claims are: in the nonstationary case, the coefficient estimator for each cross-sectional unit has dual convergence rates (T^{1/4} along the direction of the true coefficient vector and T^{3/4} orthogonal to it), the factors converge at rate t^{-\delta/2}N^{1/2} with a time-dependent limiting distribution, and collective consistency holds under T^{1/2}/N \to 0; in the cointegrated case the rates improve to (T^{1/2}, T) for coefficients and N^{1/2} for factors. The paper also proposes a factor-number selection rule, provides simulation evidence, and applies the model to daily jump arrivals for S&P 500 constituents. All proofs are deferred to a supplementary file that is not included in the submission.

Significance. If the stated results are correct, this would be the first systematic asymptotic theory for maximum likelihood estimation of high-dimensional binary factor models with nonstationary covariates and factors, extending the univariate theory of Park and Phillips (2000) to a panel/factor setting. The distinction between nonstationary and cointegrated single indices, the dual convergence rates, and the time-dependent factor rates are novel and of potential interest to econometric theory. The empirical application to jump arrivals is concrete and suggests practical relevance. However, the significance is conditional on the high-level assumptions being compatible with the paper's own data-generating process, and on the deferred proofs establishing the stated distributional claims. The absence of the supplementary proofs and the issues raised below mean that the central theoretical claims are currently not fully verifiable.

major comments (5)
  1. [Section 3.1, Assumption 4(ii)-(iii), Theorem 3.2, Remark 1] The factor CLT in Theorem 3.2 assumes that the variance of t^\delta/N \sum_i K(z0it) vanishes and that the normalized score has a Gaussian limit with deterministic covariance \Omega_{f,t}. For the paper's own DGP (Section 4.1, Case 1), z0it = \beta_{0i}'x_{it} + \lambda_{0i}'f0t contains the common I(1) factor, so for t proportional to T the cross-sectional average of K(z0it), after t^\delta scaling with \delta=1/2, converges to a non-degenerate functional of f0t/\sqrt{t}; its variance does not vanish as N grows. This is acknowledged in Remark 1, which states that the properly scaled factor Hessian converges to a stochastic limit matrix. The consequence is that the limiting law in Theorem 3.2 should be mixed normal (or the covariance should be defined conditionally on f0t), and the inversion argument in Assumption 5 requires a different normalization. The stated convergence rates may survive, but the distributional claim and the exact covariance formula are not justified as written.
  2. [Section 3.1, Assumption 4(v)] Assumption 4(v) states that for functions f in Assumptions 3(i)-(ii), (1/\sqrt{T})\sum_{t=1}^T t^\delta f(h_{0it}^{(1)}) = O_P(1), where h_{0it}^{(1)} = z0it / \|\alpha_{0i}\| is an I(1) process under Assumption 1. For a bounded integrable f with nonzero integral, the local-time approximation gives (1/\sqrt{T})\sum_{t=1}^T t^\delta f(h_{0it}^{(1)}) \approx T^\delta L_{H_{1i}}(1,0)\int f, which diverges for \delta>0. Thus Assumption 4(v), as stated, cannot hold for standard choices such as f=K^2, and it cannot be used in the proof of Theorem 3.1 without correction. The normalization in Assumption 4(v) needs to be revised or the assumption replaced by a condition that is compatible with the I(1) behavior of h_{0it}^{(1)}.
  3. [Section 3.3, Assumption 6(v)] Assumption 6(v) imposes E\|g0it\|^{4+\nu}<\infty and \|g0it\|\le C for all i,t. This is incompatible with Assumption 1, under which g0it is a (q+r)-dimensional I(1) process; in the cointegrated case the rotated component h_{0it}^{(2)}=Q_i^{(2)'}g0it is still I(1) and is unbounded. Theorems 3.5-3.7 rely on Assumption 6, so this inconsistency needs to be resolved, for example by assuming boundedness of the stationary single index z0it only and replacing the moment conditions on g0it with conditions on the stationary component h_{0it}^{(1)}.
  4. [Section 3.1, Assumption 4(iv), and Abstract] The abstract's headline sufficient condition T^{1/2}/N \to 0 for collective consistency is not the condition used by Theorem 3.1, because Assumption 4(iv) requires T^{\delta'}/N = o(1) with \delta' \ge \max(\delta, \delta/2+1/2). Since \delta > 1/4, this imposes \delta' \ge 5/8 (and \delta' \ge 3/4 when \delta=1/2, as for logit/probit), which is strictly stronger than T^{1/2}/N. The theorem, the assumptions, and the abstract must be aligned, or the proof must show that the \delta'-condition is not needed for the coefficient convergence rates.
  5. [Supplementary Material] All proofs of Theorems 3.1-3.7, Corollaries 3.1-3.4, and Proposition 1 are deferred to a Supplementary Material file that is not included in the submission. Since the paper's contribution is almost entirely asymptotic theory, the referee cannot verify the central claims. The supplement (or a complete proof appendix) should be provided for review.
minor comments (5)
  1. [Section 1] The sentence 'distinct from jump size factors (e.g., ?; Pelger 2020)' contains an unresolved citation placeholder '?' that must be fixed.
  2. [Table 1] Table 1 contains formatting and OCR artifacts (e.g., '1.924-', '201 5', '0.0 010'); the table should be carefully proofread.
  3. [Section 2.2, Step 4] The normalization step gives no tie-breaking rule when the diagonal matrix D* has repeated eigenvalues; a sentence on uniqueness would help.
  4. [Section 3.1, Corollary 3.3] The notation C_{NT,t} is used in Corollary 3.3 before it is formally defined; define it explicitly at first use.
  5. [Section 3.3, Assumption 6(v)] There is a typo in 'max_{i\ge,t\ge1}' that should read 'max_{i,t}'; more importantly, the moment condition should be stated for the stationary component, not for the full I(1) vector.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the rates follow from primitive link-function and Hessian assumptions, not from self-citation or fitted targets.

full rationale

I walked the claimed derivation chain. The main rates (T^{1/4}, T^{3/4} for alpha_i, and t^{-delta/2} N^{1/2} for f_t) are obtained from Taylor expansions of the score equations (eqs. 5-6), local-time asymptotics for nonlinear functions of integrated processes (Assumption 3, based on the external Park-Phillips theory), and the Hessian-decay condition Assumption 4(i), which specifies E(K(z)) ~ sigma^{-2 delta}. That assumption is a primitive condition on the link kernel, not an assumption of the estimator's rate; the factor rate inherits delta from it, and Assumption 4(iv) then yields the sufficient condition T^delta/N -> 0. This is a transparent assumption-to-result derivation, not a fitted parameter renamed as a prediction. There is no load-bearing self-citation chain: the cited works (Park and Phillips, Bai, Chen et al., Gao et al.) are external, and this paper extends rather than renames their results. The rotated coordinate system is explicitly a proof device: 'the rotation serves only as a tool for deriving the asymptotic theory for the proposed estimators.' Two non-circular concerns were noted. First, Remark 1 states that the normalized factor Hessian 'converges weakly to a stochastic limit matrix,' which is in tension with the deterministic Omega_{f,t} in Theorem 3.2; this is a correctness and assumptions-verification risk about Assumption 4(ii)-(iii), not an equality of input and output. Second, the manuscript contains a missing citation marker ('e.g., ?; Pelger 2020'), an editorial completeness issue. Neither constitutes circularity, so no circular step is reported.

Assumptions & free parameters 2 free parameters · 7 assumptions · 0 invented entities

The central theory rests on a chain of high-level assumptions. The most fragile is Assumption 4(i), whose exponent delta controls the factor rate and consistency condition. No new particles or entities are introduced.

free parameters (2)
  • delta (decay exponent for link function K) = Not fixed; simulations suggest roughly 0.5 for logit and probit (Supplementary Material)
    Assumption 4(i) assumes E(K(z)) is of order sigma^{-2 delta}; the factor consistency condition and factor rate depend directly on delta, so the theory is parameterized by an unestimated exponent.
  • pi_NT (factor number threshold) = In simulations, pi = hat_sigma_{N,1} (C_NT^2 T^{-1/2})^{-1/3}; must satisfy pi -> 0 and pi C_NT^2 T^{-1/2} -> infinity
    Threshold for selecting the number of factors in Theorem 3.4; no data-driven rule is derived, only a consistency condition.
assumptions (7)
  • domain assumption Assumption 1: covariates and factors are I(1) linear processes with nonsingular long-run matrices, finite 8+ moments, and functional CLT to Brownian motion H_i(s).
    Standard nonstationary regression setup, Section 2.1, needed for local time asymptotics.
  • domain assumption Assumption 2: loadings are bounded, factors satisfy FCLT with positive definite covariance (no cointegration among factor components), and cross-sectional loading matrix converges to a diagonal matrix with distinct positive limits.
    Strong factor identification and factor ordering, Section 2.1; needed for rotational identification under F'F/T^2=I and Lambda'Lambda/N diagonal.
  • domain assumption Assumption 3: link function is three times differentiable with K^2 regular and a list of integrability and boundedness conditions on M and K.
    Imposes enough smoothness and integrability for the score and Hessian expansions in Section 3.1; satisfied by logit and probit.
  • ad hoc to paper Assumption 4(i): E(K(z)) is comparable to sigma^{-2 delta} for delta in (1/4,3/4), and related moments decay at the same rate.
    This decay law is the load-bearing premise for the t-dependent factor convergence rate; delta is not derived from the link function in the main text.
  • ad hoc to paper Assumption 4(ii)-(v): variance decay of cross-sectional averages, CLT for scores, rate conditions T^{delta'}/N=o(1) and N^{delta''}/T=o(1), and a moment condition for t^{delta} f(h_{0it}).
    These are custom high-level conditions that abstract away the difficult parts of the proof; no derivation from primitive assumptions is provided in the main text.
  • domain assumption Assumption 5: normalized Hessian blocks A11, A22, A12 have bounded eigenvalues and bounded inverse eigenvalues.
    Used to invert the Hessian in the Taylor expansion; stated as mild but not verified from primitives.
  • domain assumption Assumption 6 (cointegrated case): single index lies in a bounded interval, error distribution has bounded tails away from 0 and 1, rank condition on Pi(1), alpha-mixing of z0it, and bounded g0it.
    Replaces the local-time framework with stationary-index mixing asymptotics in Section 3.3.

how reviews work

0 comments
Cite this review

Pith. "Pith review of High-Dimensional Binary Variates: Maximum Likelihood Estimation with Nonstationary Covariates and Factors." pith.science (2026). https://pith.science/paper/HGGJKY6R

@misc{pith2026250522417,
  author       = {Pith},
  title        = {Pith review of: High-Dimensional Binary Variates: Maximum Likelihood Estimation with Nonstationary Covariates and Factors},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HGGJKY6R}},
  note         = {Machine review of arXiv:2505.22417}
}
abstract

This paper introduces a high-dimensional binary variate model that accommodates nonstationary covariates and factors, and studies their asymptotic theory. This framework encompasses scenarios where single indices are nonstationary or cointegrated. For nonstationary single indices, the maximum likelihood estimator (MLE) of the coefficients has dual convergence rates and is collectively consistent under the condition $T^{1/2}/N\to0$, as both the cross-sectional dimension $N$ and the time horizon $T$ approach infinity. The MLE of all nonstationary factors is consistent when $T^{\delta}/N\to0$, where $\delta$ depends on the link function. The limiting distributions of the factors depend on time $t$, governed by the convergence of the Hessian matrix to zero. In the case of cointegrated single indices, the MLEs of both factors and coefficients converge at a higher rate of $\min(\sqrt{N},\sqrt{T})$. A distinct feature compared to nonstationary single indices is that the dual rate of convergence of the coefficients increases from $(T^{1/4},T^{3/4})$ to $(T^{1/2},T)$. Moreover, the limiting distributions of the factors do not depend on $t$ in the cointegrated case. Monte Carlo simulations verify the accuracy of the estimates. In an empirical application, we analyze jump arrivals in financial markets using this model, extract jump arrival factors, and demonstrate their efficacy in large-cross-section asset pricing.

Figures

Figures reproduced from arXiv: 2505.22417 by the authors.

Figure 1
Figure 1. Estimated number of factors for each year [PITH_FULL_IMAGE:figures/full_fig_p022_1.png] view at source ↗
Figure 2
Figure 2. Box plots of factor loadings across different industries. Notes. The three graphs correspond to the three sets of factor loadings, with the horizontal axis representing various industries. The red “+” symbols indicate outliers, and the plots display the confidence interval gaps [PITH_FULL_IMAGE:figures/full_fig_p023_2.png] view at source ↗
Figure 3
Figure 3. Estimated factors and their corresponding first-order differences. Note. The top panel displays the three estimated factors, while the bottom panel shows the corresponding first-order difference series. The top panel of [PITH_FULL_IMAGE:figures/full_fig_p023_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Canonical correlations and asset pricing results. Notes. The left panel displays the canonical correlation coefficients between the four jump arrival factors and the Fama-French-Carhart five factors. The middle panel shows the incremental R 2 from adding jump arrival f…
Figure 5
Figure 5. Figure 5: Time-variation in the percentage of explained variation for different factors. Note. This figure plots the percentage of explained variation calculated on a moving window of one year (252 trading days). To further demonstrate that the jump arrival factors have incremen…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

36 extracted references · 32 canonical work pages

  1. [1]

    Andersen, T. G., D. Dobrev, and E. Schaumburg (2012). Jump-robust volatility estimation using nearest neighbor truncation. Journal of Econometrics\/ 169\/ (1), 75--93

  2. [2]

    Bai, and K

    Ando, T., J. Bai, and K. Li (2022). Bayesian and maximum likelihood analysis of large-scale panel choice models with unobserved heterogeneity. Journal of Econometrics\/ 230\/ (1), 20--38

  3. [3]

    Bai, J. (2003). Inferential theory for factor models of large dimensions. Econometrica\/ 71\/ (1), 135--171

  4. [4]

    Bai, J. (2004). Estimating cross-section common stochastic trends in nonstationary panel data. Journal of Econometrics\/ 122\/ (1), 137--183

  5. [5]

    Bai, J. and K. Li (2012). Statistical analysis of factor models of high dimension. The Annals of Statistics\/ 40\/ (1), 436--465

  6. [6]

    Cavaliere, and L

    Barigozzi, M., G. Cavaliere, and L. Trapani (2024). Inference in heavy-tailed nonstationary multivariate time series. Journal of the American Statistical Association\/ 119\/ (545), 565--581

  7. [7]

    Bollerslev, T. and V. Todorov (2011a). Estimation of jump tails. Econometrica\/ 79\/ (6), 1727--1783

  8. [8]

    Bollerslev, T. and V. Todorov (2011b). Tails, fears, and risk premia. The Journal of Finance\/ 66\/ (6), 2165--2211

Show all 36 references
  1. [9]

    Chamberlain, G. and M. Rothschild (1983). Arbitrage, factor structure in arbitrage pricing models. Econometrica\/ 51\/ (5), 1281--1304

  2. [10]

    Chen, D., P. A. Mykland, and L. Zhang (2024). Realized regression with asynchronous and noisy high frequency and high dimensional data. Journal of Econometrics\/ 239\/ (2), 105446

  3. [11]

    Chen, L., J. J. Dolado, and J. Gonzalo (2021). Quantile factor models. Econometrica\/ 89\/ (2), 875--910

  4. [12]

    Fern \'a ndez-Val, and M

    Chen, M., I. Fern \'a ndez-Val, and M. Weidner (2021). Nonlinear factor models for network and panel data. Journal of Econometrics\/ 220\/ (2), 296--324

  5. [13]

    Gao, and B

    Dong, C., J. Gao, and B. Peng (2021). Varying-coefficient panel data models with nonstationarity and partially observed factor structure. Journal of Business & Economic Statistics\/ 39\/ (3), 700--711

  6. [14]

    Gao, and D

    Dong, C., J. Gao, and D. B. Tjostheim (2016). Estimation for single-index and partially linear single-index integrated models. Annals of Statistics\/ 44\/ (1), 425--453

  7. [15]

    Erdemlioglu, M

    Dungey, M., D. Erdemlioglu, M. Matei, and X. Yang (2018). Testing for mutually exciting jumps and financial flights in high frequency data. Journal of Econometrics\/ 202\/ (1), 18--44

  8. [16]

    Fama, E. F. and K. R. French (2015). A five-factor asset pricing model. Journal of Financial Economics\/ 116\/ (1), 1--22

  9. [17]

    Fama, E. F. and J. D. MacBeth (1973). Risk, return, and equilibrium: Empirical tests. Journal of Political Economy\/ 81\/ (3), 607--636

  10. [18]

    Liao, and M

    Fan, J., Y. Liao, and M. Mincheva (2013). Large covariance estimation by thresholding principal orthogonal complements. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 75\/ (4), 603--680

  11. [19]

    Gao, J., F. Liu, B. Peng, and Y. Yan (2023). Binary response models for heterogeneous panel data with interactive fixed effects. Journal of Econometrics\/ 235\/ (2), 1654--1679

  12. [20]

    Hansen, B. E. (1992). Convergence to stochastic integrals for dependent heterogeneous processes. Econometric Theory\/ 8\/ (4), 489--500

  13. [21]

    He, Y., Y. Hou, H. Liu, and Y. Wang (2024). Generalized principal component analysis for large-dimensional matrix factor model. arXiv preprint arXiv:2411.06423\/

  14. [22]

    He, Y., L. Li, D. Liu, and W.-X. Zhou (2025). Huber principal component analysis for large-dimensional factor models. Journal of Econometrics\/ 249 , 105993

  15. [23]

    Ma, T. F., F. Wang, and J. Zhu (2023). On generalized latent factor modeling and inference for high-dimensional binomial data. Biometrics\/ 79\/ (3), 2311--2320

  16. [24]

    Park, J. Y. and P. C. Phillips (1999). Asymptotics for nonlinear transformations of integrated time series. Econometric Theory\/ 15\/ (3), 269--298

  17. [25]

    Park, J. Y. and P. C. Phillips (2000). Nonstationary binary choice. Econometrica\/ 68\/ (5), 1249--1280

  18. [26]

    Park, J. Y. and P. C. Phillips (2001). Nonlinear regressions with integrated time series. Econometrica\/ 69\/ (1), 117--161

  19. [27]

    Pelger, M. (2019). Large-dimensional factor modeling based on high-frequency observations. Journal of Econometrics\/ 208\/ (1), 23--42

  20. [28]

    Pelger, M. (2020). Understanding systematic risk: A high-frequency approach. The Journal of Finance\/ 75\/ (4), 2179--2220

  21. [29]

    Trapani, L. (2018). A randomized sequential procedure to determine the number of factors. Journal of the American Statistical Association\/ 113\/ (523), 1341--1349

  22. [30]

    Trapani, L. (2021). Inferential theory for heterogeneity and cointegration in large panels. Journal of Econometrics\/ 220\/ (2), 474--503

  23. [31]

    Wang, F. (2022). Maximum likelihood estimation and inference for high dimensional generalized factor models with application to factor-augmented regressions. Journal of Econometrics\/ 229\/ (1), 180--200

  24. [32]

    Yuan, and J

    Xu, S., C. Yuan, and J. Guo (2025). Quasi maximum likelihood estimation for large-dimensional matrix factor models. Journal of Business & Economic Statistics\/ 43\/ (2), 439--453

  25. [33]

    Zhao, and W

    Yu, L., P. Zhao, and W. Zhou (2024). Testing the number of common factors by bootstrapped sample covariance matrix in high-dimensional factor models. Journal of the American Statistical Association\/ , 1--12

  26. [34]

    Yuan, C., Z. Gao, X. He, W. Huang, and J. Guo (2023). Two-way dynamic factor models for high-dimensional matrix-valued time series. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 85\/ (5), 1517--1537

  27. [35]

    Pan, and J

    Zhang, B., G. Pan, and J. Gao (2018). Clt for largest eigenvalues and unit root testing for high-dimensional nonstationary time series. The Annals of Statistics\/ 46\/ (5), 2186--2215

  28. [36]

    Zhou, W., J. Gao, D. Harris, and H. Kew (2024). Semi-parametric single-index predictive regression models with cointegrated regressors. Journal of Econometrics\/ 238\/ (1), 105577

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.