{"id":"93d34941-3d60-4771-9af9-e7a704229006","arxiv_id":"1908.01252","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Diversified weighted averages of panel data estimate latent factors up to an affine transformation that is consistent for factor spans even when the working number of factors R exceeds the true r, with T possibly finite.","lead":"This paper builds factor estimates by averaging many variables with chosen weights, and proves these estimates remain valid even when too many factors are used, time series are short, or factors are weak. A reader should care because it removes a fragile preprocessing step, choosing the number of factors, from many factor-model applications.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proof of Theorem 2.1 relies on Proposition A.1, which requires T → ∞; when T < R, (1/T)hat(F)'hat(F) is singular, so the advertised finite-T robustness is not established by the given argument.","rationale":"The reader identified Assumption 2.2 (rank of W'B) as the weakest assumption. I agree that this rank condition is practically hard to verify, but it is an explicit assumption of a conditional theorem; the theorem's contribution is the consequence given that condition. A more direct threat to the central claim is the mismatch between Theorem 2.1's finite-T statement and the proof's reliance on Proposition A.1, which is proved only for T → ∞. Moreover, part of the proof mechanism — invertibility of (1/T)hat(F)'hat(F) — is plainly false when T < R, a regime explicitly allowed by the theorem. This is not a subtle asymptotic edge case: it concerns the definition of P_hat(F) itself in the claimed scope. I therefore regard this as the single most load-bearing concern. I still recommend keeping the reader's CONDITIONAL verdict because the projection-error bound might be true and provable via generalized inverses or by adding the mild condition T ≥ R; the paper would then need a corrected proof or a modified statement. The concern is not that the main idea is wrong, but that the finite-T robustness is currently unsupported by the presented proof.","tokens_in":43372,"tokens_out":7927,"duration_ms":76035,"concrete_test":"Set T=2, r=1, R=3, N=1000, with f_t iid N(0,1), B iid N(0,1), u_t iid N(0,1), and W a deterministic matrix satisfying Assumptions 2.1-2.2 (e.g., first column ones, second sign-alternating, third orthogonal). By construction rank(hat(F)) ≤ T < R, so (1/T)hat(F)'hat(F) is singular for every realization; verify its minimum eigenvalue is 0, contradicting Proposition A.1(i) as invoked in the proof of Theorem 2.1. Then compute the actual projection error ||P_hat(F)P_F - P_F|| on the same simulated data; if it is already near zero at this N, the finite-T claim may be salvaged by a different argument, but the current proof does not establish it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 2.1 is stated for 'T is either finite or grows', but its proof invokes Proposition A.1, which is explicitly proved under 'T, N → ∞'. More concretely, Proposition A.1(i) asserts that (1/T)hat(F)'hat(F) has minimum eigenvalue at least c/N with probability approaching one; this is impossible when T < R because hat(F) is T × R, so rank((1/T)hat(F)'hat(F)) = rank(hat(F)) ≤ T < R, forcing the smallest eigenvalue to be identically zero. The same problem affects M'hat(F)'hat(F)M (an r × r matrix) when T < r. Thus the proof's claim that P_hat(F) and P_hat(F)M are well defined via invertible gram matrices fails exactly in the finite-T regime the paper advertises (Abstract; Section 2.4 item 4; Remark after Theorem 2.1). The projection-error bound may still hold in some finite-T cases (e.g., T=1, r=1, R=2 gives span(hat(F)) = span(F) trivially), so this is a proof gap rather than a demonstrated falsehood, but it is load-bearing because finite-T robustness is a headline selling point and the reader's strongest_claim explicitly includes 'T finite or growing'.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes estimating latent factors in an approximate factor model by cross-sectional projections of the panel onto a user-chosen weight matrix W, yielding \\hat f_t = N^{-1} W' x_t. The main theoretical claim, Theorem 2.1, is that when the working number of factors R is at least the true number r, the linear space spanned by the estimated factors consistently contains the space spanned by the true factors, at rate O_P(N^{-1/2} \\nu_{\\min}(H)^{-1}), and that this holds also for finite T. The paper further develops applications to factor-augmented forecasting, post-selection inference in high-dimensional regressions, thresholded idiosyncratic covariance estimation, a specification test for observed factors, and factor-adjusted false discovery control. Four choices of diversified weights are recommended: loading characteristics, rolling-window PCA loadings, initial transformations, and Hadamard columns. The appendix contains detailed proofs of the main estimation results and the forecast, inference, covariance, and specification-test theorems.","tokens_in":43739,"tokens_out":3665,"duration_ms":39387,"significance":"If the claims hold, the paper is a useful contribution: it offers a computationally trivial estimator of the factor space, it formalizes robustness to over-estimating the number of factors, and it covers several practically important downstream problems. The proofs are unusually detailed, and Proposition A.1 together with Theorem 2.1 provides a rigorous basis for the T, N \\to \\infty case. The paper is also honest about the fact that the estimator does not estimate the true factors themselves, only their span after an affine transformation. The main reasons I cannot endorse the manuscript as it stands are the gap between the finite-T claim and the proof, and the absence of a formal result for the FDR application.","major_comments":[{"comment":"The finite-T claim in Theorem 2.1 is not supported by the proof. Proposition A.1, on which the proof of Theorem 2.1 relies, is explicitly proved under 'T, N \\to \\infty', while Theorem 2.1 states 'T is either finite or grows'. More concretely, Proposition A.1(i) asserts that \\lambda_{\\min}((1/T)\\hat F' K \\hat F) \\ge c/N with probability approaching one, but when T < R the matrix (1/T)\\hat F' K \\hat F is T \\times T and has rank at most T < R, so its smallest eigenvalue is identically zero. The same issue affects the matrix M' \\hat F' \\hat F M used to define P_{\\hat F M} when T < r. Thus the advertised finite-T robustness, repeated in the Abstract, Section 2.4, and the remark after Theorem 2.1, is a proof gap rather than an established result. The theorem should either be restricted to T \\to \\infty, or a separate argument covering finite T should be supplied.","section":"Theorem 2.1 and Appendix A.2 / Proposition A.1"},{"comment":"The factor-adjusted false discovery control is claimed as an application ('Our theories imply the following expansion...'), but no theorem, set of regularity conditions, or proof is given for the FDR control. The expansion displayed in Section 3.5 is stated informally, and it is not shown that the resulting test statistics are sufficiently weakly dependent or that the nominal FDR level is controlled uniformly over R \\ge r. Since Section 3.5 is presented as one of the applications of the diversified-factor construction, this is a load-bearing omission for the paper's claims and must be addressed with a formal statement and proof, or the section should be explicitly labeled as heuristic.","section":"Section 3.5"},{"comment":"Assumption 2.2, requiring rank(W'B/N) = r and \\nu_{\\min}(W'B/N) \\gg N^{-1/2}, is the key condition that makes the rates in Theorem 2.1 and all later theorems non-degenerate, but it is not verifiable from the observed panel. The recommended weight choices in Section 4 are given heuristic justification only, and for the Hadamard construction in Section 4.4 no proof is provided that the resulting deterministic W satisfies Assumption 2.2 for a loading matrix B of the assumed form. Because every consistency rate in the paper depends on \\nu_{\\min}(H), the practical applicability of the method would be substantially strengthened by explicit sufficient conditions on B for at least the deterministic weight choices, or by a discussion of the consequences when Assumption 2.2 fails.","section":"Assumption 2.2 and Section 4"}],"minor_comments":[{"comment":"The caption states that Table 1 is 'computed based on one set of simulation replications'; with m = 50 forecast windows, the relative MSE entries therefore have no error bars or replication variability. Statements such as 'DP outperforms under the strong serial correlations' should be softened accordingly.","section":"Table 1"},{"comment":"The notation in Section 3.5 is inconsistent: the text defines \\bar f = (1/T)\\sum_t \\bar f_t and \\bar u = (1/T)\\sum_t \\bar u_t, but the symbols \\bar f_t and \\bar u_t are not defined before use. This should be cleaned up, for example by writing \\hat f_t and \\hat u_t.","section":"Section 3.5"},{"comment":"In the description of the moving-window weights, the periods (I) and (II) are indexed by t = 1,...,T_0 and t = T_0+1,...,T_0+T, but later the independence claim is stated for 't = m+1,...,m+T' with m undefined. The indexing should be made consistent.","section":"Section 4.2"},{"comment":"The statement that a cross-sectional CLT 'is straightforward to verify' under Assumption 2.1 is too casual: Assumption 2.1 does not itself imply a Lindeberg condition or the existence of the limit V. Since (2.2) is motivational rather than used in the main theorems, this is a presentation issue rather than a technical error.","section":"Equation (2.2)"}],"recommendation":"major_revision","confidential_remarks":"The finite-T gap in Theorem 2.1 is the most serious issue: the headline 'works even for finite length of time series' is not proven by the supplied argument. The FDR section is also not yet at the standard of the rest of the paper. If the authors can either prove a genuine finite-T statement or clearly restrict the theorem to T \\to \\infty, and provide a formal FDR result, I would be willing to reconsider."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead Fan-Liao on diversified projections. The core idea is simple and genuinely useful: estimate factors as cross-sectional weighted averages with external weights W, then prove that projections onto the estimated factor space are consistent even when R > r, provided rank(W'B) = r and the smallest singular value does not decay too fast. That stands as a real contribution. It unifies CCE-type arguments and gives a clean framework for forecasts, post-selection inference, covariance estimation, and specification tests under over-estimation.\n\nThe paper does real work. Theorem 2.1 is properly proved for T, N going to infinity, the appendix is detailed, and the simulations, though limited, support the main asymptotic claim. The four weight constructions are sensible, and one is connected to Barigozzi-Cho with proper citation.\n\nNow the soft spots, in rough order.\n\nFirst, the finite-T selling point is not actually proved. Theorem 2.1 states T finite or growing, but the proof leans on Proposition A.1, which is proved under T, N going to infinity. Proposition A.1(i) says (1/T)hat(F)'hat(F) has minimum eigenvalue at least c/N with probability approaching one; that is impossible when T < R because the matrix has rank at most T. The same issue affects M'hat(F)'hat(F)M when T < r. So the finite-T claim is load-bearing and currently unsupported. It may be salvageable—at T=1, r=1, R=2 the span result is trivial—but the given argument does not cover it. This matters because \"works even for finite length of time series\" is in the abstract.\n\nSecond, Assumption 2.2 on rank and eigen-separation of W'B/N is strong and unverifiable from the observed panel. The paper is honest about this, and the four suggested weight constructions are reasonable, but the Hadamard recommendation is asserted without a proof that the rank condition holds for general loadings.\n\nThird, Section 3.5 claims FDR control with no theorem and no proof. The expansion is sketched but no distributional result is stated. That is more than a minor omission.\n\nFourth, Table 1 is based on one simulation replication with no error bars. Fine as an illustration, not as evidence.\n\nCitation pattern is honest: prior CCE work, Barigozzi-Cho, and Moon-Weidner are acknowledged. There is no circularity: W is external and H is a transformation by construction.\n\nWho is this for? Any econometrician or statistician doing factor-augmented inference who wants to avoid factor-number selection. The paper deserves a serious referee, but it should be sent back for a real fix on finite-T or a revised statement, plus a proper FDR theorem or removal of that section. I would not desk-reject it. My own reading: the central asymptotic claim is likely right; the advertised finite-T robustness is currently a proof gap.","headline":"The over-estimation robustness result is a genuine contribution, but the advertised finite-T guarantee has a real proof gap and one application is asserted without proof.","tokens_in":44168,"tokens_out":2109,"would_cite":true,"duration_ms":22502,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H25","62M10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Estimating latent factors by pre-chosen cross-sectional weighted averages is valid even when the working number of factors exceeds the true number, so factor-augmented inference no longer requires consistently counting factors.","keywords":["large dimensions","random projections","over-estimating the number of factors","principal components","factor-augmented regression","diversified factors","post-selection inference","large covariance estimation"],"falsifier":"Take a two-factor model in which the second factor's loadings are exactly orthogonal to every column of $W$ (for instance, loadings that alternate in sign across units while $W$ is constant); then $\\nu_{\\min}(H) = 0$, so the projection error bound in Theorem 2.1 diverges rather than vanishes, and a simulation should show $\\|P_{\\hat F}M - P_F\\|$ failing to converge as $N$ grows.","tokens_in":43195,"feed_emoji":"📊","tokens_out":12892,"duration_ms":107048,"temperature":0.7,"pith_summary":"The paper tries to remove the standard requirement that the number of latent factors be consistently estimated before any factor-model inference can proceed. It claims that a simple estimator, the pre-chosen cross-sectional weighted average $\\hat f_t = (1/N)W'x_t$, spans the true factor space asymptotically whenever the working number of factors $R$ is at least the true number $r$, even when $R$ over-estimates $r$ substantially and even when the time horizon $T$ is finite. This matters because consistent factor counting demands strong factors, long and stationary series, and weak serial dependence, conditions that frequently fail in forecasting and panel-data applications. If the claim holds, forecasts, post-selection inference, large covariance estimation, and factor specification tests all remain valid without knowing $r$, and the special case $r = 0$, $R \\geq 1$ makes the procedure a safe default even when no common factors exist.","feed_headline":"Over-estimating factors no longer breaks factor-model inference","feed_subtitle":"Pre-chosen weighted averages keep forecasts, tests, and covariances valid even when the factor count is wrong.","key_machinery":"The central object is the affine transformation matrix $H = (1/N)W'B$, which links the diversified projection $\\hat f_t = H f_t + (1/N)W'u_t$ to the true latent factors. The argument's core is Proposition A.1: when $R > r$ the gram matrix $(1/T)\\hat F'\\hat F$ is invertible but its inverse is only of order $O_P(N)$, while $H'((1/T)\\hat F'\\hat F)^{-1}$ remains well behaved and $H'((1/T)\\hat F'\\hat F)^{-1}H$ converges to the generalized inverse of its population analogue, so all downstream projections $P_{\\hat F}$ behave as if the factor count were correct. The second load-bearing piece is Assumption 2.2, which requires $\\mathrm{rank}(H) = r$ and $\\nu_{\\min}(H) \\gg N^{-1/2}$ with $\\nu_{\\max}(H) \\leq C\\nu_{\\min}(H)$, meaning the user-supplied weights must be sufficiently correlated with all $r$ columns of the loading matrix while still diversifying away the idiosyncratic noise.","core_discovery":"Under Assumptions 2.1–2.4, Theorem 2.1 establishes that for every bounded $R \\geq r$ the projection error satisfies $\\|P_{\\hat F}M - P_F\\| = O_P(N^{-1/2}\\nu_{\\min}(H)^{-1})$, so the linear space spanned by the diversified-factor estimates asymptotically contains the linear space of the true factors, with $T$ either finite or growing. The estimator's clean identity $\\hat f_t = H f_t + (1/N)W'u_t$ reduces the estimation problem to an affine transformation $H = (1/N)W'B$ plus a diversifiable noise term, avoiding eigenvector analysis entirely. The paper then proves that the same $R \\geq r$ robustness carries through to factor-augmented forecasting, high-dimensional post-selection inference (including the $r = 0$ case with no factors at all), sparse thresholding estimation of the idiosyncratic covariance, and a test of whether observed factors span the latent factor space.","pith_inferences":["If the $r = 0$ insurance result is taken at face value, a practical policy recommendation follows that the paper only hints at: applied researchers should always run factor-augmented post-selection inference with at least one working factor, never pre-testing for the presence of factors.","The Hadamard deterministic weights make the estimator a fixed linear sketch of the panel with no data-dependent tuning, which connects the method to random-projection and sketching ideas in computational statistics and treats the rank condition $\\nu_{\\min}(H) \\gg N^{-1/2}$ as a coverage-type condition on the loading matrix.","Because the projection is purely cross-sectional, the estimator is a natural candidate for nonstationary or structurally broken panels, as long as weights can be learned from a pre-period, the setting the moving-window construction in Section 4.2 is designed for.","The rank condition on $H$ is effectively a demand that the user's weight directions cover the entire $r$-dimensional loading space, so a testable extension would be to check coverage empirically by comparing downstream inferences across several candidate weight matrices and looking for instability."],"forward_implications":["Out-of-sample forecasts from factor-augmented regressions achieve the rate $O_P(T^{-1/2} + N^{-1/2}\\nu_{\\min}^{-1})$ without a consistent estimator of $r$, for any bounded $R \\geq r$.","Post-selection inference on a treatment effect in a high-dimensional factor-augmented model is asymptotically normal with valid confidence intervals uniformly over all $0 \\leq r \\leq R$, including $r = 0$ where no factors exist.","The thresholded idiosyncratic covariance estimator $\\hat\\Sigma_u$ is consistent in operator norm at the rate $(\\omega_{NT})^{1-q} m_N$ for every $R \\geq r$, so over-estimating factors does not corrupt large covariance estimation.","A specification test of whether observed factors span the latent factor space has a standard normal null distribution, computed with a parametric bootstrap for the variance.","The special case $r = 0$, $R \\geq 1$ shows that extracting 'factors' when the panel is actually weakly dependent is a safe insurance procedure for factor-augmented inference."],"supporting_citations":[{"why":"Defines the benchmark principal-components factor-augmented forecasting setup the paper extends to over-estimated R.","marker":"Stock and Watson (2002)"},{"why":"Documents the inconsistency of principal-components factors under finite T and serial correlation, motivating the cross-sectional projection approach.","marker":"Bai (2003)"},{"why":"Shows eigenvectors can be inconsistent when eigenvalues are not spiked, the obstacle for over-estimating R that diversified projections avoid.","marker":"Johnstone and Lu (2009)"},{"why":"Establishes the consistent factor-number selection that the paper argues is a fragile precondition and seeks to remove.","marker":"Bai and Ng (2002)"},{"why":"Introduces common correlated effects estimation by cross-sectional averages, whose rank-condition analysis the paper refines for R greater than r.","marker":"Pesaran (2006)"},{"why":"Proposes an alternative over-estimation-robust estimator and inspires the rolling-window and block weight constructions.","marker":"Barigozzi and Cho (2018)"},{"why":"Provides the double-selection post-selection inference baseline that the factor-augmented procedure extends and compares against.","marker":"Belloni et al. (2014)"},{"why":"Supplies the thresholding estimator of the idiosyncratic covariance that Theorem 3.3 inherits for all R at least r.","marker":"Fan et al. (2013)"}],"fun_headline_variants":["Over-count factors? Robust estimator keeps inference sound","Diversified projections yield factors robust to over-specification","Factor estimation now works despite over-estimating latent count","Wrong factor number? New method keeps forecasts and tests valid","Robust factor extraction via diversified projections"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the user can supply a set of weights that are independent of the noise yet strongly enough correlated with every one of the true factor loadings, in the sense that the smallest nonzero singular value of the matrix $H = W'B/N$ does not decay faster than $N^{-1/2}$; if any true factor is nearly orthogonal to all the weights, it gets averaged away and the whole construction fails.","fun_headline_variants_meta":{"raw":{"variants":["Over-count factors? Robust estimator keeps inference sound","Diversified projections yield factors robust to over-specification","Factor estimation now works despite over-estimating latent count","Wrong factor number? New method keeps forecasts and tests valid","Robust factor extraction via diversified projections"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000581,"raw_usage":{"total_tokens":2737,"prompt_tokens":945,"completion_tokens":1792,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":1716}},"tokens_in":561,"tokens_out":1792,"duration_ms":17796,"temperature":1.0,"reasoning_tokens":1716,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:18:27.194521+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a two-factor model in which the second factor's loadings are exactly orthogonal to every column of $W$ (for instance, loadings that alternate in sign across units while $W$ is constant); then $\\nu_{\\min}(H) = 0$, so the projection error bound in Theorem 2.1 diverges rather than vanishes, and a simulation should show $\\|P_{\\hat F}M - P_F\\|$ failing to converge as $N$ grows.","supporting_citations":[{"cited_title":"and Watson, M","cited_arxiv_id":null,"evidence_quote":"Defines the benchmark principal-components factor-augmented forecasting setup the paper extends to over-estimated R."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the inconsistency of principal-components factors under finite T and serial correlation, motivating the cross-sectional projection approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows eigenvectors can be inconsistent when eigenvalues are not spiked, the obstacle for over-estimating R that diversified projections avoid."},{"cited_title":"and Ng, S","cited_arxiv_id":null,"evidence_quote":"Establishes the consistent factor-number selection that the paper argues is a fragile precondition and seeks to remove."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces common correlated effects estimation by cross-sectional averages, whose rank-condition analysis the paper refines for R greater than r."},{"cited_title":"Consistent estimation of high-dimensional factor models when the factor number is over-estimated","cited_arxiv_id":"1811.00306","evidence_quote":"Proposes an alternative over-estimation-robust estimator and inspires the rolling-window and block weight constructions."},{"cited_title":", Chernozhukov, V","cited_arxiv_id":null,"evidence_quote":"Provides the double-selection post-selection inference baseline that the factor-augmented procedure extends and compares against."},{"cited_title":", Liao, Y","cited_arxiv_id":null,"evidence_quote":"Supplies the thresholding estimator of the idiosyncratic covariance that Theorem 3.3 inherits for all R at least r."}],"review_version":1}