{"id":"95da2886-5af2-4cea-baee-ed69e84bea12","arxiv_id":"2607.06908","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"An iterative algorithm combining adaptive Marčenko–Pastur edge recalibration with a participation-ratio delocalization filter recovers weak global factors near the BBP transition in high-dimensional financial correlation matrices.","lead":"This paper proposes a method to detect hidden common factors in large financial correlation matrices by combining eigenvalue separation with an eigenvector delocalization filter. It matters because weak but real market-wide factors are often missed by standard tests near a mathematical phase-transition boundary.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The PR delocalization filter (τ=0.3) may be inert in the BH simulations: both weak-factor and idiosyncratic eigenvectors converge to PR/p→1/3>τ, so the filter passes everything that is spectrally separated, leaving the improvement attributable solely to iterative MP edge recalibration.","rationale":"The reader correctly identifies a real concern about τ transferability from synthetic to empirical data, but I believe the more load-bearing issue is upstream of that: the PR filter may not contribute anything even in the synthetic model where it is validated. The asymptotic theory (Sec. III.D, Appendix C) shows that weak factors and noise eigenvectors share the same PR/p→1/3 limit, and τ=0.3 is set below this limit. In the BH model with isotropic noise, there are no localized eigenvectors to reject, so the PR filter should pass every spectrally separated component. If this is the case, the paper's central novelty claim—combining spectral separation with eigenvector delocalization—collapses to just iterative MP edge recalibration, which is a less novel contribution. The paper does not include an ablation (τ=0 vs τ=0.3) that would settle this. The PR derivation itself (Appendix C) is correct and clean, and the iterative recalibration idea is reasonable. But without demonstrating that the PR filter adds value, the paper's framing is unsupported. The verdict remains CONDITIONAL: the method may work, but the stated mechanism is unverified. An ablation would either confirm the PR filter's contribution (strengthening the paper) or reveal it is inert (requiring reframing). I note the reader's concern about τ transferability is also valid but secondary: if the PR filter is inert in the BH model, the question of whether τ transfers to real data is moot for the simulation claims, though it could matter for the empirical application if real data has localized components the BH model lacks.","tokens_in":16163,"tokens_out":3080,"duration_ms":101918,"concrete_test":"Run the IGF algorithm on the same BH simulation (Fig. 3 setting, q=1/2, p=200…800) and on the empirical S&P 500 data (Fig. 7) with τ=0 (PR filter disabled, retaining only the iterative MP edge recalibration). Compare the detected factor counts to those obtained with τ=0.3. If the counts match within Monte Carlo error, the PR filter is inert and the paper's central novelty claim—that combining spectral separation with eigenvector delocalization improves detection—is unsupported; the improvement comes solely from the iterative recalibration of the MP edge.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central novelty claim is that combining spectral separation with eigenvector delocalization (the PR filter) improves factor detection near the BBP transition. However, the asymptotic results in Sec. III.D and Appendix C show that weak-factor directions and typical idiosyncratic eigenvectors both satisfy PR(u)/p→1/3. The operational threshold τ=0.3 is set *below* 1/3 (Sec. V, Fig. 4-5). This means that in the BH model—which has isotropic noise and dense loadings—there are no localized eigenvectors to filter out: every spectrally separated component will also satisfy the PR criterion. The PR filter is therefore expected to be inert in the BH simulations. The improvement over eigenvalue-only methods (Onatski, Tracy-Widom, MP edge) shown in Fig. 3 could be entirely due to the iterative MP edge recalibration (Sec. III.C), which adaptively lowers the noise estimate using the residual median. The paper never ablates the PR filter (e.g., by running IGF with τ=0) to demonstrate that the PR criterion contributes anything beyond the spectral recalibration. If the PR filter is inert, the paper's framing of 'combining spectral separation with eigenvector delocalization' as the key innovation is unsupported; the actual mechanism would be the iterative median-based noise recalibration alone. This is distinct from the reader's concern about τ transferability to empirical data: even if τ transfers perfectly, the PR filter may not be doing any work in the simulations that validate the method.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This manuscript proposes an iterative global factor (IGF) algorithm for detecting the number of global factors in high-dimensional correlation matrices, with a focus on the regime near the Baik–Ben Arous–Péché (BBP) phase transition. The method combines two ingredients: (i) adaptive Marčenko–Pastur (MP) edge recalibration, in which the effective noise variance is re-estimated from the residual spectrum at each step using a median-based estimator, and (ii) a participation-ratio (PR) delocalization filter that retains only eigenvectors with PR/p above a threshold τ. The author derives asymptotic PR benchmarks for the Brown–Harding (BH) factor model: the leading coherent eigenvector satisfies PR(u₁)/p → 1, while weak-factor and idiosyncratic eigenvectors satisfy PR(u)/p → 1/3. Monte Carlo simulations of the BH model show that IGF recovers the true factor count near the BBP transition where the Onatski test, Tracy–Widom criterion, and MP upper-edge criterion fail. The method is then applied to S&P 500 returns, detecting a median of 7 factors across moving windows.","tokens_in":17065,"tokens_out":1531,"duration_ms":189286,"significance":"The paper addresses a well-motivated problem at the intersection of random matrix theory and financial econometrics. The derivation of PR benchmarks for the BH factor model (Appendix C) is a clean contribution: the limits PR(u₁)/p → 1 and PR(u_α)/p → 1/3 for weak-factor and idiosyncratic directions are established rigorously for the covariance case and extended to the correlation matrix under a small-coefficient-of-variation approximation. The iterative median-based noise recalibration is a practical and sensible idea. The empirical finding of 7 median factors versus 1 from the Onatski test is potentially interesting for the econophysics community. However, the central novelty claim—that combining spectral separation with eigenvector delocalization improves detection—rests on an untested assumption that the PR filter contributes beyond the spectral recalibration, as detailed below.","major_comments":[{"comment":"The PR delocalization filter (τ = 0.3) may be inert in the BH simulations, undermining the paper's central novelty claim. The asymptotic results in Sec. III.D and Appendix C show that weak-factor directions and typical idiosyncratic eigenvectors both satisfy PR(u)/p → 1/3. The operational threshold τ = 0.3 is set below 1/3 (Sec. V, Figs. 4–5). In the BH model, which has isotropic noise and dense loadings, there are no localized eigenvectors to filter out: every spectrally separated component will also satisfy the PR criterion. The improvement over eigenvalue-only methods shown in Fig. 3 could therefore be entirely attributable to the iterative MP edge recalibration (Sec. III.C), which adaptively lowers the noise estimate using the residual median. The paper never ablates the PR filter (e.g., by running IGF with τ = 0) to demonstrate that the PR criterion contributes anything beyond the光谱","section":null},{"comment":"The calibration of τ = 0.3 involves a degree of circularity. The threshold is calibrated on synthetic BH moving-window data (Figs. 4–5) and then used to detect factors in that same synthetic model (Fig. 6) and in empirical data (Fig. 7). The synthetic recovery in Fig. 6 is somewhat circular by construction: the threshold was chosen to recover the correct count in that specific model. The empirical application is less circular, but it assumes that the empirical data-generating process sufficiently matches the BH model's isotropic noise structure for the threshold to transfer. If real financial data exhibits stronger heteroscedasticity, cross-sectional correlation in the idiosyncratic component, or non-Gaussian tails, the threshold may not be valid. The paper should discuss this transferability assumption more explicitly and, ideally, provide a robustness check with a misspecified noise结构.","section":null},{"comment":"The small-CV approximation (CVD ≈ 0.0352) underpinning the extension of PR limits to the correlation matrix (Appendix C.4) is verified only for the specific BH parameters used in the simulations. The approximation D ≈ d₀I is load-bearing for the claim that PR(u_α)/p → 1/3 holds for the correlation matrix. The paper should discuss how sensitive this approximation—and hence the τ = 0.3 threshold—is to the model parameters, and whether CVD values typical of empirical financial data would still satisfy CVD ≪ 1. Without this, the empirical claims rest on an untested assumption.","section":null}],"minor_comments":[{"comment":"Sec. V: The choice q = 1/2 is stated without much justification beyond matching the simulation regime. A brief discussion of why this particular q is appropriate for the empirical analysis, or how sensitive the results are to q, would be helpful.","section":null},{"comment":"Sec. III.E, step 2: The algorithm uses q = p/n rather than q_k = (p−k+1)/n. The text states this is because k is expected to be small relative to p, but for k = 7 and p = 417, the correction is non-negligible. A sensitivity check or justification would strengthen the presentation.","section":null},{"comment":"Fig. 3: The vertical dotted line denotes p_c ≈ 252, but Sec. IV states p_c ≈ 317 for the same parameters. These values should be reconciled.","section":null},{"comment":"The data availability statement says data are available from the corresponding author upon request. For reproducibility, depositing the code and data in a public repository would strengthen the manuscript.","section":null},{"comment":"Sec. III.D, Eq. (25): The statement that idiosyncratic eigenvectors satisfy PR(u_k)/p → 1/3 for k > K is somewhat informal, since the population-level degeneracy means individual eigenvectors are not uniquely defined. The finite-sample argument in Appendix C.3 is reasonable but could be stated more carefully in the main text.","section":null},{"comment":"Abstract: The sentence beginning 'A synthetic moving-window calibration matched to the empirical dimensions' is grammatically incomplete.","section":null},{"comment":"Sec. II: The description of the Onatski test refers to k₁ and k₂ but does not clearly state how these bounds are chosen in practice for the empirical analysis. Clarifying this would help readers assess the empirical comparison.","section":null}],"recommendation":"major_revision","confidential_remarks":"The skeptic's concern about the PR filter being inert in the BH simulations is, in my assessment, the most important issue. The author should be asked to run an ablation with τ = 0 (or equivalently, remove the PR filter) to show whether the PR criterion contributes anything in the BH setting. If it does not, the paper's framing needs to be adjusted: the PR filter's value would then be primarily as a safeguard against localized components in empirical data (where localized eigenvectors may genuinely exist), not as a contributor to the simulation results. This is a fixable issue but it goes to the heart of the paper's novelty claim. The PR limit derivations themselves are sound and represent a genuine contribution regardless of the ablation outcome."},"author_rebuttal":null,"desk_editor":{"model":"glm-5.2","letter":"Here's the short version: the paper proposes an iterative global factor (IGF) estimator that combines adaptive Marčenko–Pastur edge recalibration with a participation-ratio (PR) delocalization filter. The PR limit derivations for the Brown–Harding (BH) factor model are clean and new, and the algorithm does recover the true factor count in simulations where Onatski and Tracy–Widom fail. But there's a structural problem with the central novelty claim that needs addressing before this is publishable: the PR filter may be doing nothing in the simulations that validate the method. The asymptotic results show that weak-factor and idiosyncratic eigenvectors both satisfy PR/p→1/3. The operational threshold τ=0.3 is set below 1/3. So in the BH model — which has isotropic noise and dense loadings — every spectrally separated component will also pass the PR criterion. There are no localized eigenvectors to filter out. The improvement over eigenvalue-only methods shown in Fig. 3 could be entirely attributable to the iterative median-based noise recalibration (Sec. III.C), which adaptively lowers the noise estimate using the residual spectrum. The paper never ablates the PR filter (e.g., running IGF with τ=0) to show it contributes beyond the spectral recalibration. If the filter is inert, the framing of 'combining spectral separation with eigenvector delocalization' as the key innovation is unsupported; the actual mechanism would be the iterative recalibration alone. This is distinct from the reader's concern about τ transferability to empirical data. Even if τ transfers perfectly, the PR filter may not be doing any work in the simulations. The PR limits themselves are a legitimate theoretical contribution. The derivation in Appendix C is straightforward but correct: PR(u₁)/p→1 under strong common loadings, PR(uα)/p→1/3 for weak factors and idiosyncratic directions. The extension to correlation matrices via the small-CV approximation (CVD≈0.0352) is reasonable for the BH model but would need justification under heavier heteroscedasticity. The empirical S&P 500 application is suggestive but limited by the same threshold calibration circularity the reader flagged. The median count of 7 factors is interesting but hard to interpret without knowing whether the PR filter is actually filtering anything in real data. This deserves a serious referee. The PR limit derivation is solid and the algorithmic idea has merit. But the author needs to either (a) run the ablation with τ=0 to demonstrate the PR filter contributes beyond recalibration, or (b) reframe the contribution around the iterative MP recalibration and present the PR filter as a secondary safeguard rather than the central novelty. The current framing oversells the PR component relative to what the simulations actually demonstrate.","headline":"The PR filter may be inert in the BH simulations — the real work is the iterative MP recalibration","tokens_in":16970,"tokens_out":643,"would_cite":false,"duration_ms":69524,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Eigenvector shape, not just eigenvalue size, reveals hidden financial factors","keywords":[],"falsifier":"If empirical financial correlation matrices have eigenvector PR distributions that do not cluster near 1/3 for noise components (for example, due to sector clustering or heavy-tailed distributions), the tau = 0.3 threshold would misclassify localized noise eigenvectors as global factors or vice versa.","tokens_in":16456,"feed_emoji":"📊","tokens_out":1045,"duration_ms":169588,"temperature":0.7,"pith_summary":"In high-dimensional finance, the number of driving factors in a market is hard to pin down: when the count of assets is comparable to the count of observations, weak factors sink into the noise and standard spectral tests lose power. This paper proposes that the geometric spread of an eigenvector across the asset universe carries information that eigenvalues alone miss. The author proves that in the Brown-Harding factor model, the leading market factor spreads uniformly across all assets (participation ratio approaching the maximum), while weak factors and pure noise eigenvectors settle at a benchmark ratio of one-third. By iteratively recalibrating the noise boundary and filtering for eigenvectors that are sufficiently spread out rather than concentrated on a few assets, the proposed IGF algorithm recovers the true factor count in simulations where eigenvalue-only methods fail. Applied to S&P 500 returns, it detects a median of seven factors, far more dynamic than the single factor found by the state-of-the-art Onatski test.","feed_headline":"Eigenvector shape, not just eigenvalue size, reveals hidden financial factors","feed_subtitle":"A participation-ratio filter recovers seven market factors where standard tests find one","key_machinery":"The machinery has three moving parts. First, an iterative noise recalibration: at each step, the effective noise variance is estimated from the median of remaining eigenvalues relative to the median of the Marchenko-Pastur distribution, producing an adaptive upper spectral edge. Second, a participation-ratio filter: candidate eigenvectors must exceed a threshold (PR/p >= 0.3) to confirm they are spread across many assets rather than localized. Third, the Brown-Harding factor model provides the theoretical backbone, with Harding's sample-eigenvalue corrections describing how weak factors emerge from the noise edge above a critical dimension p_c. The combination of spectral separation plus eig","core_discovery":"The central object is the participation ratio (PR) of an eigenvector, which measures how many assets it meaningfully touches. The paper proves two asymptotic limits for the Brown-Harding factor model: the leading coherent eigenvector satisfies PR(u1)/p approaching 1 (fully extended), while weak-factor and noise eigenvectors satisfy PR(u)/p approaching 1/3 (random delocalization). These distinct limits create a separability criterion at the eigenvector level. The IGF algorithm exploits this by combining iterative noise-edge recalibration with a PR threshold at tau = 0.3, retaining only eigenvalues that both separate from the noise bulk and have extended eigenvectors. In simulations near the B","pith_inferences":["If the PR threshold transfers across asset classes or markets, the 1/3 benchmark may serve as a universal constant for real-valued noise eigenvectors, making the method applicable beyond equities to credit, options, or macroeconomic panel data.","The gap between the seven-factor IGF result and the one-factor Onatski result may partly reflect a systematic bias in eigenvalue-only tests: they are calibrated for strong factors and may have low power against the weak-factor regime that dominates real markets.","The synthetic calibration approach (matching BH model dimensions to empirical ones to set tau) could be generalized into a principled framework for parameter selection in other random-matrix-based detection problems where analytical thresholds are unavailable."],"forward_implications":["Portfolio construction could benefit from richer factor structure: if seven rather than one factor drive returns, risk models built on single-factor assumptions may systematically underestimate diversification breakdowns during stress periods.","The PR/p = 1/3 benchmark for real-delocalized eigenvectors provides a model-free diagnostic that could be applied to any high-dimensional correlation matrix, not just financial ones.","The iterative recalibration logic could be extended to detect subcritical factors that remain entirely inside the Marchenko-Pastur bulk, where no eigenvalue separation exists but eigenvector structure may still differ from pure noise.","Regime-change detection in markets could use the fluctuating factor count (4 to 14 across windows) as a real-time indicator of structural shifts in the economy."],"fun_headline_variants":["Participation ratio filter separates weak market factors from noise","Iterative eigenvector filter recovers seven dynamic factors from S&P 500 returns","Eigenvector delocalization distinguishes weak global factors near the BBP transition","Extended eigenvectors identify true financial factors where eigenvalue tests fail","An iterative PR filter recovers seven market factors near the noise edge"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The operational threshold tau = 0.3 is calibrated on a synthetic model with isotropic noise and then applied directly to real S&P 500 data. If actual financial returns have stronger cross-sectional correlations or heteroscedasticity than the model assumes, the threshold may not transfer cleanly.","fun_headline_variants_meta":{"raw":{"variants":["Participation ratio filter separates weak market factors from noise","Iterative eigenvector filter recovers seven dynamic factors from S&P 500 returns","Eigenvector delocalization distinguishes weak global factors near the BBP transition","Extended eigenvectors identify true financial factors where eigenvalue tests fail","An iterative PR filter recovers seven market factors near the noise edge","Eigenvector participation ratio distinguishes market factors from random matrix noise","Iterative noise-edge recalibration and PR filter detect weak global market factors"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":1676,"prompt_tokens":704,"completion_tokens":972,"prompt_tokens_details":null},"tokens_in":704,"tokens_out":972,"duration_ms":21405,"temperature":1.0,"reasoning_tokens":935,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T23:09:57.956717+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If empirical financial correlation matrices have eigenvector PR distributions that do not cluster near 1/3 for noise components (for example, due to sector clustering or heavy-tailed distributions), the tau = 0.3 threshold would misclassify localized noise eigenvectors as global factors or vice versa.","supporting_citations":[],"review_version":1}