{"id":"2d1675ce-3788-4fa3-b434-145aeff2cadf","arxiv_id":"2411.14016","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A spatial-sign-based multiple testing framework (SS-BH and FSS-BH) that controls the false discovery rate when selecting skilled mutual funds under heavy-tailed errors and latent factors.","lead":"This paper introduces two statistical procedures that identify mutual funds with genuinely positive 'alpha' (skill) while keeping the share of false picks low. The methods use robust spatial-sign statistics to cope with heavy-tailed returns and hidden risk factors, and the authors prove asymptotic false discovery rate control.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2's proof has a rate mismatch: the established bound for the factor-adjustment error term does not imply the claimed negligibility needed for asymptotic normality and FDR control.","rationale":"The paper's central claim is the FDR control of FSS-BH, stated as Theorem 2. The most load-bearing concern is not the external validity of the elliptical assumption but an internal gap in the proof of that theorem: the rate at which the factor-adjustment error Δ_t is bounded is too slow to justify the claimed negligibility of this term in the test statistic's asymptotic representation. This is a concrete mathematical issue that must be resolved for the central claim to hold as stated. The reader's weakest-assumption focused on the elliptical model and weak-correlation conditions; I partially agree, but the rate mismatch is more directly tied to the proof and does not depend on disagreeing with the modeling assumptions. The paper has substantial independent merit: the FSS-BH idea is well-motivated, the simulations are extensive and show good empirical performance, and the authors are transparent about relying on predecessors for parts of the analysis. The N0 definition in the theorem statements also appears to conflate the number of nulls with the number of positive-alpha funds, which is a serious notational error that should be corrected even if the underlying proof is repaired. My recommendation remains conditional acceptance, unchanged from the reader's verdict, because the identified gap is likely fixable but makes the current manuscript insufficiently rigorous for unconditional acceptance.","tokens_in":36762,"tokens_out":10143,"duration_ms":85729,"concrete_test":"Re-derive the contribution of Δ_t to the linear representation of T^{1/2}(\\breve{θ}−ωα) explicitly, without invoking the 'easy to prove' step. Compute the order of ∥T^{−1}Σ_t \\breve{r}_t^{−1}\\breve{D}^{−1/2}Δ_t∥∞ under Conditions (C5)-(C9), and verify whether it is o_p(N^{−1/2}T^{−1/2}/√log N). If the direct bound only yields o_p(N^{−1/2}/log N), then Theorem 2's proof is incomplete; either provide an additional cancellation argument (e.g., using the fact that Δ_t is orthogonal to the factor span) or adjust the theorem conditions so that the established rate suffices.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In the proof of Theorem 2 (Section 7.2), the term Δ_t = ΓW_t − \\hat{Γ}\\hat{W}_t − ΓW^⊤V_t + \\hat{Γ}\\hat{W}^⊤V_t captures the cumulative error from latent-factor estimation. The proof first derives ∥T^{−1}Σ_t Δ_t∥∞ = op(1/log N). From this, the largest bound one can directly infer for the quantity appearing in the test statistic is ∥T^{−1}Σ_t \\breve{r}_t^{−1}\\breve{D}^{−1/2}Δ_t∥∞ = op(N^{−1/2}/log N), since \\breve{r}_t = ∥\\breve{D}^{−1/2}η_t∥ is of order N^{1/2}. But earlier in the same proof the required bound is stated as op(N^{−1/2}T^{−1/2}/√log N). Under Condition (C4) log N = o(T^{1/5}), we have T^{1/2}/√log N ≫ log N, so N^{−1/2}/log N is asymptotically much larger than N^{−1/2}T^{−1/2}/√log N. Therefore op(N^{−1/2}/log N) does not imply the required op(N^{−1/2}T^{−1/2}/√log N). The proof then asserts 'it is easy to prove' the needed op(1/√log N) bound, but this assertion is not demonstrated and, as written, the exponents do not match. Without this bound, the remainder in the linear representation of T^{1/2}(\\breve{θ}−ωα) is not shown to be negligible at the 1/√log N threshold used for the Gaussian approximation, so the asymptotic normality in (11) — and hence the FDR control in Theorem 2 — is not established by the given argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two multiple testing procedures for selecting skilled mutual funds under a linear factor pricing model: SS-BH, based on spatial-sign test statistics when factors are observable, and FSS-BH, which first removes latent factors via the elliptical principal component method of He et al. (2022) and then applies spatial-sign tests. The main theoretical claims are asymptotic FDR control: Theorem 1 states FDR_{SS-BH} ≤ γ N0/N ≤ γ, and Theorem 2 states the analogous bound for FSS-BH under latent-factor conditions. The paper also reports extensive simulations with normal, t, mixture-normal, and independent-component errors, and a CRSP mutual fund application with rolling-window performance evaluation. The central claim is that the proposed procedures control FDR at the user-specified level in high-dimensional heavy-tailed settings.","tokens_in":37174,"tokens_out":6137,"duration_ms":58753,"significance":"If the asymptotic FDR theorems are correct, the paper offers a useful robust alternative to existing alpha-testing procedures such as D-BH and F-BH, particularly for heavy-tailed fund returns with latent factor structure. The use of spatial signs and elliptical PCA is well motivated, and the simulation design covers several practically relevant departures from normality. However, the paper's central contribution is not fully established: the FDR theorems rely on an imported asymptotic representation from a same-author preprint and, for Theorem 2, on a rate argument that is not demonstrated. The manuscript also has definitional inconsistencies in the theorem statements. These are load-bearing issues rather than presentational ones, so the paper needs substantial revision before its claims can be accepted.","major_comments":[{"comment":"The proof contains a rate mismatch that is not resolved. It first establishes ||T^{-1} Σ_t Δ_t||_∞ = op(1/log N). Since ||˘r_t^{-1} ˘D^{-1/2}||_∞ is of order N^{-1/2} (because ˘r_t ≍ √N and the diagonal entries of ˘D are bounded), the directly implied bound for the quantity appearing in the test statistic is ||T^{-1} Σ_t ˘r_t^{-1} ˘D^{-1/2} Δ_t||_∞ = op(N^{-1/2}/log N). However, earlier in the same proof the required bound is stated as op(N^{-1/2} T^{-1/2}/√(log N)). Under Condition (C4), log N = o(T^{1/5}), so T^{1/2}/√(log N) → ∞, and op(N^{-1/2}/log N) does not imply op(N^{-1/2} T^{-1/2}/√(log N)). The sentence \"Finally, it is easy to prove that ... = op(1/√(log N))\" asserts the needed result without proof, and the exponents do not match. Without this bound, the remainder in the linear representation of T^{1/2}(˘θ − ωα) is not shown to be negligible at the threshold required for the Gaussian approximation, so the asymptotic normality used for Theorem 2 is not established.","section":"7.2, Proof of Theorem 2"},{"comment":"The notation N0 and N1 is inconsistent in the theorem statements. In Section 2, N0 is defined as \"the true number of securities with positive alpha,\" i.e., the number of non-null hypotheses. The theorems then introduce \"the number of false null hypotheses N1 ≤ N^{ϖ}\" without defining N1, and state FDR_{SS-BH} ≤ γ N0/N ≤ γ. The proof in Section 7.1 derives FDR ≤ γ|H0|/N, where H0 is the index set of true nulls. If N0 is the number of positive-alpha funds, then the proved bound involves |H0| = N − N0, not N0, and the inequality FDR ≤ γ N0/N is generally false when non-nulls are sparse. The theorem statements should either define N0 as the number of true nulls or state the result as FDR ≤ γ(1 − N0/N) ≤ γ, with N0 denoting the number of non-nulls. This is a load-bearing definitional error in the central claim.","section":"Section 2 and Theorems 1–2"},{"comment":"The key asymptotic normality in (6) is quoted from Theorem 2.1 of Zhao et al. (2024), a same-author preprint, and the linear representation (13) is also imported from that source. The present paper does not prove or verify the conditions for these results, and the inequalities used to control the remainder terms, including max_i |C_{T,i}| = op(1/√(log N)), are stated without derivation. Since the FDR theorems depend directly on this asymptotic representation, the argument is not self-contained. Moreover, inequality (15), the uniform Gaussian approximation for the null empirical process, is only sketched; the text says it \"is essentially the Gaussian approximation\" but does not cite a theorem that yields the uniform bound over x ∈ [0, t*] under the weak-dependence conditions (C4). I ask the authors to either supply full proofs or give precise references with verification of all conditions needed for (6), (13), and (15).","section":"Equations (6) and (13), Section 7.1"},{"comment":"Theorem 2 states that it holds under Conditions (C1), (C4), and (C5)–(C8), but the proof and Lemmas 1–4 explicitly require Condition (C9), including T log N = o(N) and the bound on ||α||. Condition (C9) is not cited in the theorem statement. In addition, Condition (C5) contains a stray phrase \"m is fixed\" and Condition (C7) duplicates the notation η_t = v_t L V_t already used in (C3), while Condition (C5) defines η_t through an elliptical representation with A; the relationship between these conditions should be clarified. The theorem statement and the condition list need to be aligned so that the reader can verify which assumptions are actually used.","section":"Theorem 2 statement and Conditions (C5)–(C9)"}],"minor_comments":[{"comment":"In the sentence defining p-values for FSS-BH, the manuscript writes \"the corresponding p-value for H0i versus H1i is ps_i = 1 − Φ(T s_i)\" but the statistic just defined is T f_i; the subscript should be f, not s.","section":"Section 3, paragraph after equation (11)"},{"comment":"The estimator bω defined before equation (6) is written with a hat, but later in the same paragraph bω is used in the expression bωα, while the population quantity ω is introduced as the limit of T^{-1} Σ_t ϑ_t. Please distinguish clearly between the estimator and its limit throughout the proofs.","section":"Notation, Section 2"},{"comment":"The phrase \"elliptical principle component method\" should be \"elliptical principal component method.\" Similar spelling issues appear in Figure 7's caption (\"repectively\") and in the text (\"Kendall’ tau\").","section":"Abstract and Introduction"},{"comment":"The paper does not mention whether code or data are available for reproducing the simulations and the empirical application. Given the computational nature of the proposed procedures, a statement on code availability would improve reproducibility.","section":"Simulation and real-data sections"},{"comment":"In Scenario III, the error term is generated as ε_it = 0.5 z_{1t} + ζ_i z_{2t} + ε_it, where z_{2t} is t(3)/√3; this does not satisfy the elliptical representation assumed in Conditions (C5) and (C7). The text acknowledges this on the next page, but it would help to state explicitly which of the theoretical guarantees are expected to remain valid under such misspecification.","section":"Section 4, Scenario III"}],"recommendation":"major_revision","confidential_remarks":"The paper's central FDR theorems are currently supported by a chain of imported results from a same-author preprint (Zhao et al. 2024) and by a rate argument in Section 7.2 that appears to contain a genuine mismatch. The issue is not merely presentation: the claimed asymptotic normality of the FSS-BH statistic is not established by the given proof. I would advise the editor that acceptance should require the authors to provide a complete proof of the rate-Negligibility bound, correct the N0/N1 definitions, and either prove or precisely verify all imported asymptotic results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read: a workmanlike extension of spatial-sign multiple testing to fund selection with latent factors. SS-BH and FSS-BH are genuinely new combinations, the simulations are fairly thorough and honestly reported, and the method looks useful for heavy-tailed fund data. The catch is that Theorem 2, the headline FDR result, is not established as written. The stress-test note is right that the stated rates do not connect: the intermediate bound is written as op(1/log N), which would only give op(N^{-1/2}/log N) for the term that actually enters the test statistic, far short of the required op(N^{-1/2}T^{-1/2}/sqrt(log N)). The closing 'easy to prove' bound is also claimed at the wrong strength.\n\nWhat makes this repairable rather than fatal: on a close read, the two displayed term bounds in Section 7.2 are each op(1/(sqrt(T) sqrt(log N))) once the factor of T^{-1} is accounted for. So the derivation actually delivers the stronger intermediate ||T^{-1} sum Delta_t||_inf = op(1/(sqrt(T) sqrt(log N))), and multiplying by r_t^{-1} = O_p(N^{-1/2}) gives exactly the required rate. The machinery is there; the rates need to be stated correctly and the last step written out. A referee should demand the corrected version before accepting the theorem.\n\nWhat the paper does well: the combination of elliptical-PCA factor extraction with spatial-sign statistics is sensible; the simulation design covers normal, t, mixture-normal, and nonsymmetric errors, with and without latent factors; and FSS-BH controls FDP in most settings with better power than D-BH and F-BH under heavy tails. The authors also concede the nonsymmetric case where their FDP slips, which is honest.\n\nOther soft spots, in order: Theorem 1 leans on the same group's preprint (Zhao et al. 2024) for the linear representation and on a Liu-Shao-style Gaussian approximation that is sketched rather than proved; the N0 notation is defined as the positive-alpha count but used as the null count; Theorem 2's condition list omits (C9), which its own lemmas require; and the abstract's 'superiority of FSS-BH' overstates Figure 9, where SS-BH and FSS-BH are nearly identical, one path each, no error bars.\n\nWho it is for: empirical asset-pricing and high-dimensional-testing readers who want a robust alternative to Giglio-Liao-Xiu or Lan-Du. It deserves a serious referee. I would send it out and ask for the rates fixed, the imports verified, the notation cleaned up, and the real-data claims scaled back.","headline":"Useful, honestly-simulated spatial-sign fund-selection paper whose FSS-BH theorem, as written, has a rate gap in the proof that looks repairable on close reading, but the headline claim is not currently established.","tokens_in":37691,"tokens_out":24883,"would_cite":false,"duration_ms":192174,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H15","62G10","62H25","62G35"],"pacs":[],"model":"deepseek-v4-flash","headline":"A spatial-sign-based multiple testing procedure controls the false discovery rate when screening mutual funds for positive alpha, even under heavy-tailed returns and hidden factors.","keywords":["False discovery rate","Mutual fund selection","Spatial sign","Heavy-tailed distribution","Latent factor model","Elliptical distribution","Principal component analysis","Multiple testing"],"falsifier":"Run FSS-BH on simulated data with $T=60$, $N=200$, errors from the asymmetric independent-component model $(3-\\chi^2_3)/\\sqrt6$, and a latent factor with strong loadings; repeat enough times to estimate FDP. If the empirical FDP systematically exceeds $\\gamma=0.1$ when the $\\alpha$ signal is strong, the asymptotic FDR control is failing outside the elliptical assumption.","tokens_in":36536,"feed_emoji":"📈","tokens_out":7752,"duration_ms":67914,"temperature":0.7,"pith_summary":"Fund performance is usually measured by alpha in a linear factor pricing model, and the practical question is which of thousands of funds have genuinely positive alpha rather than luck. This paper proposes replacing the classical t-statistic with a spatial-sign statistic built on the spatial median, and feeding its normal-tail p-values into the Benjamini-Hochberg procedure. The resulting SS-BH procedure is shown to control the false discovery rate asymptotically under an elliptical error model with weak cross-fund dependence; when latent factors are present, an FSS-BH variant first removes them via a robust principal component method based on the spatial Kendall's tau matrix. If the theorems are correct, investors can run a fund screen with a preset FDR level and be protected against a majority of selected 'skilled' funds being lucky, without assuming normally distributed returns.","feed_headline":"Spatial-sign tests tame fund-picking false discoveries","feed_subtitle":"Spatial-sign screening keeps fund-selection FDR at target even with hidden factors and fat tails.","key_machinery":"The load-bearing object is the spatial sign $U(x)=x/\\|x\\|$ and the spatial median estimator $\\hat\\theta$ defined by the estimating equations $\\frac1T\\sum_t U(D^{-1/2}(Z_t-\\theta))=0$ together with a diagonal scaling constraint. This replaces the sample-mean location estimate, making the test statistic robust to heavy tails; the asymptotic covariance of $T^{1/2}\\hat D^{-1/2}(\\hat\\theta-\\omega\\alpha)$ is proportional to the shape matrix $R=D^{-1/2}\\Sigma D^{-1/2}$. For the latent-factor version, the spatial Kendall's tau matrix $K_Z=\\frac{2}{T(T-1)}\\sum_{i<j}U(Z_i-Z_j)U(Z_i-Z_j)^\\top$ is eigen-decomposed to estimate factors and loadings, avoiding the moment constraints that ordinary PCA needs. The FDR proof uses the Storey et al. (2004) equivalence between BH rejections and a threshold $\\hat t$, then bounds the empirical distribution of the null spatial-sign statistics by the Gaussian tail uniformly up to the critical value, using the weak-dependence condition (C4).","core_discovery":"The central claim is that the spatial-sign statistic $T_i^s=T^{1/2}\\varsigma^{1/2}\\hat\\theta_i/\\hat d_i$, whose p-value is $1-\\Phi(T_i^s)$, plugged into the BH procedure, keeps FDR at or below the target $\\gamma$ asymptotically, with $FDR_{SS-BH}\\le \\gamma N_0/N \\le \\gamma$. For the latent-factor case, factors and loadings are estimated from the spatial Kendall's tau matrix instead of the sample covariance, and the same spatial-sign construction is applied to the residualized returns, giving $FDR_{FSS-BH}\\le \\gamma N_0/N \\le \\gamma$ under conditions (C1), (C4), (C5)-(C8). The practical meaning is that the expected fraction of selected funds that are actually unskilled can be kept below the user-chosen level even when returns are heavy-tailed and common variation is driven by unobserved factors. This is achieved without assuming normality, at the cost of an elliptical, weakly dependent error model and a sparse-signal condition.","pith_inferences":["An adaptive or Storey-type threshold could plausibly be layered onto FSS-BH to raise power when the null proportion is high; inequality (15) in the proof is the uniform Gaussian approximation such an extension would need.","Because the FDR bound relies on weak residual dependence after factor removal, a stress test with block-correlated residuals or a non-elliptical skewed error distribution would be the natural next robustness check beyond the paper's scenarios.","The same spatial-sign machinery could be applied to other high-dimensional asset pricing screens, such as testing alphas of individual stocks or factor-specific performance, wherever the elliptical error assumption is plausible.","The sparsity condition $N_1\\le N^\\varpi$ means the method is designed for a small number of true skilled funds; in a market with many moderately skilled funds the procedure may still control FDR but lose power to identify them all."],"forward_implications":["A preset FDR level $\\gamma$ can be used to screen large fund universes, with the expected share of falsely selected 'lucky' funds bounded by $\\gamma N_0/N$ as sample size grows.","FSS-BH extends FDR-controlled alpha testing to settings with latent factors and heavy-tailed residuals, where PCA-based factor adjustment (F-BH) loses power or fails to control FDP.","In the paper's simulations, SS-BH and FSS-BH achieve higher true discovery proportions than D-BH and F-BH under $t$, mixture-normal, and independent-component errors, while controlling FDP around the nominal level.","The CRSP application suggests that portfolios built from funds selected by FSS-BH at $\\gamma=0.1$ outperform the S&P 500 and Sharpe-ratio-selected funds over the 1987-2017 period."],"supporting_citations":[{"why":"Supplies the asymptotic normality and global alpha test for spatial-sign statistics that SS-BH builds on.","marker":"[Zhao et al. (2024)]"},{"why":"Provides the moment-free elliptical principal component estimation via spatial Kendall's tau used to extract latent factors in FSS-BH.","marker":"[He et al. (2022)]"},{"why":"Defines the BH step that converts the p-values into FDR-controlled rejection sets.","marker":"[Benjamini and Hochberg, 1995]"},{"why":"Establishes the 'lucky fund' problem and the FDR framework for mutual fund alpha testing that this paper robustifies.","marker":"[Barras et al., 2010]"},{"why":"Gives the PCA-based factor-adjusted multiple testing procedure (F-BH) that FSS-BH extends and compares against.","marker":"[Lan and Du (2019)]"},{"why":"Provides the threshold characterization of BH rejections used in the FDR proof.","marker":"[Storey et al., 2004]"},{"why":"Used for the uniform Gaussian approximation of null statistics under weak dependence.","marker":"[Liu and Shao (2014)]"},{"why":"Introduced the multivariate-sign-based location estimation that the spatial median algorithm is adapted from.","marker":"[Feng et al. (2016)]"}],"fun_headline_variants":["Spatial-sign tests keep fund-picking FDR at target","Fund selection with FDR control via spatial signs","Spatial signs tame false discoveries in fund picks","Robust fund picking with FDR guarantee","Spatial-sign method controls fund-pick false discoveries"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The FDR bound depends on the idiosyncratic errors being elliptically distributed with independent components and only weak cross-fund correlation, and on the truly skilled funds being sparse with at least a few strong signals; if the error distribution is non-elliptical or residual correlation across funds is strong, the stated FDR guarantees are not proven.","fun_headline_variants_meta":{"raw":{"variants":["Spatial-sign tests keep fund-picking FDR at target","Fund selection with FDR control via spatial signs","Spatial signs tame false discoveries in fund picks","Robust fund picking with FDR guarantee","Spatial-sign method controls fund-pick false discoveries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000248,"raw_usage":{"total_tokens":1517,"prompt_tokens":889,"completion_tokens":628,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":551}},"tokens_in":505,"tokens_out":628,"duration_ms":5553,"temperature":1.0,"reasoning_tokens":551,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:37:37.167350+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run FSS-BH on simulated data with $T=60$, $N=200$, errors from the asymmetric independent-component model $(3-\\chi^2_3)/\\sqrt6$, and a latent factor with strong loadings; repeat enough times to estimate FDP. If the empirical FDP systematically exceeds $\\gamma=0.1$ when the $\\alpha$ signal is strong, the asymptotic FDR control is failing outside the elliptical assumption.","supporting_citations":[{"cited_title":"and Wermers, R","cited_arxiv_id":null,"evidence_quote":"Establishes the 'lucky fund' problem and the FDR framework for mutual fund alpha testing that this paper robustifies."},{"cited_title":"Double Robust high dimensional alpha test for linear factor pricing model","cited_arxiv_id":"2408.06612","evidence_quote":"Gives the PCA-based factor-adjusted multiple testing procedure (F-BH) that FSS-BH extends and compares against."}],"review_version":1}