REVIEW 2 major objections 4 minor 18 references
Empirical utility maximization under-controls the prioritized class on average because of over-optimism; a CAP threshold correction restores second-order unbiased control.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 01:19 UTC pith:ZR5V3RD6
load-bearing objection Clean cube-root diagnosis of EUM under-control as over-optimism, with usable CAP-corrected thresholds for both CiE and CiP plus training-data accuracy inference. the 2 major comments →
Tightening Control in Neyman--Pearson Linear Classification
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The EUM classifier’s empirical class-0 accuracy is second-order positively biased relative to its true accuracy at rate n^{2/3}; consequently the classifier under-controls the prioritized class on average. A CAP-corrected threshold restores second-order unbiasedness for both control-in-expectation and control-in-probability, and the same CAP construction yields consistent training-data predictors of the resulting class-specific accuracies.
What carries the argument
Cube-root asymptotic expansion of the empirical utility process together with the repeated two-fold cross-audit projection (CAP) bias estimator that rescales the half-sample over-optimism by the n^{-2/3} rate to produce the corrected threshold ˆτ_c.
Load-bearing premise
The conditional distribution of each linear score, given a suitable linear transformation of the features, must possess a bounded second derivative near the oracle threshold; without that smoothness the n^{2/3} expansions and the CAP correction fail.
What would settle it
In a simulation with the same sample sizes and biomarker distributions used in the paper, compute the average true sensitivity of the uncorrected EUM classifier and of the CAP-corrected cEUM classifier; if the former is not systematically below 0.95 while the latter is unbiased, or if the CAP accuracy bounds fail to cover at the claimed rate, the central claim is false.
If this is right
- Any EUM-style Neyman–Pearson linear classifier can be made second-order control-unbiased by a single CAP threshold adjustment without re-estimating the combination coefficients.
- Practitioners can obtain asymptotically valid lower bounds on sensitivity and specificity from the training sample alone, removing the need for a held-out validation set for performance reporting.
- The same CAP correction extends immediately to the control-in-probability formulation by replacing the nominal level ρ with the binomial quantile ρ_n, yielding a classifier whose control probability converges to the pre-specified δ.
- Because the bias diagnosis rests only on the cube-root geometry of the empirical process, analogous over-optimism corrections should apply to other non-smooth utility maximizers that share the same rate.
Where Pith is reading between the lines
- The CAP correction may remain useful even when the linear-score densities are only Hölder continuous, provided the bias rate can still be estimated by half-sample resampling; that would enlarge the practical scope beyond the paper’s smoothness conditions.
- If feature selection is later grafted onto the EUM step, the same second-order bias will reappear and will again be removable by CAP, suggesting a modular pipeline for high-dimensional Neyman–Pearson screening.
- The training-data accuracy bounds could be turned into sequential monitoring rules that decide when enough samples have been collected for a target control level, a use not discussed in the paper.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies linear Neyman–Pearson classification, where one class accuracy is constrained at a nominal level ρ and the other is maximized. It shows that the classical empirical utility maximization (EUM) classifier of Huang & Sanda (2022) systematically under-controls the prioritized class in finite samples. Using cube-root asymptotics, Proposition 2 establishes that the second-order (n^{2/3}) term in the empirical class-0 accuracy is positively biased, so the EUM threshold is second-order negatively biased for the oracle threshold (Theorem 3). The authors adapt the cross-audit projection (CAP) method to produce a corrected threshold τ̂_c that restores second-order unbiasedness, yielding the cEUM classifier for control-in-expectation and the cEUM_p classifier for control-in-probability (via a binomial quantile adjustment of the control level). Parallel CAP procedures give training-data estimators and accuracy bounds for the resulting class-specific accuracies. Simulations (four biomarker scenarios, n=100–500, 1000 replications) and a breast-cancer illustration support the theory.
Significance. The work supplies a precise asymptotic diagnosis of a practically important defect of EUM under the Neyman–Pearson constraint and a computationally feasible correction that restores second-order control under both expectation and high-probability criteria. The CAP-based performance prediction and inference procedures are a useful by-product for deployment decisions when independent validation data are unavailable. The theory rests on standard cube-root expansions (Kim & Pollard) under explicit Conditions 1–4 and is corroborated by extensive simulations; these are genuine strengths. The contribution is incremental relative to Huang & Sanda (2022) and Huang (2026) but fills a clear gap for constrained linear classification.
major comments (2)
- [§4, Corollary 7; Table 2] Corollary 7 and the surrounding discussion in §4 claim that both EUM_p and cEUM_p achieve the target probability δ asymptotically, yet the text explicitly states that a formal finite-sample (or even asymptotic second-order) advantage of cEUM_p over EUM_p “remains to be established.” Table 2 shows a clear finite-sample gap. Either a second-order expansion analogous to Proposition 4 should be supplied for the probability criterion, or the claim should be restricted to the first-order statement already proved and the finite-sample superiority of cEUM_p presented as empirical only.
- [Condition 4; §2.1–2.2] Condition 4 (bounded second derivative of the conditional distribution of the linear score) is load-bearing for the cube-root expansions, the Gaussian-process limits, and the sign of E{W_d(U)} that drive the bias diagnosis and the CAP correction. The paper does not discuss the practical consequences when this fails (e.g., discrete or multimodal biomarkers near the threshold). A short remark on robustness, or a simulation with a discontinuous density, would strengthen the applicability claim.
minor comments (4)
- [Figure 1; §1] Figure 1 caption and the surrounding text refer to “prediction sensitivity” and “empirical sensitivity”; a one-sentence reminder that “prediction” means conditional expectation on future data would help readers unfamiliar with the earlier CAP papers.
- [Eq. (13); §3.2] The transformation Q in the CAP estimator (13) is taken to be the normal quantile function “for range preservation.” A brief justification or sensitivity check would be useful, since the asymptotic theory only requires differentiability at the true value.
- [Table 3; Remark 1] In Table 3 the combination coefficients for the two control strategies differ slightly; a short note explaining that the EUM_p combination is re-optimized under ρ_n (Remark 1) would avoid confusion.
- [References; header] Typographical inconsistencies appear in the arXiv header (e.g., “arXiv:2607.03590v1 [math.ST] 3 Jul 2026”) and in a few reference entries (missing italics, incomplete page ranges). These should be cleaned before final submission.
Circularity Check
Minor self-citation of author's concurrent CAP theory and prior EUM asymptotics; NP-specific bias diagnosis and corrections remain independently derived under stated conditions.
specific steps
-
self citation load bearing
[Proposition 2 (and its proof sketch in Appendix)]
"Finally, n^{2/3}[ ˆψ_0{τ(ˆβββ), ˆβββ} − ˆψ_0{τ(βββ),βββ}] ⇝ W_0(U) and n^{2/3}[ ˆψ_1{τ(ˆβββ), ˆβββ} − ψ_1{τ(ˆβββ), ˆβββ}− ˆψ_1{τ(βββ),βββ}+ψ_1{τ(βββ),βββ}] ⇝W_1(U), where U = arg max_g Z(g) and E{W_d(U)}>0 for d=0,1. ... can be established by adapting the proof of Huang (2026, theorem 3)."
The positivity of the second-order bias (the over-optimism that motivates the entire refinement) is obtained by adapting a theorem from the author's concurrent CAP paper rather than being re-derived from scratch under the NP criterion. The adaptation is legitimate and the surrounding expansions are self-contained, yet the load-bearing sign result is imported via self-citation.
full rationale
The paper's central chain (decomposition of empirical accuracy bias in Prop. 1, cube-root weak convergence and positive bias of the second term for EUM in Prop. 2, threshold expansion in Thm. 3, CAP correction yielding second-order unbiasedness in Prop. 4/Cor. 5, and CiP extension in Cor. 6-7) is developed from first-order and higher-order expansions under Conditions 1-4. The n^{2/3} rate and Gaussian-process limits rest on the external Kim-Pollard cube-root theory plus extensions of the author's own prior EUM results (Huang-Sanda 2022); the CAP bias-correction device is adapted from the author's concurrent CAP paper (Huang 2026). These self-citations supply technical lemmas and the sign of E{W_d(U)}, but the NP-specific under-control claim, the construction of cEUM/cEUM_p, and the training-data accuracy predictors are not forced by definition, by a fitted free parameter, or by a uniqueness theorem that forbids alternatives. No quantity is fitted to the target control level and then re-labeled a prediction; r=16 is a fixed algorithmic choice. Simulations and the breast-cancer illustration supply external checks. The self-citation is therefore real but not load-bearing in the circular sense; score 2 is proportionate.
Axiom & Free-Parameter Ledger
free parameters (1)
- number of CAP repetitions r =
16
axioms (5)
- domain assumption Condition 1: n1/n0 converges to a finite positive constant
- domain assumption Condition 2: unique maximizer of φ(b) with strict separation
- domain assumption Condition 3: density f0{τ(β),β} exists and is strictly positive
- domain assumption Condition 4: existence of linear transformation Bd such that the conditional distribution of the score has bounded second derivative near the oracle
- standard math Cube-root asymptotics of empirical processes (Kim & Pollard 1990)
invented entities (1)
-
cEUM / cEUM_p classifiers (CAP-corrected thresholds)
no independent evidence
read the original abstract
Neyman--Pearson classification prioritizes one class by constraining its accuracy above a prespecified level, and then takes the accuracy of the other class as the utility objective. This paradigm is well suited for disease screening and diagnosis, among other applications. Statistical learning under this framework is complicated since classifier performance determines its acceptability. Furthermore, no learned classifier that is consistent for the oracle classifier can guarantee satisfaction of the control constraint in finite samples. Classical learning theory targets a control-relaxed empirical utility maximization (EUM) classifier. However, even the EUM classifier fails to achieve the desired control level on average. We conjecture that this under-control phenomenon is a manifestation of the over-optimism bias well known in standard statistical learning, and develop asymptotic theory to confirm it. Motivated by this insight, we propose refined learning procedures under two accuracy control strategies for the prioritized class: one controlling accuracy in expectation and the other with high probability. We further develop training-data-based methods to predict and infer class-specific accuracies of the resulting classifiers. Simulation studies demonstrate favorable finite-sample performance, and we illustrate the proposed methods with an application to cancer detection.
Figures
Reference graph
Works this paper leans on
-
[1]
, Howse, J
Cannon, A. , Howse, J. and Scovel, C. (2002). Learning with the Neyman--Pearson and min-max criteria. Los Alamos National Laboratory Technical Report LA-UR-02-2951
2002
-
[2]
Catalona, W. J. , Partin, A. W. , Slawin, K. M. , Brawer, M. K , Flanigan, R. C. , Patel, A. , Richie, J. P. , deKernion, J. B. , Walsh, P. C. , Scardino, P. T. , Lange, P. H. , Subong, E. N. , Parson, R. E. , Gasior, G. H. , Loveland, K. G. and Southwick, P. C. (1998). Use of the percentage of free prostate-specific antigen to enhance differentiation of ...
1998
-
[3]
Greenhouse, S. W. and Mantel, N. (1950). The evaluation of diagnostic tests. Biometrics 6 399--412
1950
-
[4]
Huang, Y. (2026). Cross-audit projection for model risk prediction. arXiv 2607.02328
Pith/arXiv arXiv 2026
-
[5]
, Parakati, I
Huang, Y. , Parakati, I. , Patil, D. H. and Sanda, M. G. (2023). Interval estimation for operating characteristic of continuous biomarkers with controlled sensitivity or specificity. Stat. Sin. 33 193--214
2023
-
[6]
and Sanda, M
Huang, Y. and Sanda, M. G. (2022). Linear biomarker combination for constrained classification. Ann. Statist. 50 2793--2815
2022
-
[7]
and Pollard, D
Kim, J. and Pollard, D. (1990). Cube root asymptotics. Ann. Statist. 18 191--219
1990
-
[8]
, Carone, M
Meisner, A. , Carone, M. , Pepe, M. S. and Kerr, K. F. (2021). Combining biomarkers by maximizing the true positive rate for a fixed false positive rate. Biom. J. 63 1223--1240
2021
-
[9]
, Pereira, J
Patricio, M. , Pereira, J. , Crisostomo, J. et al. (2018). Using resistin, glucose, age, and BMI to predict the presence of breast cancer. BMC Cancer 18 29
2018
-
[10]
Pepe, M. S. , Cai, T. and Longton, G. (2006). Combining predictors for classification using the area under the receiver operating characteristic curve. Biometrics 62 221--229
2006
-
[11]
and Tong, X
Rigollet, P. and Tong, X. (2011). Neyman--Pearson classification, convexity and stochastic constraints. J. Mach. Learn. Res. 12 2831--2855
2011
-
[12]
Sanda, M. G. , Feng, Z. , Howard, D. H. , Tomlins, S. A. , Sokoll, L. J. , Chan, D. W. , Regan, M. M. , Groskopf, J. , Chipman, J. , Patil, D. H. , Salami, S. S. , Scherr, D. S. , Kagan, J. , Srivastava, S. , Thompson, I. M., Jr , Siddiqui, J. , Fan, J. , Joon, A. Y. , Bantis, L. E. , Rubin, M. A. , Chinnayian, A. M. , Wei, J. T. and the EDRN-PCA3 Study G...
2017
-
[13]
and Nowak, R
Scott, C. and Nowak, R. (2005). A Neyman--Pearson approach to statistical learning. IEEE Trans. Inf. Theory 51 3806--3819
2005
-
[14]
Skates, S. J. , Horick, N. , Yu, Y. , Xu, F.-J. , Berchuck, A. , Havrilesky, L. J. , de Bruijn, H. W. , van der Zee, A. G. , Woolas, R. P. , Jacobs, I. J. , Zhang, Z. and Bast, R. C. Jr (2004). Preoperative sensitivity and specificity for early-stage ovarian cancer when combining cancer antigen CA-125II, CA 15-3, CA 72-4, and macrophage colony-stimulating...
2004
-
[15]
, Feng, Y
Tong, X. , Feng, Y. and Li, J. J. (2018). Neyman--Pearson classification algorithms and NP receiver operating characteristics. Sci. Adv. 4 eaao1659
2018
-
[16]
, Feng, Y
Tong, X. , Feng, Y. and Zhao, A. (2016). A survey on Neyman--Pearson classification and suggestions for future research. Wiley Interdiscip. Rev. Comput. Stat. 8 64--81
2016
-
[17]
, Xia, L
Tong, X. , Xia, L. , Wang, J. and Feng, Y. (2020). Neyman--Pearson classification: parametrics and sample size requirement. J. Mach. Learn. Res. 21 1--18
2020
-
[18]
Wang, J. , Xia, L. , Bao, Z. and Tong, X. (2022). Non-splitting Neyman--Pearson classifiers. arXiv 2112.00329
Pith/arXiv arXiv 2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.