Pith. sign in

REVIEW 2 major objections 4 minor 18 references

Empirical utility maximization under-controls the prioritized class on average because of over-optimism; a CAP threshold correction restores second-order unbiased control.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 01:19 UTC pith:ZR5V3RD6

load-bearing objection Clean cube-root diagnosis of EUM under-control as over-optimism, with usable CAP-corrected thresholds for both CiE and CiP plus training-data accuracy inference. the 2 major comments →

arxiv 2607.03590 v1 pith:ZR5V3RD6 submitted 2026-07-03 math.ST stat.MEstat.MLstat.TH

Tightening Control in Neyman--Pearson Linear Classification

classification math.ST stat.MEstat.MLstat.TH MSC 62H3062G2062C12
keywords Neyman–Pearson classificationempirical utility maximizationover-optimism biascube-root asymptoticscross-audit projectioncontrol-in-expectationcontrol-in-probabilitysecond-order asymptotics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Neyman–Pearson linear classification keeps one class’s accuracy above a fixed level (e.g., 95 % sensitivity) while maximizing the other class’s accuracy. The classical empirical-utility-maximization (EUM) procedure matches the control level on the training sample but systematically falls short of it on new data. The paper shows that this under-control is second-order positive bias of order n^{-2/3} that arises from the same over-optimism that inflates training performance in ordinary statistical learning. By adapting a cross-audit projection (CAP) correction to the threshold, the authors construct two refined classifiers—cEUM (control in expectation) and cEUM_p (control with high probability)—whose class-0 accuracy is second-order unbiased for the target. The same CAP machinery supplies training-data estimates and confidence bounds for the class-specific accuracies of the corrected classifiers, so practitioners can assess deployment risk without a separate validation set. Simulations and a breast-cancer example confirm that the bias is real, that the correction restores the nominal control level, and that the accuracy predictions track true performance closely.

Core claim

The EUM classifier’s empirical class-0 accuracy is second-order positively biased relative to its true accuracy at rate n^{2/3}; consequently the classifier under-controls the prioritized class on average. A CAP-corrected threshold restores second-order unbiasedness for both control-in-expectation and control-in-probability, and the same CAP construction yields consistent training-data predictors of the resulting class-specific accuracies.

What carries the argument

Cube-root asymptotic expansion of the empirical utility process together with the repeated two-fold cross-audit projection (CAP) bias estimator that rescales the half-sample over-optimism by the n^{-2/3} rate to produce the corrected threshold ˆτ_c.

Load-bearing premise

The conditional distribution of each linear score, given a suitable linear transformation of the features, must possess a bounded second derivative near the oracle threshold; without that smoothness the n^{2/3} expansions and the CAP correction fail.

What would settle it

In a simulation with the same sample sizes and biomarker distributions used in the paper, compute the average true sensitivity of the uncorrected EUM classifier and of the CAP-corrected cEUM classifier; if the former is not systematically below 0.95 while the latter is unbiased, or if the CAP accuracy bounds fail to cover at the claimed rate, the central claim is false.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Any EUM-style Neyman–Pearson linear classifier can be made second-order control-unbiased by a single CAP threshold adjustment without re-estimating the combination coefficients.
  • Practitioners can obtain asymptotically valid lower bounds on sensitivity and specificity from the training sample alone, removing the need for a held-out validation set for performance reporting.
  • The same CAP correction extends immediately to the control-in-probability formulation by replacing the nominal level ρ with the binomial quantile ρ_n, yielding a classifier whose control probability converges to the pre-specified δ.
  • Because the bias diagnosis rests only on the cube-root geometry of the empirical process, analogous over-optimism corrections should apply to other non-smooth utility maximizers that share the same rate.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The CAP correction may remain useful even when the linear-score densities are only Hölder continuous, provided the bias rate can still be estimated by half-sample resampling; that would enlarge the practical scope beyond the paper’s smoothness conditions.
  • If feature selection is later grafted onto the EUM step, the same second-order bias will reappear and will again be removable by CAP, suggesting a modular pipeline for high-dimensional Neyman–Pearson screening.
  • The training-data accuracy bounds could be turned into sequential monitoring rules that decide when enough samples have been collected for a target control level, a use not discussed in the paper.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper studies linear Neyman–Pearson classification, where one class accuracy is constrained at a nominal level ρ and the other is maximized. It shows that the classical empirical utility maximization (EUM) classifier of Huang & Sanda (2022) systematically under-controls the prioritized class in finite samples. Using cube-root asymptotics, Proposition 2 establishes that the second-order (n^{2/3}) term in the empirical class-0 accuracy is positively biased, so the EUM threshold is second-order negatively biased for the oracle threshold (Theorem 3). The authors adapt the cross-audit projection (CAP) method to produce a corrected threshold τ̂_c that restores second-order unbiasedness, yielding the cEUM classifier for control-in-expectation and the cEUM_p classifier for control-in-probability (via a binomial quantile adjustment of the control level). Parallel CAP procedures give training-data estimators and accuracy bounds for the resulting class-specific accuracies. Simulations (four biomarker scenarios, n=100–500, 1000 replications) and a breast-cancer illustration support the theory.

Significance. The work supplies a precise asymptotic diagnosis of a practically important defect of EUM under the Neyman–Pearson constraint and a computationally feasible correction that restores second-order control under both expectation and high-probability criteria. The CAP-based performance prediction and inference procedures are a useful by-product for deployment decisions when independent validation data are unavailable. The theory rests on standard cube-root expansions (Kim & Pollard) under explicit Conditions 1–4 and is corroborated by extensive simulations; these are genuine strengths. The contribution is incremental relative to Huang & Sanda (2022) and Huang (2026) but fills a clear gap for constrained linear classification.

major comments (2)
  1. [§4, Corollary 7; Table 2] Corollary 7 and the surrounding discussion in §4 claim that both EUM_p and cEUM_p achieve the target probability δ asymptotically, yet the text explicitly states that a formal finite-sample (or even asymptotic second-order) advantage of cEUM_p over EUM_p “remains to be established.” Table 2 shows a clear finite-sample gap. Either a second-order expansion analogous to Proposition 4 should be supplied for the probability criterion, or the claim should be restricted to the first-order statement already proved and the finite-sample superiority of cEUM_p presented as empirical only.
  2. [Condition 4; §2.1–2.2] Condition 4 (bounded second derivative of the conditional distribution of the linear score) is load-bearing for the cube-root expansions, the Gaussian-process limits, and the sign of E{W_d(U)} that drive the bias diagnosis and the CAP correction. The paper does not discuss the practical consequences when this fails (e.g., discrete or multimodal biomarkers near the threshold). A short remark on robustness, or a simulation with a discontinuous density, would strengthen the applicability claim.
minor comments (4)
  1. [Figure 1; §1] Figure 1 caption and the surrounding text refer to “prediction sensitivity” and “empirical sensitivity”; a one-sentence reminder that “prediction” means conditional expectation on future data would help readers unfamiliar with the earlier CAP papers.
  2. [Eq. (13); §3.2] The transformation Q in the CAP estimator (13) is taken to be the normal quantile function “for range preservation.” A brief justification or sensitivity check would be useful, since the asymptotic theory only requires differentiability at the true value.
  3. [Table 3; Remark 1] In Table 3 the combination coefficients for the two control strategies differ slightly; a short note explaining that the EUM_p combination is re-optimized under ρ_n (Remark 1) would avoid confusion.
  4. [References; header] Typographical inconsistencies appear in the arXiv header (e.g., “arXiv:2607.03590v1 [math.ST] 3 Jul 2026”) and in a few reference entries (missing italics, incomplete page ranges). These should be cleaned before final submission.

Circularity Check

1 steps flagged

Minor self-citation of author's concurrent CAP theory and prior EUM asymptotics; NP-specific bias diagnosis and corrections remain independently derived under stated conditions.

specific steps
  1. self citation load bearing [Proposition 2 (and its proof sketch in Appendix)]
    "Finally, n^{2/3}[ ˆψ_0{τ(ˆβββ), ˆβββ} − ˆψ_0{τ(βββ),βββ}] ⇝ W_0(U) and n^{2/3}[ ˆψ_1{τ(ˆβββ), ˆβββ} − ψ_1{τ(ˆβββ), ˆβββ}− ˆψ_1{τ(βββ),βββ}+ψ_1{τ(βββ),βββ}] ⇝W_1(U), where U = arg max_g Z(g) and E{W_d(U)}>0 for d=0,1. ... can be established by adapting the proof of Huang (2026, theorem 3)."

    The positivity of the second-order bias (the over-optimism that motivates the entire refinement) is obtained by adapting a theorem from the author's concurrent CAP paper rather than being re-derived from scratch under the NP criterion. The adaptation is legitimate and the surrounding expansions are self-contained, yet the load-bearing sign result is imported via self-citation.

full rationale

The paper's central chain (decomposition of empirical accuracy bias in Prop. 1, cube-root weak convergence and positive bias of the second term for EUM in Prop. 2, threshold expansion in Thm. 3, CAP correction yielding second-order unbiasedness in Prop. 4/Cor. 5, and CiP extension in Cor. 6-7) is developed from first-order and higher-order expansions under Conditions 1-4. The n^{2/3} rate and Gaussian-process limits rest on the external Kim-Pollard cube-root theory plus extensions of the author's own prior EUM results (Huang-Sanda 2022); the CAP bias-correction device is adapted from the author's concurrent CAP paper (Huang 2026). These self-citations supply technical lemmas and the sign of E{W_d(U)}, but the NP-specific under-control claim, the construction of cEUM/cEUM_p, and the training-data accuracy predictors are not forced by definition, by a fitted free parameter, or by a uniqueness theorem that forbids alternatives. No quantity is fitted to the target control level and then re-labeled a prediction; r=16 is a fixed algorithmic choice. Simulations and the breast-cancer illustration supply external checks. The self-citation is therefore real but not load-bearing in the circular sense; score 2 is proportionate.

Axiom & Free-Parameter Ledger

1 free parameters · 5 axioms · 1 invented entities

The central claims rest on four regularity conditions (sample-size balance, identifiability, quantile smoothness, conditional-distribution smoothness) that are standard for cube-root asymptotics of empirical processes, plus the algorithmic choice of CAP repetitions. No free parameters are fitted to the target accuracies; the binomial quantile for CiP is a known function of the nominal level and sample size.

free parameters (1)
  • number of CAP repetitions r = 16
    Fixed at r = 16 in all numerical work; affects the variance of the bias estimate but is not optimized to the data.
axioms (5)
  • domain assumption Condition 1: n1/n0 converges to a finite positive constant
    Standard balanced-sample assumption used throughout the asymptotic expansions (Section 2).
  • domain assumption Condition 2: unique maximizer of φ(b) with strict separation
    Identifiability needed for consistency of ˆβ (Section 2).
  • domain assumption Condition 3: density f0{τ(β),β} exists and is strictly positive
    Ensures the quantile map is differentiable (Section 2).
  • domain assumption Condition 4: existence of linear transformation Bd such that the conditional distribution of the score has bounded second derivative near the oracle
    Key smoothness for cube-root expansions and Gaussian-process limits (Section 2).
  • standard math Cube-root asymptotics of empirical processes (Kim & Pollard 1990)
    Background result used to obtain the n^{2/3} rates and argmax limits.
invented entities (1)
  • cEUM / cEUM_p classifiers (CAP-corrected thresholds) no independent evidence
    purpose: Second-order bias-corrected versions of the EUM classifier under CiE and CiP criteria
    Defined by the explicit correction formula (12) and its ρ_n analogue; no independent physical existence claimed.

pith-pipeline@v1.1.0-grok45 · 21434 in / 2733 out tokens · 18909 ms · 2026-07-12T01:19:53.672570+00:00 · methodology

0 comments
read the original abstract

Neyman--Pearson classification prioritizes one class by constraining its accuracy above a prespecified level, and then takes the accuracy of the other class as the utility objective. This paradigm is well suited for disease screening and diagnosis, among other applications. Statistical learning under this framework is complicated since classifier performance determines its acceptability. Furthermore, no learned classifier that is consistent for the oracle classifier can guarantee satisfaction of the control constraint in finite samples. Classical learning theory targets a control-relaxed empirical utility maximization (EUM) classifier. However, even the EUM classifier fails to achieve the desired control level on average. We conjecture that this under-control phenomenon is a manifestation of the over-optimism bias well known in standard statistical learning, and develop asymptotic theory to confirm it. Motivated by this insight, we propose refined learning procedures under two accuracy control strategies for the prioritized class: one controlling accuracy in expectation and the other with high probability. We further develop training-data-based methods to predict and infer class-specific accuracies of the resulting classifiers. Simulation studies demonstrate favorable finite-sample performance, and we illustrate the proposed methods with an application to cancer detection.

Figures

Figures reproduced from arXiv: 2607.03590 by Yijian Huang.

Figure 1
Figure 1. Figure 1: Simulation results for the EUM classifier obtained under Sce [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

18 extracted references · 2 linked inside Pith

  1. [1]

    , Howse, J

    Cannon, A. , Howse, J. and Scovel, C. (2002). Learning with the Neyman--Pearson and min-max criteria. Los Alamos National Laboratory Technical Report LA-UR-02-2951

  2. [2]

    Catalona, W. J. , Partin, A. W. , Slawin, K. M. , Brawer, M. K , Flanigan, R. C. , Patel, A. , Richie, J. P. , deKernion, J. B. , Walsh, P. C. , Scardino, P. T. , Lange, P. H. , Subong, E. N. , Parson, R. E. , Gasior, G. H. , Loveland, K. G. and Southwick, P. C. (1998). Use of the percentage of free prostate-specific antigen to enhance differentiation of ...

  3. [3]

    Greenhouse, S. W. and Mantel, N. (1950). The evaluation of diagnostic tests. Biometrics 6 399--412

  4. [4]

    Huang, Y. (2026). Cross-audit projection for model risk prediction. arXiv 2607.02328

  5. [5]

    , Parakati, I

    Huang, Y. , Parakati, I. , Patil, D. H. and Sanda, M. G. (2023). Interval estimation for operating characteristic of continuous biomarkers with controlled sensitivity or specificity. Stat. Sin. 33 193--214

  6. [6]

    and Sanda, M

    Huang, Y. and Sanda, M. G. (2022). Linear biomarker combination for constrained classification. Ann. Statist. 50 2793--2815

  7. [7]

    and Pollard, D

    Kim, J. and Pollard, D. (1990). Cube root asymptotics. Ann. Statist. 18 191--219

  8. [8]

    , Carone, M

    Meisner, A. , Carone, M. , Pepe, M. S. and Kerr, K. F. (2021). Combining biomarkers by maximizing the true positive rate for a fixed false positive rate. Biom. J. 63 1223--1240

  9. [9]

    , Pereira, J

    Patricio, M. , Pereira, J. , Crisostomo, J. et al. (2018). Using resistin, glucose, age, and BMI to predict the presence of breast cancer. BMC Cancer 18 29

  10. [10]

    Pepe, M. S. , Cai, T. and Longton, G. (2006). Combining predictors for classification using the area under the receiver operating characteristic curve. Biometrics 62 221--229

  11. [11]

    and Tong, X

    Rigollet, P. and Tong, X. (2011). Neyman--Pearson classification, convexity and stochastic constraints. J. Mach. Learn. Res. 12 2831--2855

  12. [12]

    Sanda, M. G. , Feng, Z. , Howard, D. H. , Tomlins, S. A. , Sokoll, L. J. , Chan, D. W. , Regan, M. M. , Groskopf, J. , Chipman, J. , Patil, D. H. , Salami, S. S. , Scherr, D. S. , Kagan, J. , Srivastava, S. , Thompson, I. M., Jr , Siddiqui, J. , Fan, J. , Joon, A. Y. , Bantis, L. E. , Rubin, M. A. , Chinnayian, A. M. , Wei, J. T. and the EDRN-PCA3 Study G...

  13. [13]

    and Nowak, R

    Scott, C. and Nowak, R. (2005). A Neyman--Pearson approach to statistical learning. IEEE Trans. Inf. Theory 51 3806--3819

  14. [14]

    Skates, S. J. , Horick, N. , Yu, Y. , Xu, F.-J. , Berchuck, A. , Havrilesky, L. J. , de Bruijn, H. W. , van der Zee, A. G. , Woolas, R. P. , Jacobs, I. J. , Zhang, Z. and Bast, R. C. Jr (2004). Preoperative sensitivity and specificity for early-stage ovarian cancer when combining cancer antigen CA-125II, CA 15-3, CA 72-4, and macrophage colony-stimulating...

  15. [15]

    , Feng, Y

    Tong, X. , Feng, Y. and Li, J. J. (2018). Neyman--Pearson classification algorithms and NP receiver operating characteristics. Sci. Adv. 4 eaao1659

  16. [16]

    , Feng, Y

    Tong, X. , Feng, Y. and Zhao, A. (2016). A survey on Neyman--Pearson classification and suggestions for future research. Wiley Interdiscip. Rev. Comput. Stat. 8 64--81

  17. [17]

    , Xia, L

    Tong, X. , Xia, L. , Wang, J. and Feng, Y. (2020). Neyman--Pearson classification: parametrics and sample size requirement. J. Mach. Learn. Res. 21 1--18

  18. [18]

    , Xia, L

    Wang, J. , Xia, L. , Bao, Z. and Tong, X. (2022). Non-splitting Neyman--Pearson classifiers. arXiv 2112.00329