{"id":"8b4f215f-1ffe-427b-b837-39ae144292c4","arxiv_id":"2607.03590","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"Cube-root asymptotics explain EUM under-control in Neyman–Pearson linear classification; CAP-based threshold corrections restore control-in-expectation and control-in-probability with training-data accuracy inference.","lead":"The paper shows that standard empirical utility maximization for Neyman–Pearson linear classifiers systematically under-controls the prioritized class because of over-optimism bias, and supplies corrected thresholds plus training-data accuracy prediction under both expectation and high-probability control. This matters for medical screening and diagnosis, where one error type must be tightly constrained.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly isolates Condition 4 as the weakest regularity assumption and correctly notes that it is standard for cube-root asymptotics. No stronger load-bearing flaw appears: the bias diagnosis, the CAP correction, and the control-in-probability extension are all consequences of the same expansions once that smoothness holds. The recommended verification simply re-checks the key expansion that underpins every subsequent claim; if it holds, the paper’s central argument stands. Hence the CONDITIONAL verdict with high confidence remains appropriate and needs no adjustment.","tokens_in":17375,"tokens_out":432,"duration_ms":4250,"concrete_test":"Independently re-derive the second-order expansion of ˆ\tau(ˆeta) – \tau(ˆeta) in Theorem 3 starting from the empirical-process representation of ˆ\tau(b) given in Huang & Sanda (2022, Thm 3.4) and the weak-convergence statement of Proposition 2, without invoking the CAP construction; verify that the bias term is exactly –f_0{\tau(eta),eta}^{–1} W_0(U) + o_p(n^{–2/3}).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (second-order positive bias of EUM class-0 empirical accuracy at the n^{2/3} rate, and restoration of second-order unbiasedness by the CAP-corrected threshold) is internally consistent under the stated conditions. Proposition 2, Theorem 3, Proposition 4 and Corollary 7 follow from standard cube-root expansions once Conditions 1–4 hold; the Gaussian-process limits and the sign of E{W_d(U)} are obtained by the usual argmax continuous-mapping arguments. Condition 4 is the natural smoothness hypothesis for this literature and is not an internal contradiction. Finite-sample claims that remain only empirical (advantage of cEUM_p over EUM_p) are already flagged by the reader and do not undermine the asymptotic diagnosis.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies linear Neyman–Pearson classification, where one class accuracy is constrained at a nominal level ρ and the other is maximized. It shows that the classical empirical utility maximization (EUM) classifier of Huang & Sanda (2022) systematically under-controls the prioritized class in finite samples. Using cube-root asymptotics, Proposition 2 establishes that the second-order (n^{2/3}) term in the empirical class-0 accuracy is positively biased, so the EUM threshold is second-order negatively biased for the oracle threshold (Theorem 3). The authors adapt the cross-audit projection (CAP) method to produce a corrected threshold τ̂_c that restores second-order unbiasedness, yielding the cEUM classifier for control-in-expectation and the cEUM_p classifier for control-in-probability (via a binomial quantile adjustment of the control level). Parallel CAP procedures give training-data estimators and accuracy bounds for the resulting class-specific accuracies. Simulations (four biomarker scenarios, n=100–500, 1000 replications) and a breast-cancer illustration support the theory.","tokens_in":17575,"tokens_out":1059,"duration_ms":8738,"significance":"The work supplies a precise asymptotic diagnosis of a practically important defect of EUM under the Neyman–Pearson constraint and a computationally feasible correction that restores second-order control under both expectation and high-probability criteria. The CAP-based performance prediction and inference procedures are a useful by-product for deployment decisions when independent validation data are unavailable. The theory rests on standard cube-root expansions (Kim & Pollard) under explicit Conditions 1–4 and is corroborated by extensive simulations; these are genuine strengths. The contribution is incremental relative to Huang & Sanda (2022) and Huang (2026) but fills a clear gap for constrained linear classification.","major_comments":[{"comment":"Corollary 7 and the surrounding discussion in §4 claim that both EUM_p and cEUM_p achieve the target probability δ asymptotically, yet the text explicitly states that a formal finite-sample (or even asymptotic second-order) advantage of cEUM_p over EUM_p “remains to be established.” Table 2 shows a clear finite-sample gap. Either a second-order expansion analogous to Proposition 4 should be supplied for the probability criterion, or the claim should be restricted to the first-order statement already proved and the finite-sample superiority of cEUM_p presented as empirical only.","section":"§4, Corollary 7; Table 2"},{"comment":"Condition 4 (bounded second derivative of the conditional distribution of the linear score) is load-bearing for the cube-root expansions, the Gaussian-process limits, and the sign of E{W_d(U)} that drive the bias diagnosis and the CAP correction. The paper does not discuss the practical consequences when this fails (e.g., discrete or multimodal biomarkers near the threshold). A short remark on robustness, or a simulation with a discontinuous density, would strengthen the applicability claim.","section":"Condition 4; §2.1–2.2"}],"minor_comments":[{"comment":"Figure 1 caption and the surrounding text refer to “prediction sensitivity” and “empirical sensitivity”; a one-sentence reminder that “prediction” means conditional expectation on future data would help readers unfamiliar with the earlier CAP papers.","section":"Figure 1; §1"},{"comment":"The transformation Q in the CAP estimator (13) is taken to be the normal quantile function “for range preservation.” A brief justification or sensitivity check would be useful, since the asymptotic theory only requires differentiability at the true value.","section":"Eq. (13); §3.2"},{"comment":"In Table 3 the combination coefficients for the two control strategies differ slightly; a short note explaining that the EUM_p combination is re-optimized under ρ_n (Remark 1) would avoid confusion.","section":"Table 3; Remark 1"},{"comment":"Typographical inconsistencies appear in the arXiv header (e.g., “arXiv:2607.03590v1 [math.ST] 3 Jul 2026”) and in a few reference entries (missing italics, incomplete page ranges). These should be cleaned before final submission.","section":"References; header"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a natural and carefully executed sequel to the author’s 2022 Annals paper and the concurrent CAP paper. The technical overlap is properly cited; I see no novelty or citation issues. Fit for a solid statistics journal is good; the main limitation is that the finite-sample superiority of cEUM_p is left empirical, which is already flagged and does not warrant major revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real contribution is the second-order asymptotic diagnosis: EUM’s empirical class-0 accuracy is positively biased at the n^{2/3} rate (Proposition 2), so the classifier under-controls the prioritized class on average. That matches the simulation pattern everyone has seen and turns the over-optimism intuition into a precise expansion. From there Huang builds CAP-corrected thresholds that restore second-order unbiasedness for both control-in-expectation (cEUM) and control-in-probability (cEUM_p), plus matching training-data estimators and accuracy bounds (Theorem 3, Prop 4, Corollaries 5–7).\n\nWhat works well: the higher-order theory is careful and sits on standard cube-root machinery (Kim–Pollard plus the author’s earlier CAP paper). Conditions 1–4 are explicit; the weak-convergence statements are clean. Simulations (1 000 reps, four scenarios, n = 100–500) and the breast-cancer illustration line up with the asymptotics. The comparison with eLDA is fair and shows the nonparametric correction is more stable under misspecification. Citation pattern is appropriate—no padding.\n\nSoft spots are minor and already flagged. Condition 4 (conditional second-derivative smoothness near the threshold) is the usual regularity for this literature; if densities are only Hölder the rates fail, but that is not a hidden flaw. The finite-sample edge of cEUM_p over EUM_p is still only empirical. No code is shipped. None of these undercut the central expansion or the corrected procedures.\n\nThis is for people who actually build or evaluate constrained linear classifiers in diagnostics or screening. It is not a broad theory paper, but it solves a concrete, observed problem with usable algorithms and honest inference. The math and data are solid enough that a serious editor should send it to referees. I would cite the bias diagnosis and the CAP correction if I were working in the area.","headline":"Clean cube-root diagnosis of EUM under-control as over-optimism, with usable CAP-corrected thresholds for both CiE and CiP plus training-data accuracy inference.","tokens_in":18160,"tokens_out":486,"would_cite":true,"duration_ms":4751,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H30","62G20","62C12"],"pacs":[],"model":"grok-4.5","headline":"Empirical utility maximization under-controls the prioritized class on average because of over-optimism; a CAP threshold correction restores second-order unbiased control.","keywords":["Neyman–Pearson classification","empirical utility maximization","over-optimism bias","cube-root asymptotics","cross-audit projection","control-in-expectation","control-in-probability","second-order asymptotics"],"falsifier":"In a simulation with the same sample sizes and biomarker distributions used in the paper, compute the average true sensitivity of the uncorrected EUM classifier and of the CAP-corrected cEUM classifier; if the former is not systematically below 0.95 while the latter is unbiased, or if the CAP accuracy bounds fail to cover at the claimed rate, the central claim is false.","tokens_in":18272,"feed_emoji":"🎯","tokens_out":1053,"duration_ms":8450,"temperature":0.7,"pith_summary":"Neyman–Pearson linear classification keeps one class’s accuracy above a fixed level (e.g., 95 % sensitivity) while maximizing the other class’s accuracy. The classical empirical-utility-maximization (EUM) procedure matches the control level on the training sample but systematically falls short of it on new data. The paper shows that this under-control is second-order positive bias of order n^{-2/3} that arises from the same over-optimism that inflates training performance in ordinary statistical learning. By adapting a cross-audit projection (CAP) correction to the threshold, the authors construct two refined classifiers—cEUM (control in expectation) and cEUM_p (control with high probability)—whose class-0 accuracy is second-order unbiased for the target. The same CAP machinery supplies training-data estimates and confidence bounds for the class-specific accuracies of the corrected classifiers, so practitioners can assess deployment risk without a separate validation set. Simulations and a breast-cancer example confirm that the bias is real, that the correction restores the nominal control level, and that the accuracy predictions track true performance closely.","feed_headline":"EUM under-controls; CAP threshold fixes the bias","feed_subtitle":"Second-order over-optimism is diagnosed and corrected so that sensitivity stays on target","key_machinery":"Cube-root asymptotic expansion of the empirical utility process together with the repeated two-fold cross-audit projection (CAP) bias estimator that rescales the half-sample over-optimism by the n^{-2/3} rate to produce the corrected threshold ˆτ_c.","core_discovery":"The EUM classifier’s empirical class-0 accuracy is second-order positively biased relative to its true accuracy at rate n^{2/3}; consequently the classifier under-controls the prioritized class on average. A CAP-corrected threshold restores second-order unbiasedness for both control-in-expectation and control-in-probability, and the same CAP construction yields consistent training-data predictors of the resulting class-specific accuracies.","pith_inferences":["The CAP correction may remain useful even when the linear-score densities are only Hölder continuous, provided the bias rate can still be estimated by half-sample resampling; that would enlarge the practical scope beyond the paper’s smoothness conditions.","If feature selection is later grafted onto the EUM step, the same second-order bias will reappear and will again be removable by CAP, suggesting a modular pipeline for high-dimensional Neyman–Pearson screening.","The training-data accuracy bounds could be turned into sequential monitoring rules that decide when enough samples have been collected for a target control level, a use not discussed in the paper."],"forward_implications":["Any EUM-style Neyman–Pearson linear classifier can be made second-order control-unbiased by a single CAP threshold adjustment without re-estimating the combination coefficients.","Practitioners can obtain asymptotically valid lower bounds on sensitivity and specificity from the training sample alone, removing the need for a held-out validation set for performance reporting.","The same CAP correction extends immediately to the control-in-probability formulation by replacing the nominal level ρ with the binomial quantile ρ_n, yielding a classifier whose control probability converges to the pre-specified δ.","Because the bias diagnosis rests only on the cube-root geometry of the empirical process, analogous over-optimism corrections should apply to other non-smooth utility maximizers that share the same rate."],"fun_headline_variants":["EUM under-controls from second-order bias; CAP threshold corrects it","CAP restores second-order unbiased class-0 control for EUM","Over-optimism makes EUM under-control; CAP calibrates the threshold","CAP-corrected thresholds fix finite-sample under-control in NP learning","Training CAP predictors recover unbiased class accuracies after EUM"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The conditional distribution of each linear score, given a suitable linear transformation of the features, must possess a bounded second derivative near the oracle threshold; without that smoothness the n^{2/3} expansions and the CAP correction fail.","fun_headline_variants_meta":{"raw":{"variants":["EUM under-controls from second-order bias; CAP threshold corrects it","CAP restores second-order unbiased class-0 control for EUM","Over-optimism makes EUM under-control; CAP calibrates the threshold","CAP-corrected thresholds fix finite-sample under-control in NP learning","Training CAP predictors recover unbiased class accuracies after EUM"]},"model":"grok-4.5","effort":"low","cost_usd":0.00427,"raw_usage":{"total_tokens":1281,"prompt_tokens":758,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":42700000,"prompt_tokens_details":{"text_tokens":758,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":446,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":758,"tokens_out":77,"duration_ms":3659,"temperature":1.0,"reasoning_tokens":446,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T01:19:53.672570+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"In a simulation with the same sample sizes and biomarker distributions used in the paper, compute the average true sensitivity of the uncorrected EUM classifier and of the CAP-corrected cEUM classifier; if the former is not systematically below 0.95 while the latter is unbiased, or if the CAP accuracy bounds fail to cover at the claimed rate, the central claim is false.","supporting_citations":[],"review_version":1}