{"id":"cf1f5469-d734-4734-8816-29b5997446f8","arxiv_id":"2411.13763","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An active K-step subsampling estimator for high-dimensional individualized thresholds achieves the parametric rate (s log d/N)^(1/2) under a smoothness condition, matching the new N-budget minimax lower bound up to a logarithmic factor.","lead":"The paper proposes an active subsampling algorithm that spends a label budget on patients closest to the estimated decision threshold, and proves this can recover an individualized threshold at the parametric rate with far fewer labels than passive sampling. A generalist might care because the method gives a formal answer to which patients to label in expensive EHR or clinical studies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RSC verification in §A.7 uses condition (A.76) that is incompatible with Theorem 2's tuning for β in ((1+√3)/2, (3+√15)/2), so Assumption 3.5 is not established in that regime.","rationale":"The reader's weakest assumption is Assumption 3.5, and the reader noted it is only verified for the conditional mean model. I agree that this is the load-bearing condition, but I push further: even the conditional mean verification fails under the paper's own tuning. The algebraic check of (A.76) shows a divergence for β in a large part of the phase-transition regime, so the only attempt to ground Assumption 3.5 in primitive conditions is unsuccessful. This is not a manufactured objection; it is an internal consistency problem between Proposition A.4 and Theorem 2. It does not disprove the theorems, which are stated conditionally on Assumption 3.5, but it does mean the paper has not demonstrated that its central claim applies to the models it advertises. The verdict remains CONDITIONAL: the theory is coherent if Assumption 3.5 is granted, but the support for that assumption is incomplete and, in the identified β range, demonstrably broken. A targeted numerical check can settle whether the sparse curvature actually collapses, which would confirm or resolve the concern.","tokens_in":60013,"tokens_out":18048,"duration_ms":164488,"concrete_test":"For the conditional mean model of §A.7 with β = 2, s = 10, d = 200, N = 2000, n = 20000, use the Theorem 2 tuning: δ_k = c(s log d / N)^{1/(2β)} and radius R = (s log d / N)^{β/(2β+1)}. Generate the active-sampled data at iteration k = 2, compute the empirical Hessian ∇² R^{D_k}_{δ_k}(θ) at θ = θ̂_{k−1} + R v for sparse unit vectors v, and estimate its minimum sparse eigenvalue over v with ∥v∥₀ ≤ Cs across 100 replications. If the minimum eigenvalue is not bounded below by a positive constant times c_{n,k} = N/(n P(S_k)), then Assumption 3.5 is violated in the regime where Theorem 2 asserts the parametric rate, demonstrating the proof gap is real.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All upper-bound rates pass through Assumption 3.5, the RSC/RSM condition on the sampled smoothed risk. The only verification of this assumption, given in §A.7, concerns the conditional mean model (X = θ*ᵀZ + μY + u). Proposition A.4 requires condition (A.76): s M_n³ √s / δ_k³ · (s log d / N)^{β/(2β+1)} = o(1). But Theorem 2 sets δ_k ≍ (s log d / N)^{1/(2β)} for k ≥ 2. Substituting gives s^{3/2} M_n³ (s log d / N)^{β/(2β+1) − 3/(2β)}. The exponent β/(2β+1) − 3/(2β) is negative for every β < (3+√15)/2 ≈ 3.436, which includes the entire range (1+√3)/2 ≈ 1.366 < β ≤ 3.436 covered by Theorem 2. Since s log d = o(N), the term tends to infinity as N grows, so (A.76) fails. Consequently, the appendix does not establish RSC/RSM for the conditional mean model under the theorem's own tuning for a substantial part of the claimed phase-transition regime, leaving the central parametric-rate claim without a valid verification of its load-bearing condition.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies measurement-constrained estimation of a high-dimensional individualized threshold parameter θ* in the M-estimation problem (1.2), where only N of n available units can be labeled. It proposes a K-step active subsampling algorithm that uses the current estimator to define an 'active set' of observations with X close to the current threshold, and focuses label acquisition there. The main theoretical claim is a phase transition in the Hölder smoothness β of the conditional density of X given (Y,Z): for β > (1+√3)/2, a two-step version attains the parametric l2 rate (s log d / N)^{1/2}, faster than the passive i.i.d. minimax rate (s log d / N)^{β/(2β+1)} and matching the paper's N-budget minimax lower bound up to logarithmic factors; for 1 < β ≤ (1+√3)/2 the same rate requires a finite K > 2 depending on β, and for β = 1 only a near-parametric rate with K = O(log log N) is obtained. The paper also gives implementation details, a data-driven version with cross-validation, simulations, and a diabetes readmission application.","tokens_in":60353,"tokens_out":20893,"duration_ms":165286,"significance":"If the results hold, the phase transition is significant: it shows that active subsampling can overcome the slow non-regular rate of threshold M-estimators and achieve parametric accuracy in high dimension under a label budget, with a matching lower bound in a carefully defined class of adaptive sampling mechanisms. The proposed algorithm is concrete and computationally practical, and the N-budget minimax framework (the class Q_N(P(β,s))) is a useful formalization for measurement-constrained active estimation. The paper ships no code, but the simulations and real-data analysis illustrate the potential practical value.","major_comments":[{"comment":"The verification of Assumption 3.5 in Section A.7 is incompatible with the tuning of Theorem 2 over a substantial part of the claimed range. Proposition A.4 requires condition (A.76), namely s M_n^3 √s / δ_k^3 · (s log d / N)^{β/(2β+1)} = o(1). Theorem 2 sets δ_k = c_1 (s log d / N)^{1/(2β)} for every k ≥ 2. Substituting gives s^{3/2} M_n^3 (s log d / N)^{β/(2β+1) - 3/(2β)}. The exponent is negative for all β < (3+√15)/2 ≈ 3.436, which includes the entire interval (1+√3)/2 < β ≤ 3.436 covered by Theorem 2. Since s log d = o(N), the left-hand side diverges as N grows, so (A.76) fails. This matters because Theorem 1 (and hence Theorem 2) passes through Assumption 3.5, and the appendix's verification is the only concrete evidence that this load-bearing condition holds for a nontrivial model. The paper therefore does not currently establish the parametric-rate claim for the conditional mean model on the advertised range of β, and the phase-transition threshold (1+√3)/2 is not supported by the provided verification.","section":"Section A.7, Eq. (A.76); Theorem 2"}],"minor_comments":[{"comment":"The notation is nonstandard: ⌊β⌋ is defined as the greatest integer strictly less than β. In standard usage, floor(β) is the greatest integer ≤ β. For integer β this changes the kernel order l from β to β-1. If the intended definition is l = ⌈β⌉-1, please state that explicitly.","section":"Section 1.3 / Definition 3.1"},{"comment":"The statement of Lemma A.3 writes θ_j = c√s (s log(d/s)/N)^{1/2} ω_j, but the proof and equation (A.33) imply θ_j = c (s log(d/s)/N)^{1/2} ω_j / √s. This is a typo, but it affects the displayed form of the hypotheses in the lower-bound construction and should be corrected.","section":"Lemma A.3"},{"comment":"The text states that for 1 < β ≤ (1+√3)/2 the required number of iterations K is strictly greater than 2. At the endpoint β = (1+√3)/2, the formula in Theorem 3 gives K = ⌈log_{β/(2β+1)}(1 - (β+1)/(2β^2))⌉ + 1 = 2, so the statement 'strictly greater than 2' is false at that boundary. Please qualify the statement to reflect the endpoint behavior.","section":"Section 1"},{"comment":"The simulations set δ1 = δ2 = 1 and use a Gaussian kernel, whereas the theory (Theorems 2–4) assumes a compactly supported kernel of order l with bandwidths δ_k → 0 that are specific functions of s, d, N, and n. The text explains the practical choice, but it would be helpful to comment explicitly on the gap between the implemented bandwidth and the theoretical regime, and on whether the simulation results should be interpreted as supporting the theoretical rates or only the algorithm's practical performance.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper's algorithmic idea and lower-bound framework are solid and likely publishable after revision. The central issue is the unverified RSC/RSM condition for β in (1.366, 3.436): the only concrete verification, in Section A.7, fails for that range under the theorem's own bandwidth choice. I would ask the authors to repair this verification (e.g., by a different localization argument or a different choice of δ_k that still yields the parametric rate) or to restrict the main upper-bound theorem to the range in which the verification is valid. This is a fixable major gap rather than a fatal flaw, since the theorem is conditional on Assumption 3.5 and the lower bound is independent of that condition."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis paper has a genuinely new idea: active subsampling for measurement-constrained high-dimensional threshold estimation, with a phase transition in the smoothness β of the conditional density. The two-step estimator achieving a parametric rate for smooth β, and the N-budget minimax lower bound, are real contributions. The bias-variance explanation is clear and honest.\n\nThe soft spot is load-bearing. All the upper-bound rates pass through Assumption 3.5, the RSC/RSM condition on the sampled smoothed risk. The only verification of that condition, in §A.7 for the conditional mean model, requires (A.76). Under Theorem 2's own tuning, δ_k = (s log d/N)^{1/(2β)}, the left side of (A.76) diverges for every β < (3+√15)/2 ≈ 3.436, which includes the interval ((1+√3)/2, (3+√15)/2) that the abstract highlights. So the appendix does not establish RSC/RSM in the claimed regime. The same δ_k scaling appears in the final step of Theorems 3 and 4, so the gap is not confined to Theorem 2. This is not a minor nuisance: the central parametric-rate claim is unproven for all but quite smooth densities. It may be fixable, but as written the paper does not support its main advertised result.\n\nTwo lesser points. The verification is only for the conditional mean model; the binary response class is asserted without proof. And the practical recipe (K=2, δ=1, cross-validation) is not tied to the theory, though that is typical for this literature. No code is provided for the simulations.\n\nWho it's for: researchers in non-regular high-dimensional estimation and active learning. The phase-transition idea and the N-budget minimax framework are worth discussing even with the gap. I would send it to peer review, asking the referee to focus on the RSC verification. Net: deserves referee time, but not acceptance in current form.","headline":"A promising phase-transition result undermined by a failing RSC verification for most β in the claimed range.","tokens_in":60851,"tokens_out":5705,"would_cite":false,"duration_ms":49625,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62C20","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Active label selection can push individualized threshold estimation to the parametric rate, beating the passive minimax rate when the conditional density is sufficiently smooth.","keywords":["active subsampling","measurement-constrained estimation","individualized threshold","high-dimensional M-estimation","kernel smoothing","phase transition","minimax optimality","sparsity"],"falsifier":"Check Assumption 3.5 numerically for a logistic-regression threshold model with $\\beta=1.5$: compute the minimum sparse eigenvalue of $\\nabla^2 R^{D_k}_{\\delta_k}(\\theta)$ over $\\theta$ in the ball $\\{\\theta:\\|\\theta-\\hat\\theta_{k-1}\\|_2\\le R_{k-1}\\}$; if it is not proportional to $c_{n,k}$ with high probability, the rate theorem's foundation fails. Alternatively, simulate the two-step algorithm with increasing $N$ and verify empirically whether the $\\ell_2$ error tracks $(s\\log d/N)^{1/2}$ rather than the passive rate.","tokens_in":59801,"feed_emoji":"🎯","tokens_out":9876,"duration_ms":87815,"temperature":0.7,"pith_summary":"This paper asks which data points to label when labels are scarce but covariates are cheap, for the problem of estimating the optimal individualized threshold $\\theta^*$ in a linear decision rule $X \\gtrless \\theta^{T}Z$. The authors propose a $K$-step active subsampling algorithm that starts with a uniformly sampled regularized M-estimator, then repeatedly samples only the observations whose $X-\\hat\\theta_{k-1}^TZ$ lies near zero---the most informative for the threshold---and refits. The central claim is a phase transition in $\\beta$, the H\\\"older smoothness of the conditional density of $X$ given $(Y,Z)$: for $\\beta>(1+\\sqrt{3})/2$ the two-step estimator reaches the parametric rate $O_p((s\\log d/N)^{1/2})$, strictly faster than the passive minimax rate $(s\\log d/N)^{\\beta/(2\\beta+1)}$, and matches the newly formulated $N$-budget minimax lower bound up to a log factor. For $1<\\beta\\le(1+\\sqrt{3})/2$ the same rate needs a fixed finite number of steps $>2$, and for $\\beta=1$ only a near-parametric rate is attainable. A sympathetic reader should care because in settings like electronic health record studies the bottleneck is the cost of chart review, not the availability of covariates; the paper shows this bottleneck can be partially broken by adaptive label selection.","feed_headline":"Two-step active labeling hits parametric rate for thresholds","feed_subtitle":"Adaptively labeling points near the estimated cutpoint beats the passive minimax rate when the density is smooth.","key_machinery":"The argument is carried by a zoom-in sampling rule. At each step $k\\ge2$, the algorithm defines the active set $S_k=\\{(X,Z): -b_{k-1}\\le (X-\\hat\\theta_{k-1}^TZ)/\\sqrt{1+\\|\\hat\\theta_{k-1}\\|_2^2}\\le b_{k-1}\\}$ and samples only from this thin band around the current estimated threshold, with probability $c_{n,k}$ proportional to the available budget. The loss is a smoothed surrogate $L_\\delta$ built from a kernel of order $\\lfloor\\beta\\rfloor$, whose kernel smoothing creates a bias of order $c_{n,k}\\delta_k^\\beta$ and a stochastic error of order $\\sqrt{c_{n,k}K\\log d/(n\\delta_k)}$; balancing these gives $\\delta_k\\asymp(b_{k-1}s\\log d/N)^{1/(2\\beta+1)}$ and the per-step rate $\\|\\hat\\theta_k-\\theta^*\\|_2\\lesssim(b_{k-1}s\\log d/N)^{\\beta/(2\\beta+1)}$. Since the active set has probability $\\asymp b_{k-1}$, each iteration multiplies the rate by a power of the band width $b_{k-1}$, and the stability condition $b_{k-1}\\ge C\\max\\{\\delta_k,\\|\\hat\\theta_{k-1}-\\theta^*\\|_2\\sqrt{\\log(N/(s\\log d))}\\}$ determines how small $b_{k-1}$ may be chosen. The phase transition occurs because for $\\beta>(1+\\sqrt{3})/2$ the first-step estimator already lands in the \"fast convergence region\" where $\\|\\hat\\theta_1-\\theta^*\\|_2$ is at most order $\\delta_2$, so one more step reaches the parametric rate; for smaller $\\beta$ several steps are needed to reach that region, and for $\\beta=1$ it is never reached.","core_discovery":"On its own terms, the paper establishes that active label selection can convert a non-regular high-dimensional estimation problem into one that behaves like a regular parametric problem. Concretely, when the conditional density of $X$ given $(Y,Z)$ is $\\beta$-smooth with $\\beta>(1+\\sqrt{3})/2$, the estimator produced by two iterations of the proposed algorithm satisfies $\\|\\hat\\theta_K-\\theta^*\\|_2 = O_p((s\\log d/N)^{1/2})$ with high probability, while the best passive estimator using $N$ i.i.d. labels has rate $(s\\log d/N)^{\\beta/(2\\beta+1)}$. The same parametric rate is achieved for $1<\\beta\\le(1+\\sqrt{3})/2$ by running $K=\\lceil \\log_{\\beta/(2\\beta+1)}(1-(\\beta+1)/(2\\beta^2))\\rceil+1$ iterations, and for $\\beta=1$ only $\\|\\hat\\theta_K-\\theta^*\\|_2 = O_p((\\log(N/(Ks\\log d)))^{1/4}(Ks\\log d/N)^{1/2})$ is obtained with $K=\\lceil\\log_3(\\log N)\\rceil$. The paper also defines an $N$-budget minimax risk over label-sampling distributions and proves the lower bound $(s\\log(d/s)/N)^{1/2}$, showing the active estimator is minimax optimal up to logarithmic factors and that unlabeled covariates do not improve the rate.","pith_inferences":["An immediate practical reading: when analysts cannot verify that the conditional density is smoother than $1.37$-H\\\"older, running $K=2$ is still a safe default because the dominant improvement happens between the first and second iterate, and the second iterate's tuning is simpler.","The rate improvement implies a direct cost translation: for $\\beta>(1+\\sqrt{3})/2$, the same $\\ell_2$ error as passive sampling with $N$ labels is reached with roughly $(s\\log d)^{(1-1/(2\\beta))}N^{1-1/(2\\beta)}$-style fewer labels; in EHR chart-review budgets this converts into concrete dollar savings, though the paper does not quantify this.","The proof's reliance on a region-sampling class suggests a natural stress test: allow sampling probabilities that depend on $Z$ beyond the bounded-probability and sparse-eigenvalue constraints, and see whether the lower bound still holds; one would suspect it does, because the Fano construction already chooses $Z$ uniform on $[-1,1]$.","A testable extension the authors mention but do not pursue is Lepski-type adaptation to unknown $\\beta$; the paper's own simulations fix $K=2$ and cross-validate the bandwidth, so an empirical study measuring achieved rates under unknown $\\beta$ would tell whether the phase-transition recommendation is robust."],"forward_implications":["If the density is $\\beta$-smooth with $\\beta>(1+\\sqrt{3})/2$, two labeling rounds with budget split $N_1=N/8$, $N_2=7N/8$ give the same $\\ell_2$ accuracy as a parametric estimator using $N$ labels; additional rounds do not improve the rate.","The $N$-budget minimax lower bound $(s\\log(d/s)/N)^{1/2}$ implies that no sampling scheme in the permitted class can beat the proposed algorithm's rate by more than a log factor, and that the unlabeled pool adds no rate benefit once labels are actively selected.","For intermediate smoothness $1<\\beta\\le(1+\\sqrt{3})/2$, the practical protocol should budget for more than two rounds; the required number of rounds is fixed and finite, independent of $N$.","When only Lipschitz smoothness ($\\beta=1$) holds, the achievable rate carries an extra $(\\log N)^{1/4}$ factor even with $K=\\lceil\\log_3(\\log N)\\rceil$ rounds, so the gain over passive sampling is only logarithmic."],"supporting_citations":[{"why":"Supplies the smoothed surrogate loss, the path-following algorithm, and the passive minimax rate $(s\\log d/N)^{\\beta/(2\\beta+1)}$ that the active estimator must beat; also gives the RSC/RSM template for the first step.","marker":"Feng et al. (2022)"},{"why":"Provides the Fano-type minimax lower bound theorem (Theorem 2.7) used to prove the $N$-budget lower bound.","marker":"Tsybakov (2008)"},{"why":"Gives the comparable multistage sampling rates in the univariate $\\beta=1$ case against which the paper benchmarks its iterate rates.","marker":"Mallik et al. (2020)"},{"why":"Background reference for the restricted strong convexity and smoothness conditions that Assumption 3.5 relies on.","marker":"Bühlmann and Van De Geer (2011)"},{"why":"Defines the individualized minimal clinically important difference objective (1.1) that this paper estimates.","marker":"Zhou et al. (2020)"}],"fun_headline_variants":["Active label selection hits parametric rate for high-dimensional thresholds","Smart subsampling beats passive minimax for threshold estimation","Adaptive labeling reaches parametric rate in M-estimation","Budget-efficient active learning for individualized thresholds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Assumption 3.5 — that the smoothed, selectively sampled risk is strongly convex and smooth on shrinking balls around each previous estimate — is the load-bearing condition; the paper verifies it only for a conditional mean model with Gaussian noise and assumes it for the general binary response and conditional mean classes used in the main theorems.","fun_headline_variants_meta":{"raw":{"variants":["Active label selection hits parametric rate for high-dimensional thresholds","Smart subsampling beats passive minimax for threshold estimation","Adaptive labeling reaches parametric rate in M-estimation","Budget-efficient active learning for individualized thresholds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000268,"raw_usage":{"total_tokens":1684,"prompt_tokens":1079,"completion_tokens":605,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":695,"completion_tokens_details":{"reasoning_tokens":545}},"tokens_in":695,"tokens_out":605,"duration_ms":6665,"temperature":1.0,"reasoning_tokens":545,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:54:39.854372+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check Assumption 3.5 numerically for a logistic-regression threshold model with $\\beta=1.5$: compute the minimum sparse eigenvalue of $\\nabla^2 R^{D_k}_{\\delta_k}(\\theta)$ over $\\theta$ in the ball $\\{\\theta:\\|\\theta-\\hat\\theta_{k-1}\\|_2\\le R_{k-1}\\}$; if it is not proportional to $c_{n,k}$ with high probability, the rate theorem's foundation fails. Alternatively, simulate the two-step algorithm with increasing $N$ and verify empirically whether the $\\ell_2$ error tracks $(s\\log d/N)^{1/2}$ rather than the passive rate.","supporting_citations":[{"cited_title":", Ning, Y","cited_arxiv_id":null,"evidence_quote":"Supplies the smoothed surrogate loss, the path-following algorithm, and the passive minimax rate $(s\\log d/N)^{\\beta/(2\\beta+1)}$ that the active estimator must beat; also gives the RSC/RSM template for the first step."},{"cited_title":", Banerjee, M","cited_arxiv_id":null,"evidence_quote":"Gives the comparable multistage sampling rates in the univariate $\\beta=1$ case against which the paper benchmarks its iterate rates."},{"cited_title":", Zhao, J","cited_arxiv_id":null,"evidence_quote":"Defines the individualized minimal clinically important difference objective (1.1) that this paper estimates."}],"review_version":1}