{"id":"cad58ff6-f4a4-48ad-b96e-7342f59c61f6","arxiv_id":"2607.09250","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Under Gaussian design with n ≍ d, the empirical distribution of leave-one-out influences for convex M-estimators converges to the pushforward of a four-dimensional Gaussian through an explicit nonlinear map built from resolvent summary statistics.","lead":"Leave-one-out influences of training points on high-dimensional convex M-estimators converge to an explicit limiting distribution when n is proportional to d. This distribution shows that points near the decision boundary are typically most influential, supporting a common active-learning heuristic.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader’s strongest claim matches Theorem 2.1 exactly, and the identified weakest assumption (A.1 + Gaussian design) is indeed the place where the argument is most model-dependent. That assumption is load-bearing in the technical sense that every concentration and residual-limit step relies on it, yet it is not a soft spot that undermines the mathematical result as stated: the paper never claims universality beyond the listed conditions, and the proofs are written under those conditions. Real-data figures and the active-learning discussion are presented as qualitative illustrations, not as part of the formal claim. Consequently no adjustment to the ACCEPT / high-confidence verdict is warranted. The concrete test above is a straightforward independent verification that would still be worth running before camera-ready, but a failure would point to an implementation or transcription error rather than a conceptual flaw in the argument.","tokens_in":62313,"tokens_out":587,"duration_ms":6630,"concrete_test":"Independently recompute the four summary statistics Q^(k), V^(k) for logistic loss at α=2, λ=0.05 by solving the self-consistent equations (14)–(16) numerically, then sample the pushforward φ_IF ♯ N(0_4,Q) and overlay it on a fresh Monte-Carlo histogram of exact leave-one-out IF_i (d=2000, n=4000). Agreement of the first two moments and Kolmogorov–Smirnov distance <0.05 would reconfirm the claim; a clear mismatch would indicate an error in the map or the resolvent equations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (weak convergence of the marginal IF distribution to φ_IF ♯ N(0_4,Q) under the stated assumptions) is supported by a complete, self-contained argument. The leave-one-out expansion (8), residual Gaussian limit (Lemma C.2), concentration of resolvents/summary statistics (Appendices B–C), and the final pushforward (Theorem 2.1 / D.6) form a coherent chain. The reader correctly flags the strongest modeling hypotheses (strong convexity + O(polylog n) derivative bounds + isotropic Gaussian design). These are used throughout, but they are standard for the high-dimensional M-estimation literature the paper builds on, are stated explicitly, and are not hidden. The square-loss case is handled by a transparent mollification remark; the elliptical/noisy extension is correctly labeled a conjecture. No internal inconsistency or gap that would invalidate the asymptotic characterization under the paper’s own hypotheses was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies leave-one-out influence diagnostics for ridge-regularized convex M-estimation under isotropic Gaussian design in the proportional high-dimensional regime n ≍ d. The central results are: (i) weak convergence of the marginal law of the leave-one-out test-error influence IF_i to the pushforward φ_IF ♯ N(0_4, Q), where φ_IF is an explicit nonlinear map built from resolvent summary statistics Q^(k), V^(k) characterized by self-consistent equations (Theorem 2.1); (ii) concentration in probability of the empirical distribution of DFBETAs to a related two-dimensional pushforward (Proposition 2.2). The proofs combine leave-one-out expansions, Efron–Stein and Lipschitz concentration of resolvents, and Gaussian residual limits. Section 3 then conditions the limiting law on geometric descriptors (margin, ground-truth projection) and argues that influential points tend to lie near the decision boundary, linking the asymptotics to a standard active-learning heuristic, with supporting synthetic and real-data experiments.","tokens_in":62503,"tokens_out":1230,"duration_ms":27196,"significance":"Classical influence-function theory is low-dimensional; the high-dimensional regime n ≍ d is the relevant one for many modern estimators, yet the joint dependence of each sample’s influence on the full training set has remained largely uncharacterized. The paper supplies a sharp, closed-form limiting description for two standard diagnostics under the standard Gaussian-design M-estimation model, with fully explicit summary-statistic equations and careful treatment of the 1/n versus 1/(n−1) normalization. The technical development (leave-one-out approximation, pointwise and integral concentration of resolvents, deterministic equivalents) is self-contained and builds cleanly on the El Karoui / Donoho–Montanari / Thrampoulidis line. The active-learning discussion is appropriately cautious (“evidence,” “making contact”) and is backed by both the conditional limiting law and real-data visualizations. Explicit characterizations, numerical histogram matches, and a clearly labeled conjectural extension to elliptical/noisy designs are genuine strengths.","major_comments":[{"comment":"Abstract and §2.3 state that “the distribution of the influences across the training set converges to a limiting measure.” Theorem 2.1 only establishes weak convergence of the marginal law of IF_{i_n} for a sequence of indices (plus an O(polylog n / n^{1/4}) error for Lipschitz test functions). Empirical-measure concentration is proved only for DFBETA (Proposition 2.2) and is explicitly left as a conjecture for IF. The abstract/introduction wording should be aligned with the theorem statements, or the stronger empirical claim for IF should be flagged as conjectural in the same places where the main result is announced.","section":null},{"comment":"Assumption A.1 requires strong convexity and derivatives up to order four bounded by O(polylog n). The square-loss case (Remark 2.3, Fig. 2) is recovered only by a mollification argument that is not written out. Because the square-loss formulae are used for intuition and for the population-limit comparison with classical influence functions (Remark 2.4), a short, self-contained justification that the limiting distribution is continuous under the mollification (or a direct argument for quadratic loss) would make that part of the paper fully rigorous rather than heuristic.","section":null}],"minor_comments":[{"comment":"The encoding of arrows and asymptotics in the provided text (e.g., “n/∫hortrightarrow∞”) is garbled; ensure the arXiv/source version uses standard LaTeX arrows throughout.","section":null},{"comment":"Notation for leave-one-out objects is dense (ˆw_{(i)}, ˆw_{\\i}, ˜w_i, H_{(i)}, H_{\\i}, ˜H_i). A short “notation table” early in §2 or Appendix A would help readers navigate Appendices B–D.","section":null},{"comment":"Fig. 1 (right) and Conjecture F.1: the real-data histogram is compared to the conjectural elliptical/noisy formula. State more clearly in the caption that the red curve is not covered by Theorem 2.1.","section":null},{"comment":"Section 3.2: the conditional densities ν_{IF|ω} are obtained by conditioning the four-dimensional Gaussian and pushing forward; a one-line formula for the conditional mean μ_{IF|ω} (or a pointer to the corresponding Gaussian conditional) would make the insets of Figs. 3–4 easier to reproduce.","section":null},{"comment":"Appendix G: the MNIST / chest X-ray protocol synthesizes points in span(β̂, ŵ) with oracle labels sign(⟨β, z⟩). A brief remark that this is an idealized probe of the (β, ŵ) plane, not a practical active-learning algorithm, would avoid over-interpretation.","section":null},{"comment":"Typos / polish: “succintly” → “succinctly” (p. 10); “on the other hand is the understanding” (p. 3); occasional missing articles. None affect correctness.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The technical core is solid and the contribution is real for the high-dimensional M-estimation community. The only soft spot is the slight overstatement of “distribution across the training set” versus marginal weak convergence; once that is tightened, the paper is close to accept. Scope fits a serious stat/ML theory venue."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is that this paper finally gives the limiting law of leave-one-out influences for convex M-estimation when n and d grow together. Prior high-dimensional work already used leave-one-out as a proof device for risk asymptotics; here the influences themselves become the object, and the answer is an explicit four-dimensional Gaussian pushforward through a nonlinear map built from resolvent summary statistics. That is new, and it is cleanly done.\n\nWhat works: the leave-one-out expansion, the residual Gaussian limit, the Efron–Stein/Lipschitz concentration of the resolvents, and the deterministic equivalents form a coherent chain. Assumptions (strong convexity, polylog-bounded derivatives, isotropic Gaussian) are stated up front and are the usual ones in this literature; the square-loss case is handled by a transparent mollification remark, and the elliptical/noisy extension is correctly labeled a conjecture. Numerics on synthetic data match the theory tightly; the real-data histograms and the MNIST/X-ray spatial plots are only qualitative, but they are presented as such and still make the margin–influence connection visible. The conditional analysis that recovers the “samples near the decision boundary are more influential” heuristic is a genuine payoff, not an afterthought.\n\nSoft spots are minor and proportional. The real-data experiments are illustrative only, and the code is referenced rather than permanently archived; neither undercuts the asymptotic claim. The derivative bounds and Gaussian design are load-bearing, but they are not hidden and they match the setting the paper claims to treat. I see no circularity or internal contradiction.\n\nThis is for people who work on high-dimensional M-estimation, influence diagnostics, or active learning under proportional asymptotics. It deserves a serious referee. I would bring it to reading group and I would cite the limiting law when I next need a high-dimensional influence calculation. Send it out.","headline":"Solid, first sharp asymptotics for leave-one-out influences in the n~d regime; the math holds under standard assumptions and the active-learning link is a clean payoff.","tokens_in":63097,"tokens_out":475,"would_cite":true,"duration_ms":8442,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J07","62F12","60B20"],"pacs":[],"model":"grok-4.5","headline":"In the high-dimensional regime n comparable to d, leave-one-out influences of training samples on convex M-estimators converge to an explicit limiting distribution given by a nonlinear pushforward of a four-dimensional Gaussian.","keywords":["high-dimensional M-estimation","leave-one-out influence","DFBETA","resolvent equations","active learning","Gaussian design","proportional asymptotics"],"falsifier":"Generate synthetic Gaussian data for logistic or ridge regression at moderate n = d = 1000–2000, compute the empirical histogram of leave-one-out test-error influences, and check whether it matches the predicted density φ_IF ♯ N(0_4, Q) within sampling error; a systematic mismatch would falsify the claimed weak convergence.","tokens_in":63204,"feed_emoji":"📊","tokens_out":1039,"duration_ms":14483,"temperature":0.7,"pith_summary":"Classical leave-one-out influence diagnostics become random and interdependent when the number of samples n is comparable to the dimension d. For convex ridge-regularized M-estimation under isotropic Gaussian design, the paper proves that the distribution of these influences (on test error and on parameter change) converges to a sharply characterized limiting measure. That measure is the pushforward of a low-dimensional Gaussian whose covariance is built from a handful of summary statistics of the estimator and its Hessian; the statistics themselves solve self-consistent resolvent equations. The same characterization shows that points with large influence tend to lie near the decision boundary, supplying a rigorous footing for a common active-learning heuristic. The results therefore give both an exact asymptotic description of sample importance and a practical geometric cue for data selection in high dimensions.","feed_headline":"High-dim leave-one-out influences follow an explicit law","feed_subtitle":"When n ≈ d the distribution of sample impacts on convex M-estimators is a pushforward of a four-dimensional Gaussian","key_machinery":"The four-dimensional Gaussian N(0_4, Q) whose covariance blocks are the Cauchy integrals of the resolvents Ω(z) and V(z); these resolvents satisfy closed self-consistent equations that involve only the proximal residual of the loss and the design ratio α, and the influence map φ_IF simply evaluates a first-order expansion of the test error along the leave-one-out correction.","core_discovery":"In the proportional high-dimensional limit n, d → ∞ with fixed ratio α = n/d, the marginal law of the leave-one-out test-error influence of a random training point converges weakly to the pushforward φ_IF ♯ N(0_4, Q), where φ_IF is an explicit nonlinear map built from the proximal residual and the partial derivatives of the test-error functional, and the 4 × 4 covariance Q is assembled from the resolvent moments Q^(k) and V^(k) of the leave-one-out Hessian.","pith_inferences":["Because the limiting law is fully determined by a few scalar summary statistics, one can in principle estimate those statistics once from the full-data fit and then rank every training point’s influence without repeated leave-one-out retraining.","The same resolvent machinery should extend, with only technical changes, to subset-influence diagnostics that delete blocks of size o(n), provided the blocks remain sparse relative to the Hessian spectrum.","The observed concentration of influence near the margin supplies a quantitative justification for margin-based active learning even when labels are expensive and the model is still far from the population risk minimizer."],"forward_implications":["The average influence of a training point is a non-monotonic function of the sample complexity α = n/d and is maximized at intermediate α.","Conditional on margin, the mean influence is largest for points lying near the current decision boundary and inside the disagreement region with the ground-truth separator.","DFBETA influences concentrate even more strongly: their empirical distribution converges in probability to a simple one-dimensional pushforward of a two-dimensional Gaussian.","In the population limit α → ∞ the high-dimensional law recovers the classical influence-function χ² law, linking the two asymptotic regimes.","The same geometric picture (high influence near the boundary) is observed qualitatively on real image data after a neural feature map, suggesting the heuristic remains useful beyond pure Gaussians."],"fun_headline_variants":["High-dim leave-one-out influences converge to Gaussian pushforward","Precise asymptotics for sample impacts in high-dim M-estimation","Leave-one-out influences form explicit law when n ≈ d","Training-point effects in high-dim convex M-estimators obey 4D map","Influences across samples follow known high-dimensional limiting measure"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"The loss must be strongly convex with derivatives up to order four growing at most polylogarithmically, and the covariates must be exactly isotropic Gaussian (or elliptical in the conjectural extension).","fun_headline_variants_meta":{"raw":{"variants":["High-dim leave-one-out influences converge to Gaussian pushforward","Precise asymptotics for sample impacts in high-dim M-estimation","Leave-one-out influences form explicit law when n ≈ d","Training-point effects in high-dim convex M-estimators obey 4D map","Influences across samples follow known high-dimensional limiting measure"]},"model":"grok-4.5","effort":"low","cost_usd":0.005862,"raw_usage":{"total_tokens":1541,"prompt_tokens":750,"num_sources_used":0,"completion_tokens":96,"cost_in_usd_ticks":58620000,"prompt_tokens_details":{"text_tokens":750,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":695,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":750,"tokens_out":96,"duration_ms":6116,"temperature":1.0,"reasoning_tokens":695,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T04:23:26.825387+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Generate synthetic Gaussian data for logistic or ridge regression at moderate n = d = 1000–2000, compute the empirical histogram of leave-one-out test-error influences, and check whether it matches the predicted density φ_IF ♯ N(0_4, Q) within sampling error; a systematic mismatch would falsify the claimed weak convergence.","supporting_citations":[],"review_version":1}