{"id":"1f63d246-d23f-43fa-b57e-d57d180f3ade","arxiv_id":"2411.14665","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"With support-points sample splitting, a deep-learning and super-learner hybrid gives lower MSE and deep learning alone gives faster computation than support vector machines in the authors' simulations.","lead":"Support-points sample splitting is applied to double machine learning, and three estimators, support vector machine, deep learning, and a hybrid super learner with deep learning, are compared for average treatment effect estimation. The authors report that the hybrid is most accurate while deep learning alone is fastest in their simulations, with implications for high-dimensional causal inference practice.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never proves SPSS splits satisfy DML's cross-fitting independence; because the test fold is chosen from the full data, the root-N inference in Theorem 1 is unsupported, and no random-splitting control isolates SPSS's effect.","rationale":"The paper's stated contribution is to integrate support-point sample splitting into double machine learning and to demonstrate that DL and SDL with SPSS outperform SVM with SPSS in computational efficiency and estimation quality. For this claim to hold, two things must be true: SPSS must preserve the DML theory that justifies root-N inference, and the empirical comparisons must actually measure the effect of SPSS. The paper establishes neither. The theoretical sections restate Chernozhukov et al. (2018) assumptions and Mak and Joseph (2018) support-point convergence results, but they do not prove that a data-adaptive, energy-minimizing split satisfies the cross-fitting independence or the nuisance-estimator rate conditions needed for Theorem 1. The internal theory also shows strain: Theorem 3 assumes iid support-point random variables, whereas SPSS points are constructed to be optimally representative and are not iid; the subsequent lemmas and theorems only give weak convergence, which is insufficient for the empirical-process terms in DML. The simulation design compounds this: all methods use SPSS, and timing is compared across PCs and HPC nodes with different core counts, so the reported time advantages are not controlled comparisons. There is also no random-splitting baseline, so any observed difference among SVM, DL, and SDL could be driven by the learners rather than by SPSS. Given that the central methodological innovation is exactly the replacement of random splits by support-point splits, the absence of a direct validity check is decisive. I therefore agree with the reader's weakest-assumption diagnosis and see no reason to change the REJECT verdict: the concern is not merely that the paper disagrees with consensus, but that its central inference guarantee is unverified and its experiments cannot isolate the proposed mechanism.","tokens_in":23234,"tokens_out":3376,"duration_ms":39018,"concrete_test":"Re-run the Section 5 simulations on the same DGP (Scenario 1, p = 20, 50, 80; N = 100, 500, 1000) under three splitting schemes: random K-fold cross-fitting as in Chernozhukov et al. (2018), SPSS as used in the paper, and a repeated random holdout split (B = 200 replications), with identical SVM, DL, and SDL nuisance estimators, hyperparameters, and convergence criteria. Report bias, standard deviation, MSE, and 95% confidence-interval coverage of the ATE estimate for each scheme. If SPSS coverage is substantially below nominal, or if SPSS does not improve over random K-fold on MSE or time, the paper's central inference claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central methodological claim is that replacing random k-fold cross-fitting with support-point sample splitting (SPSS) preserves DML validity and improves efficiency. This is load-bearing because every reported comparison in Tables 2-12 uses only SPSS splits, and the theoretical sections never establish that SPSS splits satisfy the DML conditions. Chernozhukov et al. (2018) Definition 1 and Assumptions 1-2 require a random K-fold partition: the nuisance estimator eta-hat_{0,k} is built on the complement fold, and the score is evaluated on an independent fold. In SPSS, the test set is chosen by minimizing an energy-distance criterion over the full sample, so the test fold is a deterministic, data-dependent function of all N observations. The independence between eta-hat_{0,k} and the evaluation sample that cross-fitting relies on is therefore broken. Support points are repulsive and dependent, not iid, so the empirical-process remainder terms in Theorem 1 (r_N, r_N', and the OP(rho_N) term) cannot be assumed to behave as they do for iid random folds. The paper's Lemma 1 and Theorems 3-6 only establish weak convergence of the support-point empirical distribution to G; that is neither necessary nor sufficient for the Neyman-orthogonality remainder rates and the root-N asymptotic normality claimed in Theorem 1. Notably, Theorem 3 assumes an iid sequence of support-point random variables, which contradicts how SPSS points are actually generated. No proof connects SPSS to Assumptions 1-2, and no simulation includes a random-k-fold control using the same SVM/DL/SDL learners. Thus even the empirical comparisons cannot separate the effect of SPSS from the choice of learner. If this gap is real, the central claim that SPSS is an efficient and valid sample splitting technique for DML is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using support points sample splitting (SPSS), a deterministic energy-distance-based subsampling method, in place of random k-fold cross-fitting inside double machine learning (DML) for average treatment effect estimation. It applies DML with three machine-learning nuisance estimators — support vector machines (SVM), deep learning (DL), and a hybrid super learner plus deep learning (SDL) — under SPSS, and compares the three on simulated high-dimensional data and on the 401(k) real-data benchmark from Chernozhukov et al. (2018). The authors claim that SPSS provides a more representative split than random splitting, that DL with SPSS is computationally most efficient, and that SDL with SPSS achieves the lowest MSE. The theoretical sections restate the DML assumptions and support-point convergence results, and the experiments report bias, SE, MSE, and runtime across many (N, p) configurations, plus a real-data comparison against Lasso with k-fold splitting.","tokens_in":23576,"tokens_out":3293,"duration_ms":35523,"significance":"If the central claim were established — that replacing random k-fold cross-fitting with support-point splits preserves DML's root-N inference while improving efficiency — the paper would be useful for high-dimensional causal inference, because SPSS could reduce the variance of cross-fitting and improve representation of the data distribution. The paper also has practical value as a comparative study of SVM, DL, and SDL nuisance estimators in DML, and it ships a large simulation grid and a real-data benchmark on a standard dataset. However, the theoretical justification is the load-bearing part of the contribution, and this part is not delivered: the manuscript never proves that SPSS splits satisfy the cross-fitting independence and empirical-process conditions required by DML theory. The empirical comparison also lacks a random-splitting control, so the claimed advantages of SPSS are not actually tested. These are not presentation issues; they concern the validity of the paper's central methodological claim.","major_comments":[{"comment":"The paper never establishes that SPSS splits satisfy the DML cross-fitting condition of Chernozhukov et al. (2018) Definition 1. In DML, the nuisance estimator η̂_{0,k} is constructed on the complement fold and the score is evaluated on an independent random fold. In SPSS, the test fold is chosen as the minimizer of an energy-distance criterion over the full sample, so it is a deterministic, data-dependent function of all N observations. The independence between the nuisance estimator and the evaluation sample is therefore broken. Theorems 3–6 only show weak convergence of the support-point empirical distribution to G; they do not provide the Neyman-orthogonality remainder rates r_N, r_N′, and N^{1/2}λ_N that Theorem 1 requires. In addition, Theorem 3 explicitly assumes an iid sequence of support-point random variables, which contradicts the repulsive, dependent construction of support points as minimizers of the energy distance. Thus the root-N asymptotic normality claim for the SPSS-based DML estimator is unsupported.","section":"Section 3, Theorems 1 and 3–6"},{"comment":"There is no random k-fold cross-fitting control in the simulation study. Every reported comparison uses only SPSS splits, so the observed differences in MSE and time among SVM, DL, and SDL cannot be attributed to SPSS; they may simply reflect the different machine-learning estimators. To support the claim that SPSS is an efficient sample-splitting technique, the authors need to run the same estimators under random k-fold splitting and compare the two splitting schemes on identical data and replications.","section":"Section 5, Tables 2–10"},{"comment":"Tables 11 and 12 report a single 'best' MSE and total time per method, scenario, and data level, rather than aggregate statistics over the 9 (or 18) cases in each level. For example, Table 11 states that SDL achieved MSE = 0.0006 for BHD, but Table 10 shows SDL MSE ranging from 0.0006 to 0.0210 across the BHD cells, and Table 2 shows SVM MSE ranging from 0.0009 to 0.0541. Selecting the favorable cell is not a valid basis for comparing methods. Report the mean, median, and dispersion of MSE and time across all replications and cases, or provide per-case comparison plots with error bars.","section":"Tables 11 and 12"},{"comment":"The computational-time comparisons are unreliable because the runs are mixed across different hardware. The text states that simulations were operated using both personal computers (4–20 cores) and the HPC cluster, and the table notes show different machines for different methods and settings (e.g., Table 2 notes PCs, Table 3 notes HPC, Table 5 notes HPC, Table 8 notes PCs). Combining these times into total times in Table 12 confounds algorithm performance with hardware differences. Additionally, Table 13 reports SDL-SPSS real-data time of 0.0429 seconds, which is implausibly smaller than DL-SPSS (1.1610 seconds) and is not explained by the small number of covariates; this needs clarification and a consistent computing environment.","section":"Section 5, Tables 2–12 and Table 13"},{"comment":"The real-data comparison is not apples-to-apples: the benchmark is Lasso with k-fold splitting, while the proposed methods use SPSS, and the comparison reports only point estimates and SEs. Since the paper's central claim is about preserving valid inference, the authors should report confidence-interval coverage or at least repeated-split variability for each method. Without such measures, the statement that SVM-SPSS 'provided the most accurate estimation with SE = 0.0006' is not informative about inferential validity.","section":"Table 13 and Section 5.5"}],"minor_comments":[{"comment":"The abstract contains a typo: 'ABSTARCT' in the full text. The phrase 'providing an optimal sub-representation' is grammatically incomplete and should be revised.","section":"Abstract"},{"comment":"There is an inconsistency in the largest covariate dimension: Section 4 lists p = (1000, 2000, 3000), while Table 1 and Section 5 use p = 5000. Please align these numbers.","section":"Section 4, paragraph 1 and Table 1"},{"comment":"The remark attributes the proof to 'Halley's theorem'; the standard result is Helly's theorem (or Helly selection theorem). Please correct the name and give a proper citation or proof sketch.","section":"Lemma 1 and its Remark"},{"comment":"In the definition of σ̂², the second empirical moment uses ψ(Z; β̂0, η̂0,k) while the rest of the expression uses β̃0. This appears to be a typo; it should be β̃0 consistently.","section":"Theorem 2"},{"comment":"The notation G_n, Ǧ_n, and G is used without a clear distinction between the support-point empirical distribution, the ordinary empirical distribution, and the limit. Define each symbol in one place to avoid confusion.","section":"Theorems 3–6"},{"comment":"The note says 'Time = the computer running time for simulation' while Table 2's note says 'Time = the running time of computing.' Please use consistent terminology and specify units (seconds vs. minutes).","section":"Section 5, Table 12 note"}],"recommendation":"reject","confidential_remarks":"The paper addresses an interesting practical question, but the central methodological claim — that SPSS preserves DML validity — is neither proved nor tested against random cross-fitting. The theory provided (Theorems 3–6) concerns convergence of the support-point empirical distribution, not the cross-fitting independence or Neyman-orthogonality remainder rates needed for Theorem 1. The simulations lack a random-splitting control, and the timing results are confounded by mixed PC/HPC usage. These are load-bearing gaps that cannot be repaired by local revisions; they require new theoretical work and a redesigned empirical study. The manuscript would need substantial new results before it could be considered for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper is a benchmark comparison of three ML learners inside DML with support-point sample splitting, but the central claim—that SPSS preserves DML's validity—is never established, and the experiments can't isolate SPSS's effect.\n\nWhat's genuinely new is the three-way comparison (SVM, DL, and SDL) under SPSS; that specific comparison isn't in the prior work they cite. The simulation grid is broad (N=100–1000, p=20–5000) and they are transparent about using both PCs and HPC, which is more than many papers do.\n\nThe theory is the big problem. Chernozhukov et al.'s cross-fitting relies on the nuisance estimator being built on a random complement fold independent of the evaluation fold. SPSS chooses the test fold by minimizing energy distance over the full sample, so the test set is a deterministic function of all N observations. The paper never shows the needed independence or that the Neyman-orthogonality remainder rates hold for repulsive, dependent support points. Theorems 3–6 only give weak convergence of the empirical distribution to G, which is neither necessary nor sufficient for the root-N normality in Theorem 1. Worse, Theorem 3 explicitly assumes the support points are iid, which contradicts how SPSS is actually generated.\n\nThe empirical side also has issues. No random k-fold control means you can't tell whether SPSS helps or hurts relative to standard cross-fitting. Timing comparisons mix PC and HPC runs, so the speed rankings in Tables 11–12 are not reliable. The summary tables pick the best MSE and best time across many settings rather than reporting aggregates. And the real data results contradict the abstract: Table 13 shows SDL is fastest, not DL, and SVM has the smallest SE. This is not a fatal flaw by itself, but it makes the paper feel rushed.\n\nIf the authors reframed this as an empirical benchmark with a random-splitting control, and either proved or explicitly dropped the claim that SPSS preserves DML theory, the comparison could be useful. As it stands, the load-bearing theoretical claim is unsupported. I'd send it to peer review rather than desk-reject, because the underlying question—whether non-random splits can preserve DML validity—is important and the empirical work is extensive enough to be salvageable with major revision. But I would not cite it in its current form.","headline":"A broad but flawed benchmark: the SPSS-for-DML claim lacks theoretical support and the experiments lack a random-splitting control.","tokens_in":24162,"tokens_out":2561,"would_cite":false,"duration_ms":25994,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62G05","62G20","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"Support points sample splitting can replace random k-fold cross-fitting inside double machine learning, making average treatment effect estimation faster and more accurate.","keywords":["double machine learning","support points sample splitting","causal inference","average treatment effect","energy distance","super learner","deep learning","high-dimensional data"],"falsifier":"A paired simulation that runs the same DML estimator, such as SDL, with SPSS and with random k-fold cross-fitting on the same data-generating processes, checking 95% confidence interval coverage and bias as N grows: if SPSS-DML coverage drops below nominal or bias does not shrink at the $\\sqrt{N}$ rate, the paper's central claim is refuted.","tokens_in":23047,"feed_emoji":"📊","tokens_out":10100,"duration_ms":80654,"temperature":0.7,"pith_summary":"This paper argues that the sample-splitting step inside double machine learning (DML) can be improved by replacing random k-fold cross-fitting with support points sample splitting (SPSS), a subsampling method that picks a fixed-size set of points minimizing energy distance to the full data. The authors embed SPSS in DML with three estimators—support vector machines, deep learning, and a hybrid super learner plus deep learning—and compare them against the random k-fold benchmarks of Chernozhukov et al. (2018). Their simulations across 54 settings (sample sizes 100 to 1000, covariate counts 20 to 5000) find that deep learning with SPSS is the most computationally efficient, while the hybrid super learner with SPSS has the lowest mean squared error. On the 401(k) real data, SVM with SPSS gives the smallest standard error and the hybrid gives the fastest runtime among the methods compared. If correct, the paper establishes SPSS as a practical replacement for random splitting that preserves the gains of Neyman-orthogonal estimation while reducing computational cost.","feed_headline":"Support-point splits speed up double machine learning","feed_subtitle":"Deep learning with support-point splitting wins on speed; the hybrid super learner wins on accuracy.","key_machinery":"The central machinery is support points sample splitting (SPSS): a subsampling method that selects $n$ points from the full sample by minimizing energy distance, $ED = \\frac{2}{n}\\sum_i \\mathbb{E}\\|\\mathbf{v}_i - \\mathbf{V}\\|_2 - \\frac{1}{n^2}\\sum_i\\sum_j \\|\\mathbf{v}_i - \\mathbf{v}_j\\|_2 - \\mathbb{E}\\|\\mathbf{V} - \\mathbf{V}'\\|_2$, so the chosen test set is an optimal representative of the data distribution rather than a random draw. The paper couples SPSS with double machine learning, which uses Neyman-orthogonal moment scores to make the target estimator insensitive to nuisance estimation error, and with three learners: SVM, deep learning, and a super learner ensemble of deep learners. The support-point convergence results (Theorems 3-6) are meant to show that the subsample converges in distribution to the original data, so nuisance functions trained on the complement inherit the needed rates.","core_discovery":"The central discovery claimed is that support points sample splitting (SPSS) is an efficient sample-splitting technique for double machine learning, providing an optimal sub-representation of the underlying data generating distribution. In the structural causal model $Y = T\\beta_0 + g_0(X) + U$ with $\\mathbb{E}[U \\mid X, T] = 0$, and $T = m_0(X) + V$ with $\\mathbb{E}[V \\mid X] = 0$, the authors estimate the average treatment effect $\\beta_0$ using Neyman-orthogonal moment scores, that is, score functions constructed so that errors in the helper regression functions $g_0$ and $m_0$ do not bias the treatment estimate. They report that DML with deep learning and SPSS outperforms SVM with SPSS in computational efficiency, and the hybrid super learner with deep learning and SPSS outperforms in estimation quality, across low-, moderate-, and big-high-dimensional settings. On the 401(k) plan data, DML with SVM-SPSS gives the smallest standard error (0.0006) and the hybrid SDL-SPSS runs fastest (0.0429 seconds), both beating the Lasso k-fold benchmark with standard error 0.0071 and time 3.487 seconds.","pith_inferences":["Editorial inference: The paper's central comparison would be stronger if it included a random-splitting control for the same three learners; without that control, the reported gains conflate the splitting method with the choice of machine learner.","Editorial inference: If SPSS preserves nuisance convergence rates, the same subsampling idea should carry over to other debiased estimators, such as doubly debiased lasso or quantile treatment effects, where random cross-fitting is currently standard.","Editorial inference: Because SPSS minimizes energy distance, it may disproportionately help small-sample and rare-stratum settings where random splits can accidentally omit important covariate regions; that hypothesis is testable by stratifying simulation results by covariate count and sample size.","Editorial inference: A practical takeaway the authors do not state is that the best choice of learner depends on the budget: SDL for accuracy, DL for speed, and SVM only in low-dimensional regimes."],"forward_implications":["For estimating average treatment effects, the hybrid super learner with deep learning and SPSS is the recommended estimator when estimation quality, measured by MSE, matters most.","Deep learning with SPSS is the recommended estimator when computational time is the constraint.","Support vector machines with SPSS are not competitive for high-dimensional causal estimation, in either accuracy or speed.","Support-point splitting can be dropped into the DML pipeline without changing the Neyman-orthogonal score framework, yielding faster or more accurate ATE estimates than random k-fold splitting.","On the 401(k) data, SPSS-based methods match or beat the Lasso k-fold benchmark on both standard error and runtime."],"supporting_citations":[{"why":"Supplies the DML framework, Neyman-orthogonal score conditions, and the random k-fold benchmark results on the 401(k) data that the paper compares against.","marker":"Chernozhukov et al. (2018)"},{"why":"Defines support points as minimizers of energy distance and supplies the convergence properties that support-point subsamples inherit.","marker":"Mak and Joseph (2018)"},{"why":"Establishes support-point-based splitting as an optimal data-splitting method, the empirical basis for SPSS.","marker":"Joseph and Vakayil (2021)"},{"why":"Provides the super learner oracle inequality that justifies constructing the hybrid SDL estimator.","marker":"Van der Laan et al. (2007)"},{"why":"Supplies the simulation nuisance functions g0 and m0 and the double machine learning implementation used in the experiments.","marker":"Bach et al. (2021)"},{"why":"Gives the energy distance definition and its properties that support points minimize.","marker":"Székely and Rizzo (2013)"},{"why":"Provides theoretical foundations for causal inference and treatment effects with high-dimensional data under Neyman orthogonality.","marker":"Belloni et al. (2017)"}],"fun_headline_variants":["Support-point splitting speeds and sharpens DML","Deep learning with SPSS: fastest DML, hybrid most accurate","SPSS beats random k-fold for causal DML","Optimal data subsampling improves double machine learning","Support points: smarter splits for faster causal ML"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that splitting data with support points instead of by random folds still gives double machine learning the statistical guarantees it needs—in particular, that the helper models' estimation errors shrink quickly enough and do not bias the treatment-effect estimate—without proving that this combined procedure satisfies those conditions.","fun_headline_variants_meta":{"raw":{"variants":["Support-point splitting speeds and sharpens DML","Deep learning with SPSS: fastest DML, hybrid most accurate","SPSS beats random k-fold for causal DML","Optimal data subsampling improves double machine learning","Support points: smarter splits for faster causal ML"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000589,"raw_usage":{"total_tokens":2818,"prompt_tokens":1049,"completion_tokens":1769,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":665,"completion_tokens_details":{"reasoning_tokens":1694}},"tokens_in":665,"tokens_out":1769,"duration_ms":16376,"temperature":1.0,"reasoning_tokens":1694,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:03:03.737379+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A paired simulation that runs the same DML estimator, such as SDL, with SPSS and with random k-fold cross-fitting on the same data-generating processes, checking 95% confidence interval coverage and bias as N grows: if SPSS-DML coverage drops below nominal or bias does not shrink at the $\\sqrt{N}$ rate, the paper's central claim is refuted.","supporting_citations":[],"review_version":1}