{"id":"4699caf4-7264-4441-8b60-a784ff828191","arxiv_id":"2608.08089","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Antithetic randomization with correlation -1/(K-1) is necessary and sufficient to keep the randomization variance of cross-validation bounded in the small-bias limit, and the jointly normal antithetic scheme is minimax optimal within a general class.","lead":"This paper proves that, for randomized cross-validation in a normal means setting, the only way to keep the randomization part of the variance bounded as the bias shrinks is to use antithetic randomization with pairwise correlation -1/(K-1). It also shows that among a general family of antithetic schemes, the jointly normal one is minimax optimal, and that a control variate restores bounded variance for non-smooth estimators.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.1's minimax optimality covers only the constructed class M; whether a non-Gaussian scheme outside M can beat the normal scheme under Assumption 2.1 remains open.","rationale":"The reader's weakest_assumption correctly identifies the scope limitation of Theorem 4.1. My independent reading of the proofs confirms that Theorem 3.1 is sound: the R_VAR decomposition in Appendix A.1, after correcting the typographical coefficient 2√α to 2/√α in the definition of T_α (the variance calculation uses 4/α, which is the correct coefficient from Eq. (3)), yields the claimed O(1) versus Θ(1/α) dichotomy. The proof of Theorem 4.1 is internally consistent within M: the variance decomposition in Lemma A.2, the lower bound via g⋆, and the characterization of equality via D=0 ⇒ deterministic G_M ⇒ joint normality are all valid. The only substantive gap is the one the reader names: the minimax theorem and its unique-attainment result are proven for the conditional-Gaussian mixture class M, not for the full class of exchangeable normal-marginal zero-sum schemes in Assumption 2.1. Because the abstract and Section 1 explicitly say 'within this class,' this gap does not falsify any stated theorem, so the ACCEPT verdict stands unchanged. The concrete check proposed would settle whether the gap is real (a counterexample outside M) or merely apparent (a lower bound that extends to all of Assumption 2.1).","tokens_in":16761,"tokens_out":26617,"duration_ms":268382,"concrete_test":"Construct or rule out a candidate outside M: for example, take an exchangeable vector (ω_1,...,ω_K) with N(0,σ^2 I_n) marginals, pairwise correlation −1/(K−1), and Σ_k ω_k = 0 almost surely, but with a non-Gaussian copula (for K=3, n=1, embed a non-Gaussian 2D law with standard normal marginals and identity covariance into the sum-zero plane). Evaluate the asymptotic reducible variance for the linear test function g⋆(y)=y/√n+b, which equals (4/(K^2 n)) Var(Σ_k ||ω_k||_2^2), and compare it with the jointly normal value 8σ^4/(K−1). If any such scheme gives a smaller value, Theorem 4.1's optimality does not extend beyond M; if an analytical lower bound shows the value is always at least 8σ^4/(K−1), the gap is closed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 1 poses the question 'which joint law is optimal,' but Theorem 4.1 (Section 4) proves minimax optimality only over the class M of schemes from Proposition 4.1, where ω_k = M_k Z with random co-isometric M_k independent of Z and sum-zero. The lower-bound proof in Appendix A.3 uses this representation essentially: D = M_1 M_2^⊤ + I_n/(K−1), and for the linear test function g⋆(y)=y/√n+b, any deviation from D=0 adds 8(K−1)σ^4/(K n) E||D||²_F ≥ 0. This gives no bound for general exchangeable, equicorrelated normal-marginal zero-sum schemes that are not conditional-Gaussian mixtures. Thus the paper leaves open whether a non-Gaussian scheme outside M can have strictly smaller limiting reducible variance. The abstract's 'within this class' makes the theorem true as stated, but the motivating global-optimality question is not fully settled. (Separately, in Appendix A.1, T_α is printed as 2√α⟨g(Y)−Y,ω̄⟩; Eq. (3) and the subsequent 4/α variance formula require 2/√α. This is a typographical slip, not a substantive error.)","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the cross-validation estimator CV_α defined in Equation (3) for the normal means problem, where K train--test views are formed by adding and subtracting scaled randomization variables ω_k. Under Assumption 2.1 (exchangeable, marginally normal, equicorrelated ω_k), the bias of CV_α depends only on the marginal law, while the reducible variance depends on the joint law. Theorem 3.1 proves that E[Var(CV_α | Y)] is O(1) as α↓0 if and only if ρ = -1/(K-1), and Θ(1/α) otherwise. Proposition 4.1 constructs a class M of antithetic schemes via random co-isometric matrices, and Theorem 4.1 proves that within M the jointly normal scheme uniquely attains the minimax value 8σ^4/(K-1). Theorems 5.1 and 5.2 extend the analysis to piecewise-smooth estimators, where antithetic randomization improves the rate of the reducible variance to O(α^{-1/2}) and a control variate restores boundedness when discontinuities are known. The proofs are detailed and the numerical section illustrates the predicted rates.","tokens_in":16978,"tokens_out":16923,"duration_ms":167723,"significance":"If the results are correct, the paper provides a fairly complete characterization under Assumption 2.1 of which pairwise correlation keeps the randomization-induced variance bounded, and it gives a sharp minimax optimality result within the constructive class M. The uniqueness statement for the jointly normal scheme is a strong and nontrivial contribution, and the paper is careful to state Theorem 4.1 as a result 'within this class'. The proofs are explicit, the technical lemmas are stated cleanly, and code for the numerical experiments is provided. The main caveat is that Theorem 4.1 does not settle the global question of optimality over all schemes satisfying Assumption 2.1: M is a restricted family, and the paper does not show that every exchangeable zero-sum scheme with normal marginals belongs to M. This is a scope limitation rather than a mathematical error, but it should be made prominent in the final version.","major_comments":[],"minor_comments":[{"comment":"The Introduction frames the question as 'which joint law is optimal', but Theorem 4.1 proves minimax optimality only over the class M of Proposition 4.1. Since Assumption 2.1 admits exchangeable zero-sum schemes with normal marginals that need not be representable as ω_k = M_k Z with independent co-isometric M_k (for example, symmetrized mixtures of singular normal distributions), the global question remains open. I recommend adding an explicit remark after Theorem 4.1 stating that M is a strict subclass and that optimality over the full Assumption 2.1 class is not claimed.","section":"Section 1 and Theorem 4.1"},{"comment":"In the proof of Theorem 3.1, the display 'T_α := 2√α {g(Y)−Y}^T ω̄' is inconsistent with Equation (3): the correct factor is 2/√α, as the subsequent variance calculation E[Var(T_α | Y)] = 4σ^2/(αK) {1+(K−1)ρ} E‖g(Y)−Y‖^2_2 confirms. Please correct the displayed definition.","section":"Appendix A.1"}],"recommendation":"minor_revision","confidential_remarks":"The paper is well within the scope of math.ST and the core theorems appear sound. The main point to monitor in revision is the presentation of Theorem 4.1's scope: the final version should avoid any impression that the minimax optimality covers all of Assumption 2.1. The typo in Appendix A.1 should also be fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here’s my take. The paper’s main contribution is Theorem 3.1: among exchangeable schemes with normal marginals, reducible variance stays bounded iff ρ = -1/(K-1). That’s a clean necessity result, and it strengthens Liu et al. by not requiring joint normality. The proof is careful and the appendix checks out modulo a typo. Theorem 4.1 is also interesting but narrower than the introduction suggests: it shows that the jointly normal scheme is minimax within the class M of schemes generated by random co-isometric matrices. The lower bound uses the representation ω_k = M_k Z essentially, so it does not rule out a non-Gaussian scheme outside M doing better. The paper explicitly says 'within this class', so it is not a false claim, but the motivating question 'which joint law is optimal' remains open. That is a scope caveat, not a fatal flaw.\n\nMinor issue: in Appendix A.1, T_α is printed as 2√α⟨g(Y)-Y, ω̄⟩; the variance calculation needs 2/√α. The result survives, but the typo should be fixed.\n\nThe non-smooth results are a nice addition. The control variate is clearly derived and the experiments match the predicted rates. The code is available and the simulations are simple but convincing.\n\nOverall: solid paper, honestly written, with a real new theorem. The minimax claim is more modest than the framing suggests, but the necessity result alone justifies publication.\n\nFor a reader: anyone working on randomized cross-validation, data fission, or variance reduction via antithetic sampling will want to see this. Worth sending to a serious referee. I’d cite it, mostly for Theorem 3.1.","headline":"A clean necessity result for antithetic randomization, plus a minimax optimality claim that is honest but narrower than the intro implies.","tokens_in":17539,"tokens_out":2453,"would_cite":true,"duration_ms":53621,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F40","62G09","62C20"],"pacs":[],"model":"deepseek-v4-flash","headline":"For smooth estimators, cross-validation variance stays bounded only when the randomization folds are antithetic; jointly normal antithetic folds are the unique minimax-optimal scheme in the constructed class.","keywords":["antithetic sampling","cross-validation","normal means problem","variance reduction","randomized risk estimation","minimax optimality","control variates"],"falsifier":"Fix a smooth estimator $g(y)=\\frac{1}{\\sqrt n}y+b$ and a randomization scheme satisfying Assumption 2.1 with $\\rho=0$; the paper's calculation predicts $E[\\mathrm{Var}(CV_\\alpha\\mid Y)] = c/\\alpha + O(\\alpha^{-1/2})$ for a computable positive constant $c$. A direct simulation or exact moment calculation across decreasing $\\alpha$ that does not show this divergence would refute Theorem 3.1. Conversely, if an exchangeable antithetic scheme with normal marginals but a non-joint-normal law outside $\\mathcal{M}$ attains a minimax value below $8\\sigma^4/(K-1)$ on the same test function, Theorem 4.1's uniqueness would be refuted.","tokens_in":16545,"feed_emoji":"📉","tokens_out":8720,"duration_ms":87853,"temperature":0.7,"pith_summary":"This paper asks which joint distribution of the randomization variables used to build train–test folds gives the best cross-validation estimator in the normal means problem. Because bias depends only on the marginal normal law, every scheme in the paper's class has the same bias; the joint law is a free knob that controls reducible variance. The paper proves that, for smooth estimators, the reducible variance stays bounded as the perturbation $\\alpha$ goes to zero if and only if the folds are antithetic, meaning the randomization vectors sum to zero almost surely, which forces pairwise correlation $\\rho=-1/(K-1)$. It then constructs a general class of antithetic schemes and shows the jointly normal one is the unique minimax-optimal member. For non-smooth estimators, antithetic randomization still improves the divergence rate, and a control variate restores bounded variance when the jump boundaries are known.","feed_headline":"Only antithetic randomization keeps cross-validation variance bounded","feed_subtitle":"Only antithetic correlation keeps the randomization-induced variance bounded as bias vanishes; joint normality is minimax.","key_machinery":"The central object is the stacked randomization vector $(\\omega_1,\\dots,\\omega_K)$ whose marginals are $N(0,\\sigma^2 I_n)$ and whose joint law is exchangeable with equicorrelation $\\rho$. Antithetic randomization is the zero-sum constraint $\\sum_k \\omega_k=0$, equivalent to $\\rho=-1/(K-1)$. The argument's engine is a variance decomposition for quadratic forms: conditioning on $Y$, the limit of the reducible variance is $\\frac{8\\sigma^4}{K-1}E\\|H_g(Y)\\|_F^2 + \\frac{8(K-1)\\sigma^4}{K}E\\operatorname{tr}(H_g(Y) D H_g(Y)D^\\top)$, where $H_g$ is the symmetrized Jacobian of the estimator and $D=M_1M_2^\\top + I_n/(K-1)$ measures the deviation of a scheme in the construction class from joint normality. Joint normality makes $D=0$, killing the nonnegative second term, and a linear test function shows this is the best possible worst case. For non-smooth estimators with finitely many jumps, the same decomposition is replaced by a small-noise increment bound giving $O(\\alpha^{-1/2})$ under antithetic randomization versus $\\Theta(\\alpha^{-1})$ otherwise.","core_discovery":"In the normal means problem, the cross-validation estimator $CV_\\alpha$ constructed from $K$ normal randomization vectors has bias determined only by the common marginal law of the vectors, while its conditional variance given the data decomposes into an irreducible sampling term and a reducible randomization term $E[\\mathrm{Var}(CV_\\alpha\\mid Y)]$. The paper's central theorems characterize this reducible term. For weakly differentiable estimators, Theorem 3.1 shows $E[\\mathrm{Var}(CV_\\alpha\\mid Y)]$ is $O(1)$ as $\\alpha\\downarrow0$ exactly when $\\rho=-1/(K-1)$, and $\\Theta(1/\\alpha)$ for every larger equicorrelation. Theorem 4.1 then fixes the value of the best worst-case limit: over the class of antithetic schemes generated by a sum-zero co-isometric construction, $\\inf_M\\sup_g \\lim_{\\alpha\\downarrow0} E[\\mathrm{Var}(CV_\\alpha\\mid Y)] = 8\\sigma^4/(K-1)$, and a scheme attains this value if and only if the stacked randomization vector is jointly normal. The paper reads this as a complete answer within its construction class: antithetic correlation is necessary for stability, and joint normality is the minimax choice.","pith_inferences":["The paper's minimax theorem is stated within its own construction class $\\mathcal{M}$; if a future example shows an exchangeable antithetic scheme with normal marginals that lies outside $\\mathcal{M}$ and beats $8\\sigma^4/(K-1)$, the global optimality question would reopen, since the paper does not claim to close that case.","The same proof technology suggests a practical diagnostic: if a cross-validation run shows reducible variance growing like $1/\\alpha$, the randomization scheme is not truly antithetic, or the estimator has unknown jump discontinuities; inspecting the empirical variance across $\\alpha$ could detect either problem.","The control variate requires knowing the thresholds, but the hard-thresholded ridge example is fully analytic; similar closed-form conditional expectations should be derivable for any estimator whose jump boundaries are affine functions of the data, which would enlarge the class of practical targets.","Since only the marginal law matters for bias, one could mix or approximate the optimal joint law without changing the estimator's expectation, which may be useful when exact zero-sum randomization is hard to enforce, for example with odd $K$ or streaming data."],"forward_implications":["Any user of randomized data splitting for cross-validation can read off the paper's rule: choose randomization vectors whose average is exactly zero; any positive equicorrelation makes the randomization-induced variance diverge as the bias is removed.","The jointly normal antithetic scheme is the safe default among the constructed schemes: no other scheme in the class has a smaller worst-case reducible variance, and the paper identifies joint normality as the unique equality case.","For non-smooth estimators such as thresholded or sparse predictors, antithetic randomization still converts a $\\Theta(\\alpha^{-1})$ divergence into $O(\\alpha^{-1/2})$, purely from the dependence among folds and without knowing the jump locations.","When the estimator's discontinuity boundaries are known in closed form, the paper's control variate makes the reducible variance bounded again, so smooth-level performance is achievable for piecewise-smooth estimators.","Because the same marginal law fixes the bias, all schemes compared in the paper are interchangeable in expectation; the theorems isolate variance as the sole criterion for choosing the joint law."],"supporting_citations":[{"why":"Proposes the antithetic Gaussian randomization scheme for cross-validation whose variance behavior this paper extends; also supplies the plug-in extension to asymptotically normal sufficient statistics.","marker":"Liu et al. (2026)"},{"why":"Introduces the coupled bootstrap estimator $CV_\\alpha$ and proves its bias vanishes as $\\alpha\\downarrow0$, providing the baseline estimator whose reducible variance is analyzed here.","marker":"Oliveira et al. (2024)"},{"why":"Identifies the small-$\\alpha$ limit of $E(CV_\\alpha\\mid Y)$ as Stein's unbiased risk estimate, placing the randomized estimator in the SURE framework.","marker":"Stein (1981)"}],"fun_headline_variants":["Antithetic correlation is the only stable cross-validation","Cross-validation worst-case variance fixed by antithetic scheme","Antithetic randomization is minimax for cross-validation","Stable cross-validation demands antithetic folds","Joint normality wins in antithetic cross-validation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The minimax conclusion is load-bearing only for schemes generated by the paper's specific sum-zero co-isometric construction; if the best possible antithetic joint law lives outside that construction class, the claimed optimality could miss it.","fun_headline_variants_meta":{"raw":{"variants":["Antithetic correlation is the only stable cross-validation","Cross-validation worst-case variance fixed by antithetic scheme","Antithetic randomization is minimax for cross-validation","Stable cross-validation demands antithetic folds","Joint normality wins in antithetic cross-validation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000225,"raw_usage":{"total_tokens":1474,"prompt_tokens":968,"completion_tokens":506,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":434}},"tokens_in":584,"tokens_out":506,"duration_ms":6464,"temperature":1.0,"reasoning_tokens":434,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:25:52.982570+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fix a smooth estimator $g(y)=\\frac{1}{\\sqrt n}y+b$ and a randomization scheme satisfying Assumption 2.1 with $\\rho=0$; the paper's calculation predicts $E[\\mathrm{Var}(CV_\\alpha\\mid Y)] = c/\\alpha + O(\\alpha^{-1/2})$ for a computable positive constant $c$. A direct simulation or exact moment calculation across decreasing $\\alpha$ that does not show this divergence would refute Theorem 3.1. Conversely, if an exchangeable antithetic scheme with normal marginals but a non-joint-normal law outside $\\mathcal{M}$ attains a minimax value below $8\\sigma^4/(K-1)$ on the same test function, Theorem 4.1's uniqueness would be refuted.","supporting_citations":[{"cited_title":"Journal of the Royal Statistical Society Series B: Statistical Methodology , pages =","cited_arxiv_id":null,"evidence_quote":"Proposes the antithetic Gaussian randomization scheme for cross-validation whose variance behavior this paper extends; also supplies the plug-in extension to asymptotically normal sufficient statistics."}],"review_version":1}