{"id":"100f7cf6-89ea-4c4f-bc63-33dffc474158","arxiv_id":"2607.24472","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"The Riesz representer in automatic DML is identified precisely when it uniquely optimizes a quadratic functional, enabling Riesz regression for endogenous first steps and nonlinear shape constraints.","lead":"This paper gives conditions under which the Riesz representer in debiased machine learning is identified, and shows it is exactly the unique optimizer of a quadratic functional. That characterization yields a general Riesz-regression estimator that handles endogenous nuisances and economic shape constraints such as monotonicity and convexity.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The load-bearing soft spot is not coercivity of Υ₀ (automatic in the paper's own NPIV example, Υ₀(a,b)=−E[a(W)b(W)] on L²(W)) but Assumption 3.5(ii): the ϱₙ-weighted operator-norm rate for the estimated conditional-expectation operator Π̂ₙ, which carries the entire \"endogenous first steps\" extension","rationale":"I read the main proofs in detail. Theorems 3.1–3.2 are a clean, correct application of Lax–Milgram plus Zeidler Theorem 22.A; Theorem S.2.1's necessity argument (via (S.2.2)–(S.2.4) and the bounded inverse theorem) checks out. The rate proofs (A.3)–(A.17) and the peeling argument in Theorem 3.4 are standard and internally consistent, including the derivation (A.19)–(A.22) showing that the curvature condition in Assumption 3.5(iv) forces the first-order condition to hold on an ϵ-enlargement of A₀—a subtle but legitimate step, and one verified exactly in Example 2.3 via (S.1.3)–(S.1.4), where α₀=ν̄_m is the unconstrained global maximizer over all of L²(W). Lemma S.2.2's Céa-type bound is fine. So I find no derivation error, and my concern is not a correctness objection but a gap between what is assumed and what is deliverable: the strongest claim advertises \"rates for endogenous... first steps,\" and the unique new ingredient that the endogenous extension requires beyond the exogenous theory is Assumption 3.5(ii). That condition is high-level, unverified for any concrete Π̂ₙ, stronger (operator norm over a growing sieve, inflated by ϱₙ) than the regression-rate literature the paper cites, and it sits directly upstream of the o_p(n^{−1/4}) requirement in Assumption 3.8(iii). This is why I only partially agree with the reader: coercivity is a real case-by-case condition but is honestly flagged by the authors and is automatic in their leading endogenous example; the sharper, less-examined pillar is the Π̂ₙ op-norm rate. Because the paper already concedes the open status of practical verification and the simulations display the symptom (residual undercoverage), this concern supports rather than worsens the reader's CONDITIONAL verdict: the mathematics stands, but the endogenous-rate claim's practical applicability is conditional on an assumption no one has yet verified. The proposed test is cheap (the DGP's Π₀ is essentially closed-form) and would settle whether the concern lands in the paper's own flagship design.","tokens_in":52021,"tokens_out":7547,"duration_ms":253915,"concrete_test":"In the NPIV DGP of §4.2.2 (d_c=d_f=5, η=0.5, ρ=0.1), Π₀(δ)=E[δ(D,Z)|S,W] is computable to arbitrary accuracy by numerical quadrature/Monte Carlo because all primitives are Φ-transforms of jointly Gaussian variables. For the adaptive-basis ridge Π̂ₙ actually used, estimate ϱₙ‖Π̂ₙ−Π₀‖op,n by maximizing ‖Π̂ₙ(δ)−Π₀(δ)‖_{L²(S,W)} over a fine grid of sieve directions δ∈∆Γₙ with ‖δ‖_H≤1, at n=1000 and, say, n=4000. If this term decays slower than n^{−1/4} (or slower than δₙ), Assumption 3.8(iii) fails through the α̂ channel and Theorem 3.4's rate, though correct, does not deliver the inference requirement in this regime—explaining the residual undercoverage in Figure 5 as Π̂ₙ error rather than anything shape constraints could fix. If it is comfortably o(n^{−1/4}), the concern does not land and the undercoverage must be attributed to γ̂ or to first-step influence error instead.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader locates the weakest assumption in coercivity of Υ₀ (Eq. 11). That is defensible in general (coercivity fails if Υ₀(a,a)=−E[νρ(X)a(W)²] with νρ approaching zero, e.g., quantile residuals with vanishing conditional density), but it is not the sharpest soft spot for this paper's strongest claim, because (i) the paper is transparent that coercivity is necessary in the strengthened sense (Remark 3.1, Theorem S.2.1, proof checked and correct), and (ii) in the flagship endogenous Example 2.3, Υ₀(a,b)=−E[a(W)b(W)] with A=L²(W), so coercivity holds trivially with κ=1 regardless of ill-posedness. The ill-posedness has been deliberately relocated, not eliminated: it lands on (a) the Severini–Tripathi range condition (existence of ν̄_m with E[ν̄_m(W)|Z]=ν_m(Z)), which the paper states is necessary for √n-estimability anyway, and (b) Assumption 3.5(ii), which requires an estimator Π̂ₙ of the unknown map Π₀(δ)=E[δ(Z)|W] satisfying ϱₙ‖Π̂ₙ−Π₀‖op,n=o_p(1), where ‖·‖op,n is the operator norm over the unit ball of a growing sieve ∆Γₙ and ϱₙ≥1 is the (possibly diverging) sieve bound of Assumption 3.3(i). This term enters the rate formula (29) directly, and through Theorem 3.4 it must be o_p(n^{−1/4}) for Proposition 3.1's inference theory (Assumption 3.8(iii)) to apply. The paper offers no construction with a verified op-norm rate: it says Π̂ₙ \"may be constructed by nonparametric regression methods\" and cites Schmidt-Hieber (2020), Farrell et al. (2021), Kohler and Langer (2021), but those deliver L² regression rates for a single target function, not sup-norm-over-sieve-directions operator rates amplified by a diverging ϱₙ in a high-dimensional instrument space (W has dimension up to 11 in the simulations). Uniform control of Π̂ₙ over all directions of an expanding, weakly bounded sieve is qualitatively harder and is precisely where the ill-posedness of the endogenous problem re-enters. Consistent with this, the paper's own NPIV simulations show coverage of only ~50–80% even","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper studies a GMM parameter θ0 depending on a high-dimensional nuisance γ0, with orthogonalization achieved through a Riesz representer α0 entering linearly in the first-step influence function. Assumption 3.1 decomposes the relevant directional derivatives through maps Π0, Ψ0, and a bilinear form Υ0. Theorem 3.1 uses Lax–Milgram to identify α0 under coercivity, and Theorem 3.2 shows, under symmetry and a definite sign, that this is equivalent to uniquely optimizing a quadratic functional. Section 3.2 develops penalized sieve Riesz regression, including the case where the operator Π0 associated with endogeneity must itself be estimated, while Section 3.3 gives a cross-fitted debiased GMM central limit theorem. Simulations and two applications examine monotonicity and convexity restrictions.","tokens_in":52539,"tokens_out":12180,"duration_ms":483471,"significance":"If the requested scope issues are addressed, this is a valuable unifying contribution. The identification/quadratic-characterization results clarify the foundations of automatic DML rather than merely proposing another learner. Particular strengths are the necessary-and-sufficient coercivity analysis in Theorem S.2.1, the clean separation of analytic and probabilistic assumptions, complete Lax–Milgram/Céa-type arguments, and rate results formulated through moduli of continuity and critical radii. The accommodation of nonlinear shape sets, estimated Π0, and partially identified first steps is potentially important. The numerical work compares against relevant unconstrained and Callaway–Sant’Anna benchmarks and supplies useful evidence, although it does not yet substitute for the missing concrete endogenous-rate verification.","major_comments":[{"comment":"The endogenous extension rests on ϱn‖Π̂n−Π0‖op,n entering Eq. (29), and Proposition 3.1 requires the resulting α̂ error to be op(n−1/4) via Assumption 3.8(iii). For the flagship NPIV case, the paper only says that Π̂n may be obtained by nonparametric regression and cites pointwise/regression-rate papers. Those rates do not directly give an operator-norm rate uniform over the growing sieve ∆Γn; the adaptively learned basis used in §4.2.2 adds further dependence. Please provide primitive conditions and a verified rate for at least one concrete Π̂n, preferably the ridge/adaptive-basis construction simulated, or explicitly present this as an unverified feasibility condition and qualify the endogenous-scope claim. This does not appear to make Theorem 3.4 false, but it leaves its main new application conditional.","section":"§3.2, Assumption 3.5(ii), Eq. (29); §4.2.2"},{"comment":"The theory requires ∆Γn⊂lin(Γ−Γ) and A0,n=Π0(∆Γn). Thus, when Γ is a monotone or convex class, α generally lies in a difference of two constrained functions, not in the original constrained class. Sections 4.2.1–4.2.2 say γ0 and α0 are estimated using the same partially monotonic/convex architectures, while Eq. (57) correctly optimizes over Γ⌣−Γ⌣. A single constrained network need not satisfy Assumption 3.3(ii) for α0. Please specify the exact computational parameterization—e.g., α=α+−α− with both components constrained, followed by Π0 or Π̂n—and establish that the reported SDML estimates use it. In leading examples Γ−Γ may also be dense in L2, so the precise sense in which α is regularized should be clarified.","section":"§3.2, definition of A0,n; §4.2.1–4.2.2; §5.2, Eq. (57)"},{"comment":"The introduction and conclusion motivate shape constraints as improving precision and mitigating the curse of dimensionality, but Theorems 3.3–3.4 are sieve-neutral: they do not compare constrained and unconstrained approximation errors, entropy integrals, or critical radii. No worked high-dimensional example verifies that the constrained neural sieves used in §4 achieve the required n−1/4 nuisance rates, and the NPIV coverage at pc=5 in Figures 5–6 remains well below nominal. Either add a concrete constrained-versus-unconstrained rate calculation supporting the curse-of-dimensionality language, or state the narrower claim that the theory permits shape constraints and that simulations provide finite-sample regularization evidence.","section":"§1 and §6; Theorems 3.3–3.4; Figures 5–6"}],"minor_comments":[{"comment":"The aggregate pre-treatment estimates are significantly nonzero at several leads, including event times −9 and −8. Although this does not affect the econometric theory, the text should explicitly caution that these diagnostics are in tension with the conditional parallel-trends assumption in Eq. (47) and avoid a strong causal interpretation of the post-treatment contrasts.","section":"§5.1.2, Figures 7–9"},{"comment":"The definitions of U0,n(δ) contain an unusual lower bound of order δ² and are later related to peeling shells. A short sentence explaining the role of the puncturing and its relation to Sn,j,M would make the proofs easier to follow.","section":"§3.2, Eqs. (25) and (27)"},{"comment":"The numerical description gives broad choices (AdamW, dropout, early stopping, ℓ2 penalties) but not enough detail to reproduce the reported λn/Jn calibration, learned NPIV basis, or constrained-head construction. Replication code or a detailed implementation appendix would substantially strengthen the computational contribution.","section":"§4.1–§4.2"},{"comment":"Please explain why the numbers receiving positive weight differ from the cleaned sample sizes, and whether this trimming is common across DML and SDML. It would also help to report the normalized weight definition explicitly.","section":"§5.2, Table 3"},{"comment":"In the literature paragraph, “Chernozhukov et al. (2024a,b) Singh (2024)” is missing a conjunction or punctuation. Some figure captions should also repeat the number of Monte Carlo replications and mark the nominal 95% coverage line.","section":"§1; Figures 3–6"}],"recommendation":"major_revision","confidential_remarks":"The main identification and estimation arguments appear technically sound, and I would not treat coercivity itself as an undisclosed circularity problem: the paper is transparent about its necessity in a strengthened sense, and it is automatic in the leading NPIV specialization. The revision is instead about demonstrated scope. The advertised endogenous and shape-constrained applications currently depend on high-level operator-rate and sieve-approximation conditions that are not instantiated for the paper's own constructions. These points should be fixable, but they require more than stylistic revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline: this paper closes the identification gap Chernozhukov et al. left open for the Riesz representer, and it does so in a way that actually covers endogenous first steps and nonlinear shape constraints. Theorems 3.1–3.2 are the core—Lax–Milgram plus an iff link to the quadratic C₀—and the proofs check out.\n\nWhat is new is not another Riesz regression recipe for the exogenous linear case. It is (i) conditions under which α₀ is identified from the orthogonalized influence equations when Γ can be nonlinear, (ii) the precise equivalence to unique optimization of C₀ when Υ₀ is coercive, symmetric, and definite, and (iii) rate results for the empirical maximizer that allow unknown link operators Π₀ and sieves including nets. The shape-constraint embedding is clean: constraints on γ₀ induce the right space for α₀ without forcing Γ itself to be linear. Simulations and the Medicaid/GMAT applications are honest illustrations; gains from monotonicity/convexity show up, especially in the NPIV designs where unconstrained coverage is poor.\n\nSoft spots, in proportion. Coercivity is high-level and must be checked case by case, but the paper is transparent about necessity (Remark 3.1, Thm S.2.1), and in their flagship NPIV example Υ₀(a,b)=−E[a(W)b(W)] so κ=1 is free. The sharper load for the endogenous extension is Assumption 3.5(ii): a ϱₙ-weighted operator-norm rate for Π̂ₙ over a growing sieve. They point to DNN regression rates; those control a single target in L², not sup over sieve directions amplified by a diverging bound. That is exactly where ill-posedness re-enters, and it lines up with residual undercoverage when d_c=d_f=5. Also open: data-driven λ_n, and the empirics are not a full horse race. None of this breaks the identification theory.\n\nWho it is for: people doing automatic DML, NPIV functionals, or shape-restricted causal work. Math and citations look solid; no circular construction. I would bring it to reading group, cite it when I need the identification backbone or shape-constrained Riesz regression, and send it to referees. Engage.","headline":"Solid identification foundation for automatic DML under endogeneity and shape constraints; the real soft spot is operator-norm rates for Π̂, not coercivity.","tokens_in":52407,"tokens_out":586,"would_cite":true,"duration_ms":19122,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"The Riesz representer in debiased machine learning is identified exactly when it uniquely optimizes a quadratic functional, unlocking automatic estimation under endogeneity and shape constraints.","keywords":["debiased machine learning","Riesz representer","Riesz regression","shape constraints","deep learning","nonparametric instrumental variables","orthogonal moments"],"falsifier":"In a known endogenous or shape-constrained design where the bilinear form fails coercivity (or is only weakly coercive), check whether multiple distinct representers satisfy the orthogonalized equations and whether Riesz regression recovers a unique, rate-consistent estimator; if uniqueness or the claimed rates still hold, the characterization is wrong.","tokens_in":51973,"feed_emoji":"📐","tokens_out":922,"duration_ms":23323,"temperature":0.7,"pith_summary":"Debiased machine learning needs a second nuisance—the Riesz representer—to cancel first-step estimation bias when the target parameter depends on a high-dimensional nuisance. This paper shows when that representer is identified from the orthogonalized moment conditions, and proves that identification holds if and only if the representer is the unique optimizer of a known quadratic functional. That characterization yields a general Riesz-regression estimator that works for endogenous first steps (such as nonparametric IV) as well as exogenous ones, and that can use classical sieves or deep nets. Shape restrictions on the first-step nuisance are folded into a (possibly nonlinear) parameter space so that the same theory covers them and can shrink estimation error in high dimensions. Simulations and two empirical applications illustrate that the constraints improve precision and coverage when they are economically motivated.","feed_headline":"Riesz representer identified by a quadratic optimum","feed_subtitle":"Automatic debiased ML now covers endogenous first steps and economic shape constraints with rates","key_machinery":"The quadratic functional C₀ (and its sample Riesz-regression counterpart). It depends only on the original moment and the first-step influence function, not on a closed form for α₀, so optimizing it automatically recovers the representer once coercivity, symmetry, and definiteness hold.","core_discovery":"Under standard smoothness and a linearity condition on the first-step influence function, the Riesz representer α₀ is the unique solution to the orthogonalized influence equations over shape-respecting perturbations precisely when a bilinear form Υ₀ is coercive; when Υ₀ is also symmetric and definite, that uniqueness is equivalent to α₀ uniquely maximizing or minimizing the quadratic functional C₀(α) = ½Υ₀(α,α) + Ψ₀(α). This equivalence is the foundation for automatic Riesz regression with rates that cover endogenous and shape-constrained first steps.","pith_inferences":["The same quadratic characterization may extend to other linear-in-α influence structures (generated regressors, dynamic discrete choice, support-function inference) once coercivity is checked.","Shape constraints act as a form of regularization that can shrink partially identified sets for the first step, potentially stabilizing selection of a representative γ₀ before debiasing.","Data-driven choice of the Riesz-regression penalty and network depth remains open; the theory only controls rates once those tunings satisfy the critical-radius conditions."],"forward_implications":["Automatic DML can be run for functionals of NPIV and other endogenous first steps without deriving a closed-form representer.","Monotonicity, convexity, and other economic shape restrictions can be built into the sieve or network for both the first step and the induced space for the representer, with the same identification and rate theory.","Cross-fit debiased GMM based on the estimated representer is asymptotically normal under the usual faster-than-n^{-1/4} nuisance rates (or product rates under double robustness).","Simulation and application evidence indicate that adding more valid shape constraints tends to cut bias, variance, and interval length while preserving coverage."],"fun_headline_variants":["Riesz representer is unique quadratic optimum","Quadratic functional pins down Riesz representer for DML","Riesz regression identifies α₀ under shape constraints","Unique quadratic maximizer yields automatic Riesz representer","Endogenous first steps allowed when Υ₀ is coercive"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The bilinear form that appears in the first-step influence must be coercive: its quadratic form stays bounded away from zero by a positive multiple of the squared norm of the representer; without that, uniqueness of the representer fails.","fun_headline_variants_meta":{"raw":{"variants":["Riesz representer is unique quadratic optimum","Quadratic functional pins down Riesz representer for DML","Riesz regression identifies α₀ under shape constraints","Unique quadratic maximizer yields automatic Riesz representer","Endogenous first steps allowed when Υ₀ is coercive"]},"model":"grok-4.5","effort":"low","cost_usd":0.004345,"raw_usage":{"total_tokens":1290,"prompt_tokens":740,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":43448000,"prompt_tokens_details":{"text_tokens":740,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":489,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":740,"tokens_out":61,"duration_ms":8183,"temperature":1.0,"reasoning_tokens":489,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T13:53:40.076745+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"In a known endogenous or shape-constrained design where the bilinear form fails coercivity (or is only weakly coercive), check whether multiple distinct representers satisfy the orthogonalized equations and whether Riesz regression recovers a unique, rate-consistent estimator; if uniqueness or the claimed rates still hold, the characterization is wrong.","supporting_citations":[],"review_version":1}