{"id":"39d5ba87-0331-4794-94a7-08d87f15ae66","arxiv_id":"2602.22083","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Discretizing a continuous mediator in causal functionals induces first-order approximation bias; a within-bin mean correction reduces it to second order.","lead":"Using bins to discretize a continuous mediator in causal effect estimation changes the target parameter and creates bias that only shrinks linearly with bin width. The paper shows a simple correction—evaluating the outcome model at the average mediator value inside each bin—reduces this approximation error to second order.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"EIF derivation in Theorem 5.2 omits the pathwise derivative of the kernel normalizing denominator; one-step asymptotic-linearity claim is unproven.","rationale":"The paper's central population-level rate results (Lemmas 3.1 and 4.1) are mathematically sound: the debiased functional replaces the within-bin conditional expectation by the outcome regression at the within-bin mean, leaving a Jensen gap bounded by curvature times within-bin variance. The invalid inequality in Lemma 3.1's proof can be repaired by replacing |mk(a1,c)-mk(a0,c)| with the bin width, so it is not the most load-bearing issue. The twice-differentiability assumption is explicit and standard, though it is violated in the applied B_PROUD example where the mediator has a point mass at zero; that is an applicability limitation rather than an internal inconsistency. The most serious concrete error is in the influence-function derivation for the smoothed one-step estimator, which is a claimed contribution and affects the paper's practical guidance about valid inference. This warrants the same conditional verdict the reader reached: the bias-reduction core can stand, but the one-step claims need correction and re-validation.","tokens_in":22107,"tokens_out":17238,"duration_ms":176292,"concrete_test":"Numerically verify the Gateaux derivative of ψ~_{h,b}(Q) along a parametric submodel that tilts the conditional density of M|A=a1,C by (1+εφ(M)) with ∫φ dP=0. Compare d/dε ψ~_{h,b}(Q_ε) at ε=0 with E_{Q0}[ϕ~_{h,b}(Q0) S(O)] using the paper's ϕ~ and the corrected (Y-μ)ω version. If the paper's formula fails the identity, Theorem 5.2 is invalidated and the one-step estimator requires re-derivation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The smoothed debiased functional in Eq. (15) defines μ_{b,k}(a1,c) using weights ω_{b,k}(m|a1,c)=K_b(m-m_k(a0,c))/E[K_b(M-m_k)|A=a1,C=c]. The denominator D_k(c) is a functional of the conditional law of M|A=a1,C=c. In Appendix B.6, Part (I), the influence function for μ_{b,k} is derived as Dfix_μb,k = I(A=a1)/π(a1|C){Yω_{b,k} - μ_{b,k}}, treating ω as fixed. But along a submodel that tilts P_{Y,M|A=a1,C} with score s, the derivative of μ_{b,k}=N/D is E[(Y-μ_{b,k})K_b/D · s]/D = E[(Y-μ_{b,k})ω_{b,k} · s]. The correct gradient is therefore (Y-μ_{b,k})ω_{b,k}, not Yω-μ. The difference μ_{b,k}(ω_{b,k}-1) has conditional mean zero but is not orthogonal to the tangent space; e.g., if Y is independent of M, μ_{b,k} is constant but their IF can yield a nonzero derivative under a tilt of M. The later chain-rule correction through m_k in Eq. (19) does not fix this, since m_k is a functional of P_{M|A=a0,C}, while the denominator depends on P_{M|A=a1,C}. Consequently, Theorem 5.2's EIF is incorrect, and the asymptotic linearity claim for the one-step estimator is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the population-level approximation error induced by discretizing a continuous mediator in causal functionals of the form θ(Q)(c)=∫ µ(m,a₁,c) f_{M|A,C}(m|a₀,c) dm, as used in mediation and front-door estimands. It defines the naive coarsened functional θ_h, shows under one-time differentiability that its coarsening error is O(w_max,K) (Lemma 3.1), and proposes a debiased coarsened functional θ~_h that evaluates the outcome regression at the within-bin conditional mean under A=a₀, achieving O(w²_max,K) under two-time differentiability (Lemma 4.1). A smoothed variant θ~_{h,b} is introduced to restore pathwise differentiability, and an influence-function-based one-step estimator is derived in Theorem 5.2. Simulations and a stroke-data application compare plug-in and one-step estimators for the naive and debiased functionals.","tokens_in":22553,"tokens_out":7505,"duration_ms":78848,"significance":"If the main approximation results are correct, the proposed debiased functional is a simple and practically valuable correction: binning a continuous mediator can be made nearly bias-free by a within-bin-mean evaluation of the outcome regression, reducing coarsening error from first order to second order in bin width. The population-level decomposition (coarsening error versus estimation error) is clearly articulated, and the simulation design in Section 6 usefully isolates the population coarsening error by using a large Monte Carlo sample. The paper does not provide machine-checked proofs or reproducible code for all experiments, but the Taylor-expansion arguments behind Lemmas 4.1 and 5.1 are transparent. However, the influence-function derivation in Theorem 5.2 contains a load-bearing error, and the statistical-estimation contribution is therefore not reliable as written.","major_comments":[{"comment":"The EIF for μ_{b,k}(a₁,c) is incorrect because the denominator D_k(c)=E[K_b(M−m_k(a₀,c))|A=a₁,C=c] is itself a functional of P_{M|A=a₁,C}. In Part (I) the derivation treats ω_{b,k} as fixed and then adds a chain-rule correction only for m_k(a₀,c), but that correction does not account for the pathwise variation of the normalizing denominator. Along a submodel with score s for the conditional law under A=a₁, the correct gradient is E[(Y−μ_{b,k})ω_{b,k}·s], so the EIF term should be (Y−μ_{b,k})ω_{b,k}, not Yω_{b,k}−μ_{b,k}. The difference μ_{b,k}(ω_{b,k}−1) has conditional mean zero but is not orthogonal to the tangent space; for example, if Y is independent of M, tilting the law of M can make this term contribute a nonzero pathwise derivative while the true functional is unchanged. Consequently, Theorem 5.2's displayed EIF is not the efficient influence function, and the asymptotic-lineari","section":"§5.2, Theorem 5.2 and Appendix B.6, Part (I)"},{"comment":"The inequality |µ_k(a₁,c)−µ_{k,a₁}(a₀,c)| ≤ L(c)|m_k(a₁,c)−m_k(a₀,c)| is not valid in general. Two distributions can have the same conditional mean inside a bin while giving different expectations of a function with bounded derivative; e.g., with a tent-shaped µ on [0,1] (slope ±L), P₀ putting mass 1/2 at 0 and 1/2 at 1, and P₁ a point mass at 1/2, both means are 1/2 but the expectations differ by L/2. This is a step in the proof of Lemma 3.1, although the resulting O(w_max,K) bound is recoverable by replacing the inequality with a bound such as |µ_k(a₁,c)−µ_{k,a₁}(a₀,c)| ≤ 2L(c)w_k(c). The proof should be corrected.","section":"Appendix B.2, Eq. (37)"},{"comment":"The paper derives a one-step estimator for the smoothed functional θ~_{h,b} (Theorem 5.2) but never simulates this estimator. The one-step estimators used in Section 6, in particular ψ~⁺_{h2} obtained by replacing θ(Q̂) with θ~_h(Q̂) in Eq. (20), target the non-smoothed debiased functional θ~_h, which Section 5 explicitly states is not pathwise differentiable in the nonparametric model. Thus the theoretical guarantees of Section 5 do not cover the debiased one-step estimator whose finite-sample performance is reported. Either the simulations should use the smoothed estimator whose theory is developed, or the claims about one-step estimation for the non-smoothed functional should be clearly labeled as heuristic without asymptotic justification.","section":"§5.2 and §6"}],"minor_comments":[{"comment":"The phrase 'when m_k(a₀,k) is fixed' appears to contain a typo; it should read 'm_k(a₀,c)'.","section":"§5.2, Theorem 5.2 statement"},{"comment":"The final sentence states '∆_h(Q)(c)=O(1/K²)'; this should be '˜∆_h(Q)(c)=O(1/K²)' for the debiased functional.","section":"Lemma 4.1, final sentence"},{"comment":"The theoretical scaling O(1/K) and O(1/K²) is derived under equal-width bins, but the simulations use equal-frequency bins. Equal-frequency bins need not have equal widths, especially under skewed mediator distributions; the text should clarify why the equal-width theory is expected to apply or provide a separate argument.","section":"§6, Simulation #1 and #2"},{"comment":"The one-step formulas for ψ⁺ and ψ⁺_h are stated without derivation. Since these estimators play a key role in the simulation comparisons, a reference to the standard derivation or a brief appendix entry would improve readability.","section":"§6, Eqs. (20)–(21)"},{"comment":"The outcome mRS is ordinal but treated as continuous. A sentence acknowledging this simplification and its potential impact on the front-door estimand would be appropriate.","section":"§7, real data application"}],"recommendation":"major_revision","confidential_remarks":"The central coarsening-bias result (Lemmas 3.1 and 4.1) appears defensible and likely fixable, but the influence-function derivation in Theorem 5.2 is not merely a presentation issue: the stated EIF is wrong because the pathwise variation of the normalizing denominator is omitted. The fix is conceptually straightforward, but it changes the displayed estimator and any downstream variance estimator. I would send the paper back for a substantive revision of Section 5.2 and Appendix B.6, and I would ask the authors to either simulate the smoothed one-step estimator or sharply delimit the claims for the non-smoothed version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core message is worth knowing: coarsening a continuous mediator introduces a first-order population-level bias, and evaluating the outcome regression at within-bin conditional means knocks that down to second order. That is the midpoint rule, but it has not been formalized in this causal-inference context, and the paper gives clean rates and a sensible practical estimator. The plug-in version is simple, works in simulations, and is likely what applied researchers will actually use.\n\nThe soft spots are real but localized. The proof of Lemma 3.1 has an invalid inequality: it bounds the difference between conditional expectations by the difference in conditional means times L, which is not generally true. The bound still holds by a Lipschitz argument, so this is fixable. The bigger problem is Theorem 5.2 and the one-step estimator for the smoothed functional. The stress test is correct: the derivation treats the weight omega as fixed, but omega contains the normalizing denominator D which depends on the conditional law of M given A=a1,C. The correct gradient for mu_b,k is (Y - mu_b,k)omega, not Y omega - mu_b,k. The extra term mu_b,k(omega - 1) has conditional mean zero but is not orthogonal to the tangent space, so the EIF is wrong. The chain-rule correction through m_k does not fix this because m_k lives on the a0 side while the denominator is on the a1 side. Concretely, if Y is independent of M, the true derivative is zero but the proposed IF gives a nonzero derivative under a tilt of M. The smoothed estimator is also never simulated, so the one-step claims are unsupported as written.\n\nThere are also strong assumptions about Op(n^{-1/2}) estimation error for fixed h, which do not generally hold with nonparametric ML. That is more of a limitation than a flaw.\n\nBottom line: the debiased plug-in functional is a solid, useful contribution for applied semiparametric work. The influence-function section needs correction or should be downgraded to a heuristic. The paper deserves a serious referee, but it should not be accepted until the EIF is fixed or the one-step claims are removed. I would cite the plug-in result if I worked on discretization in mediation/front-door settings.","headline":"Useful, simple bias correction for discretized causal functionals; the influence-function part has a real gap and needs repair.","tokens_in":22901,"tokens_out":3699,"would_cite":true,"duration_ms":35590,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62G05","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Discretizing a continuous mediator in causal functionals induces a first-order coarsening bias, and evaluating the outcome regression at within-bin conditional means removes that leading term.","keywords":["coarsening bias","mediator discretization","causal functionals","within-bin conditional means","second-order approximation","influence function","front-door functional","mediation analysis"],"falsifier":"Generate data with a known mediator–outcome regression that is continuous but not differentiable at a bin boundary (e.g., μ(m) = |m| or μ(m) = max(0,m) inside a bin), fit the debiased coarsened functional for K=2,4,8,..., and measure its error against the exact integral. If the error decays at first order, or if it does not decay at all, the claim that the within-bin-mean correction eliminates the leading term fails.","tokens_in":22022,"feed_emoji":"📉","tokens_out":9223,"duration_ms":83677,"temperature":0.7,"pith_summary":"The paper studies what happens when a continuous mediator is binned before computing causal functionals such as the mediation functional E(Y(a1, M(a0))) or the front-door functional E(Y(a0)). It shows that the naive discretized version of the integral is a different population parameter: its bias relative to the true functional is first order in the bin width, O(w_max), even when identification and nuisance estimation are perfect. The proposed fix replaces, in each bin, the within-bin outcome mean with the outcome regression evaluated at the conditional mean of the mediator under the treatment-referent level, m_k(a0,c). A Taylor expansion removes the leading term, leaving a second-order error O(w_max^2) — O(1/K^2) for equal-width bins — under twice-differentiability of the outcome regression. A kernel-smoothed variant remains second-order in bin width and bandwidth while restoring pathwise differentiability, so one-step estimators and confidence intervals become available.","feed_headline":"Within-bin means turn first-order discretization bias into second order","feed_subtitle":"A bin-level fix removes the leading coarsening bias in mediation and front-door estimates.","key_machinery":"The key object is the within-bin conditional mean m_k(a,c) = E(M | A=a, C=c, bin k), used as the expansion point in a Taylor series of the outcome regression μ(·, a1, c). The exact coarsening-error identity, Δ_h(Q)(c) = Σ_k {μ_k(a1,c) − μ_{k,a1}(a0,c)} g_k(a0,c), isolates the bias; the debiased functional replaces μ_k(a1,c) with μ(m_k(a0,c), a1,c), so the difference becomes a centered second-order remainder bounded by the second derivative of μ and the within-bin variance, which is at most w_k^2/4. A kernel-smoothed local average around m_k(a0,c) restores pathwise differentiability, enabling influence-function-based estimation.","core_discovery":"The paper establishes that the coarsening error of the naive discretized functional, Δ_h(Q)(c) = Σ_k {μ_k(a1,c) − μ_{k,a1}(a0,c)} g_k(a0,c), is first order in the bin width, and that replacing the within-bin outcome mean μ_k(a1,c) by μ(m_k(a0,c), a1,c) — the outcome regression evaluated at the treatment-a0 within-bin conditional mediator mean — removes the leading term. The remaining error is bounded by half the sup curvature of μ times the within-bin variance, giving O(w_max^2); with equal-width bins this is O(1/K^2). The paper further shows that a kernel-smoothed version of the corrected functional has combined error O(w_max^2 + b^2) and is pathwise differentiable, so one-step estimators c","pith_inferences":["The same within-bin-mean correction should apply to any causal functional that integrates a smooth regression against a reference conditional distribution — for example, other path-specific effects or g-computation formulas that currently discretize continuous covariates — as long as the reference-level bin means are estimable.","The covariance view of the bias (Remark 3.2) suggests a practical diagnostic: estimate the within-bin covariance between the outcome surface and the treatment-induced density ratio; a large nonzero covariance predicts that the naive coarsened estimate will be materially biased and the debiased version is needed.","A testable extension: in an applied dataset, compute both naive and debiased coarsened estimates at several bin counts and compare how the difference shrinks; the paper's simulations show the predicted 1/K versus 1/K^2 decay, which practitioners can reproduce to decide whether binning is safe.","Flagged from the text: the rate example for K_n stated after the smoothed estimator appears to have the inequality direction reversed — the condition 1/K_n^2 = o(n^{-1/2}) requires K_n to grow faster than n^{1/4}, not slower."],"forward_implications":["For equal-width bins, the same precision requires roughly the square root of the number of bins: approximation error drops from 1/K to 1/K^2, so K=10 gives about the error that naive binning needs K=100 to reach.","One-step estimators built on the influence function correct statistical estimation bias but not discretization bias; the target functional itself must be corrected, which is what the debiased functional does.","With the corrected functional, the smoothed one-step estimator is asymptotically equivalent to the original undiscretized functional provided nuisance estimators converge and the bin width and bandwidth go to zero fast enough.","The construction carries over to multiple mediators with a bound in terms of each mediator's maximum bin width, so the correction is not limited to univariate binning.","Simulations with correctly specified nuisance models show the debiased plug-in estimator already achieves near-nominal coverage and small MSE, so the main benefit is available without implementing influence functions."],"fun_headline_variants":["Bin-level conditional means eliminate leading discretization bias","First-order coarsening error removed by within-bin means","Debiased coarsened functional reduces binning bias to second order","Within-bin regression removes first-order binning bias"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claim collapses if the outcome regression is not twice continuously differentiable in the mediator with uniformly bounded second derivative within each bin (Lemma 4.1), and the statistical claims further assume O_p(n^{-1/2}) nuisance estimation error for fixed bins (stated before Eq. 9).","fun_headline_variants_meta":{"raw":{"variants":["Bin-level conditional means eliminate leading discretization bias","First-order coarsening error removed by within-bin means","Debiased coarsened functional reduces binning bias to second order","Within-bin regression removes first-order binning bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001358,"raw_usage":{"total_tokens":5358,"prompt_tokens":767,"completion_tokens":4591,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":4525}},"tokens_in":511,"tokens_out":4591,"duration_ms":28450,"temperature":1.0,"reasoning_tokens":4525,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T20:48:01.934664+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate data with a known mediator–outcome regression that is continuous but not differentiable at a bin boundary (e.g., μ(m) = |m| or μ(m) = max(0,m) inside a bin), fit the debiased coarsened functional for K=2,4,8,..., and measure its error against the exact integral. If the error decays at first order, or if it does not decay at all, the claim that the within-bin-mean correction eliminates the leading term fails.","supporting_citations":[],"review_version":1}