{"id":"cee7b68b-c5fd-4652-be1c-17ab3e82881e","arxiv_id":"2510.07235","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Smoothing inverse-probability-weighted CDFs with Bernstein polynomials gives monotone, boundary-adaptive distribution estimates under missingness, with smaller variance when propensities are estimated.","lead":"This paper shows how to estimate a distribution curve from data with missing values by smoothing the standard missing-data correction with Bernstein polynomials. The method produces smooth, valid distribution estimates and can beat a kernel competitor when sample sizes are moderate to large.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 7's variance-reduction proof drops a term whose stated bound is only O_L2(n^{-1/2}); the central feasible-estimator variance claim is not established as written.","rationale":"The reader's verdict is conditional, and the proof gap in Proposition 7 is the same load-bearing concern: the step from (7.15) to (7.17) drops a term whose stated bound is O_L2(n^{-1/2}), not O_L2(n^{-1}). I also checked the Taylor-expansion/empty-cell issue; it is real but secondary, since A1-A2 make the bad event exponentially rare and conditioning would repair it. I would not escalate to REJECT because the variance-reduction phenomenon itself is plausible (e.g., the one-cell constant-propensity calculation gives the claimed n^{-1}C(y) reduction), so the result may be true but the proof as written is incomplete. The proposed analytic check settles the order of the dropped term. Thus the correct disposition remains CONDITIONAL; the reader's verdict needs no change.","tokens_in":23328,"tokens_out":15748,"duration_ms":127771,"concrete_test":"Re-derive the first term of (7.15) in the one-cell, constant-propensity model (X≡0, delta_i ~ Bernoulli(pi), Y_i ~ F). There T1 = (Delta/pi²) W with Delta = bar(delta)-pi and W = n^{-1} sum_i (delta_i - pi) 1{Y_i ≤ y}. Compute the limit n Var(T1) as n → ∞. A zero limit means the dropped term is o_{L2}(n^{-1}) and Prop. 7 can be repaired by conditioning on nonempty cells and a sharper U-statistic argument; a positive limit means the variance formula in Prop. 7 is incomplete and the stated variance reduction is incorrect. This isolates exactly the step skipped between (7.15) and (7.17).","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 7's variance expansion is the paper's key finding, but its proof contains an unclosed control of a dropped term. After decomposing the second term of (7.9) as in (7.15), the first term on the RHS, call it T1, is bounded by Cauchy-Schwarz as E[T1^2] ≲ n^{-1}, i.e. T1 = O_{L2}(n^{-1/2}) in the paper's notation; nevertheless (7.16)-(7.17) absorb T1 into an O_{L2}(n^{-1}) remainder. If Var(T1) is truly of order n^{-1}, it alters the leading variance n^{-1}ν² and the claimed n^{-1}C(y) reduction is not established. The displayed bound is too crude: it does not exploit cancellation between the i and j indices in T1, so a Hoeffding/U-statistic projection is required to decide whether the variance is o(n^{-1}). A second, fixable issue is the infinite Taylor expansion of 1/hat(pi)_i around 1/pi_i in (7.9) and 1/hat(p) around 1/p, which are undefined on the event that an occupied cell has no observed response; A1-A2 make this event exponentially small, but the conditioning is never stated. These gaps leave the central variance-reduction claim not fully proven.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Bernstein-polynomial smoothing of the inverse-probability-weighted (IPW) empirical CDF for missing-at-random (MAR) data with discrete auxiliary covariates. Two estimators are studied: a pseudo estimator using known propensities and a feasible estimator with propensities estimated nonparametrically from cell proportions. The main theoretical claims are pointwise bias and variance expansions, optimal Bernstein degree m ~ n^{2/3} with respect to MSE/MISE, asymptotic normality, and — for the feasible estimator — an explicit variance reduction relative to the pseudo estimator by a nonnegative term n^{-1}C(y). A Monte Carlo study and an NHANES application illustrate finite-sample performance. The proofs are collected in Section 7. The central finding is Proposition 7, which asserts that estimating the propensities reduces the asymptotic variance.","tokens_in":23664,"tokens_out":17171,"duration_ms":126776,"significance":"If the proof gaps are repaired, the paper makes a useful contribution: it provides a shape-preserving, boundary-adaptive CDF estimator in a practically important missing-data setting, with explicit expansions and a theoretically grounded variance reduction for estimated propensities. The availability of reproducible code and the use of published combinatorial lemmas are strengths. The variance-reduction result is the most interesting finding, but it rests on a proof step that is currently insufficient as written. The paper is within the scope of the journal and would be of interest to researchers working on smoothing methods and missing-data inference.","major_comments":[{"comment":"The proof of Proposition 7 bounds the first term on the right-hand side of (7.15), call it T1, by E[T1^2] ≪ n^{-1}, i.e., T1 = O_{L2}(n^{-1/2}), but then (7.16)–(7.17) absorb T1 into an O_{L2}(n^{-1}) remainder. A term of O_{L2}(n^{-1/2}) cannot be dropped from an O_{L2}(n^{-1}) representation. The displayed Cauchy–Schwarz bound does not exploit cancellation between the i and j indices, so it does not establish that Var(T1) = o(n^{-1}). A U-statistic-type projection is needed to decide whether T1 contributes to the leading variance. Without this, the claimed variance expansion Var( bF_{n,m}(y)) = Var( eF_{n,m}(y)) - n^{-1}C(y) + O(n^{-1}m^{-1}) is not established.","section":"§7.2, Eqs. (7.15)–(7.17)"},{"comment":"The infinite Taylor expansions of 1/\\hatπ_i around 1/π_i and of 1/\\hat p around 1/p are undefined on the event that a covariate cell has no observed response (or no observation at all). This event has exponentially small probability under Assumptions A1–A2, but the proofs of Propositions 6 and 7 never condition on its complement. The expansions and the bounds in (7.13)–(7.14) need to be stated conditional on the good event, with the bad event handled separately. This is a fixable but required step.","section":"§7.2, Eqs. (7.9)–(7.10)"},{"comment":"The step 'by the law of large numbers in L2' that replaces the double-sum term by n^{-1}∑_i (δ_i - π_i)/π_i F_{Y_i|X_i}(k/m) b_{m,k}(y) is not shown. This replacement is the crux of the influence-function representation (7.17) and deserves a detailed justification, including the treatment of the diagonal, off-diagonal, and cell-specific terms. The current proof relies on an unstated projection argument.","section":"§7.2, Eq. (7.16)"}],"minor_comments":[{"comment":"The abstract (and the summary at the top of the file) state that 'for small to moderate sample sizes, the Bernstein-smoothed pseudo and feasible estimators outperform ... the integrated version of the IPW kernel density estimator', while the full-text abstract and Section 4.3 state that the feasible estimator outperforms at moderate to large n and that the pseudo Bernstein estimator is actually dominated by the I-IPW KDE at all sample sizes. Please harmonize these claims.","section":"Abstract / Section 4.3"},{"comment":"The notation π(X_i) is used interchangeably with π_i(X_i) in the same display. Please be consistent.","section":"Eq. (7.15)"},{"comment":"The leave-one-out version bF^{(-i)}_{n,m} is computed using the full-sample weights cW_i rather than propensities estimated without the ith observation. This is an approximation to the true leave-one-out criterion; it would be helpful to state this explicitly or justify that the difference is asymptotically negligible.","section":"§4.2, Eq. (4.1)"},{"comment":"The second-order Taylor expansion writes the remainder as O(|k/m - y|^3), which formally requires F to be C^3 rather than the stated C^2. Under C^2 the remainder is o(|k/m - y|^2), which still yields the claimed o(m^{-1}) after summing against the binomial weights. Please adjust the wording or the assumption.","section":"§7.1, proof of Proposition 1"}],"recommendation":"major_revision","confidential_remarks":"The main gap is technical and likely fixable: the variance-reduction result is plausible and consistent with semiparametric efficiency intuition, but the proof as written drops a term of nominal order O_{L2}(n^{-1/2}) from an O_{L2}(n^{-1}) representation. The self-citation of Ouimet (2021) is appropriate because that lemma is published. The paper fits the journal's scope. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real contribution, not a repackaging. Combining the Bernstein operator with the IPW empirical CDF under MAR is a natural extension of Leblanc and Dubnicka, but the paper works out the details: bias and variance expansions for both pseudo and feasible estimators, the optimal degree m ~ n^{2/3}, asymptotic normality, an LSCV degree selector, simulations, and an NHANES application. The pseudo-estimator results follow the standard Bernstein-CDF route and look solid. The variance-reduction formula for the feasible estimator is the genuinely new piece, and I suspect the conclusion is true, but the proof as written does not establish it. In Eqs. (7.15)-(7.17), the first term is bounded only as O_{L2}(n^{-1/2}) and then absorbed into an O_{L2}(n^{-1}) remainder. That is not legitimate on the displayed bounds; a U-statistic/Hoeffding projection is needed to show the variance contribution is o(n^{-1}). The same section expands 1/hat(pi) and 1/hat(p) in infinite Taylor series without conditioning on the event that every occupied cell has at least one observed response. A1-A2 make that event exponentially small, so the issue is fixable, but it should be stated explicitly. Minor: the abstract overstates the simulations—their own tables show the I-IPW KDE wins for small n, and the Bernstein feasible estimator only becomes competitive for moderate to large n. Also, the promised code link is a GitHub user page rather than a clearly verifiable repository. The self-citation to Ouimet (2021) is fine; it is an externally published lemma and not a circular reliance. Overall, the central idea is sound and worth publishing after a revision that closes the Proposition 7 gap and cleans up the Taylor-expansion conditioning. Send it to a serious referee; do not desk reject.","headline":"Sensible, clearly-written extension of Bernstein CDF smoothing to MAR/IPW data; the pseudo-estimator results look right, but the paper's key variance-reduction claim for the feasible estimator has a proof gap that needs fixing before acceptance.","tokens_in":24148,"tokens_out":3267,"would_cite":true,"duration_ms":29072,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62E20","62G08","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Bernstein smoothing of inverse-probability-weighted empirical CDFs yields monotone, [0,1]-valued estimates of the true distribution under missing-at-random, and estimating the propensities rather than knowing them reduces the variance.","keywords":["Bernstein polynomial","cumulative distribution function","inverse probability weighting","missing at random","nonparametric estimation","asymptotic normality","variance reduction","cross-validation"],"falsifier":"Simulate or resample a finite-sample design where a covariate cell has positive probability p but, with non-negligible frequency, contains zero observed Y's (e.g., p ≈ n^{-1/2} with small n per cell). Under that design, check whether the feasible estimator's variance equals Var(F̃_{n,m}) − n^{-1}C(y) + o(n^{-1}) and whether n^{1/2}(F̂_{n,m} − F) is asymptotically normal. A systematic discrepancy would falsify the proof's claim.","tokens_in":23201,"feed_emoji":"📈","tokens_out":7038,"duration_ms":52588,"temperature":0.7,"pith_summary":"The paper sets out to show that applying Bernstein polynomial smoothing to the inverse-probability-weighted empirical CDF yields a smooth, monotone, [0,1]-valued estimator of the population CDF when data are missing at random. It derives exact asymptotic bias and variance expansions for two versions: one with known propensities and one with propensities estimated from discrete covariates. A central claim is that the feasible estimator has smaller variance than the pseudo (oracle) estimator by an explicit nonnegative term. The paper also establishes optimal polynomial degree selection, asymptotic normality, and a practical cross-validation procedure, with simulations and a health-survey application.","feed_headline":"Feasible smoothing beats oracle for missing-data CDFs","feed_subtitle":"Bernstein-smoothed IPW CDFs stay monotone, adapt to boundaries, and gain extra accuracy from estimating propensities.","key_machinery":"The Bernstein operator B_m(φ)(y)=Σ_{k=0}^m φ(k/m) binom(m,k) y^k (1−y)^{m−k}—a binomial-weighted average against the empirical CDF—is the engine. It turns any step function on [0,1] into a smooth polynomial that stays monotone and within [0,1] and adapts to the boundaries. The bias expansion follows from the binomial variance identity Σ(k/m−y)^2 b_{m,k}(y)=y(1−y)/m; the variance reduction follows from a known double-sum expansion Σ_{k,ℓ}((k∧ℓ)/m−y)b_{m,k}b_{m,ℓ}= −m^{-1/2}√(y(1−y)/π)+o_y(m^{-1/2}). For the feasible estimator, a Taylor expansion of 1/π̂ around 1/π yields the correction C(y)=E[(1−π_1(X_1))/π_1(X_1) F_{Y_1|X_1}(y)^2].","core_discovery":"The paper claims that, under missing-at-random with a bounded response, the Bernstein-smoothed IPW empirical CDF, F̃_{n,m} = B_m(F̃_n), has pointwise bias m^{-1}B(y)+o(m^{-1}) and variance n^{-1}σ²(y)−n^{-1}m^{-1/2}V(y)+o(n^{-1}m^{-1/2}), where B(y)=½ y(1−y)f′(y). The feasible version F̂_{n,m}, which uses propensities estimated nonparametrically from discrete covariates, has the same leading bias but variance reduced by n^{-1}C(y) with C(y)≥0, so estimating the propensity improves efficiency. The optimal degree m scales as n^{2/3} in MSE and MISE, and both estimators are asymptotically normal. The estimators are genuine CDFs—monotone, [0,1]-valued, boundary-adaptive—and correct for MAR missi","pith_inferences":["Editorial inference: the variance reduction is proved for discrete X; for continuous covariates one would need additional smoothness and a stochastic equicontinuity condition, but the qualitative effect (estimation lowers variance) is likely to carry over.","Editorial inference: the authors' advice to trim or stabilize extreme weights corresponds to the proof's need for π bounded away from 0; trimming should let the variance expansion hold with modified constants.","Editorial inference: because the estimator is boundary-adaptive, it may yield quantile estimators near 0 and 1 with lower bias than kernel-based CDF estimators, an implication not tested here.","Editorial inference: the proof's Taylor expansion of 1/π̂ is only valid when |π̂−π|<π; a careful reader will want to see the argument conditioned on non-empty cells before applying the variance formula to very small samples."],"forward_implications":["The optimal degree m ∝ n^{2/3} yields an MISE improvement from n^{-1} to n^{-4/3} in the second-order term.","Estimating propensities from discrete covariates never inflates the asymptotic variance: it removes n^{-1}C(y) with C(y)≥0.","The asymptotic normality results justify pointwise confidence intervals for F(y) when the bias is negligible (n^{1/2}/m → 0).","The leave-one-out LSCV selection rule, computable in O(m²+nm), provides a data-driven degree that performs well in simulations.","The estimator is always a proper CDF, so it needs no monotonicity or boundary post-processing."],"fun_headline_variants":["Feasible Bernstein CDFs beat oracle for missing data","Estimating propensities sharpens Bernstein-smoothed CDFs","Feasible estimation beats oracle for missing-data CDFs","Bernstein IPW CDFs: estimated propensities reduce variance","Feasible beats oracle in Bernstein smoothed CDFs"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The variance-reduction proof expands 1/π̂ in an infinite Taylor series around 1/π and requires |π̂−π| < π in every covariate cell, but it never conditions on the event that a cell contains at least one observed response; if that event fails, the expansion is invalid and the claimed variance reduction is not proved.","fun_headline_variants_meta":{"raw":{"variants":["Feasible Bernstein CDFs beat oracle for missing data","Estimating propensities sharpens Bernstein-smoothed CDFs","Feasible estimation beats oracle for missing-data CDFs","Bernstein IPW CDFs: estimated propensities reduce variance","Feasible beats oracle in Bernstein smoothed CDFs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001136,"raw_usage":{"total_tokens":4607,"prompt_tokens":850,"completion_tokens":3757,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":3678}},"tokens_in":594,"tokens_out":3757,"duration_ms":16123,"temperature":1.0,"reasoning_tokens":3678,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T11:01:39.587839+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate or resample a finite-sample design where a covariate cell has positive probability p but, with non-negligible frequency, contains zero observed Y's (e.g., p ≈ n^{-1/2} with small n per cell). Under that design, check whether the feasible estimator's variance equals Var(F̃_{n,m}) − n^{-1}C(y) + o(n^{-1}) and whether n^{1/2}(F̂_{n,m} − F) is asymptotically normal. A systematic discrepancy would falsify the proof's claim.","supporting_citations":[],"review_version":1}