{"id":"2af9378a-8d53-4f38-bf71-571d03823fad","arxiv_id":"2603.07447","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"An inverse-probability-weighted Dirichlet kernel density estimator on the simplex is asymptotically normal under MAR missingness, with bias matching full-data Dirichlet KDE and variance inflated by a propensity factor.","lead":"This paper builds a density estimator for compositional data on the simplex when some compositions are missing at random, using inverse-probability weights and a Dirichlet kernel. It gives asymptotic theory, simulations, and an NHANES leukocyte example that finds a modal immune profile.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader’s strongest claim accurately restates Theorems 4.4 and 4.8, and the weakest-assumption note already isolates the two genuine restrictions (p < d for the feasible CLT, and π bounded away from zero). No deeper inconsistency, circularity, or unstated rate failure appears in the expansions or proofs. The Dirichlet-mixture simulation design favors the proposed kernel, but the paper qualifies the comparison as holding “for certain target densities,” and the NHANES analysis is correctly labeled descriptive. Because the mathematics holds under the conditions claimed, the ACCEPT verdict and low correctness risk stand.","tokens_in":31305,"tokens_out":531,"duration_ms":27843,"concrete_test":"Independently re-derive the leading variance term of Proposition 4.2 starting from the total-variance split (8.1), inserting only the local L2 bound of Lemma 9.1 and Bayes’ rule for E[κ^{2}_{s,b}(Y)|X]; confirm that the factor (1 + ζ(s)) appears exactly and that the o_s(n^{-1}b^{-d/2}) remainder is controlled by Assumption (A6).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims (bias/variance expansions and asymptotic normality for efn,b, and inheritance by the feasible ˆfn,b when p < d) are internally consistent under the stated assumptions. Bias of the pseudo-estimator matches the full-data Dirichlet KDE exactly (Prop. 4.1), variance inflates by the standard IPW factor (1 + ζ(s)) via the law-of-total-variance decomposition (8.1)–(8.3), and Lindeberg holds by the local bound of Lemma 9.2. For the feasible estimator the Taylor expansion of 1/ˆπ, the second-order NW moment bounds, and the L2 comparison n^{1/2}b^{d/4}(ˆfn,b − efn,b) → 0 when p < d (Thm. 4.8 / Remark 1) contain no hidden gaps. The p < d rate condition and π ≥ π_min > 0 (A5) are explicit, standard limitations of nonparametric IPW rather than unacknowledged soft spots; simulations and the NHANES illustration operate in regimes consistent with the theory or are presented only as finite-sample evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper develops inverse-probability-weighted Dirichlet kernel density estimators for compositional responses on the simplex under MAR missingness. It studies a pseudo estimator with known propensities and a feasible estimator that plugs in Nadaraya–Watson propensity estimates, derives pointwise bias/variance expansions, MSE rates, and asymptotic normality (with the feasible limit requiring p < d), and supports the theory with Monte Carlo experiments (two Dirichlet mixtures, n up to 800, missing rates up to 40%, 1000 replications) and an NHANES leukocyte-composition illustration. Simulations indicate that the IPW Dirichlet KDE outperforms IPW alr- and ilr-based KDEs for the chosen targets, and the application identifies a modal immune profile near (0.57, 0.32, 0.11).","tokens_in":31600,"tokens_out":1198,"duration_ms":20997,"significance":"If the asymptotics hold as stated, the paper cleanly extends Dirichlet kernel density estimation to incomplete compositional data without imputation, preserving nonnegativity and boundary behavior on the simplex. The bias of the pseudo estimator matches the full-data Dirichlet KDE by design of the Horvitz–Thompson weights, while MAR enters the variance through the explicit factor 1+ζ(s); the feasible estimator’s second-order variance reduction −n^{-1}ξ(s) is a useful efficiency observation. Strengths include full pointwise expansions and normality theorems, an IPW-adapted LSCV bandwidth criterion, a reproducible GitHub code repository, and a transparent real-data illustration. The contribution is incremental but well-executed for the nonparametric compositional-data literature and is of practical interest in microbiome, geochemistry, and survey settings where MAR is plausible.","major_comments":[{"comment":"Section 5.1 sets d=2 and p=2 (bivariate X and Y∈S²), so the Monte Carlo design operates at p=d. Theorem 4.8 and Remark 1 establish first-order asymptotic normality of the feasible estimator ˆfn,b only under p<d (so that the NW propensity rate is o of the Dirichlet rate). Finite-sample ISE comparisons remain informative, but the paper should either (i) add at least one configuration with p<d, (ii) explicitly flag that the reported simulations lie outside the regime of Theorem 4.8, or (iii) invoke the higher-order-kernel/Hölder extension sketched in Remark 1. Without one of these, the link between the feasible-estimator theory and the simulation design is incomplete.","section":null},{"comment":"Assumption (A5) requires π≥π_min>0 on the support of X, and Section 7 recommends flooring or stabilizing extreme weights in practice. The simulation design (logistic MAR up to 40% missing) can produce small estimated propensities, yet the manuscript does not report whether a floor, truncation, or stabilized weights were used when computing ˆfn,b or LSCV. Because inverse-probability weights drive both the estimator and the bandwidth criterion (5.2), a short statement of the numerical safeguards (or confirmation that none were needed) is load-bearing for reproducibility of Tables 1–2 and Figures 6–8.","section":null}],"minor_comments":[{"comment":"Abstract and §5.4 claim outperformance “for certain target densities.” The two Dirichlet mixtures are reasonable but narrow; a brief caveat that the ranking may reverse for densities with strong boundary mass or multimodality would temper the claim.","section":null},{"comment":"Figures 1–3 and 9 use placeholder-style glyphs in the manuscript text (e.g., boxes for axis labels). Ensure final production figures have readable axis labels, legends, and color scales; contour levels in Figures 2–3 would aid comparison of mode location and height.","section":null},{"comment":"Notation: κs,b is introduced after the full-data estimator; a one-line reminder that it is the Dirichlet kernel with parameters s/b+1 and (1−∥s∥1)/b+1 would help readers less familiar with [44].","section":null},{"comment":"Assumption (C1) on q(x)=E[κs,b(Y1)|X1=x] is used in the feasible-estimator proofs but is not discussed in the main text; a short remark that it is a standard smoothness condition on the smoothed regression of the kernel would improve transparency.","section":null},{"comment":"NHANES analysis (§6) correctly notes that survey design weights and clustering are ignored. Consider adding one sentence on how the modal composition might change under design-based weighting, or flag this more prominently as a limitation for population inference.","section":null},{"comment":"Typos/style: “Fr´ ed´ eric” and similar accented names appear with spacing artifacts in the author list; “H¨ older” in Remark 1; ensure consistent use of efn,b vs. ˆfn,b in the abstract and introduction.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The theory and proofs appear carefully done and reduce cleanly to known full-data Dirichlet expansions plus standard IPW/NW arguments; I did not find internal inconsistencies. The p=d simulation design is the main narrative gap relative to Theorem 4.8, but it is fixable without new theory. Fit for a solid methods journal is good; novelty is incremental (IPW + Dirichlet KDE) rather than conceptual, which is fine if the journal values complete asymptotic packages with code and applications."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean methods paper that does exactly what it claims: take the existing Dirichlet kernel density estimator on the simplex, put Horvitz–Thompson weights on it for MAR responses, and derive the pointwise bias, variance, MSE rates, and asymptotic normality for both the known-π and estimated-π versions.\n\nWhat is new is the combination and the full first-order theory. Bias of the pseudo-estimator is identical to the full-data Dirichlet KDE (same φ(s) term); variance picks up the usual IPW inflation 1 + ζ(s). For the feasible estimator they expand 1/π̂, control the NW remainders, and show the estimation error is negligible when p < d. That rate condition is explicit in Remark 1, not buried. The proofs reduce cleanly to Ouimet–Tolosana-Delgado plus standard total-variance and Lindeberg arguments; I do not see a gap in the stress-test sense.\n\nThey also ship an IPW-adapted LSCV, 1000-rep simulations against alr/ilr IPW competitors (Dirichlet wins on their two mixture models), and a transparent NHANES leukocyte illustration with code on GitHub. That is more than many theory-plus-simulation papers deliver.\n\nSoft spots are real but proportionate. The p < d restriction for normality of the feasible estimator is the main theoretical limitation; if covariates are high-dimensional you need higher-order kernels or a parametric propensity model. π bounded away from zero is assumed, so near-certain nonresponse regions are out. Simulations stay in d = 2, p = 2 (borderline for the theorem) and the NHANES analysis treats the sample as iid, which they flag. None of these break the central claims under the stated conditions.\n\nThis is for people who already work with compositional densities or MAR nonparametric estimation and want a ready-to-use simplex-respecting tool with asymptotics. It is not a foundational breakthrough, but it is honest incremental work that a methods journal should send to referees. I would cite it if I needed density estimation on the simplex with missing compositions, and I would bring the theory section to a reading group if we were covering asymmetric kernels or IPW.\n\nRecommendation: accept for peer review; the math and evidence are in good enough shape that referees can focus on polish and scope rather than repair.","headline":"Solid, usable IPW Dirichlet KDE on the simplex under MAR; asymptotics check out and the p < d caveat is stated honestly.","tokens_in":32203,"tokens_out":585,"would_cite":true,"duration_ms":6984,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G07","62E20","62G05","62G08","62G20","62H12"],"pacs":[],"model":"grok-4.5","headline":"Inverse-probability-weighted Dirichlet kernels estimate densities on the simplex under missing-at-random sampling without imputation.","keywords":["Dirichlet kernel","compositional data","simplex","missing at random","inverse probability weighting","kernel density estimation","Nadaraya–Watson","asymmetric kernel"],"falsifier":"Generate data from a known Dirichlet mixture on the 2-simplex with a logistic MAR mechanism whose propensity is bounded away from zero, compute the IPW Dirichlet estimator with LSCV bandwidth, and check whether the Monte-Carlo mean integrated squared error tracks the predicted n^{-4/(d+4)} rate and whether the studentized estimator is approximately standard normal; systematic failure of either check would falsify the central asymptotic claim.","tokens_in":32210,"feed_emoji":"📊","tokens_out":1031,"duration_ms":8667,"temperature":0.7,"pith_summary":"Compositional data live on the simplex and often have missing parts whose chance of being observed depends on fully observed covariates. This paper shows that you can estimate the density of those compositions by reweighting each observed point with the inverse of its observation probability and smoothing with an adaptive Dirichlet kernel that stays nonnegative and well-behaved near the boundary. When the observation probabilities are unknown they are estimated by ordinary Nadaraya–Watson regression. The resulting estimator has the same leading bias as the complete-data Dirichlet kernel density estimator; missingness only multiplies the variance by a factor that depends on the propensity score. Under standard smoothness conditions the estimator is asymptotically normal at the usual rate, provided the covariate dimension is smaller than the simplex dimension. Simulations and a leukocyte-composition example from NHANES illustrate that the procedure is competitive with log-ratio competitors and recovers a biologically plausible modal immune profile.","feed_headline":"Dirichlet kernels recover simplex densities under MAR missingness","feed_subtitle":"Inverse-probability weights keep the bias of the full-data estimator while only inflating variance","key_machinery":"The feasible IPW Dirichlet kernel estimator ˆf_n,b(s) = n^{-1} ∑ (δ_i / ˆπ_i) κ_{s,b}(Y_i), where κ_{s,b} is the adaptive Dirichlet kernel centered at s with bandwidth b and ˆπ_i is the Nadaraya–Watson estimate of the propensity score; its bias–variance expansions and asymptotic normality are derived by Taylor expansion of 1/ˆπ around 1/π together with the known L^{2} asymptotics of the Dirichlet kernel.","core_discovery":"Under a missing-at-random mechanism the inverse-probability-weighted Dirichlet kernel density estimator on the simplex has the same first-order bias expansion as the full-data Dirichlet estimator, while its variance is inflated only by the factor 1 + ζ(s) that encodes the variability of the inverse propensity weights; when the propensities themselves are estimated by Nadaraya–Watson regression the same leading asymptotics continue to hold whenever the covariate dimension is strictly smaller than the simplex dimension.","pith_inferences":["The p < d restriction suggests that, for high-dimensional metadata, practitioners will need either dimension reduction or a parametric propensity model before the asymptotic normality guarantee applies.","The second-order variance reduction term −n^{-1}ξ(s) that appears when propensities are estimated hints that mild misspecification of π may still be tolerable at the rates considered here.","Extending the same IPW-Dirichlet construction to structural zeros (common in microbiome data) would require only a zero-handling pre-step and would inherit the same bias expansion."],"forward_implications":["Density estimation for microbiome or geochemical compositions can proceed by inverse-probability weighting without first imputing missing taxa or assays.","The leading bias term is identical to the complete-data case, so existing bandwidth rules for Dirichlet kernels remain asymptotically valid under MAR.","When covariates are low-dimensional relative to the simplex, nonparametric propensity estimation does not degrade the first-order rate of the density estimator.","The same weighting-plus-Dirichlet construction immediately supplies a density estimate whose mode can be read as a typical compositional profile (as done for NHANES leukocytes)."],"fun_headline_variants":["IPW Dirichlet kernels match full-data bias on simplex under MAR","Dirichlet KDE on simplex keeps bias via inverse-propensity weights","Same bias expansion for IPW Dirichlet density estimator under MAR","Inverse weights inflate only variance of simplex Dirichlet KDE","Nadaraya-Watson propensities preserve Dirichlet KDE asymptotics"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The observation probability must stay bounded away from zero, and the number of continuous covariates must be smaller than the dimension of the simplex, otherwise the error from estimating the propensities swamps the density estimate.","fun_headline_variants_meta":{"raw":{"variants":["IPW Dirichlet kernels match full-data bias on simplex under MAR","Dirichlet KDE on simplex keeps bias via inverse-propensity weights","Same bias expansion for IPW Dirichlet density estimator under MAR","Inverse weights inflate only variance of simplex Dirichlet KDE","Nadaraya-Watson propensities preserve Dirichlet KDE asymptotics"]},"model":"grok-4.5","effort":"low","cost_usd":0.004666,"raw_usage":{"total_tokens":1305,"prompt_tokens":743,"num_sources_used":0,"completion_tokens":90,"cost_in_usd_ticks":46660000,"prompt_tokens_details":{"text_tokens":743,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":472,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":743,"tokens_out":90,"duration_ms":3653,"temperature":1.0,"reasoning_tokens":472,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T13:15:16.930017+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Generate data from a known Dirichlet mixture on the 2-simplex with a logistic MAR mechanism whose propensity is bounded away from zero, compute the IPW Dirichlet estimator with LSCV bandwidth, and check whether the Monte-Carlo mean integrated squared error tracks the predicted n^{-4/(d+4)} rate and whether the studentized estimator is approximately standard normal; systematic failure of either check would falsify the central asymptotic claim.","supporting_citations":[],"review_version":1}