Pith. sign in

REVIEW 2 major objections 6 minor 56 references

Dirichlet kernel density estimation on the simplex with missing data

T0 review · 2 major / 6 minor · reviewed 2026-07-15 · grok-4.5

Pith's one-line read Inverse-probability-weighted Dirichlet kernels estimate densities on the simplex under missing-at-random sampling without imputation.

desk verdict Solid, usable IPW Dirichlet KDE on the simplex under MAR; asymptotics check out and the p < d caveat is stated honestly. read the letter →

arxiv 2603.07447 v2 pith:3PBZE7MD submitted 2026-03-08 stat.ME math.STstat.APstat.TH

classification stat.MEmath.STstat.APstat.TH MSC 62G0762E2062G0562G0862G2062H12
keywords DirichletkernelcompositionaldatasimplexmissingatrandominverseprobabilityweightingdensityestimationNadaraya–Watsonasymmetric
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Compositional data live on the simplex and often have missing parts whose chance of being observed depends on fully observed covariates. This paper shows that you can estimate the density of those compositions by reweighting each observed point with the inverse of its observation probability and smoothing with an adaptive Dirichlet kernel that stays nonnegative and well-behaved near the boundary. When the observation probabilities are unknown they are estimated by ordinary Nadaraya–Watson regression. The resulting estimator has the same leading bias as the complete-data Dirichlet kernel density estimator; missingness only multiplies the variance by a factor that depends on the propensity score. Under standard smoothness conditions the estimator is asymptotically normal at the usual rate, provided the covariate dimension is smaller than the simplex dimension. Simulations and a leukocyte-composition example from NHANES illustrate that the procedure is competitive with log-ratio competitors and recovers a biologically plausible modal immune profile.

What carries the argument

The feasible IPW Dirichlet kernel estimator ˆf_n,b(s) = n^{-1} ∑ (δ_i / ˆπ_i) κ_{s,b}(Y_i), where κ_{s,b} is the adaptive Dirichlet kernel centered at s with bandwidth b and ˆπ_i is the Nadaraya–Watson estimate of the propensity score; its bias–variance expansions and asymptotic normality are derived by Taylor expansion of 1/ˆπ around 1/π together with the known L^{2} asymptotics of the Dirichlet kernel.

What would settle it

Generate data from a known Dirichlet mixture on the 2-simplex with a logistic MAR mechanism whose propensity is bounded away from zero, compute the IPW Dirichlet estimator with LSCV bandwidth, and check whether the Monte-Carlo mean integrated squared error tracks the predicted n^{-4/(d+4)} rate and whether the studentized estimator is approximately standard normal; systematic failure of either check would falsify the central asymptotic claim.

Watch

Extended reading notes

Core claim

Under a missing-at-random mechanism the inverse-probability-weighted Dirichlet kernel density estimator on the simplex has the same first-order bias expansion as the full-data Dirichlet estimator, while its variance is inflated only by the factor 1 + ζ(s) that encodes the variability of the inverse propensity weights; when the propensities themselves are estimated by Nadaraya–Watson regression the same leading asymptotics continue to hold whenever the covariate dimension is strictly smaller than the simplex dimension.

Load-bearing premise

The observation probability must stay bounded away from zero, and the number of continuous covariates must be smaller than the dimension of the simplex, otherwise the error from estimating the propensities swamps the density estimate.

Editorial extensions

If this is right

  • Density estimation for microbiome or geochemical compositions can proceed by inverse-probability weighting without first imputing missing taxa or assays.
  • The leading bias term is identical to the complete-data case, so existing bandwidth rules for Dirichlet kernels remain asymptotically valid under MAR.
  • When covariates are low-dimensional relative to the simplex, nonparametric propensity estimation does not degrade the first-order rate of the density estimator.
  • The same weighting-plus-Dirichlet construction immediately supplies a density estimate whose mode can be read as a typical compositional profile (as done for NHANES leukocytes).

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The p < d restriction suggests that, for high-dimensional metadata, practitioners will need either dimension reduction or a parametric propensity model before the asymptotic normality guarantee applies.
  • The second-order variance reduction term −n^{-1}ξ(s) that appears when propensities are estimated hints that mild misspecification of π may still be tolerable at the rates considered here.
  • Extending the same IPW-Dirichlet construction to structural zeros (common in microbiome data) would require only a zero-handling pre-step and would inherit the same bias expansion.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper develops inverse-probability-weighted Dirichlet kernel density estimators for compositional responses on the simplex under MAR missingness. It studies a pseudo estimator with known propensities and a feasible estimator that plugs in Nadaraya–Watson propensity estimates, derives pointwise bias/variance expansions, MSE rates, and asymptotic normality (with the feasible limit requiring p < d), and supports the theory with Monte Carlo experiments (two Dirichlet mixtures, n up to 800, missing rates up to 40%, 1000 replications) and an NHANES leukocyte-composition illustration. Simulations indicate that the IPW Dirichlet KDE outperforms IPW alr- and ilr-based KDEs for the chosen targets, and the application identifies a modal immune profile near (0.57, 0.32, 0.11).

Significance. If the asymptotics hold as stated, the paper cleanly extends Dirichlet kernel density estimation to incomplete compositional data without imputation, preserving nonnegativity and boundary behavior on the simplex. The bias of the pseudo estimator matches the full-data Dirichlet KDE by design of the Horvitz–Thompson weights, while MAR enters the variance through the explicit factor 1+ζ(s); the feasible estimator’s second-order variance reduction −n^{-1}ξ(s) is a useful efficiency observation. Strengths include full pointwise expansions and normality theorems, an IPW-adapted LSCV bandwidth criterion, a reproducible GitHub code repository, and a transparent real-data illustration. The contribution is incremental but well-executed for the nonparametric compositional-data literature and is of practical interest in microbiome, geochemistry, and survey settings where MAR is plausible.

major comments (2)
  1. Section 5.1 sets d=2 and p=2 (bivariate X and Y∈S²), so the Monte Carlo design operates at p=d. Theorem 4.8 and Remark 1 establish first-order asymptotic normality of the feasible estimator ˆfn,b only under p<d (so that the NW propensity rate is o of the Dirichlet rate). Finite-sample ISE comparisons remain informative, but the paper should either (i) add at least one configuration with p<d, (ii) explicitly flag that the reported simulations lie outside the regime of Theorem 4.8, or (iii) invoke the higher-order-kernel/Hölder extension sketched in Remark 1. Without one of these, the link between the feasible-estimator theory and the simulation design is incomplete.
  2. Assumption (A5) requires π≥π_min>0 on the support of X, and Section 7 recommends flooring or stabilizing extreme weights in practice. The simulation design (logistic MAR up to 40% missing) can produce small estimated propensities, yet the manuscript does not report whether a floor, truncation, or stabilized weights were used when computing ˆfn,b or LSCV. Because inverse-probability weights drive both the estimator and the bandwidth criterion (5.2), a short statement of the numerical safeguards (or confirmation that none were needed) is load-bearing for reproducibility of Tables 1–2 and Figures 6–8.
minor comments (6)
  1. Abstract and §5.4 claim outperformance “for certain target densities.” The two Dirichlet mixtures are reasonable but narrow; a brief caveat that the ranking may reverse for densities with strong boundary mass or multimodality would temper the claim.
  2. Figures 1–3 and 9 use placeholder-style glyphs in the manuscript text (e.g., boxes for axis labels). Ensure final production figures have readable axis labels, legends, and color scales; contour levels in Figures 2–3 would aid comparison of mode location and height.
  3. Notation: κs,b is introduced after the full-data estimator; a one-line reminder that it is the Dirichlet kernel with parameters s/b+1 and (1−∥s∥1)/b+1 would help readers less familiar with [44].
  4. Assumption (C1) on q(x)=E[κs,b(Y1)|X1=x] is used in the feasible-estimator proofs but is not discussed in the main text; a short remark that it is a standard smoothness condition on the smoothed regression of the kernel would improve transparency.
  5. NHANES analysis (§6) correctly notes that survey design weights and clustering are ignored. Consider adding one sentence on how the modal composition might change under design-based weighting, or flag this more prominently as a limitation for population inference.
  6. Typos/style: “Fr´ ed´ eric” and similar accented names appear with spacing artifacts in the author list; “H¨ older” in Remark 1; ensure consistent use of efn,b vs. ˆfn,b in the abstract and introduction.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: IPW identities and standard kernel expansions yield the asymptotics; self-citations supply independent full-data Dirichlet lemmas, not the missing-data claims.

full rationale

The derivation chain is self-contained and non-circular. Under MAR, E[efn,b(s)] = E[κs,b(Y1)] by conditional independence and E[δ|X]=π(X), so Bias[efn,b] equals the full-data Dirichlet bias by Horvitz–Thompson construction (Prop. 4.1 / §8.1), not by fitting f. Variance follows from the law of total variance (8.1)–(8.3), producing the explicit IPW inflation factor 1+ζ(s); asymptotic normality uses a standard Lindeberg argument with the local kernel bound. For the feasible estimator, the Taylor expansion of 1/ˆπ, NW moment bounds, and the L2 comparison n^{1/2}b^{d/4}(ˆfn,b−efn,b)→0 when p<d are ordinary nonparametric calculations (Props. 4.5–4.6, Thm. 4.8). Citations to Ouimet–Tolosana-Delgado and related Dirichlet-kernel papers supply full-data bias/variance expansions and technical lemmas (L2 norm, max bound); those results do not include the MAR/IPW target and are not used as uniqueness or ansatz smuggling. Bandwidth LSCV and propensity estimation are data-driven tools for practice, not inputs that force the asymptotic statements. No self-definitional loop, fitted-as-prediction, or load-bearing uniqueness import appears.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

The central asymptotic claims rest on standard nonparametric regularity (smooth density, kernel moments, bandwidth rates), the MAR factorization π(x)=P(δ=1|X=x), positivity of π, and dimension/rate conditions that make propensity estimation negligible. No new physical entities are introduced; free parameters are the usual smoothing bandwidths chosen by rate or LSCV.

free parameters (3)
  • Dirichlet bandwidth b
    Smoothing parameter for κ_{s,b}; theory uses b ∼ c n^{-2/(d+4)}; practice selects b by IPW LSCV over a discrete grid B={0.01,...,0.35}.
  • Propensity bandwidth h
    Nadaraya–Watson bandwidth for ˆπ; theory uses h ∼ κ n^{-1/(p+4)}; simulations use Silverman's rule of thumb.
  • LSCV candidate grid and ISE grid (res, ε)
    Implementation choices (res=40 for LSCV integral, res=300 for ISE, ε=0.01) affect reported finite-sample performance though not the asymptotic theorems.
assumptions (6)
  • domain assumption Responses are MAR: P(δ=1|Y,X)=P(δ=1|X)=π(X).
    Section 2; load-bearing for unbiasedness of IPW reconstruction of the full-data density.
  • domain assumption π is bounded away from zero on {g>0} (A5).
    Controls inverse weights; used throughout variance and Lindeberg arguments.
  • standard math Target density f is twice continuously differentiable on S_d (A3) for bias; Lipschitz (A2) for variance.
    Section 3; standard for second-order kernel bias expansions.
  • domain assumption Covariate density g has bounded support and is bounded below on its support (A4); π and g have continuous bounded second derivatives (B2).
    Needed for NW propensity rates and remainder control in Propositions 4.5–4.6.
  • ad hoc to paper For asymptotic normality of ˆfn,b, p < d so n^{-4/(p+4)}=o(n^{-4/(d+4)}).
    Theorem 4.8 and Remark 1; authors note higher-order kernels could relax this but omit that case.
  • standard math Classical kernel K* is bounded, symmetric, with finite second moments (B3).
    Standard NW assumptions for propensity estimation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Dirichlet kernel density estimation on the simplex with missing data." pith.science (2026). https://pith.science/paper/3PBZE7MD

@misc{pith2026260307447,
  author       = {Pith},
  title        = {Pith review of: Dirichlet kernel density estimation on the simplex with missing data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3PBZE7MD}},
  note         = {Machine review of arXiv:2603.07447}
}
read the original abstract

Nonparametric density estimation for compositional data supported on the simplex is examined under a missing at random mechanism. Rather than imputing missing values and estimating the density from a completed data set, we adopt a strategy based on inverse probability weighting. The proposed estimator uses an adaptive Dirichlet kernel, which ensures nonnegativity on the simplex and favorable behavior near the boundary. When the observation probabilities are unknown, they are estimated through a Nadaraya-Watson regression step. The large-sample properties of the estimator are derived, including pointwise bias and variance expansions, optimal smoothing rates, and asymptotic normality. A simulation study investigates its finite-sample performance under varying sample sizes and missing rates. Simulations show our method outperforms inverse-probability-weighted kernel density estimators based on additive and isometric log-ratio transformations of the data for certain target densities. The methodology is further illustrated through an application to leukocyte composition data from the National Health and Nutrition Examination Survey (NHANES), which allows for the identification of the modal immune profile in the sampled population.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 41 canonical work pages

  1. [1]

    Aboubacar and C

    A. Aboubacar and C. C. Kokonendji. Asymptotic results for recursive multivariate associated-kernel estimators of the probability density mass function of a data stream.Comm. Statist. Theory Methods, 54(7):2109–2129, 2025.DOI:10.1080/03610926.2024.2360041

  2. [2]

    K. G. Abraham, A. Maitland, and S. M. Bianchi. Nonresponse in the american time use survey: Who is missing from the data and how much does it matter?Public Opin. Q., 70(5):676–703, 2006. DOI:10.1093/poq/nfl037

  3. [3]

    Aitchison

    J. Aitchison. The statistical analysis of compositional data.J. R. Stat. Soc. Ser. B. Stat. Methodol., 44(2):139–160, 1982.DOI:10.1111/j.2517-6161.1982.tb01195.x

  4. [4]

    Aitchison.The statistical analysis of compositional data

    J. Aitchison.The statistical analysis of compositional data. Monographs on Statistics and Applied Probability. Chapman & Hall, London, 1986. ISBN 0-412-28060-4.DOI:10.1007/978-94-009-410 9-0

  5. [5]

    Aitchison and I

    J. Aitchison and I. J. Lauder. Kernel density estimation for compositional data.J. R. Stat. Soc. Ser. C. Appl. Stat., 34(2):129–137, 1985.DOI:10.2307/2347365

  6. [6]

    Alwadeai, S

    A. Alwadeai, S. Khardani, S. Bouzebda, and W. Jedidi. Beta kernel density estimation with missing data.Preprint, 2026

  7. [7]

    Bertin, C

    K. Bertin, C. Genest, N. Klutchnikoff, and F. Ouimet. Minimax properties of Dirichlet kernel density estimators.J. Multivariate Anal., 195:Paper No. 105158, 16 pp., 2023.DOI:10.1016/j.jmva.202 3.105158

  8. [8]

    Bouzebda, A

    S. Bouzebda, A. Nezzal, and I. Elhattab. Limit theorems for nonparametric conditionalU-statistics smoothed by asymmetric kernels.AIMS Math., 9(9):26195–26282, 2024.DOI:10.3934/math.20241 280. 29

Show all 56 references
  1. [9]

    Buonacera, B

    A. Buonacera, B. Stancanelli, M. Colaci, and L. Malatino. Neutrophil to lymphocyte ratio: An emerging marker of the relationships between the immune system and diseases.Int. J. Mol. Sci., 23 (7):3636, 2022.DOI:10.3390/ijms23073636

  2. [10]

    E. J. M. Carranza. Analysis and mapping of geochemical anomalies using logratio-transformed stream sediment data with censored values.J. Geochem. Explor., 110(2):167–185, 2011.DOI: 10.1016/j.gexplo.2011.05.007

  3. [11]

    J. E. Chac´ on, G. Mateu-Figueras, and J. A. Mart´ ın-Fern´ andez. Gaussian kernels for density estima- tion with compositional data.Comput. Geosci., 37(5):702–711, 2011.DOI:10.1016/j.cageo.2009 .12.011

  4. [12]

    S. F. M. Chastin, J. Palarea-Albaladejo, M. L. Dontje, and D. A. Skelton. Combined effects of time spent in physical activity, sedentary behaviors and sleep on obesity and cardio-metabolic health markers: A novel compositional data analysis approach.PLoS One, 10(10):e0139984, ...

  5. [13]

    Daayeb and F

    H. Daayeb and F. Ouimet. DirichletKDEMissingData, 2026. GitHub repository available online at https://github.com/FredericOuimetMcGill/DirichletKDEMissingData

  6. [14]

    Daayeb, C

    H. Daayeb, C. Genest, S. Khardani, N. Klutchnikoff, and F. Ouimet. A comparison of Dirichlet kernel regression methods on the simplex.arXiv e-prints, 2025.DOI:10.48550/arXiv.2502.08461

  7. [15]

    Daayeb, S

    H. Daayeb, S. Khardani, and F. Ouimet. Dirichlet kernel density estimation for strongly mixing sequences on the simplex.To appear in Math. Methods Statist., 2026.DOI:10.48550/arXiv.2506. 08816

  8. [16]

    S. R. Dubnicka. Kernel density estimation with missing data and auxiliary variables.Aust. N. Z. J. Stat., 51(3):247–270, 2009.DOI:10.1111/j.1467-842X.2009.00541.x

  9. [17]

    J. J. Egozcue, V. Pawlowsky-Glahn, G. Mateu-Figueras, and C. Barcel´ o-Vidal. Isometric logratio transformations for compositional data analysis.Math. Geol., 35(3):279–300, 2003.DOI:10.1023/A: 1023818214614

  10. [18]

    Endres, L

    C. Endres, L. Ale, R. Gentleman, and D. Sarkar. nhanesA: NHANES Data Retrieval, 2025. R package version 1.4. doi:10.32614/CRAN.package.nhanesA

  11. [19]

    Ertefaie, N

    A. Ertefaie, N. S. Hejazi, and M. J. van der Laan. Nonparametric inverse-probability-weighted estimators based on the highly adaptive lasso.Biometrics, 79(2):1029–1041, 2023.DOI:10.1111/bi om.13719

  12. [20]

    Esstafa, C

    Y. Esstafa, C. C. Kokonendji, and T.-B.-T. Ngˆ o. Asymptotic properties of continuous associated- kernel density estimators.Comm. Statist. Theory Methods, 55(5):1568–1588, 2026.DOI:10.1080/ 03610926.2025.2530135

  13. [21]

    T. H. I. Fakhouri, C. B. Martin, T. C. Chen, L. J. Akinbami, C. L. Ogden, R. Paulose-Ram, M. K. Riddles, W. Van de Kerckhove, S. B. Roth, J. Clark, L. K. Mohadjer, and R. E. Fay. An investigation of nonresponse bias and survey location variability in the 2017–2018 national hea...

  14. [22]

    Forget, C

    P. Forget, C. Khalifa, J.-P. Defour, D. Latinne, M.-C. Van Pel, and M. De Kock. What is the normal value of the neutrophil-to-lymphocyte ratio?BMC Res. Notes, 10(1):12, 2017.DOI:10.1186/s131 04-016-2335-5

  15. [23]

    Funke and R

    B. Funke and R. Kawka. Nonparametric density estimation for multivariate bounded data using two non-negative multiplicative bias correction methods.Comput. Statist. Data Anal., 92:148–162, 2015. DOI:10.1016/j.csda.2015.07.006

  16. [24]

    G. Geenens. Explicit formula for asymptotic higher moments of the Nadaraya-Watson estimator. Sankhya A, 76(1):77–100, 2014.DOI:10.1007/s13171-013-0035-y

  17. [25]

    Genest and F

    C. Genest and F. Ouimet. Local linear smoothing for regression surfaces on the simplex using Dirichlet kernels.Statist. Papers, 66(4):97, 2025.DOI:10.1007/s00362-025-01708-8. 30

  18. [26]

    Gharbi, W

    R. Gharbi, W. Jedidi, S. Khardani, and F. Ouimet. A Bernstein polynomial approach for the estimation of cumulative distribution functions in the presence of missing data.arXiv e-prints, 2025. DOI:10.48550/arXiv.2510.07235

  19. [27]

    G. B. Gloor, J. M. Macklaim, V. Pawlowsky-Glahn, and J. J. Egozcue. Microbiome datasets are compositional: And this is not optional.Front. Microbiol., 8:6 pp., 2017.DOI:10.3389/fmicb.20 17.02224

  20. [28]

    R. M. Groves and E. Peytcheva. The impact of nonresponse rates on nonresponse bias: A meta- analysis.Public Opin. Q., 72(2):167–189, 2008.DOI:10.1093/poq/nfn011

  21. [29]

    E. C. Grunsky and P. de Caritat. State-of-the-art analysis of geochemical data for mineral explo- ration.Geochem.: Explor. Environ. Anal., 20(2):217–232, 2020.DOI:10.1144/geochem2019-031

  22. [30]

    Gueorguieva, R

    R. Gueorguieva, R. Rosenheck, and D. Zelterman. Dirichlet component regression and its applications to psychiatric data.Comput. Statist. Data Anal., 52(12):5344–5355, 2008.DOI:10.1016/j.csda.2 008.05.030

  23. [31]

    D. G. Horvitz and D. J. Thompson. A generalization of sampling without replacement from a finite universe.J. Amer. Statist. Assoc., 47(260):663–685, 1952.DOI:10.1080/01621459.1952.10483446

  24. [32]

    K. Hron, M. Templ, and P. Filzmoser. Imputation of missing values for compositional data using classical and robust methods.Comput. Statist. Data Anal., 54(12):3095–3107, 2010.DOI:10.1016/ j.csda.2009.11.023

  25. [33]

    Z. Hu, D. A. Follmann, and J. Qin. Semiparametric dimension reduction estimation for mean response with missing data.Biometrika, 97(2):305–319, 2010.DOI:10.1093/biomet/asq005

  26. [34]

    Structure, function and diversity of the healthy human microbiome.Nature, 486:207–214, 2012.DOI:10.1038/nature11234

    Human Microbiome Project Consortium. Structure, function and diversity of the healthy human microbiome.Nature, 486:207–214, 2012.DOI:10.1038/nature11234

  27. [35]

    Jiang, W

    R. Jiang, W. V. Li, and J. J. Li. mbImpute: an accurate and robust imputation method for microbiome data.Genome Biol., 22:192, 2021.DOI:10.1186/s13059-021-02400-4

  28. [36]

    C. C. Kokonendji and S. M. Som´ e. On multivariate associated kernels to estimate general density functions.J. Korean Statist. Soc., 47(1):112–126, 2018.DOI:10.1016/j.jkss.2017.10.002

  29. [37]

    C. C. Kokonendji and S. M. Som´ e. Bayesian bandwidths in semiparametric modelling for nonnegative orthant data with diagnostics.Stats, 4(1):162–183, 2021.DOI:10.3390/stats4010013

  30. [38]

    Kratz, M

    A. Kratz, M. Ferraro, P. M. Sluss, and K. B. Lewandrowski. Normal reference laboratory values.N. Engl. J. Med., 351(15):1548–1563, 2004.DOI:10.1056/NEJMcpc049016

  31. [39]

    M. L. C. Leite. Applying compositional data methodology to nutritional epidemiology.Stat. Methods Med. Res., 25(6):3057–3065, 2016.DOI:10.1177/0962280214560047

  32. [40]

    R. J. A. Little and D. B. Rubin.Statistical Analysis with Missing Data. John Wiley & Sons, 3rd edition, 2019. ISBN 9780470526798.DOI:10.1002/9781119482260

  33. [41]

    J. A. Mart´ ın-Fern´ andez, C. Barcel´ o-Vidal, and V. Pawlowsky-Glahn. Dealing with zeros and missing values in compositional data sets using nonparametric imputation.Math. Geol., 35(3):253–278, 2003. DOI:10.1023/A:1023866030544

  34. [42]

    F. Ouimet. Asymptotic properties of Bernstein estimators on the simplex.J. Multivariate Anal., 185:Paper No. 104784, 20 pp., 2021.DOI:10.1016/j.jmva.2021.104784

  35. [43]

    F. Ouimet. On the boundary properties of Bernstein estimators on the simplex.Open Stat., 3(1): 48–62, 2022.DOI:10.1515/stat-2022-0111

  36. [44]

    Ouimet and R

    F. Ouimet and R. Tolosana-Delgado. Asymptotic properties of Dirichlet kernel density estimators. J. Multivariate Anal., 187:Paper No. 104832, 25 pp., 2022.DOI:10.1016/j.jmva.2021.104832

  37. [45]

    Pal and C

    S. Pal and C. Heumann. Clustering compositional data using Dirichlet mixture model.PLoS One, 17(5):e0268438, 2022.DOI:10.1371/journal.pone.0268438

  38. [46]

    Palarea-Albaladejo and J

    J. Palarea-Albaladejo and J. A. Mart´ ın-Fern´ andez. Values below detection limit in compositional chemical data.Anal. Chim. Acta, 764:32–43, 2013.DOI:10.1016/j.aca.2012.12.029. 31

  39. [47]

    Pawlowsky-Glahn, J

    V. Pawlowsky-Glahn, J. Egozcue, and R. Tolosana-Delgado.Modeling and Analysis of Compositional Data. John Wiley & Sons, Ltd, Chichester, UK, 2015. ISBN 978-1-118-44306-4.DOI:10.1002/9781 119003144

  40. [48]

    Peleg and E

    O. Peleg and E. Borenstein. Interpolation of microbiome composition in longitudinal data sets. mBio, 15(9):e01150–24, 2024.DOI:10.1128/mbio.01150-24

  41. [49]

    J. M. Robins. Correcting for non-compliance in randomized trials using structural nested mean models.Comm. Statist. Theory Methods, 23(8):2379–2412, 1994.DOI:10.1080/03610929408831393

  42. [50]

    D. B. Rubin. Inference and missing data.Biometrika, 63(3):581–592, 1976.DOI:10.1093/biomet /63.3.581

  43. [51]

    S. R. Seaman and I. R. White. Review of inverse probability weighting for dealing with missing data.Stat. Methods Med. Res., 22(3):278–295, 2013.DOI:10.1177/0962280210395740

  44. [52]

    Tenbusch

    A. Tenbusch. Two-dimensional Bernstein polynomial density estimators.Metrika, 41:233–253, 1994. DOI:10.1007/BF01895321

  45. [53]

    A. A. Tsiatis.Semiparametric Theory and Missing Data. Springer, 2006.DOI:10.1007/0-387-373 45-4

  46. [54]

    Vega-G´ amez and P

    F. Vega-G´ amez and P. J. Alonso-Gonz´ alez. How likely is it to beat the target at different investment horizons: an approach using compositional data in strategic portfolios.Financ Innov, 10:125, 2024. DOI:10.1186/s40854-023-00601-3

  47. [55]

    L. Wang. Dimension reduction for kernel-assisted M-estimators with missing response at random. Ann. Inst. Stat. Math., 71(4):889–910, 2019.DOI:10.1007/s10463-018-0664-y

  48. [56]

    J. Xie, Y. Wang, and E. Garc´ ıa-Portugu´ es. Density estimation for compositional data using non- parametric mixtures.arXiv e-prints, 2025.DOI:10.48550/arXiv.2510.07608. 32

Pith tools

Reviewed July 15, 2026 · model on record in the stance chip above.