Pith. sign in

REVIEW 2 minor 37 references

Dimension reduction of multivariate densities in Bayes spaces

T0 review · 0 major / 2 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read Multivariate densities in Bayes space decompose orthogonally into independent geometric marginals and an interactive component, making FPCA equivalent to separate multivariate analyses on the parts.

desk verdict The paper extends Bayes-space FPCA to multivariate densities by proving equivalence between direct and decomposed FPCA plus PCA-optimality of the independent-interactive variance split. read the letter →

arxiv 2606.19011 v1 pith:KMLSM236 submitted 2026-06-17 stat.ME

classification stat.ME
keywords Bayesspacemultivariatedensitiescentredlogratiotransformationfunctionalprincipalcomponentanalysisdimensionreductionorthogonaldecompositionvariancegeometricmarginals
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper shows that the Bayes space structure, via the centred logratio transformation, lets multivariate probability densities be split into mutually orthogonal parts: geometric marginals that capture independent variation and a remaining interactive component. This split decomposes the total variance in a way that is optimal for principal component analysis, so the eigenfunctions and scores from FPCA on the full densities break down cleanly into contributions from each part. The equivalence means one can run FPCA directly on the densities or on the decomposed pieces and obtain matching results. A reader would care because the approach turns the constrained, relative nature of density data into a geometrically natural setting where dimension reduction reveals separate sources of variation rather than mixing them.

What carries the argument

The orthogonal decomposition of multivariate densities into independent geometric marginals and interactive component, enabled by the centred logratio (clr) transformation that gives an isometric isomorphism to an L² subspace.

What would settle it

If the eigenfunctions and scores from direct FPCA on multivariate densities fail to match the decomposed versions up to the claimed additive structure, or if the variance explained by the parts is not maximal among all orthogonal splits, the optimality and equivalence would not hold.

Watch

Extended reading notes

Core claim

Embedding multivariate PDFs in the Bayes space enables an orthogonal decomposition into independent and interactive components, with the independent part further split into mutually orthogonal geometric marginals. The centred logratio transformation maps this structure isometrically to a subspace of L², so functional principal component analysis applies directly. The resulting variance decomposition is optimal in the PCA sense, and applying FPCA to the original densities is equivalent to multivariate FPCA on the decomposed form, with eigenfunctions and scores decomposing accordingly.

Load-bearing premise

The centred logratio transformation establishes an isometric isomorphism between the Bayes space and a subspace of L² space.

Editorial extensions

If this is right

  • The decomposition of total variance is optimal in a PCA sense, so eigenfunctions and scores from FPCA have a direct interpretation in terms of independent and interactive contributions.
  • FPCA applied directly to multivariate densities produces results equivalent to multivariate FPCA performed on the decomposed independent and interactive parts.
  • Eigenfunctions and scores obtained from the full densities decompose additively according to the independent and interactive split.
  • The decomposition applied to empirical housing and geological data yields interpretable components that separate sources of variation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same orthogonal split could be used with other functional data techniques such as functional regression or clustering on density data.
  • Fields that routinely work with joint distributions, such as compositional data or spatial statistics, might adopt the geometric marginals as a standard way to separate marginal and dependence effects.
  • Simulated examples with known independent and dependence structures could be used to check whether the PCA optimality holds numerically beyond the theoretical proof.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The manuscript develops dimension reduction techniques for multivariate probability density functions within the Bayes space framework. It utilizes the centred logratio (clr) transformation to establish an isometric isomorphism with a subspace of L², allowing the application of functional principal component analysis (FPCA). The key contributions include an orthogonal decomposition of multivariate densities into independent and interactive components, with the independent part further decomposed into orthogonal geometric marginals. The paper proves that this variance decomposition is optimal in a PCA sense and demonstrates the equivalence of applying FPCA directly to the densities versus to their decomposed form, with corresponding decomposition of eigenfunctions and scores. The theoretical results are illustrated with applications to housing and geological data.

Significance. If the results hold, this provides a significant advancement in the analysis of multivariate density data by offering a structured way to decompose and interpret variance sources. The reliance on the standard clr isometry ensures the framework is built on solid Hilbert space foundations, and the optimality and equivalence results could influence how FPCA is applied and interpreted in compositional data analysis. The empirical applications demonstrate practical utility. The use of an established isometric isomorphism and the focus on reproducible theoretical structure are strengths.

minor comments (2)
  1. Abstract: the phrase 'equivalent in a certain sense' is imprecise; a brief clarification of the precise sense of equivalence (e.g., with respect to the inner product or the resulting scores) would improve readability without altering the claim.
  2. The manuscript would benefit from an explicit statement early in the introduction of how the geometric marginals are defined and why they are mutually orthogonal under the Bayes-space inner product.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the positive summary, significance assessment, and recommendation of minor revision. No major comments were listed in the report.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; derivation self-contained via standard clr isometry

full rationale

The paper's central results—the PCA-optimality of the independent/interactive variance decomposition and the equivalence between direct and decomposed FPCA—follow from the isometric isomorphism property of the clr transformation, which is invoked as an established fact from the Bayes-space literature rather than derived or fitted within the manuscript. No equation reduces a claimed prediction to a self-defined quantity, no load-bearing uniqueness theorem is imported from the authors' own prior work, and the decomposition is presented as a direct consequence of the Hilbert-space inner product supplied by clr. The framework therefore contains no self-referential steps that collapse the claimed results to their inputs by construction.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central claim rests on the isometric properties of the clr transformation and the existence of an orthogonal decomposition into independent and interactive components; these are treated as given by the Bayes-space framework rather than derived anew in the paper.

assumptions (1)
  • domain assumption The centred logratio (clr) transformation establishes an isometric isomorphism between the Bayes space and a subspace of L² space.
    This property is invoked to justify applying FPCA and the subsequent orthogonal decomposition to multivariate densities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Dimension reduction of multivariate densities in Bayes spaces." pith.science (2026). https://pith.science/paper/KMLSM236

@misc{pith2026260619011,
  author       = {Pith},
  title        = {Pith review of: Dimension reduction of multivariate densities in Bayes spaces},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KMLSM236}},
  note         = {Machine review of arXiv:2606.19011}
}
abstract

The Bayes space provides a Hilbert space structure for analysing probability density functions (PDFs), equipping them with a geometry that reflects their relative and constrained nature. A key tool in this framework is the centred logratio (clr) transformation, which establishes an isometric isomorphism between the Bayes space and (a subspace of) the classical $L^2$ space. This makes it possible to apply functional data analysis (FDA) techniques, particularly functional principal component analysis (FPCA), to both univariate and multivariate density data in the context of dimension reduction. For multivariate PDFs, embedding them in the Bayes space enables an orthogonal decomposition into independent and interactive components. Furthermore, the independent part can be decomposed into mutually orthogonal geometric marginals. This structure provides more profound insights into the sources of variation in multivariate densities. We show that this decomposition of the total variance is optimal in a PCA sense, impacting the interpretation of the eigenfunctions and scores resulting from FPCA. We demonstrate that applying FPCA directly to multivariate densities is equivalent in a certain sense to applying multivariate FPCA to their decomposed form, with the resulting eigenfunctions and scores decomposing accordingly. The unique decomposition based on these theoretical results is applied to housing and geological empirical data respectively, demonstrating the interpretability and practical value of this approach.

Figures

Figures reproduced from arXiv: 2606.19011 by the authors.

Figure 1
Figure 1. New York state: The first row shows the raw data, the density estimate and the clr-transformed density. The second row shows the orthogonal decomposition of this density on clr level (the two geometric marginals and the interactive part). bottom row show the orthogonal decomposition – the geometric marginals for price and size, respectively, and the interactive part (all in L 2 0 ). The clr-transformed densities (cl… view at source ↗
Figure 2
Figure 2. Housing scores and their decomposition with colours according to the four main US regions and Puerto Rico, which does not belong to any of these regions. and therefore have the smallest influence on the structure of the total scores. While the loadings corresponding to the geometric marginals are rather easily interpretable, the interaction loadings indicate a more complex dependence structure, i.e., stronger intera… view at source ↗
Figure 3
Figure 3. Eigenfunctions of US housing data for the first two functional principal components and their orthogonal decomposition on clr level. The percentages show how much variance is explained by each eigenfunction or its part. The percentages in brackets show how much variance of the functional principal component are explained by the particular parts of the eigenfunction. squared norm of the price marginal. Moreover, it i… view at source ↗
Figures from the paper (23 more)
Figure 4
Figure 4. Figure 4: Scree plot for the housing data. Together, the first two principal components explain 39.11 % of variance in the data. 14 [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Squared norms of the centered bivariate densities and their decomposition for all the states. 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 interaction size (y) price (x) Alabama Alaska Arizona Arkansas California Colorado Connecticut Delaware District of Columbia Florida Georgia Ha…
Figure 6
Figure 6. Figure 6: Relative norms of the orthogonal parts of the centered bivariate housing densities. also the most informative ones. The variance decomposition suggests that the reason is that they contain the majority of the variance in the density data. Note that while in Section 4 w…
Figure 7
Figure 7. Figure 7: Decomposition of FPCA scores depending on the imputation strategy. 16 [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Loadings: 1 × mean 8 10 14 18 6 7 8 9 10 11 29.28 % (100 %) log(price) log(size) 8 10 14 18 6 7 8 9 10 11 9.69 % (100 %) log(price) log(size) 8 10 14 18 6 7 8 9 10 11 9.58 % (100 %) log(price) log(size) 8 10 14 18 6 7 8 9 10 11 17.39 % (59.38 %) log(price) log(size) 8 …
Figure 9
Figure 9. Figure 9: Loadings: 0.9 × mean 17 [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]
Figure 10
Figure 10. Figure 10: Loadings: 0.8 × mean 8 10 14 18 6 7 8 9 10 11 29.15 % (100 %) log(price) log(size) 8 10 14 18 6 7 8 9 10 11 9.59 % (100 %) log(price) log(size) 8 10 14 18 6 7 8 9 10 11 9.49 % (100 %) log(price) log(size) 8 10 14 18 6 7 8 9 10 11 17.39 % (59.68 %) log(price) log(size)…
Figure 11
Figure 11. Figure 11: Loadings: 0.7 × mean 18 [PITH_FULL_IMAGE:figures/full_fig_p018_11.png]
Figure 12
Figure 12. Figure 12: Loadings: 0.6 × mean 8 10 14 18 6 7 8 9 10 11 28.97 % (100 %) log(price) log(size) 8 10 14 18 6 7 8 9 10 11 9.6 % (100 %) log(price) log(size) 8 10 14 18 6 7 8 9 10 11 9.26 % (100 %) log(price) log(size) 8 10 14 18 6 7 8 9 10 11 17.39 % (60.04 %) log(price) log(size) …
Figure 13
Figure 13. Figure 13: Loadings: 0.5 × mean 19 [PITH_FULL_IMAGE:figures/full_fig_p019_13.png]
Figure 14
Figure 14. Figure 14: Percentage of explained variance by the first three components depending on the imputation constant (multiple of the default value). 1.0 0.9 0.8 0.7 0.6 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 variance 1.0 0.9 0.8 0.7 0.6 0.5 0.1 0.2 0.3 0.4 0.5 0.6 relative variance x−margin…
Figure 15
Figure 15. Figure 15: Decomposition of the total variance of the housing density data depending on the imputation constant (multiple of the default value). Top panel: the variance of each part in the orthogonal decomposition; bottom panel: the proportion of variance (relative variance) for…
Figure 16
Figure 16. Figure 16: Example data for the Louny district: The first row shows the raw data, the density estimate, and the clr-transformed density. The second row shows the orthogonal decomposition of this density on clr level (the two geometric marginals and the interactive part). percent…
Figure 17
Figure 17. Figure 17: Scores of geological data and their decomposition. The colours corresponds to the selected districts where concentrations of Cu and Zn are unrelated (red), related (green) and where wine or hops are grown (blue). The districts that were not assigned to any of the thre…
Figure 18
Figure 18. Figure 18: Geological data eigenfunctions and their orthogonal decomposition corresponding to the first two functional principal components on clr level. The percentages show how much variance is explained by each eigenfunction or its part. The percentages in brackets show how m…
Figure 19
Figure 19. Figure 19: Scree plot for the geological data. The first two principal components together explain 41.9 % of variance in the data [PITH_FULL_IMAGE:figures/full_fig_p023_19.png]
Figure 20
Figure 20. Figure 20: Squared norms of the centered bivariate densities and their decomposition for all the Czech districts. 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 interaction Zn (y) Cu (x) Benesov Beroun Blansko Breclav Brno−mesto Brno−venkov Bruntal Ceska Lipa Ceske Budejovice Cesky Krumlov Dec…
Figure 21
Figure 21. Figure 21: Relative norms of the orthogonal density parts of the geological data. principal components with close percentages of explained variance like in the case of the housing data, such that small changes in the data cannot lead to any changes in the order of components [P…
Figure 22
Figure 22. Figure 22: Decomposition of FPCA scores depending on the imputation strategy. 1 2 3 4 2 3 4 5 23.97 % (100 %) log(Cu) log(Zn) 1 2 3 4 2 3 4 5 17.93 % (100 %) log(Cu) log(Zn) 1 2 3 4 2 3 4 5 17 % (70.92 %) log(Cu) log(Zn) 1 2 3 4 2 3 4 5 3.04 % (16.96 %) log(Cu) log(Zn) 1 2 3 4 2…
Figure 23
Figure 23. Figure 23: Loadings: 1 × (imputation value) 25 [PITH_FULL_IMAGE:figures/full_fig_p025_23.png]
Figure 24
Figure 24. Figure 24: Loadings: 0.7 × (imputation value) 1 2 3 4 2 3 4 5 23.86 % (100 %) log(Cu) log(Zn) 1 2 3 4 2 3 4 5 17.77 % (100 %) log(Cu) log(Zn) 1 2 3 4 2 3 4 5 16.95 % (71.06 %) log(Cu) log(Zn) 1 2 3 4 2 3 4 5 2.97 % (16.74 %) log(Cu) log(Zn) 1 2 3 4 2 3 4 5 3.17 % (13.27 %) log(C…
Figure 25
Figure 25. Figure 25: Loadings: 0.5 × (imputation value) 26 [PITH_FULL_IMAGE:figures/full_fig_p026_25.png]
Figure 26
Figure 26. Figure 26: Percentage of explained variance by the first two components depending on the imputation constant (multiple of the default value). 1.0 0.9 0.8 0.7 0.6 0.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 variance 1.0 0.9 0.8 0.7 0.6 0.5 0.25 0.30 0.35 0.40 0.45 relative variance x−margina…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

37 extracted references

  1. [1]

    K. G. van den Boogaart, J. J. Egozcue, V . Pawlowsky-Glahn, Bayes Hilbert spaces, Australian & New Zealand Journal of Statistics 56 (2014) 171–194

  2. [2]

    K. G. van den Boogaart, R. Tolosana-Delgado, Multivariate Bayes Spaces and Compositions, in: C. Thomas- Agnan, V . Pawlowsky-Glahn (Eds.), Proceedings of the 9th International Workshop on Compositional Data Analysis (CoDaWork 2022), CoDa Association, Toulouse, 2022

  3. [3]

    Delicado, Dimensionality reduction when data are density functions, Computational Statistics & Data Analysis 55 (2011) 401–420

    P. Delicado, Dimensionality reduction when data are density functions, Computational Statistics & Data Analysis 55 (2011) 401–420

  4. [4]

    Eckardt, J

    M. Eckardt, J. Mateu, S. Greven, Generalized functional additive mixed models with (functional) compositional covariates for areal Covid-19 incidence curves, Journal of the Royal Statistical Society Series C: Applied Statis- tics 73 (2024) 880–901

  5. [5]

    J. J. Egozcue, J. L. Díaz-Barrero, V . Pawlowsky-Glahn, Hilbert space of probability density functions based on Aitchison geometry, Acta Mathematica Sinica 22 (2006) 1175–1182

  6. [6]

    Filzmoser, K

    P. Filzmoser, K. Hron, A. Menafoglio, Logratio approach to distributional modeling, in: A. Daouia, A. Ruiz- Gazen (Eds.), Advances in Contemporary Statistics and Econometrics: Festschrift in Honor of Christine Thomas–Agnan, Springer International Publishing, Cham, 2021, pp. 451–470

  7. [7]

    Genest, K

    C. Genest, K. Hron, J. G. Nešlehová, Orthogonal decomposition of multivariate densities in Bayes spaces and relation with their copula–based representation, Journal of Multivariate Analysis 198 (2023) 105228

  8. [8]

    C. Happ, S. Greven, Multivariate functional principal component analysis for data observed on different (dimen- sional) domains, Journal of the American Statistical Association 113 (2018) 649–659

Show all 37 references
  1. [9]

    K. Hron, J. Machalová, A. Menafoglio, Bivariate densities in Bayes spaces: orthogonal decomposition and spline representation, Statistical Papers 64 (2023) 1629–1667

  2. [10]

    K. Hron, A. Menafoglio, M. Templ, K. Hr˚ uzová, P. Filzmoser, Simplical principal component analysis for density functions in Bayes spaces, Computational Statistics and Data Analysis 94 (2016) 330–350. 28

  3. [11]

    Johnson, D

    R. Johnson, D. Wichern, Applied Multivariate Statistical Analysis, Prentice Hall, Upper Saddle River, 6th edi- tion, 2007

  4. [12]

    Kneip, K

    A. Kneip, K. J. Utikal, Inference for density families using functional principal component analysis, Journal of the American Statistical Association 96 (2001) 519–542

  5. [13]

    Kokoszka, M

    P. Kokoszka, M. Reimherr, Introduction to Functional Data Analysis, CRC Press, Boca Raton, 2017

  6. [14]

    Kutta, A

    T. Kutta, A. Jach, M. Haddad, P. Kokoszka, H. Wang, Detection and localization of changes in a panel of densities, Journal of Multivariate Analysis 205 (2025) 105374

  7. [15]

    X. Lei, Z. Chen, H. Li, Functional outlier detection for density-valued data with application to robustify distribution-to-distribution regression, Technometrics 65 (2023) 351–362

  8. [16]

    Y . Ma, X. Zhou, W. Wu, A stochastic process representation for time warping functions, Computational Statistics and Data Analysis 194 (2024) 107941

  9. [17]

    Maier, A

    E.-M. Maier, A. Fottner, S. Greven, A. Stöcker, Additive density regression, 2025

  10. [18]

    Maier, A

    E.-M. Maier, A. Stöcker, B. Fitzenberger, S. Greven, Additive density-on-scalar regression in Bayes Hilbert spaces with an application to gender economics, Annals of Applied Statistics 19 (2025) 680–700

  11. [19]

    Matys Grygar, U

    T. Matys Grygar, U. Radoji ˇci´c, I. Pavl˚ u, S. Greven, J. Nešlehová, Š. T˚ umová, K. Hron, Exploratory functional data analysis of multivariate densities for the identification of agricultural soil contamination by risk elements, Journal of Geochemical Exploration 259 (2024) 107416

  12. [20]

    Menafoglio, M

    A. Menafoglio, M. Grasso, P. Secchi, B. Colosimo, Monitoring of probability density functions via simplicial functional pca with application to image data, Technometrics 60 (2018) 497–510

  13. [21]

    Menafoglio, M

    A. Menafoglio, M. Grasso, P. Secchi, B. M. Colosimo, A class-kriging predictor for functional compositions with application to particle-size curves in heterogeneous aquifers, Mathematical Geosciences 48 (2016) 463– 485

  14. [22]

    Menafoglio, A

    A. Menafoglio, A. Guadagnini, P. Secchi, A kriging approach based on Aitchison geometry for the characteriza- tion of particle-size curves in heterogeneous aquifers, Stochastic Environmental Research and Risk Assessment 28 (2014) 1835–1851

  15. [23]

    Murph, J

    A. Murph, J. Strait, K. Moran, J. Hyman, P. Stauffer, Visualisation and outlier detection for probability density function ensembles, Stat 13 (2024) e662

  16. [24]

    Pavl˚ u, J

    I. Pavl˚ u, J. Machalová, R. Tolosana-Delgado, K. Hron, K. Bachmann, K. G. van den Boogaart, Principal com- ponent analysis for distributions observed by samples in bayes spaces, Mathematical Geosciences 56 (2024) 1641–1669

  17. [25]

    Pawlowsky-Glahn, J

    V . Pawlowsky-Glahn, J. J. Egozcue, R. Tolosana-Delgado, Modeling and Analysis of Compositional Data, Wi- ley, Chichester, 2015

  18. [26]

    Petersen, C

    A. Petersen, C. Zhang, P. Kokoszka, Modeling probability density functions as data objects, Econometrics and Statistics 21 (2022) 159–178

  19. [27]

    Podlešáková, J

    E. Podlešáková, J. N ˇemeˇcek, G. Halová, Proposal of soil contamination limits for persistent organic xenobiotic substances in the Czech Republic, Rostlinná výroba 42 (1996) 49–54

  20. [28]

    Poláková, K

    Š. Poláková, K. Hutarová, D. Reininger, L. Kubík, Registr kontaminovaných ploch 2 M HNO3 (1990–2009), Technical report, Ústˇrední kontrolní a zkušební ústav zemˇedˇelský v Brnˇe, Brno, Czech Republic, 2011

  21. [29]

    J. Qiu, X. Dai, Z. Zhu, Nonparametric estimation of repeated densities with heterogeneous sample sizes, Journal of the American Statistical Association 119 (2024) 176–188. 29

  22. [30]

    Ramsay, B

    J. Ramsay, B. W. Silverman, Functional Data Analysis, Springer, New York, 2 edition, 2005. [31]RCore Team,R: A Language and Environment for Statistical Computing,RFoundation for Statistical Comput- ing, Vienna, Austria, 2025

  23. [31]

    A. S. Sakib, USA real estate dataset, 2022. [accessed 2025-10-30]

  24. [32]

    Škor ˇna, J

    S. Škor ˇna, J. Machalová, J. Burkotová, K. Hron, S. Greven, Approximation of bivariate densities with composi- tional splines, 2024

  25. [33]

    Steyer, S

    L. Steyer, S. Greven, Principal component analysis in Bayes spaces for sparsely sampled density functions, 2023

  26. [34]

    Talská, A

    R. Talská, A. Menafoglio, K. Hron, J. J. Egozcue, J. Palarea-Albaladejo, Weighting the domain of probability densities in functional data analysis, Stat 9 (2020) e283

  27. [35]

    Talská, A

    R. Talská, A. Menafoglio, J. Machalová, K. Hron, E. Fišerová, Compositional regression with functional re- sponse, Computational Statistics & Data Analysis 123 (2018) 66–85

  28. [36]

    H. Wang, L. Shangguan, R. Guan, L. Billard, Principal component analysis for compositional data vectors, Computational Statistics 30 (2015) 1079–1096

  29. [37]

    Zbíral, I

    J. Zbíral, I. Honsa, S. Malý, D. ˇCižmár, Soil Analysis III, Central Institute for Supervising and Testing in Agriculture, Brno, Czech Republic, 2004. 30

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.