REVIEW 3 major objections 5 minor 9 references
On the Spherical Dirichlet Distribution: Corrections and Results
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that the Spherical-Dirichlet Distribution, built as the square-root of a Dirichlet vector, correctly models unit vectors on the positive orthant of the hypersphere, with closed-form moments and mode and workable MLE and…
desk verdict Solid corrected math for a known object, but the applied claims outrun the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the square-root map $x_i=\sqrt{z_i}$ from the $(p-1)$-dimensional Dirichlet simplex to the positive orthant of the unit sphere, together with the surface-measure element $d\omega_{p-1}(x)=dz/(2^{p-1}\sqrt{z_1\cdots z_p})$ supplied by the paper's reference [3]. This factor, not the full Jacobian of the $p$-dimensional transformation, determines the density's $2^{p-1}\prod_i x_i$ term and all subsequent moment and mode formulas. The inference machinery is the exponential-family form of the log-likelihood with sufficient statistics $\sum_i\log x_{ij}$, leading to moment and likelihood equations solved numerically by L-BFGS-B.
What would settle it
Integrate the claimed density over the positive orthant of $\mathbb{S}^{p-1}$ for a concrete case, e.g. $p=3$, $\alpha=(2,2,2)$, using the standard spherical surface element; if the result is not exactly 1 and the second moment does not reproduce $\alpha_i/\alpha_0$ for a simulated sample of vectors drawn uniformly from the orthant with $\alpha_i=1/2$, the measure factor in the density is mistaken.
Extended reading notes
Core claim
The central claim is that the distribution with density $f_{\mathrm{SDir}}(x;\alpha) = \frac{2^{p-1}\Gamma(\alpha_0)}{\prod_{i=1}^p \Gamma(\alpha_i)}\prod_{i=1}^p x_i^{2\alpha_i-1}$ on the positive orthant of the unit sphere $\mathbb{S}^{p-1}_+$, obtained from $x_i=\sqrt{z_i}$ where $z$ is Dirichlet, is correctly normalized. The paper further claims that its first moment is $E(x_i)=\Gamma(\alpha_i+\tfrac12)\Gamma(\alpha_0)/(\Gamma(\alpha_i)\Gamma(\alpha_0+\tfrac12))$, its second moment is $E(x_i^2)=\alpha_i/\alpha_0$, and its mode is $\sqrt{(2\alpha_i-1)/(2\alpha_0-p)}$, and that these formulas make the distribution a practical parametric family for vectors on the positive orthant. The corrected surface-measure factor $2^{p-1}\prod x_i$ is what makes the density integrate to one.
Load-bearing premise
The load-bearing assumption is that the surface-measure formula $d\omega_{p-1}(x)=dz/(2^{p-1}\sqrt{z_1\cdots z_p})$, stated without proof, is correct; if that factor is wrong, every derived density, moment, mode, and estimator shifts.
Editorial extensions
If this is right
- Text-mining and gene-expression data, once normalized to unit length, can be modeled with a single parametric distribution that assigns zero mass outside the positive orthant.
- The MLE system solves to accurate estimates, with the reported simulations showing MLE converging in tens of iterations and lower error than MOM.
- Setting all $\alpha_i=1/2$ recovers the uniform distribution on the positive orthant, giving a baseline model with no fitted parameters.
- As the common $\alpha$ grows, the SDD concentrates to a point mass at the mean direction rather than approaching a von Mises or normal distribution on the sphere.
- Because $E(x_i^2)=\alpha_i/\alpha_0$, the second empirical moments directly give a quick method-of-moments estimate of the parameter proportions.
Reading between the lines
- A direct testable extension: for any sample of non-negative vectors, the SDD's fit could be checked by verifying that the sample mean of squared coordinates approximates $\hat\alpha_i/\hat\alpha_0$, a diagnostic the paper does not explicitly propose.
- The same square-root transformation could be applied to other simplex-based models (e.g., logistic-normal or generalized Dirichlet) to generate new sphere-supported distributions with the same surface-measure correction.
- The paper's log-shift transformation $\ln(1.10+x)$ for zero counts is an ad hoc step; a dedicated zero-inflated SDD or a Bayesian treatment of zeros would make the model directly applicable to sparse term-frequency matrices.
- Because the surface measure formula is imported without derivation, a careful check in the literature on the induced measure of the square-root map would settle whether the density's factor is $2^{p-1}$ or something else; the editorial instinct is to verify this before using the moments.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript corrects and extends the author's earlier treatment of the Spherical-Dirichlet Distribution (SDD), defined by taking a Dirichlet random vector z on the simplex and setting x_i = sqrt(z_i), so that x lies in the positive orthant of the unit sphere. The paper derives the SDD density with respect to surface measure, computes first and second moments, variance, covariance, mode, and some limiting cases, and compares the SDD with the von Mises and Fisher-Bingham distributions. It then proposes method-of-moments (MOM) and maximum-likelihood (MLE) estimators, with an L-BFGS-B implementation, and reports a simulation study and a text-mining application using transformed term-frequency vectors. The abstract and introduction claim that the SDD is a useful new model for positive-orthant directional data in text mining and gene-expression analysis.
Significance. The theoretical core of the paper is largely sound and useful. The moment formulas in Eqs. (16), (20), and (21) are transparent, are derived by kernel recognition, and agree with the corresponding Dirichlet moments after the transformation z_i = x_i^2, which is a genuine check on the normalizing constant. The simulation study is a useful sanity check for the proposed MOM and MLE procedures, and the correction of an earlier published error is of value to users of this distribution. However, the paper's applied claim that the SDD is a useful model for text data is not supported by the real-data section, which contains no goodness-of-fit assessment or baseline comparison. The manuscript therefore needs revision before the applied claims can be accepted.
major comments (3)
- [§Basic Properties, display preceding Eq. (3)] The intermediate display f_SDir(x;alpha) = f_Dir(x^2) 2^{p-1} / sqrt(x_1^2 ... x_p^2) has the surface-measure factor inverted. Combining Eq. (2) as d omega = dz / (2^{p-1} sqrt(z_1...z_p)) with z_i = x_i^2 gives dz = 2^{p-1} sqrt(z_1...z_p) d omega = 2^{p-1}(x_1...x_p) d omega, so the SDD density must be f_SDir = f_Dir(x^2) 2^{p-1}(x_1...x_p), not f_Dir(x^2) 2^{p-1}/(x_1...x_p). The printed denominator would produce exponents 2 alpha_i - 3 and would contradict the correctly normalized density in Eq. (3). Please correct this display and state explicitly that dz in Eq. (2) is the (p-1)-dimensional Lebesgue measure on the simplex.
- [§Text Mining Example, Table 3] The text-mining section reports that MOM and MLE produce similar alpha estimates, but this does not provide evidence that the SDD actually fits the transformed normalized term-frequency vectors. Because the SDD is equivalent to a Dirichlet distribution on z = x^2, a natural and inexpensive check would be to test whether the squared, normalized data follow Dirichlet(alpha + 1/2), using marginal beta QQ plots, second-moment comparisons, or a likelihood-ratio comparison against a more flexible model. Without such a diagnostic, the close MOM/MLE agreement only shows internal consistency of the two estimation algorithms, not the distributional adequacy on which the paper's applied claim rests. The ad hoc transformation x_transf = ln(1.10 + x) also needs justification and some sensitivity analysis, since the conclusion may depend on the choice 1.10.
- [§Mode, Eq. (28)] The interior-mode formula (28) is derived under the condition alpha_i > 1/2 for all i, and when any alpha_i <= 1/2 the mode should be analyzed on the boundary of the positive orthant. The text-mining estimates in Table 3 include several alpha_i below 1/2, so the fitted model in the application lies outside the regime where Eq. (28) applies. The paper should either restrict the mode discussion to the interior case and note its inapplicability to the fitted model, or develop the boundary-mode analysis. This is not a contradiction, but it is a gap between the theoretical section and the application.
minor comments (5)
- [§Probability Density Function and Normalizing Constant, Eq. (2)] The measure dz in Eq. (2) is not defined; please state that it is the flat Lebesgue measure on the (p-1)-dimensional simplex and add a reference or a short proof for the surface-measure formula, since it is load-bearing for the density derivation.
- [§Simulation Example, Table 1] The phrase "percentage error based on vector norm ratios" is not defined; please specify the formula used to compute the reported percentage errors.
- [§Limiting Behavior, Eq. (35)] Equation (35) appears to contain typesetting or algebraic errors: the term (1 - mu_alpha^2 / alpha) is dimensionally inconsistent with the variance expression (17), and the following display does not clearly follow. Please rewrite the symmetric covariance matrix carefully.
- [§Abstract and Introduction] The abstract calls the SDD a "novel probability distribution," but the paper is a correction and extension of Guardiola (2020); please phrase the novelty claim as a corrected and extended treatment rather than a fully new distribution.
- [§Mode and Relationship with the Mean] The section title promises a relationship between mode and mean, but the only result is an asymptotic equality as alpha tends to infinity; consider renaming the section or adding a finite-sample comparison for the symmetric case.
Circularity Check
No significant circularity: the SDD derivations follow from the Dirichlet density and standard calculus, with no fitted parameter or load-bearing self-citation.
full rationale
The derivation chain is self-contained given the Dirichlet density and the stated surface-measure change of variables. The SDD density (3) is obtained by the transformation z_i = x_i^2, with the surface element dω_{p-1}(x) = dz/(2^{p-1}√(z_1...z_p)) cited from Gupta; this is an external reference, not a self-citation. Moments (7), (15), (19), and (20) are computed by recognizing the integrand as the kernel of the same SDD with adjusted parameters and using the already-established normalizing constant; this is a standard calculus identity, not a circular definition. The mode (28) follows from maximizing the log-density with a Lagrange multiplier. The MOM and MLE estimators are derived from these moment identities and the likelihood; no fitted value is fed back into the derivation. The only self-referential text is the abstract's statement that the note corrects 'Guardiola (2020)', which is the prior paper being corrected and is not used to justify any mathematical claim. The text-mining application estimates parameters and observes that MOM and MLE agree, but it does not provide a goodness-of-fit test; this is a substantive support/validity limitation rather than a circularity, because the fitted estimates do not enter the derivations of the density or estimators.
Assumptions & free parameters
assumptions (4)
- standard math The starting point is the standard Dirichlet density on the (p-1)-simplex with parameters alpha_i > 0.
- standard math The surface measure of the positive orthant sphere satisfies d omega_{p-1}(x) = dz / (2^(p-1) sqrt(z_1 ... z_p)) for z_i = x_i^2.
- standard math Frame's asymptotic Gamma(x+a) / Gamma(x) ~ x^a as x tends to infinity.
- domain assumption The text-mining data follow an SDD after the ln(1.10 + x) transformation and L2 normalization.
invented entities (1)
-
Spherical-Dirichlet Distribution (SDD)
Cite this review
Pith. "Pith review of On the Spherical Dirichlet Distribution: Corrections and Results." pith.science (2026). https://pith.science/paper/AQLQBKB2
@misc{pith2026250604441,
author = {Pith},
title = {Pith review of: On the Spherical Dirichlet Distribution: Corrections and Results},
year = {2026},
howpublished = {\url{https://pith.science/paper/AQLQBKB2}},
note = {Machine review of arXiv:2506.04441}
}
read the original abstract
This note corrects a technical error in Guardiola (2020, Journal of Statistical Distributions and Applications), presents updated derivations, and offers an extended discussion of the properties of the spherical Dirichlet distribution. Today, data mining and gene expressions are at the forefront of modern data analysis. Here we introduce a novel probability distribution that is applicable in these fields. This paper develops the proposed Spherical-Dirichlet Distribution designed to fit vectors located at the positive orthant of the hypersphere, as it is often the case for data in these fields, avoiding unnecessary probability mass. Basic properties of the proposed distribution, including normalizing constants and moments are developed. Relationships with other distributions are also explored. Estimators based on classical inferential statistics, such as method of moments and maximum likelihood estimators are obtained. Two applications are developed: the first one uses simulated data, and the second uses a real text mining example. Both examples are fitted using the proposed Spherical-Dirichlet Distribution and their results are discussed.
Reference graph
Works this paper leans on
-
[1]
arXiv e-prints, 1605–00316 (2016)
Suvrit, S.: Directional statistics in machine learning: a brief review. arXiv e-prints, 1605–00316 (2016). 1605.00316
arXiv 2016
-
[2]
The Annals of Mathematical Statistics33(4), 1272–1280 (1962)
Olkin, I., Rubin, H.: A characterization of the wishart distribution. The Annals of Mathematical Statistics33(4), 1272–1280 (1962). doi:10.1214/aoms/1177704739
arXiv 1962
-
[3]
Chapman and Hall/CRC, Boca Raton (2000) Guardiola Page 15 of 15
Gupta, A.K., Nagar, D.K.: Matrix Variate Distributions. Chapman and Hall/CRC, Boca Raton (2000) Guardiola Page 15 of 15
work page 2000
-
[4]
The American Mathematical Monthly56(8), 529–535 (1949)
Frame, J.S.: An approximation to the quotient of gamma function. The American Mathematical Monthly56(8), 529–535 (1949)
work page 1949
-
[5]
Wiley series in probability and statistics, p
Mardia, K.V., Jupp, P.E.: Directional Statistics, 2nd edn. Wiley series in probability and statistics, p. 179. Wiley, Chichester, England (2000)
work page 2000
-
[6]
Journal of the Royal Statistical Society
Mardia, K.V.: Statistics of directional data. Journal of the Royal Statistical Society. Series B (Methodological) 37(3), 349–393 (1975)
work page 1975
-
[7]
Journal of the Royal Statistical Society: Series B (Methodological)44(1), 71–80 (1982)
Kent, J.T.: The Fisher-Bingham distribution on the sphere. Journal of the Royal Statistical Society: Series B (Methodological)44(1), 71–80 (1982). doi:10.1111/j.2517-6161.1982.tb01189.x. https://rss.onlinelibrary.wiley.com/doi/pdf/10.1111/j.2517-6161.1982.tb01189.x
arXiv 1982
-
[8]
Computers and Mathematics with Applications24(10), 11–17 (1992)
Narayanan, A.: A note on parameter estimation in the multivariate beta distribution. Computers and Mathematics with Applications24(10), 11–17 (1992). doi:10.1016/0898-1221(92)90016-B
Show all 9 references
-
[9]
https://www.cs.cmu.edu/afs/cs/project/theo-20/www/data/news20.html Accessed 2019-09-01
Lang, K.: CMU Text Learning Group Data Archives. https://www.cs.cmu.edu/afs/cs/project/theo-20/www/data/news20.html Accessed 2019-09-01
2019
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.