Pith. sign in

REVIEW 3 major objections 8 minor 94 references

High dimensional statistical inference: theoretical development to data analytics

T0 review · 3 major / 8 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This survey of high-dimensional inference claims the field is best understood through bias-corrected quadratic forms, random projections, and sparse regularization, and it organizes the main tests and estimators around those mechanisms.

desk verdict A useful but unfinished survey chapter: solid organization and coverage of high-dimensional mean/covariance testing, with a real hole where the author's own dependent-data test is presented but its published correction is never integrated. read the letter →

arxiv 1908.06600 v1 pith:2BZYCHZA submitted 2019-08-19 math.ST stat.MEstat.TH

classification math.STstat.MEstat.TH MSC 62-0262H1562H1262H10
keywords high-dimensionalinferencemeanvectortestingrandomprojectionscovariancematrixestimationgraphicallassoDirichlet-multinomialmultivariatecountmodelsasymptoticnormality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This chapter is a survey of high-dimensional statistical inference, and its central claim is that the field can be organized around a few recurring mechanisms: bias-corrected quadratic forms of the mean difference, random projections, and sparse regularization of covariance structure. This matters for genomics, metagenomics, and text mining, where the number of variables routinely exceeds the sample size and classical tools such as Hotelling's $T^2$ are undefined. The chapter compares the four standard asymptotic mean tests, the random-projection alternative, and the dependent-observation extension, and it identifies open gaps such as covariance testing by projections and feasible inference for multivariate Poisson models. The manuscript itself contains unfinished spots--a placeholder at the start of Section 2.5 and a missing citation in Section 3.2--that qualify the literal promise of completeness even though the coverage is systematic.

What carries the argument

The load-bearing objects are: (1) unbiased quadratic-form functionals such as $M_n = (\bar{X}-\bar{Y})^\top(\bar{X}-\bar{Y}) - \frac{n+m}{nm}\mathrm{tr}(S)$, whose expectation is the squared mean difference; (2) leave-out ratio-consistent variance estimators built from expressions like $\hat{\mathrm{tr}}(\Sigma_1^2)$; (3) random projection matrices derived from the Johnson-Lindenstrauss lemma; (4) $\ell_1$ penalties on the precision matrix, notably the graphical lasso; and (5) the Dirichlet-multinomial hierarchy for overdispersed count data. These carry the argument because each method is defined by which functional it uses and by which variance or penalty estimator it pairs with that functional.

What would settle it

A concrete check: simulate two samples from an exchangeable covariance model where all pairwise correlations equal a fixed $\rho>0$, run $T_{BS}$ and $T_{CQ}$, and compare empirical rejection rates to the nominal level; the chapter predicts inflated type I error because the trace-ratio condition fails, so a simulation that still controls size would refute that specific claim about the strength of covariance assumptions.

Watch

Extended reading notes

Core claim

On its own terms, the chapter establishes that high-dimensional inference now has a coherent toolkit: for mean vectors, replace the singular sample covariance inverse with bias-corrected norms of the mean difference and ratio-consistent variance estimators; for covariance matrices, impose structure through banding or $\ell_1$-regularized precision estimation; for counts, work mainly with multinomial and Dirichlet-multinomial models, where high-dimensional testing succeeds only under sparsity conditions. The four asymptotic mean tests $T_{BS}$, $T_{CQ}$, $T_{SD}$, and $T_{PA}$ are presented as successive relaxations of assumptions on distributions, covariance strength, and the relationship between dimension and sample size, and random-projection tests are presented as an exact alternative for small samples. The chapter also notes that the proofs for the dependent-observation test $T_{APR}$ required later corrections, a caveat it states in Section 2.5.

Load-bearing premise

The chapter's value as a comprehensive reference rests on the correctness of its condensed descriptions of each test, and the manuscript itself flags incomplete parts--a placeholder in Section 2.5, a missing citation in Section 3.2, and known proof corrections to the dependent-observation test--so that promise is not yet fully delivered.

Editorial extensions

If this is right

  • A practitioner comparing two high-dimensional samples can choose among $T_{BS}$, $T_{CQ}$, $T_{SD}$, and $T_{PA}$ based on whether the signal is in raw or standardized units, because the first two are orthogonal-invariant and the last two are scale-invariant.
  • Random-projection tests such as RAPTT give exact rather than asymptotic p-values when $n+m$ is small, but they require multiple projections and bootstrap calibration, making them computationally heavier than the asymptotic tests.
  • Sparse precision-matrix estimation through the graphical lasso turns covariance estimation into network discovery, but the theory is developed mainly under Gaussian assumptions.
  • High-dimensional testing of multinomial parameters is feasible only when the probability vectors are not too concentrated and the dimension grows at most linearly with sample size.
  • Multivariate Bernoulli, binomial, and Poisson models remain largely without feasible parameter inference, so count data from text mining and genomics cannot yet be handled with the same maturity as continuous data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The invariance distinction implies a testable practical rule: when variables are measured on very different scales, scale-invariant tests should detect uniform standardized shifts more often than orthogonal-invariant tests, which a simulation study could verify directly.
  • Because asymptotic growth-rate conditions like $p/n\to\delta$ cannot be confirmed from any single finite dataset, the chapter's own comparison suggests that simulation-based calibration of test choice is more defensible in practice than relying on the stated rates.
  • Extending random projections to covariance-matrix hypotheses, which the chapter identifies as open, should work for sphericity and identity tests because those hypotheses are invariant under orthogonal projection of the data.
  • For latent Dirichlet allocation, the chapter notes that statistical properties of estimators are not established, so developing hypothesis tests for comparing topic parameters across corpora is an open problem it implicitly highlights.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. This manuscript is a survey chapter, intended for the Handbook of Statistics, that promises a comprehensive overview of high-dimensional statistical inference with an emphasis on data-analytic applicability. It covers tests for the mean vector under independence (Bai-Saranadasa, Chen-Qin, Srivastava-Du, Park-Ayyala), projection-based and random-projection approaches, and a section on dependent observations; estimation and tests for covariance matrices; and discrete multivariate models (multinomial, Dirichlet-multinomial, Bernoulli, binomial, Poisson). The chapter also discusses open problems and computational considerations.

Significance. If completed and corrected, the chapter would fill a useful niche: it collects the main asymptotic mean tests, compares their invariance and assumptions, and connects high-dimensional theory to genomics and text-mining applications. The comparison of the invariance properties of TBS, TCQ, TSD, and TPA, as well as the discussion of the Dirichlet-multinomial model, are informative and would be valuable pedagogical material. However, the current text contains a literal placeholder, an unresolved citation, and, more importantly, an unintegrated correction to a presented test, so the promised 'comprehensive overview' is not yet delivered. The significance of the review will depend on these issues being fixed.

major comments (3)
  1. [Section 2.2, Eq. (14)] The claim that for k < p and a full-row-rank matrix R, R(µ1−µ2)=0 iff µ1−µ2=0 is false, because R has a (p−k)-dimensional null space. Consequently, the projected hypothesis (14) is not equivalent to the original hypothesis (1): alternatives with µ1−µ2 in the null space of R are invisible to any test based on the projected data. The power simulation in Figure 1 illustrates the loss for small k but does not repair the equivalence statement, which is load-bearing for the motivation of all projection-based tests in Sections 2.2 and 2.3.
  2. [Section 2.3, Eqs. (19)-(21)] The description of RAPTT contains an inconsistent rejection rule. The text defines ψα by P(p>ψα|H0)=1−α and then says to reject if p>ψα, while the algorithm estimates ψα as p[M(1−α)], the (1−α)-quantile of the simulated null distribution of p. Under the stated alternatives, the individual projected Hotelling p-values are stochastically smaller than uniform, so the average p is small; a test that rejects for large average p-values has power tending to zero under those alternatives. The rejection region should be the lower tail (reject if p < ψα, with ψα the α-quantile), and the two definitions of ψα should be made consistent.
  3. [Section 2.5, Eq. (33) and following paragraph] The chapter presents TAPR and assumptions (APR I)-(APR IV), then states that Cho et al. [23] identified theoretical errors in Ayyala et al. and provided corrections, but it does not say which assumptions were corrected, whether (APR II)'s M=O(n^{1/8}) and (APR III) survive, or whether the leave-out variance estimator referenced only to [6] is the corrected one. Since the chapter also refers the reader elsewhere for the estimator's exact form, a reader cannot actually learn the TAPR test from this 'comprehensive' chapter. The section opens with the placeholder '(write motivation)', which further confirms that this part of the manuscript is unfinished.
minor comments (8)
  1. [Section 2.1, below Eq. (8)] In the leave-out definitions, Y(i) and Y(i,j) are defined with denominators n−1 and n−2; these should be m−1 and m−2 to match the sample size of the Y group.
  2. [Section 2.1, paragraph after Eq. (5)] The expression 'Bn = 2n^{1−α}λmax' appears to have the exponent reversed; for p=Cn^α, the bias is of order n^{α−1}, which is what makes the subsequent statement 'diverges when α>1' correct.
  3. [Section 2.5, Eq. (31)] The second indicator inside the double sum contains an undefined index 'i'; it should presumably be 'a' or another explicitly defined index, and the expression should be checked.
  4. [Section 3.1, below Eq. (35)] The sentence saying the banded estimator is 'clearly indicating Σ̂(k) as the diagonal estimator' is imprecise; the diagonal estimator is the special case k=1, while for k>1 the estimator is banded.
  5. [Section 3.2, Eq. (42)] The citation 'John [ ? ]' is unresolved; the reference must be supplied before publication.
  6. [Section 4.2] The sentence 'While the density function is known to be globally convex, maximization can still lead us to a local maxima' is contradictory as written; if the relevant function is convex, interior local maxima are not a concern, and if the intended statement concerns the log-likelihood, concavity is the usual property.
  7. [Section 4.3.1, Eq. (58)] The displayed mass function is not readable as printed; the notation with products and subscripts should be rewritten with explicit joint probabilities.
  8. [Throughout] There are numerous typos (e.g., 'projeting', 'repsectively', 'distributiosn', 'likeliho0d', 'conjuction', 'moreknown than unkown') and at least one grammatical break ('R1 and R2 are the for notational convenience'); a careful copyedit is needed.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this survey chapter quotes and cites methods rather than deriving them; self-citations are descriptive, not load-bearing.

full rationale

No circularity found. This is a survey chapter, not an original derivation: the test statistics are presented with their stated assumptions and are cited to their original sources (e.g., Bai-Saranadasa in equation (6), Chen-Qin in equation (9), Srivastava-Du in equation (10), Park-Ayyala in equation (13)). The author's own methods appear in Sections 2.1, 2.5, and the data-analytic discussion, but those descriptions are expository and are not used as premises to prove a new result or to force a methodological choice. The T_APR passage explicitly reports that Cho et al. identified proof errors in Ayyala et al. and that corrections exist, which is an in-text limitation statement rather than a derivation step. The placeholder '(write motivation)' at the start of Section 2.5 and the incomplete 'John [ ? ]' citation in Section 3.2 are completeness and polish defects, not evidence that an input was relabeled as an output. No equation is self-definitional, no fitted parameter is renamed a prediction, and no uniqueness assertion is imported from a same-author citation. Therefore the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The chapter is a review, so it does not introduce new axioms; the entries above capture the background assumptions that the review's claims about the surveyed methods depend on.

assumptions (4)
  • domain assumption The asymptotic null distributions of the reviewed test statistics (e.g., T_BS, T_CQ, T_SD, T_PA) are correctly reproduced from the original papers.
    The review's usefulness as a reference rests on the accuracy of these stated distributions; the author does not re-derive them.
  • domain assumption The factor model X = mu + Gamma Z with i.i.d. components Z having finite fourth moments is a valid representation for the data in the mean-vector tests.
    This assumption underlies the Bai-Saranadasa, Chen-Qin, and Park-Ayyala tests described in Section 2.1.
  • domain assumption For the dependent-observations test, the data are realizations of M-dependent strictly stationary Gaussian processes with second-order stationarity.
    This is assumption (APR I) in Section 2.5, taken from Ayyala et al. (2017).
  • standard math The Johnson-Lindenstrauss lemma guarantees the existence of random projections preserving pairwise distances.
    Invoked in Section 2.3 to justify random projection dimension reduction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of High dimensional statistical inference: theoretical development to data analytics." pith.science (2026). https://pith.science/paper/2BZYCHZA

@misc{pith2026190806600,
  author       = {Pith},
  title        = {Pith review of: High dimensional statistical inference: theoretical development to data analytics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2BZYCHZA}},
  note         = {Machine review of arXiv:1908.06600}
}
read the original abstract

This article is due to appear in the Handbook of Statistics, Vol. 43, Elsevier/North-Holland, Amsterdam, edited by Arni S. R. Srinivasa Rao and C. R. Rao. In modern day analytics, there is ever growing need to develop statistical models to study high dimensional data. Between dimension reduction, asymptotics-driven methods and random projection based methods, there are several approaches developed so far. For high dimensional parametric models, estimation and hypothesis testing for mean and covariance matrices have been extensively studied. However, practical implementation of these methods are fairly limited and are primarily restricted to researchers involved in high dimensional inference. With several applied fields such as genomics, metagenomics and social networking, high dimensional inference is a key component of big data analytics. In this chapter, a comprehensive overview of high dimensional inference and its applications in data analytics is provided. Key theoretical developments and computational tools are presented, giving readers an in-depth understanding of challenges in big data analysis.

Figures

Figures reproduced from arXiv: 1908.06600 by the authors.

Figure 1
Figure 1. Figure complete data (k = p), whereas projecting into a single dimension always fails to reject H0. The smallest k for which the p-value supports rejecting H0 for δ = 0.2, 0.4 and 1 are 26, 15 and 5 respectively. The variance is kept constant for the three models, which implies the difference in results in due to δ. As δ increases, there is greater separation between the two means and hence smaller k is sufficient. … view at source ↗
Figure 2
Figure 2. Figure [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

94 extracted references · 52 canonical work pages

  1. [23]

    S. Cho, J. Lim, D. N. Ayyala, J. Park, and A. Roy. Note on Mean Vector Testing for High-Dimensional Dependent Observations. arXiv e-prints, art. arXiv:1904.09344, Apr 2019

  2. [6]

    D. N. Ayyala, J. Park, and A. Roy. Mean vector testing for high-dimensional dependent observations. Journal of Multivariate Analysis , 153:136–155, 2017. ISSN 0047-259X. doi: 10.1016/j.jmva.2016.09.012. URL http://www.sciencedirect.com/science/article/pii/S0047259X16300999

  3. [1]

    Achlioptas

    D. Achlioptas. Database-friendly random projections: Johnson-Lindenstrauss with binary coins. Journal of Computer and System Sciences , 66(4):671–687, 2003. ISSN 00220000. doi: 10.1016/S0022-0000(03) 00025-4. 31

  4. [2]

    A. C. Aitken and H. T. Gonin. XI.On Fourfold Sampling with and without Replacement. Proceedings of the Royal Society of Edinburgh , 55:114–125, 1936. doi: 10.1017/S0370164600014413

  5. [3]

    P. M. E. Altham. Two Generalizations of the Binomial Distribution. Journal of the Royal Statistical Society. Series C (Applied Statistics), 27(2):162–167, 1978. ISSN 00359254. doi: 10.2307/2346943. URL http://www.jstor.org/stable/2346943

  6. [4]

    T. W. Anderson. An Introduction to Multivariate Statistical Analysis, 3rd edition . John Wiley and Sons, 2003

  7. [5]

    D. N. Ayyala, D. E. Frankhouser, G. Marcucci, J.-O. Ganbat, P. Yan, R. Bundschuh, and S. Lin. Statistical methods for detecting differentially methylated regions based on MethylCap-seq data. Brief- ings in Bioinformatics , 17(6):926–937, 10 2015. ISSN 1467-5463. doi: 10.1093/bib/bbv089. URL https://doi.org/10.1093/bib/bbv089

  8. [7]

    Bai and H

    Z. Bai and H. Saranadasa. Effect of High Dimension: By an Example of a Two Sample Problem. Statistica Sinica, 6:311–329, 1996. ISSN 10170405

Show all 94 references
  1. [8]

    Z. Bai, D. Jiang, J. F. Yao, and S. Zheng. Corrections to LRT on large-dimensional covariance matrix by RMT. Annals of Statistics , 37(6 B):3822–3840, 2009. ISSN 00905364. doi: 10.1214/09-AOS694

  2. [9]

    Balakrishnan and L

    S. Balakrishnan and L. Wasserman. Hypothesis testing for high-dimensional multinomials: A se- lective review1. Annals of Applied Statistics , 12(2):727–749, 2018. ISSN 19417330. doi: 10.1214/ 18-AOAS1155SF

  3. [10]

    H. E. Barmi and R. L. Dykstra. Restricted multinomial maximum likelihood estimation based upon Fenchel duality. Statistics & Probability Letters , 21(2):121–130, 1994. ISSN 0167-7152. doi: 10.1016/0167-7152(94)90219-4. URL http://www.sciencedirect.com/science/article/pii/ 0167...

  4. [11]

    P. J. Bickel and E. Levina. Covariance regularization by thresholding. Annals of Statistics , 36(6): 2577–2604, 2008. ISSN 00905364. doi: 10.1214/08-AOS600

  5. [12]

    Bien and R

    J. Bien and R. J. Tibshirani. Sparse estimation of a covariance matrix. Biometrika, 98(4):807–820,

  6. [13]

    Bingham and H

    E. Bingham and H. Mannila. Random projection in dimensionality reduction: Applications to image and text data. In Proceedings of the Seventh ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’01, pages 245–250, New York, NY, USA, 2001. ACM. ISBN 1...

  7. [14]

    Biswas and J

    A. Biswas and J. S. Hwang. A new bivariate binomial distribution. Statistics and Probability Letters , 60(2):231–240, 2002. ISSN 01677152. doi: 10.1016/S0167-7152(02)00323-1

  8. [15]

    D. M. Blei, B. B. Edu, A. Y. Ng, A. S. Edu, M. I. Jordan, and J. B. Edu. technique...Latent Dirichlet Allocation. Journal of Machine Learning Research , 3:993–1022, 2003. ISSN 15324435. doi: 10.1162/ jmlr.2003.3.4-5.993. 32

  9. [16]

    P. J. Brockwell and R. A. Davis. Time Series: Theory and Methods . Springer-Verlag, Berlin, Heidelberg,

  10. [17]

    T. Cai, W. Liu, and X. Luo. A constrained 𝓁1 minimization approach to sparse precision matrix estimation. Journal of the American Statistical Association , 106(494):594–607, 2011. doi: 10.1198/jasa. 2011.tm10155. URL https://doi.org/10.1198/jasa.2011.tm10155

  11. [18]

    Liu, and Y

    Cai, TT., W. Liu, and Y. Xia. Two-sample test of high dimensional means under dependence. Journal of the Royal Statistical Society. Series B: Statistical Methodology , 76(2):349–372, 2014. ISSN 13697412. doi: 10.1111/rssb.12034

  12. [19]

    M. C. Cario and B. L. Nelson. Modeling and generating random vectors with arbitrary marginal dis- tributions and correlation matrix. Industrial Engineering, pages 1–19, 1997. URL http://citeseerx. ist.psu.edu/viewdoc/download?doi=10.1.1.48.281{&}rep=rep1{&}type=pdf

  13. [20]

    S.-O. Chan, I. Diakonikolas, P. Valiant, and G. Valiant. Optimal Algorithms for Testing Closeness of Discrete Distributions. Proceedings of the Twenty-Fifth Annual ACM-SIAM Symposium on Discrete Algorithms, pages 1193–1203, 2013. doi: 10.1137/1.9781611973402.88

  14. [21]

    Chen and H

    J. Chen and H. Li. Variable selection for sparse Dirichlet-multinomial regression with an application to microbiome data analysis. Annals of Applied Statistics , 7(1):418–442, 2013. ISSN 19326157. doi: 10.1214/12-AOAS592

  15. [22]

    S. X. Chen and Y. L. Qin. A two-sample test for high-dimensional data with applications to gene-set testing. Annals of Statistics , 38(2):808–835, 2010. ISSN 00905364. doi: 10.1214/09-AOS716

  16. [24]

    J. H. Chung and D. A. S. Fraser. Randomization tests for a multivariate two-sample problem.Journal of the American Statistical Association , 53(283):729–735, 1958. URL https://www.jstor.org/stable/ 2282050

  17. [25]

    S. A. Crossley, M. Dascalu, and D. S. Mcnamara. How important is size? An Investigation of Corpus Size and Meaning in both Latent Semantic Analysis and Latent Dirichlet Allocation. Proceedings of the Thirtieth International Florida Artificial Intelligence Research Society Confe...

  18. [26]

    B. Dai, S. Ding, and G. Wahba. Multivariate Bernoulli distribution. Bernoulli, 19(4):1465–1483,

  19. [27]

    Danaher, P

    P. Danaher, P. Wang, and D. M. Witten. The joint graphical lasso for inverse covariance estimation across multiple classes. Journal of the Royal Statistical Society. Series B: Statistical Methodology , 76 (2):373–397, 2014. ISSN 13697412. doi: 10.1111/rssb.12033

  20. [28]

    P. J. Danaher. Parameter estimation for the dirichlet-multinomial distribution using supplementary beta-binomial data. Communications in Statistics - Theory and Methods , 17(6):1777–1788, 1988. doi: 10.1080/03610928808829713

  21. [29]

    M. J. Daniels and R. E. Kass. Shrinkage estimators for covariance matrices. Biometrics, 57(4):1173– 1184, 2001. ISSN 0006341X. doi: 10.1111/j.0006-341X.2001.01173.x. 33

  22. [30]

    a. P. Dempster. A High Dimensional Two Sample Significance Test. The Annals of Mathematical Statistics, 29(4):995–1010, 1958. ISSN 0003-4851. doi: 10.1214/aoms/1177706437

  23. [31]

    J. Fan, F. Han, and H. Liu. Challenges of Big Data analysis. National Science Review , 1(2):293–314,

  24. [32]

    Fradkin and D

    D. Fradkin and D. Madigan. Experiments with random projections for machine learning. In Proceedings of the Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , KDD ’03, pages 517–522, New York, NY, USA, 2003. ACM. ISBN 1-58113-737-0. doi: 10.1145/...

  25. [33]

    Friedman, T

    J. Friedman, T. Hastie, and R. Tibshirani. Sparse inverse covariance estimation with the graphical lasso. Biostatistics, 9(3):432–441, 2008. ISSN 14654644. doi: 10.1093/biostatistics/kxm045

  26. [34]

    Goodfellow, Y

    I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning . MIT Press, 2016. http://www. deeplearningbook.org

  27. [35]

    K. B. Gregory, R. J. Carroll, V. Baladandayuthapani, and S. N. Lahiri. A Two-Sample Test for Equality of Means in High Dimension. Journal of the American Statistical Association , 110(510):837–849, 2015. ISSN 1537274X. doi: 10.1080/01621459.2014.934826

  28. [36]

    J. Guo, E. Levina, G. Michailidis, and J. Zhu. Joint estimation of multiple graphical models.Biometrika, 98(1):1–15, 2011. ISSN 00063444. doi: 10.1093/biomet/asq060

  29. [37]

    H. S. Hariharan and R. P. Velu. On estimating dirichlet parametersa comparison of initial values. Journal of Statistical Computation and Simulation , 48(1-2):47–58, 1993. ISSN 15635163. doi: 10.1080/ 00949659308811539

  30. [38]

    Hoeffding

    W. Hoeffding. Asymptotically Optimal Tests for Multinomial Distributions Author ( s ): Wassily Hoeffding Source : The Annals of Mathematical Statistics , Vol . 36 , No . 2 ( Apr ., 1965 ), pp . 369- 401 Published by : Institute of Mathematical Statistics Stable URL : ht. The Ann...

  31. [39]

    M. D. Hoffman, D. M. Blei, and F. Bach. Online Learning for Latent Dirichlet Allocation. InAdvances in Neural Information Processing Systems 23, volume 1, pages 856–864, 2010. ISBN 9781450300551. URL http://papers.nips.cc/paper/3902-online-learning-for-latent-dirichlet-allocation.pdf

  32. [40]

    Holmes, K

    I. Holmes, K. Harris, and C. Quince. Dirichlet multinomial mixtures: Generative models for microbial metagenomics. PLoS ONE, 7(2), 2012. ISSN 19326203. doi: 10.1371/journal.pone.0030126

  33. [41]

    Hotelling

    H. Hotelling. The generalization of student’s ratio. The Annals of Mathematical Statistics, 2(3):360–378, 08 1931. doi: 10.1214/aoms/1177732979. URL https://doi.org/10.1214/aoms/1177732979

  34. [42]

    W. N. Hudson, H. G. Tucker, and J. A. Veeh. Limit theorems for the multivariate binomial distribution. Journal of Multivariate Analysis , 18(1):32–45, 1986. ISSN 10957243. doi: 10.1016/0047-259X(86) 90056-4

  35. [43]

    Inouye, E

    D. Inouye, E. Yang, G. Allen, and P. Ravikumar. A Review of Multivariate Distributions for Count Data Derived from the Poisson Distribution. Wiley Interdisciplinary Review Computational Statistics , 9(3), 2017. doi: 10.1002/wics.1398.A. 34

  36. [44]

    N. P. Jewell and J. D. Kalbfleisch. Maximum likelihood estimation of ordered multinomial parameters. Biostatistics, 5(2):291–306, 2004. ISSN 14654644. doi: 10.1093/biostatistics/5.2.291

  37. [45]

    Jiang, T

    D. Jiang, T. Jiang, and F. Yang. Likelihood ratio tests for covariance matrices of high-dimensional normal distributions. Journal of Statistical Planning and Inference , 142(8):2241–2256, 2012. ISSN 03783758. doi: 10.1016/j.jspi.2012.02.057. URL http://dx.doi.org/10.1016/j.jsp...

  38. [46]

    W. B. Johnson and J. Lindenstrauss. Extensions of Lipschitz mappings into a Hilbert space. Contem- porary Mathematics, 26:189–206, 1984. doi: 10.1090/conm/026/737400

  39. [47]

    Karlis and E

    D. Karlis and E. Xekalaki. Mixed Poisson Distributions. International Statistical Review, 73(1):35–58,

  40. [48]

    A. S. Krishnamoorthy. Multivariate Binomial and Poisson Distributions. Sankhya B , 11(2):117–124,

  41. [49]

    A. Kudo. A multivariate analogue of the one-sided test. Biometrika, 50(3):403–418, 1963. URL https://www.jstor.org/stable/2333909

  42. [50]

    Ledoit and M

    O. Ledoit and M. Wolf. Some hypothesis tests for the covariance matrix when the dimension is large compared to the sample size. The Annals of Statistics , 30(4):1081–1102, 2002

  43. [51]

    T. Leonard. A Bayesian Approach to Some Multinomial Estimation and Pretesting Problems. Journal of the American Statistical Association , 72(360):869–874, 1977

  44. [52]

    B. Levin. A Representation for Multinomial Cumulative Distribution Functions. The Annals of Statis- tics, 9(5):1123–1126, 1981. URL https://www.jstor.org/stable/2240628

  45. [53]

    Li and S

    J. Li and S. X. Chen. Two sample tests for high-dimensional covariance matrices. Annals of Statistics , 40(2):908–940, 2012. ISSN 00905364. doi: 10.1214/12-AOS993

  46. [54]

    P. Li, T. J. Hastie, and K. W. Church. Very sparse random projections. In Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining - KDD ’06 , pages 287–296,

  47. [55]

    M. E. Lopes, L. J. Jacob, and M. J. Wainwright. A More Powerful Two-Sample Test in High Dimensions using Random Projection. 2015

  48. [56]

    P. J. McMurdie and S. Holmes. Waste not, want not: Why rarefying microbiome data is inadmissible. PLOS Computational Biology , 10(4):1–12, 04 2014. doi: 10.1371/journal.pcbi.1003531. URL https: //doi.org/10.1371/journal.pcbi.1003531

  49. [57]

    K. S. Miller. On the Inverse of the Sum of Matrices. Mathematics Magazine, 54(2):67–72, 1981. URL https://www.jstor.org/stable/2690437

  50. [58]

    Mimno, M

    D. Mimno, M. D. Hoffman, and D. M. Blei. Sparse Stochastic Inference for Latent Dirichlet allocation. In ICML’12 Proceedings of the 29th International Coference on International Conference on Machine Learning, pages 1515–1522, 2012. URL http://arxiv.org/abs/1206.6425

  51. [59]

    C. Morris. Central Limit Theorems for Multinomial Sums. The Annals of Statistics, 3(1):165–188, 1975. URL https://www.jstor.org/stable/2958086

  52. [60]

    R. J. Muirhead. Aspects of Multivariate Statistical Theory . John Wiley and Sons, 1982. 35

  53. [61]

    H. Nagao. On some test criteria for covariance matrix. The Annals of Statistics , 1(4):700–709, 1973

  54. [62]

    R. B. Nelson. An Introduction to Copulas . Springer Series in Statistics, second edition. ISBN 0-387- 28659-4

  55. [64]

    Park and D

    J. Park and D. N. Ayyala. A test for the mean vector in large dimension and small sam- ples. Journal of Statistical Planning and Inference , 143(5):929–943, may 2013. ISSN 0378-3758. doi: 10.1016/J.JSPI.2012.11.001. URL https://www.sciencedirect.com/science/article/pii/ S03783...

  56. [65]

    Plunkett and J

    A. Plunkett and J. Park. Two-sample test for sparse high-dimensional multinomial distributions. Test, 2018. ISSN 11330686. doi: 10.1007/s11749-018-0600-8. URL https://doi.org/10.1007/ s11749-018-0600-8

  57. [66]

    C. R. Rao. Advanced Statistical Methods in Data Science . Wiley, 1952. ISBN 02-850820-3

  58. [67]

    C. R. Rao. Maximum likelihood estimation for the multinomial distribution. Sankhy: The Indian Journal of Statistics (1933-1960) , 18(1/2):139–148, 1957. ISSN 00364452. URL http://www.jstor. org/stable/25048341

  59. [68]

    G. Ronning. Maximum likelihood estimation of dirichlet distributions. Journal of Statistical Computa- tion and Simulation , 32(4):215–221, 1989. ISSN 15635163. doi: 10.1080/00949658908811178

  60. [69]

    S., Srivastava and M

    M. S., Srivastava and M. Du. A test for the mean vector with fewer observations than the dimension. Journal of Multivariate Analysis , 99(3):386–402, 2008. ISSN 0047-259X. doi: 10.1016/j.jmva.2006.11

  61. [70]

    J. R. Schott. A test for the equality of covariance matrices when the dimension is large relative to the sample sizes. Computational Statistics and Data Analysis , 51(12):6535–6542, 2007. ISSN 01679473. doi: 10.1016/j.csda.2007.03.004

  62. [71]

    Shin and R

    K. Shin and R. Pasupathy. An algorithm for fast generation of bivariate poisson random vectors. INFORMS Journal on Computing , 22(1):81–92, 2010. ISSN 10919856. doi: 10.1287/ijoc.1090.0332

  63. [72]

    M. Sklar. Fast MLE Computation for the Dirichlet Multinomial. 2014. URL http://arxiv.org/abs/ 1405.0099

  64. [73]

    M. S. Srivastava. A test for the mean vector with fewer observations than the dimension under non- normality. Journal of Multivariate Analysis , 100(3):518–532, 2009. ISSN 0047259X. doi: 10.1016/j. jmva.2008.06.006

  65. [74]

    M. S. Srivastava. Some Tests Concerning the Covariance Matrix in High Dimensional Data. Journal of the Japan Statistical Society , 35(2):251–272, 2013. ISSN 1882-2754. doi: 10.14490/jjss.35.251

  66. [75]

    M. S. Srivastava and H. Yanagihara. Testing the equality of several covariance matrices with fewer observations than the dimension. Journal of Multivariate Analysis , 101(6):1319–1329, 2010. ISSN 0047259X. doi: 10.1016/j.jmva.2009.12.010. URL http://dx.doi.org/10.1016/j.jmva.2...

  67. [76]

    M. S. Srivastava, S. Katayama, and Y. Kano. A two sample test in high dimensional data. Journal of Multivariate Analysis , 114(1):349–358, 2013. ISSN 10957243. doi: 10.1016/j.jmva.2012.08.014. URL http://dx.doi.org/10.1016/j.jmva.2012.08.014

  68. [77]

    URL http://www.sciencedirect.com/science/article/pii/S0047259X06001990

  69. [78]

    Srivastava, P

    R. Srivastava, P. Li, and D. Ruppert. RAPTT: An Exact Two-Sample Test in High Dimensions Using Random Projections. Journal of Computational and Graphical Statistic , 25(3):954–970, 2016. doi: 10.1080/10618600.2015.1062771

  70. [79]

    J. M. Stern and S. Zacks. Testing the independence of Poisson variates under the Holgate bivariate distribution: The power of a new evidence test. Statistics and Probability Letters , 60(3):313–320, 2002. ISSN 01677152. doi: 10.1016/S0167-7152(02)00314-0

  71. [80]

    Z. Sun, T. Wang, K. Deng, X. F. Wang, R. Lafyatis, Y. Ding, M. Hu, and W. Chen. DIMM-SC: A Dirichlet mixture model for clustering droplet-based single cell transcriptomic data. Bioinformatics, 34 (1):139–146, 2018. ISSN 14602059. doi: 10.1093/bioinformatics/btx490

  72. [81]

    J. L. Teugels. Some representations of the multivariate Bernoulli and binomial distributions. Journal of Multivariate Analysis , 32(2):256–268, 1990. ISSN 10957243. doi: 10.1016/0047-259X(90)90084-U

  73. [82]

    Tibshirani

    R. Tibshirani. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical So- ciety. Series B: Statistical Methodology , 58(1):267–288, 1996. URL https://www.jstor.org/stable/ 2346178

  74. [83]

    Z. Wang, M. Gerstein, and M. Snyder. Rna-seq: a revolutionary tool for transcriptomics. Nature Review Genetics, 10(1):57–63, 2009. doi: 10.1038/nrg2484

  75. [84]

    Y. Wu, M. G. Genton, and L. A. Stefanski. A multivariate two-sample mean test for small sample size and missing data. Biometrics, 62(3):877–885, 2006. ISSN 0006341X. doi: 10.1111/j.1541-0420.2006. 00533.x

  76. [85]

    M. S. Srivastava, H. Yanagihara, and T. Kubokawa. Tests for covariance matrices in high dimension with less sample size. Journal of Multivariate Analysis , 130:289–309, 2014. ISSN 10957243. doi: 10.1016/j.jmva.2014.06.003. URL http://dx.doi.org/10.1016/j.jmva.2014.06.003

  77. [86]

    Zhong, S

    P.-S. Zhong, S. X. Chen, and M. Xu. Tests alternative to higher criticism for high-dimensional means under sparsity and column-wise dependence. Ann. Statist. , 41(6):2820–2851, 12 2013. doi: 10.1214/ 13-AOS1168. URL https://doi.org/10.1214/13-AOS1168

  78. [87]

    R. S. Zoh, A. Sarkar, R. J. Carroll, and B. K. Mallick. A Powerful Bayesian Test for Equality of Means in High Dimensions. Journal of the American Statistical Association , 113(524):1733–1741, 2018. ISSN 1537274X. doi: 10.1080/01621459.2017.1371024. URL https://doi.org/10.1080...

  79. [93]

    Zelterman

    D. Zelterman. Goodness-of-Fit Tests for Large Sparse Distributions Multinomial. Journal of the Amer- ican Statistical Association, 82(398):624–629, 2013. URL https://www.jstor.org/stable/2289474

  80. [1951]

    URL https://www.jstor.org/stable/25048072

  81. [2006]

    doi: 10.1145/1150402.1150436

    ISBN 1595933395. doi: 10.1145/1150402.1150436

  82. [2010]

    doi: 10.1111/j.1751-5823.2005.tb00250.x

  83. [2011]

    doi: 10.1093/biomet/asr054

    ISSN 00063444. doi: 10.1093/biomet/asr054

  84. [2013]

    doi: 10.3150/12-BEJSP10

    ISSN 1350-7265. doi: 10.3150/12-BEJSP10. URL http://projecteuclid.org/euclid.bj/ 1377612861

  85. [2014]

    doi: 10.1093/nsr/nwt032

    ISSN 2053714X. doi: 10.1093/nsr/nwt032

  86. [2018]

    URL http://arxiv.org/abs/1807.00930

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.