REVIEW 3 major objections 8 minor 94 references
High dimensional statistical inference: theoretical development to data analytics
T0 review · 3 major / 8 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This survey of high-dimensional inference claims the field is best understood through bias-corrected quadratic forms, random projections, and sparse regularization, and it organizes the main tests and estimators around those mechanisms.
desk verdict A useful but unfinished survey chapter: solid organization and coverage of high-dimensional mean/covariance testing, with a real hole where the author's own dependent-data test is presented but its published correction is never integrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing objects are: (1) unbiased quadratic-form functionals such as $M_n = (\bar{X}-\bar{Y})^\top(\bar{X}-\bar{Y}) - \frac{n+m}{nm}\mathrm{tr}(S)$, whose expectation is the squared mean difference; (2) leave-out ratio-consistent variance estimators built from expressions like $\hat{\mathrm{tr}}(\Sigma_1^2)$; (3) random projection matrices derived from the Johnson-Lindenstrauss lemma; (4) $\ell_1$ penalties on the precision matrix, notably the graphical lasso; and (5) the Dirichlet-multinomial hierarchy for overdispersed count data. These carry the argument because each method is defined by which functional it uses and by which variance or penalty estimator it pairs with that functional.
What would settle it
A concrete check: simulate two samples from an exchangeable covariance model where all pairwise correlations equal a fixed $\rho>0$, run $T_{BS}$ and $T_{CQ}$, and compare empirical rejection rates to the nominal level; the chapter predicts inflated type I error because the trace-ratio condition fails, so a simulation that still controls size would refute that specific claim about the strength of covariance assumptions.
Extended reading notes
Core claim
On its own terms, the chapter establishes that high-dimensional inference now has a coherent toolkit: for mean vectors, replace the singular sample covariance inverse with bias-corrected norms of the mean difference and ratio-consistent variance estimators; for covariance matrices, impose structure through banding or $\ell_1$-regularized precision estimation; for counts, work mainly with multinomial and Dirichlet-multinomial models, where high-dimensional testing succeeds only under sparsity conditions. The four asymptotic mean tests $T_{BS}$, $T_{CQ}$, $T_{SD}$, and $T_{PA}$ are presented as successive relaxations of assumptions on distributions, covariance strength, and the relationship between dimension and sample size, and random-projection tests are presented as an exact alternative for small samples. The chapter also notes that the proofs for the dependent-observation test $T_{APR}$ required later corrections, a caveat it states in Section 2.5.
Load-bearing premise
The chapter's value as a comprehensive reference rests on the correctness of its condensed descriptions of each test, and the manuscript itself flags incomplete parts--a placeholder in Section 2.5, a missing citation in Section 3.2, and known proof corrections to the dependent-observation test--so that promise is not yet fully delivered.
Editorial extensions
If this is right
- A practitioner comparing two high-dimensional samples can choose among $T_{BS}$, $T_{CQ}$, $T_{SD}$, and $T_{PA}$ based on whether the signal is in raw or standardized units, because the first two are orthogonal-invariant and the last two are scale-invariant.
- Random-projection tests such as RAPTT give exact rather than asymptotic p-values when $n+m$ is small, but they require multiple projections and bootstrap calibration, making them computationally heavier than the asymptotic tests.
- Sparse precision-matrix estimation through the graphical lasso turns covariance estimation into network discovery, but the theory is developed mainly under Gaussian assumptions.
- High-dimensional testing of multinomial parameters is feasible only when the probability vectors are not too concentrated and the dimension grows at most linearly with sample size.
- Multivariate Bernoulli, binomial, and Poisson models remain largely without feasible parameter inference, so count data from text mining and genomics cannot yet be handled with the same maturity as continuous data.
Reading between the lines
- The invariance distinction implies a testable practical rule: when variables are measured on very different scales, scale-invariant tests should detect uniform standardized shifts more often than orthogonal-invariant tests, which a simulation study could verify directly.
- Because asymptotic growth-rate conditions like $p/n\to\delta$ cannot be confirmed from any single finite dataset, the chapter's own comparison suggests that simulation-based calibration of test choice is more defensible in practice than relying on the stated rates.
- Extending random projections to covariance-matrix hypotheses, which the chapter identifies as open, should work for sphericity and identity tests because those hypotheses are invariant under orthogonal projection of the data.
- For latent Dirichlet allocation, the chapter notes that statistical properties of estimators are not established, so developing hypothesis tests for comparing topic parameters across corpora is an open problem it implicitly highlights.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey chapter, intended for the Handbook of Statistics, that promises a comprehensive overview of high-dimensional statistical inference with an emphasis on data-analytic applicability. It covers tests for the mean vector under independence (Bai-Saranadasa, Chen-Qin, Srivastava-Du, Park-Ayyala), projection-based and random-projection approaches, and a section on dependent observations; estimation and tests for covariance matrices; and discrete multivariate models (multinomial, Dirichlet-multinomial, Bernoulli, binomial, Poisson). The chapter also discusses open problems and computational considerations.
Significance. If completed and corrected, the chapter would fill a useful niche: it collects the main asymptotic mean tests, compares their invariance and assumptions, and connects high-dimensional theory to genomics and text-mining applications. The comparison of the invariance properties of TBS, TCQ, TSD, and TPA, as well as the discussion of the Dirichlet-multinomial model, are informative and would be valuable pedagogical material. However, the current text contains a literal placeholder, an unresolved citation, and, more importantly, an unintegrated correction to a presented test, so the promised 'comprehensive overview' is not yet delivered. The significance of the review will depend on these issues being fixed.
major comments (3)
- [Section 2.2, Eq. (14)] The claim that for k < p and a full-row-rank matrix R, R(µ1−µ2)=0 iff µ1−µ2=0 is false, because R has a (p−k)-dimensional null space. Consequently, the projected hypothesis (14) is not equivalent to the original hypothesis (1): alternatives with µ1−µ2 in the null space of R are invisible to any test based on the projected data. The power simulation in Figure 1 illustrates the loss for small k but does not repair the equivalence statement, which is load-bearing for the motivation of all projection-based tests in Sections 2.2 and 2.3.
- [Section 2.3, Eqs. (19)-(21)] The description of RAPTT contains an inconsistent rejection rule. The text defines ψα by P(p>ψα|H0)=1−α and then says to reject if p>ψα, while the algorithm estimates ψα as p[M(1−α)], the (1−α)-quantile of the simulated null distribution of p. Under the stated alternatives, the individual projected Hotelling p-values are stochastically smaller than uniform, so the average p is small; a test that rejects for large average p-values has power tending to zero under those alternatives. The rejection region should be the lower tail (reject if p < ψα, with ψα the α-quantile), and the two definitions of ψα should be made consistent.
- [Section 2.5, Eq. (33) and following paragraph] The chapter presents TAPR and assumptions (APR I)-(APR IV), then states that Cho et al. [23] identified theoretical errors in Ayyala et al. and provided corrections, but it does not say which assumptions were corrected, whether (APR II)'s M=O(n^{1/8}) and (APR III) survive, or whether the leave-out variance estimator referenced only to [6] is the corrected one. Since the chapter also refers the reader elsewhere for the estimator's exact form, a reader cannot actually learn the TAPR test from this 'comprehensive' chapter. The section opens with the placeholder '(write motivation)', which further confirms that this part of the manuscript is unfinished.
minor comments (8)
- [Section 2.1, below Eq. (8)] In the leave-out definitions, Y(i) and Y(i,j) are defined with denominators n−1 and n−2; these should be m−1 and m−2 to match the sample size of the Y group.
- [Section 2.1, paragraph after Eq. (5)] The expression 'Bn = 2n^{1−α}λmax' appears to have the exponent reversed; for p=Cn^α, the bias is of order n^{α−1}, which is what makes the subsequent statement 'diverges when α>1' correct.
- [Section 2.5, Eq. (31)] The second indicator inside the double sum contains an undefined index 'i'; it should presumably be 'a' or another explicitly defined index, and the expression should be checked.
- [Section 3.1, below Eq. (35)] The sentence saying the banded estimator is 'clearly indicating Σ̂(k) as the diagonal estimator' is imprecise; the diagonal estimator is the special case k=1, while for k>1 the estimator is banded.
- [Section 3.2, Eq. (42)] The citation 'John [ ? ]' is unresolved; the reference must be supplied before publication.
- [Section 4.2] The sentence 'While the density function is known to be globally convex, maximization can still lead us to a local maxima' is contradictory as written; if the relevant function is convex, interior local maxima are not a concern, and if the intended statement concerns the log-likelihood, concavity is the usual property.
- [Section 4.3.1, Eq. (58)] The displayed mass function is not readable as printed; the notation with products and subscripts should be rewritten with explicit joint probabilities.
- [Throughout] There are numerous typos (e.g., 'projeting', 'repsectively', 'distributiosn', 'likeliho0d', 'conjuction', 'moreknown than unkown') and at least one grammatical break ('R1 and R2 are the for notational convenience'); a careful copyedit is needed.
Circularity Check
No circularity: this survey chapter quotes and cites methods rather than deriving them; self-citations are descriptive, not load-bearing.
full rationale
No circularity found. This is a survey chapter, not an original derivation: the test statistics are presented with their stated assumptions and are cited to their original sources (e.g., Bai-Saranadasa in equation (6), Chen-Qin in equation (9), Srivastava-Du in equation (10), Park-Ayyala in equation (13)). The author's own methods appear in Sections 2.1, 2.5, and the data-analytic discussion, but those descriptions are expository and are not used as premises to prove a new result or to force a methodological choice. The T_APR passage explicitly reports that Cho et al. identified proof errors in Ayyala et al. and that corrections exist, which is an in-text limitation statement rather than a derivation step. The placeholder '(write motivation)' at the start of Section 2.5 and the incomplete 'John [ ? ]' citation in Section 3.2 are completeness and polish defects, not evidence that an input was relabeled as an output. No equation is self-definitional, no fitted parameter is renamed a prediction, and no uniqueness assertion is imported from a same-author citation. Therefore the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The asymptotic null distributions of the reviewed test statistics (e.g., T_BS, T_CQ, T_SD, T_PA) are correctly reproduced from the original papers.
- domain assumption The factor model X = mu + Gamma Z with i.i.d. components Z having finite fourth moments is a valid representation for the data in the mean-vector tests.
- domain assumption For the dependent-observations test, the data are realizations of M-dependent strictly stationary Gaussian processes with second-order stationarity.
- standard math The Johnson-Lindenstrauss lemma guarantees the existence of random projections preserving pairwise distances.
Cite this review
Pith. "Pith review of High dimensional statistical inference: theoretical development to data analytics." pith.science (2026). https://pith.science/paper/2BZYCHZA
@misc{pith2026190806600,
author = {Pith},
title = {Pith review of: High dimensional statistical inference: theoretical development to data analytics},
year = {2026},
howpublished = {\url{https://pith.science/paper/2BZYCHZA}},
note = {Machine review of arXiv:1908.06600}
}
read the original abstract
This article is due to appear in the Handbook of Statistics, Vol. 43, Elsevier/North-Holland, Amsterdam, edited by Arni S. R. Srinivasa Rao and C. R. Rao. In modern day analytics, there is ever growing need to develop statistical models to study high dimensional data. Between dimension reduction, asymptotics-driven methods and random projection based methods, there are several approaches developed so far. For high dimensional parametric models, estimation and hypothesis testing for mean and covariance matrices have been extensively studied. However, practical implementation of these methods are fairly limited and are primarily restricted to researchers involved in high dimensional inference. With several applied fields such as genomics, metagenomics and social networking, high dimensional inference is a key component of big data analytics. In this chapter, a comprehensive overview of high dimensional inference and its applications in data analytics is provided. Key theoretical developments and computational tools are presented, giving readers an in-depth understanding of challenges in big data analysis.
Figures
Reference graph
Works this paper leans on
-
[23]
S. Cho, J. Lim, D. N. Ayyala, J. Park, and A. Roy. Note on Mean Vector Testing for High-Dimensional Dependent Observations. arXiv e-prints, art. arXiv:1904.09344, Apr 2019
work page Pith review arXiv 1904
-
[6]
D. N. Ayyala, J. Park, and A. Roy. Mean vector testing for high-dimensional dependent observations. Journal of Multivariate Analysis , 153:136–155, 2017. ISSN 0047-259X. doi: 10.1016/j.jmva.2016.09.012. URL http://www.sciencedirect.com/science/article/pii/S0047259X16300999
-
[1]
D. Achlioptas. Database-friendly random projections: Johnson-Lindenstrauss with binary coins. Journal of Computer and System Sciences , 66(4):671–687, 2003. ISSN 00220000. doi: 10.1016/S0022-0000(03) 00025-4. 31
-
[2]
A. C. Aitken and H. T. Gonin. XI.On Fourfold Sampling with and without Replacement. Proceedings of the Royal Society of Edinburgh , 55:114–125, 1936. doi: 10.1017/S0370164600014413
- [3]
-
[4]
T. W. Anderson. An Introduction to Multivariate Statistical Analysis, 3rd edition . John Wiley and Sons, 2003
2003
-
[5]
D. N. Ayyala, D. E. Frankhouser, G. Marcucci, J.-O. Ganbat, P. Yan, R. Bundschuh, and S. Lin. Statistical methods for detecting differentially methylated regions based on MethylCap-seq data. Brief- ings in Bioinformatics , 17(6):926–937, 10 2015. ISSN 1467-5463. doi: 10.1093/bib/bbv089. URL https://doi.org/10.1093/bib/bbv089
-
[7]
Bai and H
Z. Bai and H. Saranadasa. Effect of High Dimension: By an Example of a Two Sample Problem. Statistica Sinica, 6:311–329, 1996. ISSN 10170405
1996
Show all 94 references
-
[8]
Z. Bai, D. Jiang, J. F. Yao, and S. Zheng. Corrections to LRT on large-dimensional covariance matrix by RMT. Annals of Statistics , 37(6 B):3822–3840, 2009. ISSN 00905364. doi: 10.1214/09-AOS694
2009 doi
-
[9]
Balakrishnan and L
S. Balakrishnan and L. Wasserman. Hypothesis testing for high-dimensional multinomials: A se- lective review1. Annals of Applied Statistics , 12(2):727–749, 2018. ISSN 19417330. doi: 10.1214/ 18-AOAS1155SF
2018
-
[10]
H. E. Barmi and R. L. Dykstra. Restricted multinomial maximum likelihood estimation based upon Fenchel duality. Statistics & Probability Letters , 21(2):121–130, 1994. ISSN 0167-7152. doi: 10.1016/0167-7152(94)90219-4. URL http://www.sciencedirect.com/science/article/pii/ 0167...
1994 doi
-
[11]
P. J. Bickel and E. Levina. Covariance regularization by thresholding. Annals of Statistics , 36(6): 2577–2604, 2008. ISSN 00905364. doi: 10.1214/08-AOS600
2008 doi
-
[12]
Bien and R
J. Bien and R. J. Tibshirani. Sparse estimation of a covariance matrix. Biometrika, 98(4):807–820,
-
[13]
Bingham and H
E. Bingham and H. Mannila. Random projection in dimensionality reduction: Applications to image and text data. In Proceedings of the Seventh ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’01, pages 245–250, New York, NY, USA, 2001. ACM. ISBN 1...
2001
-
[14]
Biswas and J
A. Biswas and J. S. Hwang. A new bivariate binomial distribution. Statistics and Probability Letters , 60(2):231–240, 2002. ISSN 01677152. doi: 10.1016/S0167-7152(02)00323-1
2002 doi
-
[15]
D. M. Blei, B. B. Edu, A. Y. Ng, A. S. Edu, M. I. Jordan, and J. B. Edu. technique...Latent Dirichlet Allocation. Journal of Machine Learning Research , 3:993–1022, 2003. ISSN 15324435. doi: 10.1162/ jmlr.2003.3.4-5.993. 32
2003
-
[16]
P. J. Brockwell and R. A. Davis. Time Series: Theory and Methods . Springer-Verlag, Berlin, Heidelberg,
-
[17]
T. Cai, W. Liu, and X. Luo. A constrained 𝓁1 minimization approach to sparse precision matrix estimation. Journal of the American Statistical Association , 106(494):594–607, 2011. doi: 10.1198/jasa. 2011.tm10155. URL https://doi.org/10.1198/jasa.2011.tm10155
2011 doi
-
[18]
Liu, and Y
Cai, TT., W. Liu, and Y. Xia. Two-sample test of high dimensional means under dependence. Journal of the Royal Statistical Society. Series B: Statistical Methodology , 76(2):349–372, 2014. ISSN 13697412. doi: 10.1111/rssb.12034
2014 doi
-
[19]
M. C. Cario and B. L. Nelson. Modeling and generating random vectors with arbitrary marginal dis- tributions and correlation matrix. Industrial Engineering, pages 1–19, 1997. URL http://citeseerx. ist.psu.edu/viewdoc/download?doi=10.1.1.48.281{&}rep=rep1{&}type=pdf
1997
-
[20]
S.-O. Chan, I. Diakonikolas, P. Valiant, and G. Valiant. Optimal Algorithms for Testing Closeness of Discrete Distributions. Proceedings of the Twenty-Fifth Annual ACM-SIAM Symposium on Discrete Algorithms, pages 1193–1203, 2013. doi: 10.1137/1.9781611973402.88
2013 doi
-
[21]
Chen and H
J. Chen and H. Li. Variable selection for sparse Dirichlet-multinomial regression with an application to microbiome data analysis. Annals of Applied Statistics , 7(1):418–442, 2013. ISSN 19326157. doi: 10.1214/12-AOAS592
2013 doi
-
[22]
S. X. Chen and Y. L. Qin. A two-sample test for high-dimensional data with applications to gene-set testing. Annals of Statistics , 38(2):808–835, 2010. ISSN 00905364. doi: 10.1214/09-AOS716
2010 doi
-
[24]
J. H. Chung and D. A. S. Fraser. Randomization tests for a multivariate two-sample problem.Journal of the American Statistical Association , 53(283):729–735, 1958. URL https://www.jstor.org/stable/ 2282050
1958
-
[25]
S. A. Crossley, M. Dascalu, and D. S. Mcnamara. How important is size? An Investigation of Corpus Size and Meaning in both Latent Semantic Analysis and Latent Dirichlet Allocation. Proceedings of the Thirtieth International Florida Artificial Intelligence Research Society Confe...
2017
-
[26]
B. Dai, S. Ding, and G. Wahba. Multivariate Bernoulli distribution. Bernoulli, 19(4):1465–1483,
-
[27]
Danaher, P
P. Danaher, P. Wang, and D. M. Witten. The joint graphical lasso for inverse covariance estimation across multiple classes. Journal of the Royal Statistical Society. Series B: Statistical Methodology , 76 (2):373–397, 2014. ISSN 13697412. doi: 10.1111/rssb.12033
2014 doi
-
[28]
P. J. Danaher. Parameter estimation for the dirichlet-multinomial distribution using supplementary beta-binomial data. Communications in Statistics - Theory and Methods , 17(6):1777–1788, 1988. doi: 10.1080/03610928808829713
1988 doi
-
[29]
M. J. Daniels and R. E. Kass. Shrinkage estimators for covariance matrices. Biometrics, 57(4):1173– 1184, 2001. ISSN 0006341X. doi: 10.1111/j.0006-341X.2001.01173.x. 33
2001 arXiv
-
[30]
a. P. Dempster. A High Dimensional Two Sample Significance Test. The Annals of Mathematical Statistics, 29(4):995–1010, 1958. ISSN 0003-4851. doi: 10.1214/aoms/1177706437
1958
-
[31]
J. Fan, F. Han, and H. Liu. Challenges of Big Data analysis. National Science Review , 1(2):293–314,
-
[32]
Fradkin and D
D. Fradkin and D. Madigan. Experiments with random projections for machine learning. In Proceedings of the Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , KDD ’03, pages 517–522, New York, NY, USA, 2003. ACM. ISBN 1-58113-737-0. doi: 10.1145/...
2003
-
[33]
Friedman, T
J. Friedman, T. Hastie, and R. Tibshirani. Sparse inverse covariance estimation with the graphical lasso. Biostatistics, 9(3):432–441, 2008. ISSN 14654644. doi: 10.1093/biostatistics/kxm045
2008 doi
-
[34]
Goodfellow, Y
I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning . MIT Press, 2016. http://www. deeplearningbook.org
2016
-
[35]
K. B. Gregory, R. J. Carroll, V. Baladandayuthapani, and S. N. Lahiri. A Two-Sample Test for Equality of Means in High Dimension. Journal of the American Statistical Association , 110(510):837–849, 2015. ISSN 1537274X. doi: 10.1080/01621459.2014.934826
2015
-
[36]
J. Guo, E. Levina, G. Michailidis, and J. Zhu. Joint estimation of multiple graphical models.Biometrika, 98(1):1–15, 2011. ISSN 00063444. doi: 10.1093/biomet/asq060
2011 doi
-
[37]
H. S. Hariharan and R. P. Velu. On estimating dirichlet parametersa comparison of initial values. Journal of Statistical Computation and Simulation , 48(1-2):47–58, 1993. ISSN 15635163. doi: 10.1080/ 00949659308811539
1993
-
[38]
Hoeffding
W. Hoeffding. Asymptotically Optimal Tests for Multinomial Distributions Author ( s ): Wassily Hoeffding Source : The Annals of Mathematical Statistics , Vol . 36 , No . 2 ( Apr ., 1965 ), pp . 369- 401 Published by : Institute of Mathematical Statistics Stable URL : ht. The Ann...
1965
-
[39]
M. D. Hoffman, D. M. Blei, and F. Bach. Online Learning for Latent Dirichlet Allocation. InAdvances in Neural Information Processing Systems 23, volume 1, pages 856–864, 2010. ISBN 9781450300551. URL http://papers.nips.cc/paper/3902-online-learning-for-latent-dirichlet-allocation.pdf
2010
-
[40]
Holmes, K
I. Holmes, K. Harris, and C. Quince. Dirichlet multinomial mixtures: Generative models for microbial metagenomics. PLoS ONE, 7(2), 2012. ISSN 19326203. doi: 10.1371/journal.pone.0030126
2012 doi
-
[41]
Hotelling
H. Hotelling. The generalization of student’s ratio. The Annals of Mathematical Statistics, 2(3):360–378, 08 1931. doi: 10.1214/aoms/1177732979. URL https://doi.org/10.1214/aoms/1177732979
1931
-
[42]
W. N. Hudson, H. G. Tucker, and J. A. Veeh. Limit theorems for the multivariate binomial distribution. Journal of Multivariate Analysis , 18(1):32–45, 1986. ISSN 10957243. doi: 10.1016/0047-259X(86) 90056-4
1986 doi
-
[43]
Inouye, E
D. Inouye, E. Yang, G. Allen, and P. Ravikumar. A Review of Multivariate Distributions for Count Data Derived from the Poisson Distribution. Wiley Interdisciplinary Review Computational Statistics , 9(3), 2017. doi: 10.1002/wics.1398.A. 34
2017 doi
-
[44]
N. P. Jewell and J. D. Kalbfleisch. Maximum likelihood estimation of ordered multinomial parameters. Biostatistics, 5(2):291–306, 2004. ISSN 14654644. doi: 10.1093/biostatistics/5.2.291
2004 doi
-
[45]
Jiang, T
D. Jiang, T. Jiang, and F. Yang. Likelihood ratio tests for covariance matrices of high-dimensional normal distributions. Journal of Statistical Planning and Inference , 142(8):2241–2256, 2012. ISSN 03783758. doi: 10.1016/j.jspi.2012.02.057. URL http://dx.doi.org/10.1016/j.jsp...
2012 doi
-
[46]
W. B. Johnson and J. Lindenstrauss. Extensions of Lipschitz mappings into a Hilbert space. Contem- porary Mathematics, 26:189–206, 1984. doi: 10.1090/conm/026/737400
1984 doi
-
[47]
Karlis and E
D. Karlis and E. Xekalaki. Mixed Poisson Distributions. International Statistical Review, 73(1):35–58,
-
[48]
A. S. Krishnamoorthy. Multivariate Binomial and Poisson Distributions. Sankhya B , 11(2):117–124,
-
[49]
A. Kudo. A multivariate analogue of the one-sided test. Biometrika, 50(3):403–418, 1963. URL https://www.jstor.org/stable/2333909
1963
-
[50]
Ledoit and M
O. Ledoit and M. Wolf. Some hypothesis tests for the covariance matrix when the dimension is large compared to the sample size. The Annals of Statistics , 30(4):1081–1102, 2002
2002
-
[51]
T. Leonard. A Bayesian Approach to Some Multinomial Estimation and Pretesting Problems. Journal of the American Statistical Association , 72(360):869–874, 1977
1977
-
[52]
B. Levin. A Representation for Multinomial Cumulative Distribution Functions. The Annals of Statis- tics, 9(5):1123–1126, 1981. URL https://www.jstor.org/stable/2240628
1981
-
[53]
Li and S
J. Li and S. X. Chen. Two sample tests for high-dimensional covariance matrices. Annals of Statistics , 40(2):908–940, 2012. ISSN 00905364. doi: 10.1214/12-AOS993
2012 doi
-
[54]
P. Li, T. J. Hastie, and K. W. Church. Very sparse random projections. In Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining - KDD ’06 , pages 287–296,
-
[55]
M. E. Lopes, L. J. Jacob, and M. J. Wainwright. A More Powerful Two-Sample Test in High Dimensions using Random Projection. 2015
2015
-
[56]
P. J. McMurdie and S. Holmes. Waste not, want not: Why rarefying microbiome data is inadmissible. PLOS Computational Biology , 10(4):1–12, 04 2014. doi: 10.1371/journal.pcbi.1003531. URL https: //doi.org/10.1371/journal.pcbi.1003531
2014 doi
-
[57]
K. S. Miller. On the Inverse of the Sum of Matrices. Mathematics Magazine, 54(2):67–72, 1981. URL https://www.jstor.org/stable/2690437
1981
-
[58]
Mimno, M
D. Mimno, M. D. Hoffman, and D. M. Blei. Sparse Stochastic Inference for Latent Dirichlet allocation. In ICML’12 Proceedings of the 29th International Coference on International Conference on Machine Learning, pages 1515–1522, 2012. URL http://arxiv.org/abs/1206.6425
2012 arXiv
-
[59]
C. Morris. Central Limit Theorems for Multinomial Sums. The Annals of Statistics, 3(1):165–188, 1975. URL https://www.jstor.org/stable/2958086
1975
-
[60]
R. J. Muirhead. Aspects of Multivariate Statistical Theory . John Wiley and Sons, 1982. 35
1982
-
[61]
H. Nagao. On some test criteria for covariance matrix. The Annals of Statistics , 1(4):700–709, 1973
1973
-
[62]
R. B. Nelson. An Introduction to Copulas . Springer Series in Statistics, second edition. ISBN 0-387- 28659-4
-
[64]
Park and D
J. Park and D. N. Ayyala. A test for the mean vector in large dimension and small sam- ples. Journal of Statistical Planning and Inference , 143(5):929–943, may 2013. ISSN 0378-3758. doi: 10.1016/J.JSPI.2012.11.001. URL https://www.sciencedirect.com/science/article/pii/ S03783...
2013 doi
-
[65]
Plunkett and J
A. Plunkett and J. Park. Two-sample test for sparse high-dimensional multinomial distributions. Test, 2018. ISSN 11330686. doi: 10.1007/s11749-018-0600-8. URL https://doi.org/10.1007/ s11749-018-0600-8
2018 doi
-
[66]
C. R. Rao. Advanced Statistical Methods in Data Science . Wiley, 1952. ISBN 02-850820-3
1952
-
[67]
C. R. Rao. Maximum likelihood estimation for the multinomial distribution. Sankhy: The Indian Journal of Statistics (1933-1960) , 18(1/2):139–148, 1957. ISSN 00364452. URL http://www.jstor. org/stable/25048341
1933
-
[68]
G. Ronning. Maximum likelihood estimation of dirichlet distributions. Journal of Statistical Computa- tion and Simulation , 32(4):215–221, 1989. ISSN 15635163. doi: 10.1080/00949658908811178
1989 doi
-
[69]
S., Srivastava and M
M. S., Srivastava and M. Du. A test for the mean vector with fewer observations than the dimension. Journal of Multivariate Analysis , 99(3):386–402, 2008. ISSN 0047-259X. doi: 10.1016/j.jmva.2006.11
2008 doi
-
[70]
J. R. Schott. A test for the equality of covariance matrices when the dimension is large relative to the sample sizes. Computational Statistics and Data Analysis , 51(12):6535–6542, 2007. ISSN 01679473. doi: 10.1016/j.csda.2007.03.004
2007 doi
-
[71]
Shin and R
K. Shin and R. Pasupathy. An algorithm for fast generation of bivariate poisson random vectors. INFORMS Journal on Computing , 22(1):81–92, 2010. ISSN 10919856. doi: 10.1287/ijoc.1090.0332
2010
-
[72]
M. Sklar. Fast MLE Computation for the Dirichlet Multinomial. 2014. URL http://arxiv.org/abs/ 1405.0099
2014 arXiv
-
[73]
M. S. Srivastava. A test for the mean vector with fewer observations than the dimension under non- normality. Journal of Multivariate Analysis , 100(3):518–532, 2009. ISSN 0047259X. doi: 10.1016/j. jmva.2008.06.006
2009 doi
-
[74]
M. S. Srivastava. Some Tests Concerning the Covariance Matrix in High Dimensional Data. Journal of the Japan Statistical Society , 35(2):251–272, 2013. ISSN 1882-2754. doi: 10.14490/jjss.35.251
2013 doi
-
[75]
M. S. Srivastava and H. Yanagihara. Testing the equality of several covariance matrices with fewer observations than the dimension. Journal of Multivariate Analysis , 101(6):1319–1329, 2010. ISSN 0047259X. doi: 10.1016/j.jmva.2009.12.010. URL http://dx.doi.org/10.1016/j.jmva.2...
2010 doi
-
[76]
M. S. Srivastava, S. Katayama, and Y. Kano. A two sample test in high dimensional data. Journal of Multivariate Analysis , 114(1):349–358, 2013. ISSN 10957243. doi: 10.1016/j.jmva.2012.08.014. URL http://dx.doi.org/10.1016/j.jmva.2012.08.014
2013 doi
-
[77]
URL http://www.sciencedirect.com/science/article/pii/S0047259X06001990
-
[78]
Srivastava, P
R. Srivastava, P. Li, and D. Ruppert. RAPTT: An Exact Two-Sample Test in High Dimensions Using Random Projections. Journal of Computational and Graphical Statistic , 25(3):954–970, 2016. doi: 10.1080/10618600.2015.1062771
2016
-
[79]
J. M. Stern and S. Zacks. Testing the independence of Poisson variates under the Holgate bivariate distribution: The power of a new evidence test. Statistics and Probability Letters , 60(3):313–320, 2002. ISSN 01677152. doi: 10.1016/S0167-7152(02)00314-0
2002 doi
-
[80]
Z. Sun, T. Wang, K. Deng, X. F. Wang, R. Lafyatis, Y. Ding, M. Hu, and W. Chen. DIMM-SC: A Dirichlet mixture model for clustering droplet-based single cell transcriptomic data. Bioinformatics, 34 (1):139–146, 2018. ISSN 14602059. doi: 10.1093/bioinformatics/btx490
2018 doi
-
[81]
J. L. Teugels. Some representations of the multivariate Bernoulli and binomial distributions. Journal of Multivariate Analysis , 32(2):256–268, 1990. ISSN 10957243. doi: 10.1016/0047-259X(90)90084-U
1990 doi
-
[82]
Tibshirani
R. Tibshirani. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical So- ciety. Series B: Statistical Methodology , 58(1):267–288, 1996. URL https://www.jstor.org/stable/ 2346178
1996
-
[83]
Z. Wang, M. Gerstein, and M. Snyder. Rna-seq: a revolutionary tool for transcriptomics. Nature Review Genetics, 10(1):57–63, 2009. doi: 10.1038/nrg2484
2009 doi
-
[84]
Y. Wu, M. G. Genton, and L. A. Stefanski. A multivariate two-sample mean test for small sample size and missing data. Biometrics, 62(3):877–885, 2006. ISSN 0006341X. doi: 10.1111/j.1541-0420.2006. 00533.x
2006
-
[85]
M. S. Srivastava, H. Yanagihara, and T. Kubokawa. Tests for covariance matrices in high dimension with less sample size. Journal of Multivariate Analysis , 130:289–309, 2014. ISSN 10957243. doi: 10.1016/j.jmva.2014.06.003. URL http://dx.doi.org/10.1016/j.jmva.2014.06.003
2014 doi
-
[86]
Zhong, S
P.-S. Zhong, S. X. Chen, and M. Xu. Tests alternative to higher criticism for high-dimensional means under sparsity and column-wise dependence. Ann. Statist. , 41(6):2820–2851, 12 2013. doi: 10.1214/ 13-AOS1168. URL https://doi.org/10.1214/13-AOS1168
2013 doi
-
[87]
R. S. Zoh, A. Sarkar, R. J. Carroll, and B. K. Mallick. A Powerful Bayesian Test for Equality of Means in High Dimensions. Journal of the American Statistical Association , 113(524):1733–1741, 2018. ISSN 1537274X. doi: 10.1080/01621459.2017.1371024. URL https://doi.org/10.1080...
2018
-
[93]
Zelterman
D. Zelterman. Goodness-of-Fit Tests for Large Sparse Distributions Multinomial. Journal of the Amer- ican Statistical Association, 82(398):624–629, 2013. URL https://www.jstor.org/stable/2289474
2013
-
[1951]
URL https://www.jstor.org/stable/25048072
-
[2006]
doi: 10.1145/1150402.1150436
ISBN 1595933395. doi: 10.1145/1150402.1150436
-
[2010]
doi: 10.1111/j.1751-5823.2005.tb00250.x
2005
- [2011]
-
[2013]
doi: 10.3150/12-BEJSP10
ISSN 1350-7265. doi: 10.3150/12-BEJSP10. URL http://projecteuclid.org/euclid.bj/ 1377612861
- [2014]
-
[2018]
URL http://arxiv.org/abs/1807.00930
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.