REVIEW 5 minor 300 references
Partial Differential Equation Barriers to Identifiability in Infinite Mixture Models
T0 review · 0 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A kernel annihilated by a nontrivial parameter PDE makes infinite mixtures non-identifiable, and many standard families are of this kind.
desk verdict A systematic and mostly correct treatment of PDE barriers to identifiability in infinite mixtures, worth a serious referee despite two minor gaps and a factor-2 typo. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the parameter differential operator $L_\theta$, a differential or difference-differential operator acting on the parameter $\theta$ such that $L_\theta f(x|\theta)=0$ for almost every $x$; the paper calls this a parameter PDE. For a non-trivial such operator, the formal adjoint $L^*_\theta$ applied to a compactly supported test function produces a signed density perturbation $h = L^*_\theta u$ that integrates to zero and is orthogonal to the kernel in the sense that $\int f(x|\theta)h(\theta)\,d\theta = 0$. Adding $\varepsilon h$ to the true mixing density preserves non-negativity for small $\varepsilon$ and leaves the mixture density unchanged, so the operator's null space is directly converted into a continuum of indistinguishable mixing measures. The generalization to shifts uses the same adjoint mechanism on translates of a ball.
What would settle it
Exhibit any kernel annihilated by a non-trivial parameter PDE for which the mixture operator is injective on the class of mixing measures with densities bounded below on an open ball; the theorem says none such exists, so such an example would refute it. A concrete numerical check would be to construct the perturbation $g = g^* + \varepsilon L^*_\theta u$ for the Gaussian heat-equation kernel and confirm that $p_G = p_{G^*}$ holds to floating precision, with failure indicating a gap in the adjoint argument.
Extended reading notes
Core claim
The central discovery is that non-identifiability in infinite mixture models is generated by parameter PDEs. Theorem 1 states that if the kernel f satisfies $L_\theta f(x|\theta)=0$ for a non-trivial differential operator on an open ball $B$, and the true mixing measure has a density bounded away from zero on $B$, then there is another probability measure $G \neq G^*$ with $p_G = p_{G^*}$ almost everywhere. Theorem 2 extends this to difference-differential operators with shifts, where the kernel is evaluated at shifted parameter values. The construction is explicit: integrate by parts with a compactly supported test function to build a mean-zero signed perturbation, then add a small multiple of it to the true density. Consequences are drawn for estimation: along with a minimax lower bound showing that under non-identifiability no estimator can recover the mixing measure in Wasserstein distance, even over discrete mixing measures.
Load-bearing premise
The load-bearing premise is that the true mixing measure has a density with respect to Lebesgue measure that is bounded away from zero on some open ball, and for shift operators that the shifted parameter values also lie in the parameter space; if the true measure is discrete, or its density can vanish on every open ball, the negative results do not directly apply.
Editorial extensions
If this is right
- Infinite location-scale Gaussian mixtures are non-identifiable: the heat equation $\partial_\nu f = \tfrac12 \partial_\mu^2 f$ is a parameter PDE, so distinct mixing measures over $(\mu,\nu)$ yield the same density.
- Any exponential family whose parameter dimension exceeds its sufficient-statistic dimension satisfies a non-trivial parameter PDE, making over-parameterized exponential kernels a systematic source of non-identifiability.
- When identifiability fails, the paper's minimax result implies a positive constant lower bound on the expected Wasserstein error of any estimator, uniformly over discrete mixing measures, so unconstrained estimation is impossible in the worst case.
- Identifiability is preserved for generalized translation families, such as Gaussian with fixed variance, gamma with fixed shape, and Laplace with fixed scale, and for exponential families whose sufficient statistic has dimension at least that of the parameter and a determining range.
- A concrete super-exponential kernel with sufficient statistics $x$ and $e^{x^2}$ has no non-trivial annihilating parameter PDE and an injective mixture operator, showing that the barrier is not universal.
Reading between the lines
- Because the theorem's perturbation is local and generic, the non-identifiability it produces is not a knife-edge phenomenon: any true mixing measure with a density bounded below on a small open ball is surrounded by indistinguishable alternatives, so the negative results likely extend to regularized estimators whenever the prior puts mass on continuous mixing densities.
- The minimax argument in Proposition 3 suggests that discrete priors such as Dirichlet-process mixtures inherit the worst-case obstruction even though the construction in Theorems 1–2 does not directly apply to discrete mixing measures; testing whether direct non-identifiable pairs exist for discrete measures would sharpen the practical implications.
- The over-parameterization criterion can be read as a design guide: to keep infinite mixtures identifiable, restrict kernels to parameterizations with sufficient-statistic dimension at least as large as parameter dimension, or fix shared nuisance parameters such as a common scale, before attempting nonparametric estimation of the mixing measure.
- Whether the absence of a parameter PDE is sufficient for identifiability remains open; Proposition 6's super-exponential family provides a test case where the two coincide, suggesting that growth mismatch between sufficient statistics may be the relevant condition to explore.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies identifiability of mixing measures in infinite mixture models p_G(x)=∫ f(x|θ)G(dθ). Its main negative results (Theorems 1 and 2) show that if the kernel is annihilated by a nontrivial linear differential operator, or by a difference-differential operator with shifts, on an open parameter region, and if the true mixing measure has a density bounded below on a suitable ball, then another probability measure on the parameter space produces the same mixture density almost everywhere. The authors give sufficient conditions for such annihilators (Proposition 1 for exponential families with parameter dimension exceeding the sufficient-statistic dimension; Proposition 2 for polynomial score functions), verify them on seven standard kernel families, and derive a minimax lower bound for estimation over discrete mixing measures (Proposition 3). Complementing this, Propositions 4–6 describe three kernel classes for which the mixture operator is injective. Proofs are collected in the supplementary appendix.
Significance. The paper's contribution is a clean structural explanation of non-identifiability in several widely used infinite mixture models, together with a balanced set of positive identifiability results. The constructions in Theorems 1 and 2 are explicit and the example identities I checked (Gaussian heat equation, Gamma, Beta, negative binomial) are correct; Proposition 3 gives a useful worst-case transfer from non-identifiability to estimation failure over discrete mixtures. The main limitation, namely the density-lower-bound requirement on the true mixing measure in the negative results, is explicitly acknowledged and partially addressed. If the minor proof clarifications below are made, the paper should be publishable.
minor comments (5)
- [A.3 (Proof of Proposition 2)] The final step claims that because the free coefficient is set to 1, the operator is non-trivial, but after column relabeling this coefficient may correspond to the zero-order multi-index, which would not satisfy Definition 1. The claim is true, but the proof should justify it: if the free column is the zero-order column, the top-block equation forces c_A to be nonzero whenever a solution exists, since otherwise the constant-monomial row of Φ would not be annihilated; alternatively, the authors can relabel so that a positive-order column is free. Please add this argument.
- [A.5 (Proof of Proposition 3)] The displayed lower bound evaluates to 3c/16, not 3c/32; since 3c/16 > 3c/32 the stated bound remains valid, but the constant in the display and in the statement should be made consistent.
- [A.4 (Proof of Theorem 2)] In the final display, the total variation identity should read d_TV(G,G*) = (ε/2) ∫ |h(θ)| dθ; the factor ε is missing. This does not affect the conclusion G ≠ G*.
- [Definitions 4 and A.4] The number of shift vectors is denoted J in Definition 4 but K in the proof of Theorem 2; please align the notation.
- [Abstract, Section 1.1, Section 2, A.8] Several typos need correction: 'flips side' should be 'flip side'; 'which is constitutes' should be 'which is'; 'satisifes' should be 'satisfies'; 'eqiuivalently' should be 'equivalently'; and 'is call regular closed set' should be 'is called a regular closed set'.
Circularity Check
No significant circularity: the non-identifiability construction is derived from the kernel PDE and standard adjoint/Fourier arguments; no fitted input is relabeled as a prediction and no load-bearing self-citation.
full rationale
The paper's central claim (Theorem 1) is self-contained: given a non-trivial parameter PDE L_theta f = 0, it sets h = L*_theta u for a compactly supported test function, shows integral h = 0 and integral f h = 0 by adjointness, and perturbs g* by epsilon h with epsilon = m/(2M). This is a direct construction, not an equivalent reformulation of the conclusion. Propositions 1 and 2 prove existence of annihilating operators from rank deficiency q < d2; the coefficient functions are solved by Cramer's rule, and the examples verify the PDEs by direct differentiation and recurrences (heat equation, Student-t PDE, beta/gamma/negative-binomial/noncentral-chi-squared shift relations). No parameter is fitted to data and no 'prediction' is a renamed input. The identifiability-preserving propositions rely on external, independent results (Fourier deconvolution, Muntz-Szasz density, Cartwright-Levinson logarithmic-integral lemma from Koosis), not on the authors' prior theorems. The only self-citations (Bariletto et al. 2026a,b) are contextual references in the literature review and discussion; they are not used to establish Theorem 1, Proposition 3, or any other main result. Proposition 3 is explicitly conditional, converting non-identifiability into a worst-case lower bound rather than assuming it. The density-lower-bound restriction is acknowledged in the text and is a scope condition, not a circular step.
Assumptions & free parameters
assumptions (6)
- standard math Fubini-Tonelli and dominated convergence can be applied to interchange integrals over x and theta.
- domain assumption The kernel is C^r in the parameter on the relevant open set and the coefficients of the annihilating operator are C^{|alpha|}, so formal adjoints and integrations by parts apply.
- domain assumption For shift PDEs, the translated balls B+v_j lie inside the parameter space and are disjoint, enabling the h_k components to be considered separately.
- standard math Muntz-Szasz type density of monomials y^s in C(Omega) for Proposition 5.
- standard math Paley-Wiener theorem for compactly supported measures on R^2 and Cartwright-Levinson logarithmic integral lemma for Proposition 6.
- standard math Le Cam's two-point inequality for minimax lower bounds in Proposition 3.
Cite this review
Pith. "Pith review of Partial Differential Equation Barriers to Identifiability in Infinite Mixture Models." pith.science (2026). https://pith.science/paper/Z6QJ2OHA
@misc{pith2026260808597,
author = {Pith},
title = {Pith review of: Partial Differential Equation Barriers to Identifiability in Infinite Mixture Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z6QJ2OHA}},
note = {Machine review of arXiv:2608.08597}
}
read the original abstract
We study identifiability of mixing measures in infinite mixture models. We show that, in many common cases, lack of identifiability can be characterized in terms of certain differential structures of the kernel family with respect to its parameters. In our main results, we prove that when the kernel is annihilated by a non-trivial differential or difference-differential operator over the parameter space, there exist infinitely many distinct mixing measures yielding the same mixture density. We give verifiable conditions for such operators to exist, covering many common cases, including the location-scale Gaussian, location-scale Student-t, Gamma, Beta, Dirichlet, negative binomial and non-central Chi-squared families. Furthermore, our conditions apply to any exponential family whose parameter dimension exceeds the dimension of its sufficient statistic and, more generally, to kernels with polynomially-growing score functions. We complement our results with a minimax lower bound on the estimation error for the mixing measure in the Wasserstein distance under non-identifiability. On the flips side, we describe three classes of kernels for which identifiability is preserved and nonparametric statistical inference remains possible.
Reference graph
Works this paper leans on
-
[1]
Minimax Optimal Rate for Parameter Estimation in Multivariate Deviated Models , volume =
Do, Dat and Nguyen, Huy and Nguyen, Khai and Ho, Nhat , booktitle =. Minimax Optimal Rate for Parameter Estimation in Multivariate Deviated Models , volume =
-
[2]
Attention is All you Need , volume =
Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, ukasz and Polosukhin, Illia , booktitle =. Attention is All you Need , volume =
-
[3]
On the Representation Collapse of Sparse Mixture of Experts , booktitle =
Chi, Zewen and Dong, Li and Huang, Shaohan and Dai, Damai and Ma, Shuming and Patra, Barun and Singhal, Saksham and Bajaj, Payal and Song, Xia and Mao, Xian-Ling and Huang, Heyan and Wei, Furu , editor =. On the Representation Collapse of Sparse Mixture of Experts , booktitle =
-
[4]
Sparse Mixers: Combining MoE and Mixing to build a more efficient
Lee-Thorp, James and Ainslie, Joshua , month = dec, year =. Sparse Mixers: Combining MoE and Mixing to build a more efficient. Findings of the
-
[5]
Proceedings of the ICML
On Least Square Estimation in Softmax Gating Mixture of Experts , author=. Proceedings of the ICML
-
[6]
Proceedings of the ICML
Nguyen, Huy and Akbarian, Pedram and Ho, Nhat , booktitle ="Proceedings of the ICML", year=. Is temperature sample efficient for softmax
-
[7]
Advances in Neural Information Processing Systems
FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion , author=. Advances in Neural Information Processing Systems
-
[8]
arXiv preprint arXiv:2402.02526 , year=
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition , author=. arXiv preprint arXiv:2402.02526 , year=
Show all 300 references
-
[9]
arXiv preprint arXiv:2502.03029 , year=
On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation , author=. arXiv preprint arXiv:2502.03029 , year=
-
[10]
arXiv preprint arXiv:2502.00281 , year=
Sigmoid Self-Attention is Better than Softmax Self-Attention: A Mixture-of-Experts Perspective , author=. arXiv preprint arXiv:2502.00281 , year=
-
[11]
arXiv preprint arXiv:2502.03044 , year=
RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts , author=. arXiv preprint arXiv:2502.03044 , year=
-
[12]
arXiv preprint arXiv:2503.03213 , year=
Convergence Rates for Softmax Gating Mixture of Experts , author=. arXiv preprint arXiv:2503.03213 , year=
-
[13]
arXiv preprint arXiv:2410.12258 , year=
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts , author=. arXiv preprint arXiv:2410.12258 , year=
-
[14]
arXiv preprint arXiv:2410.11222 , year=
Quadratic Gating Functions in Mixture of Experts: A Statistical Insight , author=. arXiv preprint arXiv:2410.11222 , year=
-
[15]
Huy Nguyen and Pedram Akbarian and Trang Pham and Trang Nguyen and Shujian Zhang and Nhat Ho , title =
-
[16]
The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=
Mixture of Experts Meets Prompt-Based Continual Learning , author=. The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=
-
[17]
The Thirteenth International Conference on Learning Representations , year=
Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts , author=. The Thirteenth International Conference on Learning Representations , year=
-
[18]
, journal=
Barron, A.R. , journal=. Universal approximation bounds for superpositions of a sigmoidal function , year=
-
[19]
Damai Dai and Chengqi Deng and Chenggang Zhao and R. X. Xu and Huazuo Gao and Deli Chen and Jiashi Li and Wangding Zeng and Xingkai Yu and Y. Wu and Zhenda Xie and Y. K. Li and Panpan Huang and Fuli Luo and Chong Ruan and Zhifang Sui and Wenfeng Liang , year=. Deep
-
[20]
2024 , journal =
Mixture of A Million Experts , author=. 2024 , journal =
2024
-
[21]
B. Yu. Assouad, F ano, and L e C am. Festschrift for Lucien Le Cam. 1997
1997
-
[22]
arXiv preprint arXiv:2310.14188 , author =
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts , copyright =. arXiv preprint arXiv:2310.14188 , author =
-
[23]
Nguyen, Huy and Nguyen, TrungTin and Nguyen, Khai and Ho, Nhat , title =
-
[24]
Journal of Statistical Distributions and Applications , author =
Approximations of conditional probability density functions in. Journal of Statistical Distributions and Applications , author =. 2021 , keywords =. doi:10.1186/s40488-021-00125-0 , abstract =
2021 doi
-
[25]
International Conference on Learning Representations , year=
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts , author=. International Conference on Learning Representations , year=
-
[26]
The Thirteenth International Conference on Learning Representations , year=
Theory on Mixture-of-Experts in Continual Learning , author=. The Thirteenth International Conference on Learning Representations , year=
-
[27]
and Doan, Thanh-Nam and Liu, Chenghao and Ramasamy, Savitha and Li, Xiaoli and HOI, Steven , month = dec, year =
Do, Truong Giang and Le, Huy Khiem and Nguyen, TrungTin and Pham, Quang and Nguyen, Binh T. and Doan, Thanh-Nam and Liu, Chenghao and Ramasamy, Savitha and Li, Xiaoli and HOI, Steven , month = dec, year =. Proceedings of the 2023
2023
-
[28]
Kwon, Jeongyeol and Qian, Wei and Caramanis, Constantine and Chen, Yudong and Davis, Damek , editor =. Global. Proceedings of the. 2019 , pages =
2019
-
[29]
Proceedings of the
Kwon, Jeongyeol and Caramanis, Constantine , editor =. Proceedings of the. 2020 , pages =
2020
-
[30]
Neural Networks , author =
Improved learning algorithms for mixture of experts in multiclass classification , volume =. Neural Networks , author =. 1999 , keywords =. doi:https://doi.org/10.1016/S0893-6080(99)00043-X , abstract =
1999 doi
-
[31]
Journal of the Royal Statistical Society
Iteratively. Journal of the Royal Statistical Society. Series B (Methodological) , author =. 1984 , note =
1984
-
[32]
IEEE Transactions on Pattern Analysis and Machine Intelligence , author =
Sparse multinomial logistic regression: fast algorithms and generalization bounds , volume =. IEEE Transactions on Pattern Analysis and Machine Intelligence , author =. 2005 , pages =. doi:10.1109/TPAMI.2005.127 , number =
2005 doi
-
[33]
A regression model with a hidden logistic process for feature extraction from time series , doi =
Chamroukhi, Faicel and Same, Allou and Govaert, Gerard and Aknin, Patrice , year =. A regression model with a hidden logistic process for feature extraction from time series , doi =. 2009
2009
-
[34]
Neural Networks , author =
Time series modeling by a regression approach based on a latent process , volume =. Neural Networks , author =. 2009 , note =
2009
-
[35]
Estimation and feature selection in mixtures of generalized linear experts models , journal =
Huynh, Bao Tuyen and Chamroukhi, Faicel , year =. Estimation and feature selection in mixtures of generalized linear experts models , journal =
-
[36]
The Annals of Applied Statistics , author =
A mixture of experts model for rank data with applications in election studies , volume =. The Annals of Applied Statistics , author =. 2008 , note =. doi:10.1214/08-AOAS178 , number =
2008 doi
-
[37]
, editor =
Jiang, Wenxin and Tanner, Martin A. , editor =. Hierarchical. Proceedings of the. 1999 , annote =
1999
-
[38]
Neural Computation , author =
On the. Neural Computation , author =. 1999 , pages =. doi:10.1162/089976699300016403 , abstract =
1999 doi
-
[39]
Information Retrieval , author =
Hierarchical. Information Retrieval , author =. 2002 , pages =. doi:10.1023/A:1012782908347 , abstract =
2002 doi
-
[40]
Journal of the American Statistical Association , author =
Bayesian. Journal of the American Statistical Association , author =. 1996 , pages =
1996
-
[41]
Neural Networks , author =
A. Neural Networks , author =. 1997 , keywords =
1997
-
[42]
Hierarchical mixtures of experts methodology applied to continuous speech recognition , volume =
Zhao, Ying and Schwartz, Richard and Sroka, Jason and Makhoul, John , year =. Hierarchical mixtures of experts methodology applied to continuous speech recognition , volume =. Advances in
-
[43]
Liang, Hanxue and Fan, Zhiwen and Sarkar, Rishov and Jiang, Ziyu and Chen, Tianlong and Zou, Kai and Cheng, Yu and Hao, Cong and Wang, Zhangyang , year =. M ^3
-
[44]
Advances in
Hazimeh, Hussein and Zhao, Zhe and Chowdhery, Aakanksha and Sathiamoorthy, Maheswaran and Chen, Yihua and Mazumder, Rahul and Hong, Lichan and Chi, Ed , editor =. Advances in. 2021 , pages =
2021
-
[45]
, year =
Ma, Jiaqi and Zhao, Zhe and Yi, Xinyang and Chen, Jilin and Hong, Lichan and Chi, Ed H. , year =. Modeling. Proceedings of the 24th. doi:10.1145/3219819.3220007 , abstract =
-
[46]
and Robinson, A.J
Waterhouse, S.R. and Robinson, A.J. , year =. Classification using hierarchical mixtures of experts , doi =. Proceedings of
-
[47]
and Dai, Andrew M
Zhou, Yanqi and Lei, Tao and Liu, Hanxiao and Du, Nan and Huang, Yanping and Zhao, Vincent Y. and Dai, Andrew M. and Chen, Zhifeng and Le, Quoc V. and Laudon, James , editor =. Mixture-of-. Advances in
-
[48]
and Salakhutdinov, R
Karakoulas, G. and Salakhutdinov, R. , year =. Semi-supervised mixture-of-experts classification , doi =. Fourth
-
[49]
The Annals of Statistics , author =
Optimal estimation of high-dimensional. The Annals of Statistics , author =. 2023 , note =. doi:10.1214/22-AOS2207 , number =
2023 doi
-
[50]
Regularized
Chamroukhi, Faicel and Huynh, Bao Tuyen , year =. Regularized. doi:10.1109/IJCNN.2018.8489670 , booktitle =
2018
-
[51]
Journal de la Société Française de Statistique , author =
Regularized. Journal de la Société Française de Statistique , author =. 2019 , pages =
2019
-
[52]
Neurocomputing , author =
A hidden process regression model for functional data description. Neurocomputing , author =. 2010 , pages =
2010
-
[53]
Journal of the Royal Statistical Society: Series B (Methodological) , author =
Maximum. Journal of the Royal Statistical Society: Series B (Methodological) , author =. 1977 , keywords =. doi:10.1111/j.2517-6161.1977.tb01600.x , abstract =
1977
-
[54]
Electronic Journal of Statistics , author =
Mixture of. Electronic Journal of Statistics , author =. 2014 , pages =
2014
-
[55]
Chen and J
H. Chen and J. Chen. Tests for homogeneity in normal mixtures in the presence of a structural parameter. Statistica Sinica. 2003
2003
-
[56]
Kasahara and K
H. Kasahara and K. Shimotsu. Testing the number of components in normal mixture regression models. Journal of the American Statistical Association. 2014
2014
-
[57]
van de Geer
S. van de Geer. Empirical Processes in M-estimation. 2000
2000
-
[58]
E. Polak. Optimization: Algorithms and Consistent Approximations. 1997
1997
-
[59]
R. M. Dudley. Probabilities and metrics: Convergence of laws on metric spaces, with a view to statistical testing. 1976
1976
-
[60]
Topics in Optimal Transportation
C\'edric Villani. Topics in Optimal Transportation. 2003
2003
-
[61]
Optimal transport: Old and New
C\'edric Villani. Optimal transport: Old and New. 2008
2008
-
[62]
S. T. Rachev and L. Ruschendorf. Mass transportation problems. Vol I: Theory, Vol II: Probability and its Applications. 1998
1998
-
[63]
S. T. Rachev. Probability metrics and the stability of stochastic systems. 1991
1991
-
[64]
C. Zhang. Fourier methods for estimating mixing densities and distributions. Annals of Statistics. 1990
1990
-
[65]
J. Fan. On the optimal rates of convergence for nonparametric deconvolution problems. Annals of Statistics. 1991
1991
-
[66]
R. J. Carroll and P. Hall. Optimal rates of convergence for deconvolving a density. Journal of American Statistical Association. 1988
1988
-
[67]
Convergence rates for
Ho, Nhat and Yang, Chiao-Yu and Jordan, Michael I , journal=. Convergence rates for
-
[68]
International Conference on Machine Learning , pages=
Refined convergence rates for maximum likelihood estimation under finite mixture models , author=. International Conference on Machine Learning , pages=. 2022 , organization=
2022
-
[69]
Bariletto, Nicola and Nguyen, Huy and Ho, Nhat and Rinaldo, Alessandro , journal=
-
[70]
Dedecker and B
J. Dedecker and B. Michel. Minimax rates of convergence for W asserstein deconvolution with supersmooth errors in any dimension. Arxiv manuscript. 2013
2013
-
[71]
Bickel and D
P. Bickel and D. Freedman. Some asymptotic theory for the bootstrap. Annals of Statistics. 1981
1981
-
[72]
del Barrio and J
E. del Barrio and J. Cuesta-Albertos and C. Matr\'an and J. Rodr\'iguez-Rodr\'iguez. Tests of goodness of fit based on the L_2 -Wasserstein distance. Annals of Statistics. 1999
1999
-
[73]
C. Mallows. A note on asymptotic joint normality. Annals of Mathematical Statistics. 1972
1972
-
[74]
Dobrushin
R. Dobrushin. Describing a system of random variables by conditional distributions. Theory Probab. Appl. 1970
1970
-
[75]
A. L. Gibbs and F. E. Su. On choosing and bounding probability metrics. International Statistical Review. 2002
2002
-
[76]
Ferraty and P
F. Ferraty and P. Vieu. Nonparametric Functional Data Analysis: T heory and P ractice. 2006
2006
-
[77]
J. O. Ramsay and B. W. Silverman. Applied functional data analysis: M ethods and case studies. 2002
2002
-
[78]
J. O. Ramsay and B.W. Silverman. Functional Data Analysis. 2006
2006
-
[79]
Abraham and P.A
C. Abraham and P.A. Cornillon and E. Matzner-Lober and N. Molinari. Unsupervised curve clustering using B-splines. Scand. J. Statist. 2003
2003
-
[80]
Biau and L
G. Biau and L. Devroye and G. Lugosi. On the Performance of Clustering in Hilbert Spaces. IEEE Trans. Inform. Theory. 2008
2008
-
[81]
Chiou and P.-L
J.-M. Chiou and P.-L. Li. Functional clustering and identifying substructures of longitudinal data. J. Roy. Statist. Soc. Ser. B. 2007
2007
-
[82]
J. A. Cuesta-Albertos and R. Fraiman. Impartial trimmed k-means for functional data. Comput. Statist. Data Anal. 2007
2007
-
[83]
Dabo-Niang and F
S. Dabo-Niang and F. Ferraty and P. Vieu. Mode estimation for functional random variable and its application for curves classication. Far East J. Theor. Stat. 2006
2006
-
[84]
Fraiman and A
R. Fraiman and A. Justel and M. Svarc. Selection of variables for cluster analysis and classication rules. J. Amer. Stat. Assoc. 2008
2008
-
[85]
Fraiman and G
R. Fraiman and G. Muniz. Trimmed means for functional data. Test. 2001
2001
-
[86]
G. M. James and C.A. Sugar. Clustering for sparsely sampled functional data. J. Amer. Stat. Assoc. 2003
2003
-
[87]
Ma and W
P. Ma and W. Zhong. Penalized clustering of large-scale functional data with multiple covariates. J. Amer. Statist. Assoc. 2008
2008
-
[88]
Tokushige and H
S. Tokushige and H. Yadohisa and K. Inada. Crisp and fuzzy k-means clustering algorithms for multivariate functional data. Comput. Statist. 2007
2007
-
[89]
Ishwaran and M
H. Ishwaran and M. Zarepour. Dirichlet prior sieves in finite normal mixtures. Statistica Sinica. 2002
2002
-
[90]
H. Teicher. Identifiability of finite mixtures. Ann. Math. Statist. 1963
1963
-
[91]
S. Ghosal. Dirichlet process, related priors and posterior asymptotics. Manuscript. 2007
2007
-
[92]
Ghosal and J
S. Ghosal and J. K. Ghosh and R. V. Ramamoorthi. Posterior consistency of Dirichlet mixtures in density estimation. Annals of Statistics. 1999
1999
-
[93]
Barron and M
A. Barron and M. Schervish and L. Wasserman. The consistency of posterior distributions in nonparametric problems. Ann. Statist. 1999
1999
-
[94]
Wu and S
Y. Wu and S. Ghosal. L1-Consistency of Dirichlet mixtures in multivariate Bayesian density estimation. 2009
2009
-
[95]
M.L. Stein. Interpolation of Spatial Data. 1999
1999
-
[96]
Winn and A
J. Winn and A. Criminisi and T. Minka. Object Categorization by Learned Universal Visual Dictionary. Proc. IEEE Intl. Conf. on Computer Vision (ICCV). 2005
2005
-
[97]
Brumback and J
B.A. Brumback and J. Rice. Smoothing spline models for the analysis of nested and crossed samples of curves. J. Amer. Statist. Assoc. 1998
1998
-
[98]
Sudderth and A
E. Sudderth and A. Torralba and W. Freeman and A. Willsky. Describing Visual Scenes Using Transformed Objects and Parts. International Journal of Computer Vision. 2008
2008
-
[99]
Pritchard and M
J. Pritchard and M. Stephens and P. Donnelly. Inference of populaton structure using multilocus genotype data. Genetics. 2000
2000
-
[100]
DeIorio and P
M. DeIorio and P. Muller and G.L. Rosner and S.N. MacEachern. An ANOVA model for dependent random measures. J. Amer. Statist. Assoc. 2004
2004
-
[101]
Duan and M
J. Duan and M. Guindani and A. Gelfand. Generalized spatial D irichlet processes. Biometrika. 2007
2007
-
[102]
Dunson and J.-H
D.B. Dunson and J.-H. Park. Kernel stick-breaking processes. Biometrika. 2008
2008
-
[103]
D.B. Dunson. Kernel local partition processes for functional data. 2008
2008
-
[104]
D.B. Dunson. Nonparametric Bayes local partition models for random effects. Biometrika. 2008
2008
-
[105]
R. F. MacLehose and D.B. Dunson. Nonparametric Bayes kernel-based priors for functional data analysis. Statistica Sinica. 2008
2008
-
[106]
Pillai and F
N. Pillai and F. Liang and S. Mukerjee and R. Wolpert and Q. Wu. Characterizing the function space for Bayesian kernel models. 2006
2006
-
[107]
An and C
Q. An and C. Wang and I. Shterev and E. Wang and L. Carin and D. Dunson. Hierachicial kernel stick-breaking process for multi-task image analysis. Proc. ICML. 2008
2008
-
[108]
Ferguson
T.S. Ferguson. A B ayesian analysis of some nonparametric problems. Ann. Statist. 1973
1973
-
[109]
Gelfand and A
A.E. Gelfand and A. Kottas and S.N. MacEachern. Bayesian nonparametric spatial modeling with D irichlet process mixing. J. Amer. Statist. Assoc. 2005
2005
-
[110]
Griffin and M.F
J.E. Griffin and M.F. Steel. Bayesian nonparametric spatial modeling with D irichlet process mixing. 2005
2005
-
[111]
Griffin and M.F
J.E. Griffin and M.F. Steel. Order-based dependent D irichlet processes. J. Amer. Statist. Assoc. 2006
2006
-
[112]
Ishwaran and L.F
H. Ishwaran and L.F. James. Gibbs sampling methods for stick-breaking priors. J. Amer. Statist. Assoc. 2001
2001
-
[113]
MacEachern
S.N. MacEachern. Dependent D irichlet processes. 2000
2000
-
[114]
Petrone and M
S. Petrone and M. Guidani and A.E. Gelfand. Hybrid D irichlet processes for functional data. J. Royal Stat. Soc. Series B. 2009
2009
-
[115]
Rodriguez and D
A. Rodriguez and D. Dunson and A.E. Gelfand. The nested D irichlet process. J. Amer. Statist. Assoc. to appear
-
[116]
Sethuraman
J. Sethuraman. A constructive definition of D irichlet priors. Statistica Sinica. 1994
1994
-
[117]
Teh and M.I
Y.W. Teh and M.I. Jordan and M.J. Beal and D.M. Blei. Hierarchical D irichlet processes. J. Amer. Statist. Assoc. 2006
2006
-
[118]
Blei and A.Y
D.M. Blei and A.Y. Ng and M.I. Jordan. Latent D irichlet allocation. J. Mach. Learn. Res. 2003
2003
-
[119]
Wang and E
X. Wang and E. Grimson. Spatial latent D irichlet allocation. NIPS 20. 2008
2008
-
[120]
Figueiredo and D.S
M.A. Figueiredo and D.S. Cheng and V. Murino. Clustering under prior knowledge with application to image segmentation. NIPS 19. 2007
2007
-
[121]
Fernandez and P
C. Fernandez and P. Green. Modelling spatially correlated data via mixtures: A B ayesian approach. J. Roy. Statist. Soc, Series B. 2002
2002
-
[122]
Green and S
P. Green and S. Richardson. Hidden M arkov models and desease mapping. J. Amer. Statist. Assoc. 2001
2001
-
[123]
R. R. Phelps. Convex functions, monotone operators and differentiability. 1993
1993
-
[124]
M. J. Wainwright and M. I. Jordan. Graphical models, exponential families, and variational inference
-
[125]
Bertsekas
D.P. Bertsekas. Nonlinear Programming. 1995
1995
-
[126]
Harmonic analysis on semigroups
-
[127]
Aronszajn
N. Aronszajn. Theory of reproducing kernels. Transactions of the American Mathematical Society
-
[128]
D. G. Luenberger. Optimization by Vector Space Methods
-
[129]
Rockafellar
G. Rockafellar. Convex Analysis. 1970
1970
-
[130]
Hiriart-Urruty and C
J. Hiriart-Urruty and C. Lemar\'echal. Fundamentals of Convex Analysis. 2001
2001
-
[131]
Koltchinskii and D
V. Koltchinskii and D. Panchenko. Empirical margin distributions and bounding the generalization error of combined classifiers. Annals of Statistics
-
[132]
Bousquet and A
O. Bousquet and A. Elisseeff. Stability and generalization. Journal of Machine Learning Research
-
[133]
Freund and R
Y. Freund and R. Schapire. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences. 1997
1997
-
[134]
Bartlett and S
P. Bartlett and S. Mendelson. G aussian and R ademacher complexities: R isk bounds and structural results. Journal of Machine Learning Research
-
[135]
Predd and S
J. Predd and S. Kulkarni and H. V. Poor. Consistency in Models for Communication Constrained Distributed Learning. Proceedings of the COLT
-
[136]
Reproducing Kernel Hilbert Spaces : Applications in Statistical Signal Processing. 1982
1982
-
[137]
S. Saitoh. Theory of Reproducing Kernels and its Applications. 1988
1988
-
[138]
Blackwell
D. Blackwell. Comparison of experiments. Proceeding of 2nd Berkeley Symposium on Probability and Statistics. 1951
1951
-
[139]
Blackwell
D. Blackwell. Equivalent comparisons of experiments. Annals of Statistics. 1953
1953
-
[140]
Bradt and S
R. Bradt and S. Karlin. On the design and comparison of certain dichotomous experiments. Annals of Statistics. 1956
1956
-
[141]
Goel and M
P. Goel and M. DeGroot. Comparisons of experiments and information measures. Annals of Statistics. 1979
1979
-
[142]
Chernoff
H. Chernoff. A measure of asymptotic efficiency for tests of a hypothesis based on a sum of observations. Annals of Statistics. 1952
1952
-
[143]
P. Massart. Some applications of concentration inequalities to statistics. Annales de la Facult\'e des Sciences de Toulouse. 2000
2000
-
[144]
McDiarmid
C. McDiarmid. On the method of bounded differences. Surveys in Combinatorics. 1989
1989
-
[145]
A. W. van der Vaart and J. Wellner. Weak Convergence and Empirical Processes. 1996
1996
-
[146]
A. W. van der Vaart. Asymptotic Statistics. 1998
1998
-
[147]
Le \ Cam
L. Le \ Cam. Asymptotic methods in statistical decision theory. 1986
1986
-
[148]
van de Geer
S. van de Geer. Empirical processes in M-estimation. 2000
2000
-
[149]
T. Zhang. Statistical behavior and consistency of classification methods based on convex risk minimization. Annal of Statistics
-
[150]
T. Zhang. From -entropy to KL -entropy: Analysis of Minimum Complexity Density Estimation. Ann. Statist. 2006
2006
-
[151]
T. Zhang. Information Theoretical Upper and Lower Bounds for Statistical Estimation. IEEE Trans.Info. 2006
2006
-
[152]
Steinwart
I. Steinwart. Consistency of support vector machines and other regularized kernel machines. IEEE Trans. Info. Theory. 2005
2005
-
[153]
Friedman and T
J. Friedman and T. Hastie and R. Tibshirani. Additive logistic regression: A statistical view of boosting. Annals of Statistics. 2000
2000
-
[154]
Bartlett and M
P. Bartlett and M. I. Jordan and J. D. McAuliffe. Convexity, classification and risk bounds. Journal of the American Statistical Association. 2006
2006
-
[155]
L. Breiman. Arcing classifiers. Annals of Statistics. 1998
1998
-
[156]
W. Jiang. Process consistency for Adaboost. Annals of Statistics. 2004
2004
-
[157]
Lugosi and N
G. Lugosi and N. Vayatis. On the Bayes-risk consistency of regularized boosting methods. Annals of Statistics. 2004
2004
-
[158]
Mannor and R
S. Mannor and R. Meir and T. Zhang. Greedy algorithms for classification - consistency, convergence rates and adaptivity. Journal of Machine Learning Research. 2003
2003
-
[159]
Yang and A
Y. Yang and A. Barron. Information theoretic determination of minimax rates of convergence. Annals of Statistics. 1999
1999
-
[160]
Guo and S
D. Guo and S. Shamai and S. Verdú. Mutual Information and Minimum Mean-Square Error in Gaussian Channels. IEEE Trans. Information Theory. 2005
2005
-
[161]
Cover and J
T. Cover and J. Thomas. Elements of information theory. 1991
1991
-
[162]
Csisz\'ar
I. Csisz\'ar. Information-type measures of difference of probability distributions and indirect observation. Studia Sci. Math. Hungar
-
[163]
S. M. Ali and S. D. Silvey. A general class of coefficients of divergence of one distribution from another. J. Royal Stat. Soc. Series B
-
[164]
T. Kailath. RKHS approach to detection and estimation problems--- P art I : D eterministic signals in G aussian noise. IEEE T rans. I nfo. T heory
-
[165]
Kailath and H
T. Kailath and H. V. Poor. Detection of stochastic processes. IEEE T rans. I nfo. T heory
-
[166]
T. Kailath. The Divergence and B hattacharyya Distance Measures in Signal Selection. IEEE Trans. on Communication Technology
-
[167]
Longo and T
M. Longo and T. Lookabaugh and R. Gray. Quantization for Decentralized Hypothesis Testing under Communication Contraints. IEEE Trans. on Information Theory
-
[168]
H. V. Poor and J. B. Thomas. Applications of A li- S ilvey distance measures in the design of generalized quantizers for binary decision systems. IEEE Trans. on Communications
-
[169]
R. R. Tenney and Sandell, N. R. Jr. Detection with Distributed Sensors. IEEE Trans. Aero. Electron. Sys
-
[170]
H. L. van Trees. Detection, Estimation and Modulation Theory. 1990
1990
-
[171]
Z. Luo. Universal decentralized estimation in a bandwidth constrained sensor network. 2003
2003
-
[172]
Advances in Statistical Signal Processing
J. N. Tsitsiklis , BOOKTITLE = "Advances in Statistical Signal Processing", EDITORS = "H. V. Poor and J. B. Thomas", TITLE = "Decentralized Detection", publisher =
-
[173]
Tsitsiklis
J. Tsitsiklis. Extremal properties of likelihood-ratio quantizers. IEEE Trans. on Communication. 1993
1993
-
[174]
Advances in Statistical Signal Processing
S. A. Kassam , BOOKTITLE = "Advances in Statistical Signal Processing", EDITORS = "H. V. Poor and J. B. Thomas", TITLE = "Nonparametric signal detection", publisher =
-
[175]
H. V. Poor. An introduction to signal detection and estimation. 1994
1994
-
[176]
Viswanathan and A
R. Viswanathan and A. Ansari. Distributed detection of a signal in generalized G aussian noise. IEEE Trans. Acoust., Speech, and Signal Process. 1989
1989
-
[177]
F. Topsoe. Some inequalities for information divergence and related measures of discrimination. IEEE Transactions on Information Theory. 2000
2000
-
[178]
Nasipuri and S
A. Nasipuri and S. Tantaratana. Nonparametric distributed detection using W ilcoxon statistics. Signal Processing. 1997
1997
-
[179]
R. S. Blum and S. A. Kassam and H. V. Poor. Distributed detection with multiple sensors: P art II --- Advanced Topics. Proceedings of the IEEE
-
[180]
Han and P
J. Han and P. K. Varshney and V. C. Vannicola. Some results on distributed nonparametric detection. Proc. 29th Conf. on Decision and Control. 1990
1990
-
[181]
M. M. Al-Ibrahim and P. K. Varshney. Nonparametric sequential detection based on multisensor data. Proc. 23rd Annu. Conf. on Inform. Sci. and Syst. 1989
1989
-
[182]
E. K. Hussaini and A. A. M. Al-Bassiouni and Y. A. El-Far. Decentralized CFAR signal detection. Signal Processing. 1995
1995
-
[183]
V. V. Veeravalli and T. Basar and H. V. Poor. Decentralized sequential detection with a fusion center performing the sequential test. IEEE Trans. Info. Theory. 1993
1993
-
[184]
J. F. Chamberland and V. V. Veeravalli. Decentralized detection in sensor networks. IEEE Transactions on Signal Processing. 2003
2003
-
[185]
Nguyen and M
X. Nguyen and M. J. Wainwright and M. I. Jordan. Nonparametric decentralized detection using kernel methods. IEEE Transactions on Signal Processing. 2005
2005
-
[186]
Nguyen and M
X. Nguyen and M. J. Wainwright and M. I. Jordan. Divergence measures, surrogate loss functions and experiment design. Advances in Neural Information Processing Systems 11
-
[187]
Nguyen and M
X. Nguyen and M. J. Wainwright and M. I. Jordan. On divergences, surrogate loss functions and decentralized detection
-
[188]
Nguyen and M
X. Nguyen and M. J. Wainwright and M. I. Jordan. On surrogate loss functions and f -divergences. Annals of Statistics. 2009
2009
-
[189]
Nguyen and M
X. Nguyen and M. J. Wainwright and M. I. Jordan. Estimating divergence functionals and the likelihood ratio by convex risk minimization. IEEE Transactions on Information Theory. 2010
2010
-
[190]
Jaakkola and D
T. Jaakkola and D. Haussler. Exploiting generative models in discriminative classifiers. Advances in Neural Information Processing Systems 11
-
[191]
Taskar and C
B. Taskar and C. Guestrin and D. Koller. Max- M argin M arkov N etworks. NIPS 15
-
[192]
Dougherty and R
J. Dougherty and R. Kohavi and M. Sahami. Supervised and unsupervised discretization of continuous features. Proceedings of the ICML
-
[193]
Cortes and V
C. Cortes and V. Vapnik. Support-Vector Networks. Machine Learning. 1995
1995
-
[194]
Blake and C.J
C.L. Blake and C.J. Merz. UCI Repository of machine learning databases. 1998
1998
-
[195]
Tsuda and T
K. Tsuda and T. Kin and K. Asai. Marginalized Kernels for Biological Sequences. Bioinformatics
-
[196]
Learning with Kernels. 2002
2002
-
[197]
Cristianini and J
N. Cristianini and J. Shawe-Taylor. An Introduction to Support Vector Machines (and other kernel based learning methods). 2000
2000
-
[198]
Chong and S
C. Chong and S. P. Kumar. Sensor Networks: Evolution, Opportunities, and Challenges. Proceedings of the IEEE
-
[199]
Bodik and G
P. Bodik and G. Friedman and L. Biewald and H. Levine and G. Candea and K. Patel and G. Tolle and J. Hui and A. Fox and M. I. Jordan and D. Patterson. Combining Visualization and Statistical Analysis to Improve Operator Confidence and Efficiency for Failure Detection and Local...
2005
-
[200]
Support-vector networks
Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine Learning. 1995
1995
-
[201]
J. H. Chen. Optimal rate of convergence for finite mixture models. Annals of Statistics. 1995
1995
-
[202]
S. J. Yakowitz and J. D. Spragins. On the identifiability of finite mixtures. Annals of Statistics. 1968
1968
-
[203]
X. Nguyen. Convergence of latent mixing measures in finite and infinite mixture models. Annals of Statistics. 2013
2013
-
[204]
B. Lindsay. Mixture models: Theory, geometry and applications. 1995
1995
-
[205]
G. J. McLachlan and K. E. Basford. Mixture models: Inference and Applications to Clustering. Statistics: Textbooks and Monographs. 1988
1988
-
[206]
van de Geer
S. van de Geer. Rates of convergence for the maximum likelihood estimator in mixture models. Journal of Nonparametric Statistics. 1996
1996
-
[207]
Shen and W
X. Shen and W. H. Wong. Convergence rate of sieves estimates. Annals of Statistics. 1994
1994
-
[208]
Shen and L
X. Shen and L. Wasserman. Rates of convergence of posterior distributions. Annals of Statistics. 2001
2001
-
[209]
Ghosal and J
S. Ghosal and J. K. Ghosh and A. van der Vaart. Convergence rates of posterior distributions. Annals of Statistics. 2000
2000
-
[210]
S. G. Walker, A. Lijoi and I. Prunster. On rates of convergence for posterior distributions in infinite-dimensional models. Annals of Statistics. 2007
2007
-
[211]
Ghosal and A
S. Ghosal and A. van der Vaart. Posterior convergence rates of Dirichlet mixtures at smooth densities. Annals of Statistics. 2007
2007
-
[212]
Ghosal and J
S. Ghosal and J. K. Ghosh and R. V. Ramamoorthi. Posterior consistency of D irichlet mixtures in density estimation. The Annals of Statistics. 1999
1999
-
[213]
Ghosal and A
S. Ghosal and A. van der Vaart. Entropies and rates of convergence for maximum likelihood and B ayes estimation for mixtures of normal densities. The Annals of Statistics. 2001
2001
-
[214]
On consistency of nonparametric normal mixtures for
Lijoi, Antonio and Pr. On consistency of nonparametric normal mixtures for. Journal of the American Statistical Association , volume=. 2005 , publisher=
2005
-
[215]
C. R. Genovese. and L. Wasserman. Rates of convergence for the Gaussian mixture sieve. Annals of Statistics. 2000
2000
-
[216]
C. Villani. Topics in Optimal Transportation.Graduate Studies in Mathematics. 2003
2003
-
[217]
C. Villani. Optimal Transport: Old and New. Grundlehren der Mathematischen Wissenschaften [Fundamental Principles of Mathemtical Sciences]. 2009
2009
-
[218]
H. Teicher. Identifiability of Mixtures. Annals of Statistics. 1961
1961
-
[219]
H. Teicher. On the mixture of distributions. Annals of Statistics. 1960
1960
-
[220]
H. Teicher. Identifiability of finite mixtures. Annals of Statistics. 1963
1963
-
[221]
H. Teicher. Identifiability of mixtures of product measures. Annals of Statistics. 1967
1967
-
[222]
Osiewalski and M
J. Osiewalski and M. F. J. Steel. Robust Bayesian Inference in l_ q -spherical models. Biometrika. 1993
1993
-
[223]
J. T. Kent. Identifiability of finite mixtures for directional data. Annals of Statistics. 1983
1983
-
[224]
Y. S. Hsu and M. D. Fraser and J. J. Walker. Identifiability of finite mixtures of von Mises distributions. Annals of Statistics. 1981
1981
-
[225]
K. V. Mardia. Statistics of directional data. Journal of the Royal Statistical Society. Series B(Methodological). 1975
1975
-
[226]
S. M. Ali and S. D. Silvey. A general class of coefficients of divergence of one distribution from another. Journal of the Royal Statistical Society. Series B(Methodological). 1966
1966
-
[227]
Peel and G
D. Peel and G. J. McLachlan. Robust mixture modelling using the t distribution. Statistics and Computing. 2000
2000
-
[228]
Azzalini and A
A. Azzalini and A. D. Valle. The multivariate skew-normal distribution. Biometrika. 1996
1996
-
[229]
Azzalini and A
A. Azzalini and A. Capitanio. Statistical applications of the multivariate skew-normal distribution. Journal of the Royal Statistical Society, Series B(Methodological). 1999
1999
-
[230]
E. S. Allman and C. Matias and J. A. Rhodes. Identifiability of parameters in latent structure models with many observed variables. Annals of Statistics. 2009
2009
-
[231]
Hall and X
P. Hall and X. -H. Zhou. Nonparametric estimation of component distributions in a multivariate mixture. Annals of Statistics. 2003
2003
-
[232]
Hall and A
P. Hall and A. Neeman and R. Pakyari and R. Elmore. Nonparametric inference in multivariate mixtures. Biometrika. 2005
2005
-
[233]
Elmore and P
R. Elmore and P. Hall and A. Neeman. An application of classical invariant theory to identifiability in nonparametric mixtures. Ann. Inst. Fourier (Grenoble). 2005
2005
-
[234]
Le \ Cam
L. Le \ Cam. Convergence of estimates under dimensionality reductions. Annals of Statistics. 1973
1973
-
[235]
Buchberger , title=
B. Buchberger , title=
-
[236]
Sturmfels
B. Sturmfels. Solving system of polynomial equations. 2002
2002
-
[237]
Wiper and D
M. Wiper and D. R. Insua and F. Ruggeri. Mixtures of Gamma distributions with applications. Journal of Computational and Graphical Statistics. 2001
2001
-
[238]
S. X. Lee and G. J. McLachlan. On mixtures of skew normal and skew t -distributions. Advances in Data Analysis and Classification. 2013
2013
-
[239]
Ghosal and A
S. Ghosal and A. Roy. Predicting false discovery proportion under dependence. Journal of the American Statistical Association. 2011
2011
-
[240]
Zhang and A
T. Zhang and A. Weisel and M. S. Greco. Multivariate generalized Gaussian distribution: Convexity and graphical models. IEEE Transactions on Signal Processing. 2013
2013
-
[241]
Ho and X
N. Ho and X. Nguyen. On strong identifiability and convergence rates of parameter estimation in finite mixtures. Electronic Journal of Statistics. 2016
2016
-
[242]
Ho and X
N. Ho and X. Nguyen. Convergence rates of parameter estimation for some weakly identifiable finite mixtures. Annals of Statistics. 2016
2016
-
[243]
Kulis and M
B. Kulis and M. I. Jordan. Revisiting k-means: new algorithm via Bayesian nonparametrics. Proceedings of the 29^ th International Conference on Machine Learning. 2012
2012
-
[244]
Hallin and C
M. Hallin and C. Ley. Skew-symmetric distributions and Fisher information - a tale of two densities. Bernoulli. 2012
2012
-
[245]
Hallin and C
M. Hallin and C. Ley. Skew-symmetric distributions and Fisher information: the double sin of skew-normal. Bernoulli. 2014
2014
-
[246]
M. Chiogna. A note on the asymptotic distribution of the maximum likelihood estimator for the scalar skew-normal distribution. Statistical Methods and Applications. 2005
2005
-
[247]
Azzalini and A
A. Azzalini and A. Capitanio. Distributions generated by perturbation of symmetry with emphasis on a multivariate skew t-distribution. Journal of Royal Statistical Society: Series B (Statistical Methodology). 2003
2003
-
[248]
Azzalini
A. Azzalini. Further results on a class of distributions which includes the normal ones. Statistica (Bologna). 1986
1986
-
[249]
T. J. DiCiccio and A. C. Monti. Inferential aspects of the skew-exponential power distribution. Journal of the American Statistical Association. 2004
2004
-
[250]
Azzalini and A
A. Azzalini and A. Capitanio. Statistical applications of the multivariate skew normal distribution. Journal of Royal Statistical Society: Series B (Statistical Methodology). 1999
1999
-
[251]
T. I. Lin and J. C. Lee and S. Y. Yen. Finite mixture modelling using the skew normal distribution. Statistica Sinica. 2007
2007
-
[252]
Ley and D
C. Ley and D. Paindaveine. On the singularity of multivariate skew-symmetric models. Journal of Multivariate Analysis. 2010
2010
-
[253]
X. Nguyen. Borrowing strength in hierarchical Bayes: convergence of the Dirichlet base measure. Bernoulli. 2015+
2015
-
[254]
Azzalini and A
A. Azzalini and A. Capitanio. Distributions generated by perturbation of symmetry with emphasis on a multivariate skew t distribution. Journal of Royal Statistical Society: Series B (Statistical Methodology). 2003
2003
-
[255]
T. I. Lin. Maximum likelihood estimation for multivariate skew normal mixture models. Journal of Multivariate Analysis. 2009
2009
-
[256]
T. I. Lin. Robust mixture modeling using multivariate skew t-distributions. Statistics and Computing. 2010
2010
-
[257]
T. I. Lin and J. C. Lee and W. J. Hsieh. Robust mixture modelling using the skew t-distribution. Statistics and Computing. 2007
2007
-
[258]
Lee and G
S. Lee and G. J. McLachlan. Finite mixtures of multivariate skew t-distributions: some recent and new results. Statistics and Computing. 2014
2014
-
[259]
H. J. Ho and S. Pyne and T. I. Line. Maximum likelihood inference for mixtures of skew student- t-Normal distributions through practical EM-type algorithms. Statistics and Computing. 2012
2012
-
[260]
M. O. Prates and C. R. B. Cabral and V. H. Lachos. mixsmsn: fitting finite mixture of scale mixture of skew-normal distributions. Journal of Statistical Software. 2013
2013
-
[261]
S. W. Schnatter and S. Pyne. Bayesian inference for finite mixtures of univariate and multivariate skew-normal and skew-t distributions. Biostatistics. 2009
2009
-
[262]
R. B. Arellano-Valle and L. M. Castro and M. C. Genton and H. W. Gómez. Bayesian inference for shape mixtures of skewed distributions, with application to regression analysis. Bayesian Analysis. 2008
2008
-
[263]
R. B. Arellano-Valle and M. C. Genton and R. H. Loschi. Shape mixtures of multivariate skew-normal distributions. Journal of Multivariate Analysis. 2009
2009
-
[264]
C. B. Zeller and C. R. B. Cabral and V. H. Lachos. Robust mixture regression modeling based on scale mixtures of skew-normal distributions. TEST. 2015
2015
-
[265]
Azzallini
A. Azzallini. A class of distributions which includes the normal ones. Scadinavian Journal of Statistics. 1985
1985
-
[266]
R. B. Arellano-Valle and A. Azzallini. The centred parametrization for the multivariate skew-normal distribution. Journal of Multivariate Analysis. 2008
2008
-
[267]
M. G. Genton. Skew-elliptical distributions and their applications: a journey beyond normality. 2004
2004
-
[268]
Wang and J
J. Wang and J. Boyer and M. C. Genton. A skew-symmetric representation of multivariate distribution. Statistica Sinica. 2004
2004
-
[269]
Canale and B
A. Canale and B. Scarpa. Bayesian nonparametric location-scale-shape mixtures. TEST. 2015
2015
-
[270]
Ishwaran and L
H. Ishwaran and L. F. James and J. Sun. Bayesian model selection in finite mixtures by marginal density decompositions. Journal of the American Statistical Association. 2001
2001
-
[271]
Rousseau and K
J. Rousseau and K. Mengersen. Asymptotic behaviour of the posterior distribution in overfitted mixture models. Journal of the Royal Statistical Society: Series B (Statistical Methodology). 2011
2011
-
[272]
Petralia and V
F. Petralia and V. Rao and D. B. Dunson. Repulsive mixtures. Advances in Neural Information Processing Systems (NIPS). 2012
2012
-
[273]
Rotnitzky and D
A. Rotnitzky and D. R. Cox and M. Bottai and J. Robins. Likelihood-based inference with singular information matrix. Bernoulli. 2000
2000
-
[274]
Heinrich and J
P. Heinrich and J. Kahn. Optimal rates for finite mixture estimation. Under review. 2016+
2016
-
[275]
L. F. Lee and A. Chesher. Specification testing when score test statistics are identically zero. Journal of Econometrics. 1986
1986
-
[276]
J. Chen. Consistency of the MLE under mixture models. arXiv preprint arXiv:1607.01251. 2016
2016 arXiv
-
[277]
Chen and X
J. Chen and X. Tan and R. Zhang. Inference for normal mixtures in mean and variance. Statistica Sinica. 2008
2008
-
[278]
Cox and J
D. Cox and J. Little and D. O'Shea. Ideals, Varieties, and Algorithms: An Introduction to Computational Algebraic Geometry and Commutative Algebra. 2007
2007
-
[279]
B. Stumfel. Solving systems of polynomial equations. 2002
2002
-
[280]
Xiao and G
S. Xiao and G. Zeng. Determination of the limits for multivariate rational functions. Science China Mathematics. 2014
2014
-
[281]
N. M. Kiefer. A remark on the parameterization of a model for heterogeneity. 1982
1982
-
[282]
Ho and X
N. Ho and X. Nguyen. Singularity structures and impacts on parameter estimation in finite mixtures of distributions. 2016
2016
-
[283]
Toussile and E
W. Toussile and E. Gassiat. Variable selection in model-based clustering using multilocus genotype data. Advances in Data Analysis and Classification. 2009
2009
-
[284]
Gassiat and R
E. Gassiat and R. V. Handel. The local geometry of finite mixtures. Transaction of the American Mathematical Society. 2014
2014
-
[285]
Basu and R
S. Basu and R. Pollack and M. Roy. Algorithms in real algebraic geometry. 2006
2006
-
[286]
R. A. Jacobs and M. I. Jordan and S. J. Nowlan and G. E. Hinton. Adaptive mixtures of local experts. Neural Computation. 1991
1991
-
[287]
M. I. Jordan and L. Xu. Convergence results for the EM approach to mixtures of experts architectures. Neural Networks. 1995
1995
-
[288]
Heinrich and J
P. Heinrich and J. Kahn. Strong identifiability and optimal minimax rates for finite mixture estimation. Annals of Statistics. 2017+
2017
-
[289]
Balakrishnan and M
S. Balakrishnan and M. J. Wainwright and B. Yu , Journal =. Statistical guarantees for the
-
[290]
Dwivedi and N
R. Dwivedi and N. Ho and K. Khamaru and M. J. Wainwright and M. I. Jordan and B. Yu , Journal =. Singularity, misspecification, and the convergence rate of
-
[291]
Dwivedi and N
R. Dwivedi and N. Ho and K. Khamaru and M. J. Wainwright and M. I. Jordan and B. Yu , Date-Modified =. AISTATS , Title =
-
[292]
Optimal Estimation of
Wu, Yihong and Yang, Pengkun , year =. Optimal Estimation of. The Annals of Statistics , volume =
-
[293]
The Annals of Statistics , volume=
Strong identifiability and optimal minimax rates for finite mixture estimation , author=. The Annals of Statistics , volume=
-
[294]
Jordan , title =
Nhat Ho and Chiao-Yu Yang and Michael I. Jordan , title =. Journal of Machine Learning Research , year =
-
[295]
Mixtures of Experts Unlock Parameter Scaling for Deep
Johan Samir Obando Ceron and Ghada Sokar and Timon Willi and Clare Lyle and Jesse Farebrother and Jakob Nicolaus Foerster and Gintare Karolina Dziugaite and Doina Precup and Pablo Samuel Castro , booktitle=. Mixtures of Experts Unlock Parameter Scaling for Deep
-
[296]
Electronic Journal of Statistics , volume=
Strong identifiability and parameter learning in regression with heterogeneous response , author=. Electronic Journal of Statistics , volume=. 2025 , publisher=
2025
-
[297]
On posterior contraction of parameters and interpretability in
Guha, Aritra and Ho, Nhat and Nguyen, XuanLong , journal=. On posterior contraction of parameters and interpretability in. 2021 , publisher=
2021
-
[298]
Jiang and M
W. Jiang and M. A. Tanner. On the identifiability of mixtures-of-experts. Neural Networks. 1999
1999
-
[299]
M. I. Jordan and R. A. Jacobs. Hierarchical mixtures of experts and the EM algorithm. Neural Computation. 1994
1994
-
[300]
Boyd and L
S. Boyd and L. Vandenberghe. Convex Optimization. 2004
2004
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.