Pith. sign in

REVIEW 5 minor 300 references

Partial Differential Equation Barriers to Identifiability in Infinite Mixture Models

T0 review · 0 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A kernel annihilated by a nontrivial parameter PDE makes infinite mixtures non-identifiable, and many standard families are of this kind.

desk verdict A systematic and mostly correct treatment of PDE barriers to identifiability in infinite mixtures, worth a serious referee despite two minor gaps and a factor-2 typo. read the letter →

arxiv 2608.08597 v1 pith:Z6QJ2OHA submitted 2026-08-09 math.ST stat.TH

classification math.STstat.TH MSC 62G0562G2062E1035A30
keywords mixturemodelsidentifiabilitymixingmeasuresparameterPDEexponentialfamiliesWassersteindistanceminimaxestimationoverparameterization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to show that a single differential condition on the kernel explains when the mixing measure in an infinite mixture model cannot be recovered from the mixture density. The condition is the existence of a non-trivial differential or difference-differential operator that annihilates the kernel as a function of its parameters. Under a mild regularity assumption on the true mixing measure, such an operator yields infinitely many distinct mixing measures with the same mixture density, so identifiability fails. The paper proves that this mechanism is pervasive, covering location-scale Gaussian and Student-t, Gamma, Beta, Dirichlet, negative binomial and non-central chi-squared families, and that over-parameterization is a systematic route to it. It also identifies three kernel classes where the mechanism is absent and identifiability is preserved.

What carries the argument

The load-bearing object is the parameter differential operator $L_\theta$, a differential or difference-differential operator acting on the parameter $\theta$ such that $L_\theta f(x|\theta)=0$ for almost every $x$; the paper calls this a parameter PDE. For a non-trivial such operator, the formal adjoint $L^*_\theta$ applied to a compactly supported test function produces a signed density perturbation $h = L^*_\theta u$ that integrates to zero and is orthogonal to the kernel in the sense that $\int f(x|\theta)h(\theta)\,d\theta = 0$. Adding $\varepsilon h$ to the true mixing density preserves non-negativity for small $\varepsilon$ and leaves the mixture density unchanged, so the operator's null space is directly converted into a continuum of indistinguishable mixing measures. The generalization to shifts uses the same adjoint mechanism on translates of a ball.

What would settle it

Exhibit any kernel annihilated by a non-trivial parameter PDE for which the mixture operator is injective on the class of mixing measures with densities bounded below on an open ball; the theorem says none such exists, so such an example would refute it. A concrete numerical check would be to construct the perturbation $g = g^* + \varepsilon L^*_\theta u$ for the Gaussian heat-equation kernel and confirm that $p_G = p_{G^*}$ holds to floating precision, with failure indicating a gap in the adjoint argument.

Watch

Extended reading notes

Core claim

The central discovery is that non-identifiability in infinite mixture models is generated by parameter PDEs. Theorem 1 states that if the kernel f satisfies $L_\theta f(x|\theta)=0$ for a non-trivial differential operator on an open ball $B$, and the true mixing measure has a density bounded away from zero on $B$, then there is another probability measure $G \neq G^*$ with $p_G = p_{G^*}$ almost everywhere. Theorem 2 extends this to difference-differential operators with shifts, where the kernel is evaluated at shifted parameter values. The construction is explicit: integrate by parts with a compactly supported test function to build a mean-zero signed perturbation, then add a small multiple of it to the true density. Consequences are drawn for estimation: along with a minimax lower bound showing that under non-identifiability no estimator can recover the mixing measure in Wasserstein distance, even over discrete mixing measures.

Load-bearing premise

The load-bearing premise is that the true mixing measure has a density with respect to Lebesgue measure that is bounded away from zero on some open ball, and for shift operators that the shifted parameter values also lie in the parameter space; if the true measure is discrete, or its density can vanish on every open ball, the negative results do not directly apply.

Editorial extensions

If this is right

  • Infinite location-scale Gaussian mixtures are non-identifiable: the heat equation $\partial_\nu f = \tfrac12 \partial_\mu^2 f$ is a parameter PDE, so distinct mixing measures over $(\mu,\nu)$ yield the same density.
  • Any exponential family whose parameter dimension exceeds its sufficient-statistic dimension satisfies a non-trivial parameter PDE, making over-parameterized exponential kernels a systematic source of non-identifiability.
  • When identifiability fails, the paper's minimax result implies a positive constant lower bound on the expected Wasserstein error of any estimator, uniformly over discrete mixing measures, so unconstrained estimation is impossible in the worst case.
  • Identifiability is preserved for generalized translation families, such as Gaussian with fixed variance, gamma with fixed shape, and Laplace with fixed scale, and for exponential families whose sufficient statistic has dimension at least that of the parameter and a determining range.
  • A concrete super-exponential kernel with sufficient statistics $x$ and $e^{x^2}$ has no non-trivial annihilating parameter PDE and an injective mixture operator, showing that the barrier is not universal.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the theorem's perturbation is local and generic, the non-identifiability it produces is not a knife-edge phenomenon: any true mixing measure with a density bounded below on a small open ball is surrounded by indistinguishable alternatives, so the negative results likely extend to regularized estimators whenever the prior puts mass on continuous mixing densities.
  • The minimax argument in Proposition 3 suggests that discrete priors such as Dirichlet-process mixtures inherit the worst-case obstruction even though the construction in Theorems 1–2 does not directly apply to discrete mixing measures; testing whether direct non-identifiable pairs exist for discrete measures would sharpen the practical implications.
  • The over-parameterization criterion can be read as a design guide: to keep infinite mixtures identifiable, restrict kernels to parameterizations with sufficient-statistic dimension at least as large as parameter dimension, or fix shared nuisance parameters such as a common scale, before attempting nonparametric estimation of the mixing measure.
  • Whether the absence of a parameter PDE is sufficient for identifiability remains open; Proposition 6's super-exponential family provides a test case where the two coincide, suggesting that growth mismatch between sufficient statistics may be the relevant condition to explore.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper studies identifiability of mixing measures in infinite mixture models p_G(x)=∫ f(x|θ)G(dθ). Its main negative results (Theorems 1 and 2) show that if the kernel is annihilated by a nontrivial linear differential operator, or by a difference-differential operator with shifts, on an open parameter region, and if the true mixing measure has a density bounded below on a suitable ball, then another probability measure on the parameter space produces the same mixture density almost everywhere. The authors give sufficient conditions for such annihilators (Proposition 1 for exponential families with parameter dimension exceeding the sufficient-statistic dimension; Proposition 2 for polynomial score functions), verify them on seven standard kernel families, and derive a minimax lower bound for estimation over discrete mixing measures (Proposition 3). Complementing this, Propositions 4–6 describe three kernel classes for which the mixture operator is injective. Proofs are collected in the supplementary appendix.

Significance. The paper's contribution is a clean structural explanation of non-identifiability in several widely used infinite mixture models, together with a balanced set of positive identifiability results. The constructions in Theorems 1 and 2 are explicit and the example identities I checked (Gaussian heat equation, Gamma, Beta, negative binomial) are correct; Proposition 3 gives a useful worst-case transfer from non-identifiability to estimation failure over discrete mixtures. The main limitation, namely the density-lower-bound requirement on the true mixing measure in the negative results, is explicitly acknowledged and partially addressed. If the minor proof clarifications below are made, the paper should be publishable.

minor comments (5)
  1. [A.3 (Proof of Proposition 2)] The final step claims that because the free coefficient is set to 1, the operator is non-trivial, but after column relabeling this coefficient may correspond to the zero-order multi-index, which would not satisfy Definition 1. The claim is true, but the proof should justify it: if the free column is the zero-order column, the top-block equation forces c_A to be nonzero whenever a solution exists, since otherwise the constant-monomial row of Φ would not be annihilated; alternatively, the authors can relabel so that a positive-order column is free. Please add this argument.
  2. [A.5 (Proof of Proposition 3)] The displayed lower bound evaluates to 3c/16, not 3c/32; since 3c/16 > 3c/32 the stated bound remains valid, but the constant in the display and in the statement should be made consistent.
  3. [A.4 (Proof of Theorem 2)] In the final display, the total variation identity should read d_TV(G,G*) = (ε/2) ∫ |h(θ)| dθ; the factor ε is missing. This does not affect the conclusion G ≠ G*.
  4. [Definitions 4 and A.4] The number of shift vectors is denoted J in Definition 4 but K in the proof of Theorem 2; please align the notation.
  5. [Abstract, Section 1.1, Section 2, A.8] Several typos need correction: 'flips side' should be 'flip side'; 'which is constitutes' should be 'which is'; 'satisifes' should be 'satisfies'; 'eqiuivalently' should be 'equivalently'; and 'is call regular closed set' should be 'is called a regular closed set'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the non-identifiability construction is derived from the kernel PDE and standard adjoint/Fourier arguments; no fitted input is relabeled as a prediction and no load-bearing self-citation.

full rationale

The paper's central claim (Theorem 1) is self-contained: given a non-trivial parameter PDE L_theta f = 0, it sets h = L*_theta u for a compactly supported test function, shows integral h = 0 and integral f h = 0 by adjointness, and perturbs g* by epsilon h with epsilon = m/(2M). This is a direct construction, not an equivalent reformulation of the conclusion. Propositions 1 and 2 prove existence of annihilating operators from rank deficiency q < d2; the coefficient functions are solved by Cramer's rule, and the examples verify the PDEs by direct differentiation and recurrences (heat equation, Student-t PDE, beta/gamma/negative-binomial/noncentral-chi-squared shift relations). No parameter is fitted to data and no 'prediction' is a renamed input. The identifiability-preserving propositions rely on external, independent results (Fourier deconvolution, Muntz-Szasz density, Cartwright-Levinson logarithmic-integral lemma from Koosis), not on the authors' prior theorems. The only self-citations (Bariletto et al. 2026a,b) are contextual references in the literature review and discussion; they are not used to establish Theorem 1, Proposition 3, or any other main result. Proposition 3 is explicitly conditional, converting non-identifiability into a worst-case lower bound rather than assuming it. The density-lower-bound restriction is acknowledged in the text and is a scope condition, not a circular step.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The central claim rests on standard functional analysis (formal adjoints, Fubini), explicit smoothness and boundary hypotheses on the kernel and mixing measure, and several external classical theorems. No free parameters are fitted, and no new entities are postulated.

assumptions (6)
  • standard math Fubini-Tonelli and dominated convergence can be applied to interchange integrals over x and theta.
    Used throughout Theorems 1 and 2 and Proposition 3 to transfer zero mixture integrals to zero-mass perturbations.
  • domain assumption The kernel is C^r in the parameter on the relevant open set and the coefficients of the annihilating operator are C^{|alpha|}, so formal adjoints and integrations by parts apply.
    Definition 1 and Theorem 1; all listed examples meet this.
  • domain assumption For shift PDEs, the translated balls B+v_j lie inside the parameter space and are disjoint, enabling the h_k components to be considered separately.
    Theorem 2 hypothesis on rho and the compact parameter space.
  • standard math Muntz-Szasz type density of monomials y^s in C(Omega) for Proposition 5.
    Used to conclude the span of y^s is dense in C(Omega).
  • standard math Paley-Wiener theorem for compactly supported measures on R^2 and Cartwright-Levinson logarithmic integral lemma for Proposition 6.
    Used to prove the super-exponential kernel is identifiable.
  • standard math Le Cam's two-point inequality for minimax lower bounds in Proposition 3.
    Converts separation of mixture densities into a lower bound on Wasserstein risk.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Partial Differential Equation Barriers to Identifiability in Infinite Mixture Models." pith.science (2026). https://pith.science/paper/Z6QJ2OHA

@misc{pith2026260808597,
  author       = {Pith},
  title        = {Pith review of: Partial Differential Equation Barriers to Identifiability in Infinite Mixture Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z6QJ2OHA}},
  note         = {Machine review of arXiv:2608.08597}
}
read the original abstract

We study identifiability of mixing measures in infinite mixture models. We show that, in many common cases, lack of identifiability can be characterized in terms of certain differential structures of the kernel family with respect to its parameters. In our main results, we prove that when the kernel is annihilated by a non-trivial differential or difference-differential operator over the parameter space, there exist infinitely many distinct mixing measures yielding the same mixture density. We give verifiable conditions for such operators to exist, covering many common cases, including the location-scale Gaussian, location-scale Student-t, Gamma, Beta, Dirichlet, negative binomial and non-central Chi-squared families. Furthermore, our conditions apply to any exponential family whose parameter dimension exceeds the dimension of its sufficient statistic and, more generally, to kernels with polynomially-growing score functions. We complement our results with a minimax lower bound on the estimation error for the mixing measure in the Wasserstein distance under non-identifiability. On the flips side, we describe three classes of kernels for which identifiability is preserved and nonparametric statistical inference remains possible.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

300 extracted references · 59 canonical work pages

  1. [1]

    Minimax Optimal Rate for Parameter Estimation in Multivariate Deviated Models , volume =

    Do, Dat and Nguyen, Huy and Nguyen, Khai and Ho, Nhat , booktitle =. Minimax Optimal Rate for Parameter Estimation in Multivariate Deviated Models , volume =

  2. [2]

    Attention is All you Need , volume =

    Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, ukasz and Polosukhin, Illia , booktitle =. Attention is All you Need , volume =

  3. [3]

    On the Representation Collapse of Sparse Mixture of Experts , booktitle =

    Chi, Zewen and Dong, Li and Huang, Shaohan and Dai, Damai and Ma, Shuming and Patra, Barun and Singhal, Saksham and Bajaj, Payal and Song, Xia and Mao, Xian-Ling and Huang, Heyan and Wei, Furu , editor =. On the Representation Collapse of Sparse Mixture of Experts , booktitle =

  4. [4]

    Sparse Mixers: Combining MoE and Mixing to build a more efficient

    Lee-Thorp, James and Ainslie, Joshua , month = dec, year =. Sparse Mixers: Combining MoE and Mixing to build a more efficient. Findings of the

  5. [5]

    Proceedings of the ICML

    On Least Square Estimation in Softmax Gating Mixture of Experts , author=. Proceedings of the ICML

  6. [6]

    Proceedings of the ICML

    Nguyen, Huy and Akbarian, Pedram and Ho, Nhat , booktitle ="Proceedings of the ICML", year=. Is temperature sample efficient for softmax

  7. [7]

    Advances in Neural Information Processing Systems

    FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion , author=. Advances in Neural Information Processing Systems

  8. [8]

    arXiv preprint arXiv:2402.02526 , year=

    CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition , author=. arXiv preprint arXiv:2402.02526 , year=

Show all 300 references
  1. [9]

    arXiv preprint arXiv:2502.03029 , year=

    On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation , author=. arXiv preprint arXiv:2502.03029 , year=

  2. [10]

    arXiv preprint arXiv:2502.00281 , year=

    Sigmoid Self-Attention is Better than Softmax Self-Attention: A Mixture-of-Experts Perspective , author=. arXiv preprint arXiv:2502.00281 , year=

  3. [11]

    arXiv preprint arXiv:2502.03044 , year=

    RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts , author=. arXiv preprint arXiv:2502.03044 , year=

  4. [12]

    arXiv preprint arXiv:2503.03213 , year=

    Convergence Rates for Softmax Gating Mixture of Experts , author=. arXiv preprint arXiv:2503.03213 , year=

  5. [13]

    arXiv preprint arXiv:2410.12258 , year=

    Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts , author=. arXiv preprint arXiv:2410.12258 , year=

  6. [14]

    arXiv preprint arXiv:2410.11222 , year=

    Quadratic Gating Functions in Mixture of Experts: A Statistical Insight , author=. arXiv preprint arXiv:2410.11222 , year=

  7. [15]

    Huy Nguyen and Pedram Akbarian and Trang Pham and Trang Nguyen and Shujian Zhang and Nhat Ho , title =

  8. [16]

    The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

    Mixture of Experts Meets Prompt-Based Continual Learning , author=. The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

  9. [17]

    The Thirteenth International Conference on Learning Representations , year=

    Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts , author=. The Thirteenth International Conference on Learning Representations , year=

  10. [18]

    , journal=

    Barron, A.R. , journal=. Universal approximation bounds for superpositions of a sigmoidal function , year=

  11. [19]

    Damai Dai and Chengqi Deng and Chenggang Zhao and R. X. Xu and Huazuo Gao and Deli Chen and Jiashi Li and Wangding Zeng and Xingkai Yu and Y. Wu and Zhenda Xie and Y. K. Li and Panpan Huang and Fuli Luo and Chong Ruan and Zhifang Sui and Wenfeng Liang , year=. Deep

  12. [20]

    2024 , journal =

    Mixture of A Million Experts , author=. 2024 , journal =

  13. [21]

    B. Yu. Assouad, F ano, and L e C am. Festschrift for Lucien Le Cam. 1997

  14. [22]

    arXiv preprint arXiv:2310.14188 , author =

    A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts , copyright =. arXiv preprint arXiv:2310.14188 , author =

  15. [23]

    Nguyen, Huy and Nguyen, TrungTin and Nguyen, Khai and Ho, Nhat , title =

  16. [24]

    Journal of Statistical Distributions and Applications , author =

    Approximations of conditional probability density functions in. Journal of Statistical Distributions and Applications , author =. 2021 , keywords =. doi:10.1186/s40488-021-00125-0 , abstract =

  17. [25]

    International Conference on Learning Representations , year=

    Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts , author=. International Conference on Learning Representations , year=

  18. [26]

    The Thirteenth International Conference on Learning Representations , year=

    Theory on Mixture-of-Experts in Continual Learning , author=. The Thirteenth International Conference on Learning Representations , year=

  19. [27]

    and Doan, Thanh-Nam and Liu, Chenghao and Ramasamy, Savitha and Li, Xiaoli and HOI, Steven , month = dec, year =

    Do, Truong Giang and Le, Huy Khiem and Nguyen, TrungTin and Pham, Quang and Nguyen, Binh T. and Doan, Thanh-Nam and Liu, Chenghao and Ramasamy, Savitha and Li, Xiaoli and HOI, Steven , month = dec, year =. Proceedings of the 2023

  20. [28]

    Kwon, Jeongyeol and Qian, Wei and Caramanis, Constantine and Chen, Yudong and Davis, Damek , editor =. Global. Proceedings of the. 2019 , pages =

  21. [29]

    Proceedings of the

    Kwon, Jeongyeol and Caramanis, Constantine , editor =. Proceedings of the. 2020 , pages =

  22. [30]

    Neural Networks , author =

    Improved learning algorithms for mixture of experts in multiclass classification , volume =. Neural Networks , author =. 1999 , keywords =. doi:https://doi.org/10.1016/S0893-6080(99)00043-X , abstract =

  23. [31]

    Journal of the Royal Statistical Society

    Iteratively. Journal of the Royal Statistical Society. Series B (Methodological) , author =. 1984 , note =

  24. [32]

    IEEE Transactions on Pattern Analysis and Machine Intelligence , author =

    Sparse multinomial logistic regression: fast algorithms and generalization bounds , volume =. IEEE Transactions on Pattern Analysis and Machine Intelligence , author =. 2005 , pages =. doi:10.1109/TPAMI.2005.127 , number =

  25. [33]

    A regression model with a hidden logistic process for feature extraction from time series , doi =

    Chamroukhi, Faicel and Same, Allou and Govaert, Gerard and Aknin, Patrice , year =. A regression model with a hidden logistic process for feature extraction from time series , doi =. 2009

  26. [34]

    Neural Networks , author =

    Time series modeling by a regression approach based on a latent process , volume =. Neural Networks , author =. 2009 , note =

  27. [35]

    Estimation and feature selection in mixtures of generalized linear experts models , journal =

    Huynh, Bao Tuyen and Chamroukhi, Faicel , year =. Estimation and feature selection in mixtures of generalized linear experts models , journal =

  28. [36]

    The Annals of Applied Statistics , author =

    A mixture of experts model for rank data with applications in election studies , volume =. The Annals of Applied Statistics , author =. 2008 , note =. doi:10.1214/08-AOAS178 , number =

  29. [37]

    , editor =

    Jiang, Wenxin and Tanner, Martin A. , editor =. Hierarchical. Proceedings of the. 1999 , annote =

  30. [38]

    Neural Computation , author =

    On the. Neural Computation , author =. 1999 , pages =. doi:10.1162/089976699300016403 , abstract =

  31. [39]

    Information Retrieval , author =

    Hierarchical. Information Retrieval , author =. 2002 , pages =. doi:10.1023/A:1012782908347 , abstract =

  32. [40]

    Journal of the American Statistical Association , author =

    Bayesian. Journal of the American Statistical Association , author =. 1996 , pages =

  33. [41]

    Neural Networks , author =

    A. Neural Networks , author =. 1997 , keywords =

  34. [42]

    Hierarchical mixtures of experts methodology applied to continuous speech recognition , volume =

    Zhao, Ying and Schwartz, Richard and Sroka, Jason and Makhoul, John , year =. Hierarchical mixtures of experts methodology applied to continuous speech recognition , volume =. Advances in

  35. [43]

    Liang, Hanxue and Fan, Zhiwen and Sarkar, Rishov and Jiang, Ziyu and Chen, Tianlong and Zou, Kai and Cheng, Yu and Hao, Cong and Wang, Zhangyang , year =. M ^3

  36. [44]

    Advances in

    Hazimeh, Hussein and Zhao, Zhe and Chowdhery, Aakanksha and Sathiamoorthy, Maheswaran and Chen, Yihua and Mazumder, Rahul and Hong, Lichan and Chi, Ed , editor =. Advances in. 2021 , pages =

  37. [45]

    , year =

    Ma, Jiaqi and Zhao, Zhe and Yi, Xinyang and Chen, Jilin and Hong, Lichan and Chi, Ed H. , year =. Modeling. Proceedings of the 24th. doi:10.1145/3219819.3220007 , abstract =

  38. [46]

    and Robinson, A.J

    Waterhouse, S.R. and Robinson, A.J. , year =. Classification using hierarchical mixtures of experts , doi =. Proceedings of

  39. [47]

    and Dai, Andrew M

    Zhou, Yanqi and Lei, Tao and Liu, Hanxiao and Du, Nan and Huang, Yanping and Zhao, Vincent Y. and Dai, Andrew M. and Chen, Zhifeng and Le, Quoc V. and Laudon, James , editor =. Mixture-of-. Advances in

  40. [48]

    and Salakhutdinov, R

    Karakoulas, G. and Salakhutdinov, R. , year =. Semi-supervised mixture-of-experts classification , doi =. Fourth

  41. [49]

    The Annals of Statistics , author =

    Optimal estimation of high-dimensional. The Annals of Statistics , author =. 2023 , note =. doi:10.1214/22-AOS2207 , number =

  42. [50]

    Regularized

    Chamroukhi, Faicel and Huynh, Bao Tuyen , year =. Regularized. doi:10.1109/IJCNN.2018.8489670 , booktitle =

  43. [51]

    Journal de la Société Française de Statistique , author =

    Regularized. Journal de la Société Française de Statistique , author =. 2019 , pages =

  44. [52]

    Neurocomputing , author =

    A hidden process regression model for functional data description. Neurocomputing , author =. 2010 , pages =

  45. [53]

    Journal of the Royal Statistical Society: Series B (Methodological) , author =

    Maximum. Journal of the Royal Statistical Society: Series B (Methodological) , author =. 1977 , keywords =. doi:10.1111/j.2517-6161.1977.tb01600.x , abstract =

  46. [54]

    Electronic Journal of Statistics , author =

    Mixture of. Electronic Journal of Statistics , author =. 2014 , pages =

  47. [55]

    Chen and J

    H. Chen and J. Chen. Tests for homogeneity in normal mixtures in the presence of a structural parameter. Statistica Sinica. 2003

  48. [56]

    Kasahara and K

    H. Kasahara and K. Shimotsu. Testing the number of components in normal mixture regression models. Journal of the American Statistical Association. 2014

  49. [57]

    van de Geer

    S. van de Geer. Empirical Processes in M-estimation. 2000

  50. [58]

    E. Polak. Optimization: Algorithms and Consistent Approximations. 1997

  51. [59]

    R. M. Dudley. Probabilities and metrics: Convergence of laws on metric spaces, with a view to statistical testing. 1976

  52. [60]

    Topics in Optimal Transportation

    C\'edric Villani. Topics in Optimal Transportation. 2003

  53. [61]

    Optimal transport: Old and New

    C\'edric Villani. Optimal transport: Old and New. 2008

  54. [62]

    S. T. Rachev and L. Ruschendorf. Mass transportation problems. Vol I: Theory, Vol II: Probability and its Applications. 1998

  55. [63]

    S. T. Rachev. Probability metrics and the stability of stochastic systems. 1991

  56. [64]

    C. Zhang. Fourier methods for estimating mixing densities and distributions. Annals of Statistics. 1990

  57. [65]

    J. Fan. On the optimal rates of convergence for nonparametric deconvolution problems. Annals of Statistics. 1991

  58. [66]

    R. J. Carroll and P. Hall. Optimal rates of convergence for deconvolving a density. Journal of American Statistical Association. 1988

  59. [67]

    Convergence rates for

    Ho, Nhat and Yang, Chiao-Yu and Jordan, Michael I , journal=. Convergence rates for

  60. [68]

    International Conference on Machine Learning , pages=

    Refined convergence rates for maximum likelihood estimation under finite mixture models , author=. International Conference on Machine Learning , pages=. 2022 , organization=

  61. [69]

    Bariletto, Nicola and Nguyen, Huy and Ho, Nhat and Rinaldo, Alessandro , journal=

  62. [70]

    Dedecker and B

    J. Dedecker and B. Michel. Minimax rates of convergence for W asserstein deconvolution with supersmooth errors in any dimension. Arxiv manuscript. 2013

  63. [71]

    Bickel and D

    P. Bickel and D. Freedman. Some asymptotic theory for the bootstrap. Annals of Statistics. 1981

  64. [72]

    del Barrio and J

    E. del Barrio and J. Cuesta-Albertos and C. Matr\'an and J. Rodr\'iguez-Rodr\'iguez. Tests of goodness of fit based on the L_2 -Wasserstein distance. Annals of Statistics. 1999

  65. [73]

    C. Mallows. A note on asymptotic joint normality. Annals of Mathematical Statistics. 1972

  66. [74]

    Dobrushin

    R. Dobrushin. Describing a system of random variables by conditional distributions. Theory Probab. Appl. 1970

  67. [75]

    A. L. Gibbs and F. E. Su. On choosing and bounding probability metrics. International Statistical Review. 2002

  68. [76]

    Ferraty and P

    F. Ferraty and P. Vieu. Nonparametric Functional Data Analysis: T heory and P ractice. 2006

  69. [77]

    J. O. Ramsay and B. W. Silverman. Applied functional data analysis: M ethods and case studies. 2002

  70. [78]

    J. O. Ramsay and B.W. Silverman. Functional Data Analysis. 2006

  71. [79]

    Abraham and P.A

    C. Abraham and P.A. Cornillon and E. Matzner-Lober and N. Molinari. Unsupervised curve clustering using B-splines. Scand. J. Statist. 2003

  72. [80]

    Biau and L

    G. Biau and L. Devroye and G. Lugosi. On the Performance of Clustering in Hilbert Spaces. IEEE Trans. Inform. Theory. 2008

  73. [81]

    Chiou and P.-L

    J.-M. Chiou and P.-L. Li. Functional clustering and identifying substructures of longitudinal data. J. Roy. Statist. Soc. Ser. B. 2007

  74. [82]

    J. A. Cuesta-Albertos and R. Fraiman. Impartial trimmed k-means for functional data. Comput. Statist. Data Anal. 2007

  75. [83]

    Dabo-Niang and F

    S. Dabo-Niang and F. Ferraty and P. Vieu. Mode estimation for functional random variable and its application for curves classication. Far East J. Theor. Stat. 2006

  76. [84]

    Fraiman and A

    R. Fraiman and A. Justel and M. Svarc. Selection of variables for cluster analysis and classication rules. J. Amer. Stat. Assoc. 2008

  77. [85]

    Fraiman and G

    R. Fraiman and G. Muniz. Trimmed means for functional data. Test. 2001

  78. [86]

    G. M. James and C.A. Sugar. Clustering for sparsely sampled functional data. J. Amer. Stat. Assoc. 2003

  79. [87]

    Ma and W

    P. Ma and W. Zhong. Penalized clustering of large-scale functional data with multiple covariates. J. Amer. Statist. Assoc. 2008

  80. [88]

    Tokushige and H

    S. Tokushige and H. Yadohisa and K. Inada. Crisp and fuzzy k-means clustering algorithms for multivariate functional data. Comput. Statist. 2007

  81. [89]

    Ishwaran and M

    H. Ishwaran and M. Zarepour. Dirichlet prior sieves in finite normal mixtures. Statistica Sinica. 2002

  82. [90]

    H. Teicher. Identifiability of finite mixtures. Ann. Math. Statist. 1963

  83. [91]

    S. Ghosal. Dirichlet process, related priors and posterior asymptotics. Manuscript. 2007

  84. [92]

    Ghosal and J

    S. Ghosal and J. K. Ghosh and R. V. Ramamoorthi. Posterior consistency of Dirichlet mixtures in density estimation. Annals of Statistics. 1999

  85. [93]

    Barron and M

    A. Barron and M. Schervish and L. Wasserman. The consistency of posterior distributions in nonparametric problems. Ann. Statist. 1999

  86. [94]

    Wu and S

    Y. Wu and S. Ghosal. L1-Consistency of Dirichlet mixtures in multivariate Bayesian density estimation. 2009

  87. [95]

    M.L. Stein. Interpolation of Spatial Data. 1999

  88. [96]

    Winn and A

    J. Winn and A. Criminisi and T. Minka. Object Categorization by Learned Universal Visual Dictionary. Proc. IEEE Intl. Conf. on Computer Vision (ICCV). 2005

  89. [97]

    Brumback and J

    B.A. Brumback and J. Rice. Smoothing spline models for the analysis of nested and crossed samples of curves. J. Amer. Statist. Assoc. 1998

  90. [98]

    Sudderth and A

    E. Sudderth and A. Torralba and W. Freeman and A. Willsky. Describing Visual Scenes Using Transformed Objects and Parts. International Journal of Computer Vision. 2008

  91. [99]

    Pritchard and M

    J. Pritchard and M. Stephens and P. Donnelly. Inference of populaton structure using multilocus genotype data. Genetics. 2000

  92. [100]

    DeIorio and P

    M. DeIorio and P. Muller and G.L. Rosner and S.N. MacEachern. An ANOVA model for dependent random measures. J. Amer. Statist. Assoc. 2004

  93. [101]

    Duan and M

    J. Duan and M. Guindani and A. Gelfand. Generalized spatial D irichlet processes. Biometrika. 2007

  94. [102]

    Dunson and J.-H

    D.B. Dunson and J.-H. Park. Kernel stick-breaking processes. Biometrika. 2008

  95. [103]

    D.B. Dunson. Kernel local partition processes for functional data. 2008

  96. [104]

    D.B. Dunson. Nonparametric Bayes local partition models for random effects. Biometrika. 2008

  97. [105]

    R. F. MacLehose and D.B. Dunson. Nonparametric Bayes kernel-based priors for functional data analysis. Statistica Sinica. 2008

  98. [106]

    Pillai and F

    N. Pillai and F. Liang and S. Mukerjee and R. Wolpert and Q. Wu. Characterizing the function space for Bayesian kernel models. 2006

  99. [107]

    An and C

    Q. An and C. Wang and I. Shterev and E. Wang and L. Carin and D. Dunson. Hierachicial kernel stick-breaking process for multi-task image analysis. Proc. ICML. 2008

  100. [108]

    Ferguson

    T.S. Ferguson. A B ayesian analysis of some nonparametric problems. Ann. Statist. 1973

  101. [109]

    Gelfand and A

    A.E. Gelfand and A. Kottas and S.N. MacEachern. Bayesian nonparametric spatial modeling with D irichlet process mixing. J. Amer. Statist. Assoc. 2005

  102. [110]

    Griffin and M.F

    J.E. Griffin and M.F. Steel. Bayesian nonparametric spatial modeling with D irichlet process mixing. 2005

  103. [111]

    Griffin and M.F

    J.E. Griffin and M.F. Steel. Order-based dependent D irichlet processes. J. Amer. Statist. Assoc. 2006

  104. [112]

    Ishwaran and L.F

    H. Ishwaran and L.F. James. Gibbs sampling methods for stick-breaking priors. J. Amer. Statist. Assoc. 2001

  105. [113]

    MacEachern

    S.N. MacEachern. Dependent D irichlet processes. 2000

  106. [114]

    Petrone and M

    S. Petrone and M. Guidani and A.E. Gelfand. Hybrid D irichlet processes for functional data. J. Royal Stat. Soc. Series B. 2009

  107. [115]

    Rodriguez and D

    A. Rodriguez and D. Dunson and A.E. Gelfand. The nested D irichlet process. J. Amer. Statist. Assoc. to appear

  108. [116]

    Sethuraman

    J. Sethuraman. A constructive definition of D irichlet priors. Statistica Sinica. 1994

  109. [117]

    Teh and M.I

    Y.W. Teh and M.I. Jordan and M.J. Beal and D.M. Blei. Hierarchical D irichlet processes. J. Amer. Statist. Assoc. 2006

  110. [118]

    Blei and A.Y

    D.M. Blei and A.Y. Ng and M.I. Jordan. Latent D irichlet allocation. J. Mach. Learn. Res. 2003

  111. [119]

    Wang and E

    X. Wang and E. Grimson. Spatial latent D irichlet allocation. NIPS 20. 2008

  112. [120]

    Figueiredo and D.S

    M.A. Figueiredo and D.S. Cheng and V. Murino. Clustering under prior knowledge with application to image segmentation. NIPS 19. 2007

  113. [121]

    Fernandez and P

    C. Fernandez and P. Green. Modelling spatially correlated data via mixtures: A B ayesian approach. J. Roy. Statist. Soc, Series B. 2002

  114. [122]

    Green and S

    P. Green and S. Richardson. Hidden M arkov models and desease mapping. J. Amer. Statist. Assoc. 2001

  115. [123]

    R. R. Phelps. Convex functions, monotone operators and differentiability. 1993

  116. [124]

    M. J. Wainwright and M. I. Jordan. Graphical models, exponential families, and variational inference

  117. [125]

    Bertsekas

    D.P. Bertsekas. Nonlinear Programming. 1995

  118. [126]

    Harmonic analysis on semigroups

  119. [127]

    Aronszajn

    N. Aronszajn. Theory of reproducing kernels. Transactions of the American Mathematical Society

  120. [128]

    D. G. Luenberger. Optimization by Vector Space Methods

  121. [129]

    Rockafellar

    G. Rockafellar. Convex Analysis. 1970

  122. [130]

    Hiriart-Urruty and C

    J. Hiriart-Urruty and C. Lemar\'echal. Fundamentals of Convex Analysis. 2001

  123. [131]

    Koltchinskii and D

    V. Koltchinskii and D. Panchenko. Empirical margin distributions and bounding the generalization error of combined classifiers. Annals of Statistics

  124. [132]

    Bousquet and A

    O. Bousquet and A. Elisseeff. Stability and generalization. Journal of Machine Learning Research

  125. [133]

    Freund and R

    Y. Freund and R. Schapire. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences. 1997

  126. [134]

    Bartlett and S

    P. Bartlett and S. Mendelson. G aussian and R ademacher complexities: R isk bounds and structural results. Journal of Machine Learning Research

  127. [135]

    Predd and S

    J. Predd and S. Kulkarni and H. V. Poor. Consistency in Models for Communication Constrained Distributed Learning. Proceedings of the COLT

  128. [136]

    Reproducing Kernel Hilbert Spaces : Applications in Statistical Signal Processing. 1982

  129. [137]

    S. Saitoh. Theory of Reproducing Kernels and its Applications. 1988

  130. [138]

    Blackwell

    D. Blackwell. Comparison of experiments. Proceeding of 2nd Berkeley Symposium on Probability and Statistics. 1951

  131. [139]

    Blackwell

    D. Blackwell. Equivalent comparisons of experiments. Annals of Statistics. 1953

  132. [140]

    Bradt and S

    R. Bradt and S. Karlin. On the design and comparison of certain dichotomous experiments. Annals of Statistics. 1956

  133. [141]

    Goel and M

    P. Goel and M. DeGroot. Comparisons of experiments and information measures. Annals of Statistics. 1979

  134. [142]

    Chernoff

    H. Chernoff. A measure of asymptotic efficiency for tests of a hypothesis based on a sum of observations. Annals of Statistics. 1952

  135. [143]

    P. Massart. Some applications of concentration inequalities to statistics. Annales de la Facult\'e des Sciences de Toulouse. 2000

  136. [144]

    McDiarmid

    C. McDiarmid. On the method of bounded differences. Surveys in Combinatorics. 1989

  137. [145]

    A. W. van der Vaart and J. Wellner. Weak Convergence and Empirical Processes. 1996

  138. [146]

    A. W. van der Vaart. Asymptotic Statistics. 1998

  139. [147]

    Le \ Cam

    L. Le \ Cam. Asymptotic methods in statistical decision theory. 1986

  140. [148]

    van de Geer

    S. van de Geer. Empirical processes in M-estimation. 2000

  141. [149]

    T. Zhang. Statistical behavior and consistency of classification methods based on convex risk minimization. Annal of Statistics

  142. [150]

    T. Zhang. From -entropy to KL -entropy: Analysis of Minimum Complexity Density Estimation. Ann. Statist. 2006

  143. [151]

    T. Zhang. Information Theoretical Upper and Lower Bounds for Statistical Estimation. IEEE Trans.Info. 2006

  144. [152]

    Steinwart

    I. Steinwart. Consistency of support vector machines and other regularized kernel machines. IEEE Trans. Info. Theory. 2005

  145. [153]

    Friedman and T

    J. Friedman and T. Hastie and R. Tibshirani. Additive logistic regression: A statistical view of boosting. Annals of Statistics. 2000

  146. [154]

    Bartlett and M

    P. Bartlett and M. I. Jordan and J. D. McAuliffe. Convexity, classification and risk bounds. Journal of the American Statistical Association. 2006

  147. [155]

    L. Breiman. Arcing classifiers. Annals of Statistics. 1998

  148. [156]

    W. Jiang. Process consistency for Adaboost. Annals of Statistics. 2004

  149. [157]

    Lugosi and N

    G. Lugosi and N. Vayatis. On the Bayes-risk consistency of regularized boosting methods. Annals of Statistics. 2004

  150. [158]

    Mannor and R

    S. Mannor and R. Meir and T. Zhang. Greedy algorithms for classification - consistency, convergence rates and adaptivity. Journal of Machine Learning Research. 2003

  151. [159]

    Yang and A

    Y. Yang and A. Barron. Information theoretic determination of minimax rates of convergence. Annals of Statistics. 1999

  152. [160]

    Guo and S

    D. Guo and S. Shamai and S. Verdú. Mutual Information and Minimum Mean-Square Error in Gaussian Channels. IEEE Trans. Information Theory. 2005

  153. [161]

    Cover and J

    T. Cover and J. Thomas. Elements of information theory. 1991

  154. [162]

    Csisz\'ar

    I. Csisz\'ar. Information-type measures of difference of probability distributions and indirect observation. Studia Sci. Math. Hungar

  155. [163]

    S. M. Ali and S. D. Silvey. A general class of coefficients of divergence of one distribution from another. J. Royal Stat. Soc. Series B

  156. [164]

    T. Kailath. RKHS approach to detection and estimation problems--- P art I : D eterministic signals in G aussian noise. IEEE T rans. I nfo. T heory

  157. [165]

    Kailath and H

    T. Kailath and H. V. Poor. Detection of stochastic processes. IEEE T rans. I nfo. T heory

  158. [166]

    T. Kailath. The Divergence and B hattacharyya Distance Measures in Signal Selection. IEEE Trans. on Communication Technology

  159. [167]

    Longo and T

    M. Longo and T. Lookabaugh and R. Gray. Quantization for Decentralized Hypothesis Testing under Communication Contraints. IEEE Trans. on Information Theory

  160. [168]

    H. V. Poor and J. B. Thomas. Applications of A li- S ilvey distance measures in the design of generalized quantizers for binary decision systems. IEEE Trans. on Communications

  161. [169]

    R. R. Tenney and Sandell, N. R. Jr. Detection with Distributed Sensors. IEEE Trans. Aero. Electron. Sys

  162. [170]

    H. L. van Trees. Detection, Estimation and Modulation Theory. 1990

  163. [171]

    Z. Luo. Universal decentralized estimation in a bandwidth constrained sensor network. 2003

  164. [172]

    Advances in Statistical Signal Processing

    J. N. Tsitsiklis , BOOKTITLE = "Advances in Statistical Signal Processing", EDITORS = "H. V. Poor and J. B. Thomas", TITLE = "Decentralized Detection", publisher =

  165. [173]

    Tsitsiklis

    J. Tsitsiklis. Extremal properties of likelihood-ratio quantizers. IEEE Trans. on Communication. 1993

  166. [174]

    Advances in Statistical Signal Processing

    S. A. Kassam , BOOKTITLE = "Advances in Statistical Signal Processing", EDITORS = "H. V. Poor and J. B. Thomas", TITLE = "Nonparametric signal detection", publisher =

  167. [175]

    H. V. Poor. An introduction to signal detection and estimation. 1994

  168. [176]

    Viswanathan and A

    R. Viswanathan and A. Ansari. Distributed detection of a signal in generalized G aussian noise. IEEE Trans. Acoust., Speech, and Signal Process. 1989

  169. [177]

    F. Topsoe. Some inequalities for information divergence and related measures of discrimination. IEEE Transactions on Information Theory. 2000

  170. [178]

    Nasipuri and S

    A. Nasipuri and S. Tantaratana. Nonparametric distributed detection using W ilcoxon statistics. Signal Processing. 1997

  171. [179]

    R. S. Blum and S. A. Kassam and H. V. Poor. Distributed detection with multiple sensors: P art II --- Advanced Topics. Proceedings of the IEEE

  172. [180]

    Han and P

    J. Han and P. K. Varshney and V. C. Vannicola. Some results on distributed nonparametric detection. Proc. 29th Conf. on Decision and Control. 1990

  173. [181]

    M. M. Al-Ibrahim and P. K. Varshney. Nonparametric sequential detection based on multisensor data. Proc. 23rd Annu. Conf. on Inform. Sci. and Syst. 1989

  174. [182]

    E. K. Hussaini and A. A. M. Al-Bassiouni and Y. A. El-Far. Decentralized CFAR signal detection. Signal Processing. 1995

  175. [183]

    V. V. Veeravalli and T. Basar and H. V. Poor. Decentralized sequential detection with a fusion center performing the sequential test. IEEE Trans. Info. Theory. 1993

  176. [184]

    J. F. Chamberland and V. V. Veeravalli. Decentralized detection in sensor networks. IEEE Transactions on Signal Processing. 2003

  177. [185]

    Nguyen and M

    X. Nguyen and M. J. Wainwright and M. I. Jordan. Nonparametric decentralized detection using kernel methods. IEEE Transactions on Signal Processing. 2005

  178. [186]

    Nguyen and M

    X. Nguyen and M. J. Wainwright and M. I. Jordan. Divergence measures, surrogate loss functions and experiment design. Advances in Neural Information Processing Systems 11

  179. [187]

    Nguyen and M

    X. Nguyen and M. J. Wainwright and M. I. Jordan. On divergences, surrogate loss functions and decentralized detection

  180. [188]

    Nguyen and M

    X. Nguyen and M. J. Wainwright and M. I. Jordan. On surrogate loss functions and f -divergences. Annals of Statistics. 2009

  181. [189]

    Nguyen and M

    X. Nguyen and M. J. Wainwright and M. I. Jordan. Estimating divergence functionals and the likelihood ratio by convex risk minimization. IEEE Transactions on Information Theory. 2010

  182. [190]

    Jaakkola and D

    T. Jaakkola and D. Haussler. Exploiting generative models in discriminative classifiers. Advances in Neural Information Processing Systems 11

  183. [191]

    Taskar and C

    B. Taskar and C. Guestrin and D. Koller. Max- M argin M arkov N etworks. NIPS 15

  184. [192]

    Dougherty and R

    J. Dougherty and R. Kohavi and M. Sahami. Supervised and unsupervised discretization of continuous features. Proceedings of the ICML

  185. [193]

    Cortes and V

    C. Cortes and V. Vapnik. Support-Vector Networks. Machine Learning. 1995

  186. [194]

    Blake and C.J

    C.L. Blake and C.J. Merz. UCI Repository of machine learning databases. 1998

  187. [195]

    Tsuda and T

    K. Tsuda and T. Kin and K. Asai. Marginalized Kernels for Biological Sequences. Bioinformatics

  188. [196]

    Learning with Kernels. 2002

  189. [197]

    Cristianini and J

    N. Cristianini and J. Shawe-Taylor. An Introduction to Support Vector Machines (and other kernel based learning methods). 2000

  190. [198]

    Chong and S

    C. Chong and S. P. Kumar. Sensor Networks: Evolution, Opportunities, and Challenges. Proceedings of the IEEE

  191. [199]

    Bodik and G

    P. Bodik and G. Friedman and L. Biewald and H. Levine and G. Candea and K. Patel and G. Tolle and J. Hui and A. Fox and M. I. Jordan and D. Patterson. Combining Visualization and Statistical Analysis to Improve Operator Confidence and Efficiency for Failure Detection and Local...

  192. [200]

    Support-vector networks

    Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine Learning. 1995

  193. [201]

    J. H. Chen. Optimal rate of convergence for finite mixture models. Annals of Statistics. 1995

  194. [202]

    S. J. Yakowitz and J. D. Spragins. On the identifiability of finite mixtures. Annals of Statistics. 1968

  195. [203]

    X. Nguyen. Convergence of latent mixing measures in finite and infinite mixture models. Annals of Statistics. 2013

  196. [204]

    B. Lindsay. Mixture models: Theory, geometry and applications. 1995

  197. [205]

    G. J. McLachlan and K. E. Basford. Mixture models: Inference and Applications to Clustering. Statistics: Textbooks and Monographs. 1988

  198. [206]

    van de Geer

    S. van de Geer. Rates of convergence for the maximum likelihood estimator in mixture models. Journal of Nonparametric Statistics. 1996

  199. [207]

    Shen and W

    X. Shen and W. H. Wong. Convergence rate of sieves estimates. Annals of Statistics. 1994

  200. [208]

    Shen and L

    X. Shen and L. Wasserman. Rates of convergence of posterior distributions. Annals of Statistics. 2001

  201. [209]

    Ghosal and J

    S. Ghosal and J. K. Ghosh and A. van der Vaart. Convergence rates of posterior distributions. Annals of Statistics. 2000

  202. [210]

    S. G. Walker, A. Lijoi and I. Prunster. On rates of convergence for posterior distributions in infinite-dimensional models. Annals of Statistics. 2007

  203. [211]

    Ghosal and A

    S. Ghosal and A. van der Vaart. Posterior convergence rates of Dirichlet mixtures at smooth densities. Annals of Statistics. 2007

  204. [212]

    Ghosal and J

    S. Ghosal and J. K. Ghosh and R. V. Ramamoorthi. Posterior consistency of D irichlet mixtures in density estimation. The Annals of Statistics. 1999

  205. [213]

    Ghosal and A

    S. Ghosal and A. van der Vaart. Entropies and rates of convergence for maximum likelihood and B ayes estimation for mixtures of normal densities. The Annals of Statistics. 2001

  206. [214]

    On consistency of nonparametric normal mixtures for

    Lijoi, Antonio and Pr. On consistency of nonparametric normal mixtures for. Journal of the American Statistical Association , volume=. 2005 , publisher=

  207. [215]

    C. R. Genovese. and L. Wasserman. Rates of convergence for the Gaussian mixture sieve. Annals of Statistics. 2000

  208. [216]

    C. Villani. Topics in Optimal Transportation.Graduate Studies in Mathematics. 2003

  209. [217]

    C. Villani. Optimal Transport: Old and New. Grundlehren der Mathematischen Wissenschaften [Fundamental Principles of Mathemtical Sciences]. 2009

  210. [218]

    H. Teicher. Identifiability of Mixtures. Annals of Statistics. 1961

  211. [219]

    H. Teicher. On the mixture of distributions. Annals of Statistics. 1960

  212. [220]

    H. Teicher. Identifiability of finite mixtures. Annals of Statistics. 1963

  213. [221]

    H. Teicher. Identifiability of mixtures of product measures. Annals of Statistics. 1967

  214. [222]

    Osiewalski and M

    J. Osiewalski and M. F. J. Steel. Robust Bayesian Inference in l_ q -spherical models. Biometrika. 1993

  215. [223]

    J. T. Kent. Identifiability of finite mixtures for directional data. Annals of Statistics. 1983

  216. [224]

    Y. S. Hsu and M. D. Fraser and J. J. Walker. Identifiability of finite mixtures of von Mises distributions. Annals of Statistics. 1981

  217. [225]

    K. V. Mardia. Statistics of directional data. Journal of the Royal Statistical Society. Series B(Methodological). 1975

  218. [226]

    S. M. Ali and S. D. Silvey. A general class of coefficients of divergence of one distribution from another. Journal of the Royal Statistical Society. Series B(Methodological). 1966

  219. [227]

    Peel and G

    D. Peel and G. J. McLachlan. Robust mixture modelling using the t distribution. Statistics and Computing. 2000

  220. [228]

    Azzalini and A

    A. Azzalini and A. D. Valle. The multivariate skew-normal distribution. Biometrika. 1996

  221. [229]

    Azzalini and A

    A. Azzalini and A. Capitanio. Statistical applications of the multivariate skew-normal distribution. Journal of the Royal Statistical Society, Series B(Methodological). 1999

  222. [230]

    E. S. Allman and C. Matias and J. A. Rhodes. Identifiability of parameters in latent structure models with many observed variables. Annals of Statistics. 2009

  223. [231]

    Hall and X

    P. Hall and X. -H. Zhou. Nonparametric estimation of component distributions in a multivariate mixture. Annals of Statistics. 2003

  224. [232]

    Hall and A

    P. Hall and A. Neeman and R. Pakyari and R. Elmore. Nonparametric inference in multivariate mixtures. Biometrika. 2005

  225. [233]

    Elmore and P

    R. Elmore and P. Hall and A. Neeman. An application of classical invariant theory to identifiability in nonparametric mixtures. Ann. Inst. Fourier (Grenoble). 2005

  226. [234]

    Le \ Cam

    L. Le \ Cam. Convergence of estimates under dimensionality reductions. Annals of Statistics. 1973

  227. [235]

    Buchberger , title=

    B. Buchberger , title=

  228. [236]

    Sturmfels

    B. Sturmfels. Solving system of polynomial equations. 2002

  229. [237]

    Wiper and D

    M. Wiper and D. R. Insua and F. Ruggeri. Mixtures of Gamma distributions with applications. Journal of Computational and Graphical Statistics. 2001

  230. [238]

    S. X. Lee and G. J. McLachlan. On mixtures of skew normal and skew t -distributions. Advances in Data Analysis and Classification. 2013

  231. [239]

    Ghosal and A

    S. Ghosal and A. Roy. Predicting false discovery proportion under dependence. Journal of the American Statistical Association. 2011

  232. [240]

    Zhang and A

    T. Zhang and A. Weisel and M. S. Greco. Multivariate generalized Gaussian distribution: Convexity and graphical models. IEEE Transactions on Signal Processing. 2013

  233. [241]

    Ho and X

    N. Ho and X. Nguyen. On strong identifiability and convergence rates of parameter estimation in finite mixtures. Electronic Journal of Statistics. 2016

  234. [242]

    Ho and X

    N. Ho and X. Nguyen. Convergence rates of parameter estimation for some weakly identifiable finite mixtures. Annals of Statistics. 2016

  235. [243]

    Kulis and M

    B. Kulis and M. I. Jordan. Revisiting k-means: new algorithm via Bayesian nonparametrics. Proceedings of the 29^ th International Conference on Machine Learning. 2012

  236. [244]

    Hallin and C

    M. Hallin and C. Ley. Skew-symmetric distributions and Fisher information - a tale of two densities. Bernoulli. 2012

  237. [245]

    Hallin and C

    M. Hallin and C. Ley. Skew-symmetric distributions and Fisher information: the double sin of skew-normal. Bernoulli. 2014

  238. [246]

    M. Chiogna. A note on the asymptotic distribution of the maximum likelihood estimator for the scalar skew-normal distribution. Statistical Methods and Applications. 2005

  239. [247]

    Azzalini and A

    A. Azzalini and A. Capitanio. Distributions generated by perturbation of symmetry with emphasis on a multivariate skew t-distribution. Journal of Royal Statistical Society: Series B (Statistical Methodology). 2003

  240. [248]

    Azzalini

    A. Azzalini. Further results on a class of distributions which includes the normal ones. Statistica (Bologna). 1986

  241. [249]

    T. J. DiCiccio and A. C. Monti. Inferential aspects of the skew-exponential power distribution. Journal of the American Statistical Association. 2004

  242. [250]

    Azzalini and A

    A. Azzalini and A. Capitanio. Statistical applications of the multivariate skew normal distribution. Journal of Royal Statistical Society: Series B (Statistical Methodology). 1999

  243. [251]

    T. I. Lin and J. C. Lee and S. Y. Yen. Finite mixture modelling using the skew normal distribution. Statistica Sinica. 2007

  244. [252]

    Ley and D

    C. Ley and D. Paindaveine. On the singularity of multivariate skew-symmetric models. Journal of Multivariate Analysis. 2010

  245. [253]

    X. Nguyen. Borrowing strength in hierarchical Bayes: convergence of the Dirichlet base measure. Bernoulli. 2015+

  246. [254]

    Azzalini and A

    A. Azzalini and A. Capitanio. Distributions generated by perturbation of symmetry with emphasis on a multivariate skew t distribution. Journal of Royal Statistical Society: Series B (Statistical Methodology). 2003

  247. [255]

    T. I. Lin. Maximum likelihood estimation for multivariate skew normal mixture models. Journal of Multivariate Analysis. 2009

  248. [256]

    T. I. Lin. Robust mixture modeling using multivariate skew t-distributions. Statistics and Computing. 2010

  249. [257]

    T. I. Lin and J. C. Lee and W. J. Hsieh. Robust mixture modelling using the skew t-distribution. Statistics and Computing. 2007

  250. [258]

    Lee and G

    S. Lee and G. J. McLachlan. Finite mixtures of multivariate skew t-distributions: some recent and new results. Statistics and Computing. 2014

  251. [259]

    H. J. Ho and S. Pyne and T. I. Line. Maximum likelihood inference for mixtures of skew student- t-Normal distributions through practical EM-type algorithms. Statistics and Computing. 2012

  252. [260]

    M. O. Prates and C. R. B. Cabral and V. H. Lachos. mixsmsn: fitting finite mixture of scale mixture of skew-normal distributions. Journal of Statistical Software. 2013

  253. [261]

    S. W. Schnatter and S. Pyne. Bayesian inference for finite mixtures of univariate and multivariate skew-normal and skew-t distributions. Biostatistics. 2009

  254. [262]

    R. B. Arellano-Valle and L. M. Castro and M. C. Genton and H. W. Gómez. Bayesian inference for shape mixtures of skewed distributions, with application to regression analysis. Bayesian Analysis. 2008

  255. [263]

    R. B. Arellano-Valle and M. C. Genton and R. H. Loschi. Shape mixtures of multivariate skew-normal distributions. Journal of Multivariate Analysis. 2009

  256. [264]

    C. B. Zeller and C. R. B. Cabral and V. H. Lachos. Robust mixture regression modeling based on scale mixtures of skew-normal distributions. TEST. 2015

  257. [265]

    Azzallini

    A. Azzallini. A class of distributions which includes the normal ones. Scadinavian Journal of Statistics. 1985

  258. [266]

    R. B. Arellano-Valle and A. Azzallini. The centred parametrization for the multivariate skew-normal distribution. Journal of Multivariate Analysis. 2008

  259. [267]

    M. G. Genton. Skew-elliptical distributions and their applications: a journey beyond normality. 2004

  260. [268]

    Wang and J

    J. Wang and J. Boyer and M. C. Genton. A skew-symmetric representation of multivariate distribution. Statistica Sinica. 2004

  261. [269]

    Canale and B

    A. Canale and B. Scarpa. Bayesian nonparametric location-scale-shape mixtures. TEST. 2015

  262. [270]

    Ishwaran and L

    H. Ishwaran and L. F. James and J. Sun. Bayesian model selection in finite mixtures by marginal density decompositions. Journal of the American Statistical Association. 2001

  263. [271]

    Rousseau and K

    J. Rousseau and K. Mengersen. Asymptotic behaviour of the posterior distribution in overfitted mixture models. Journal of the Royal Statistical Society: Series B (Statistical Methodology). 2011

  264. [272]

    Petralia and V

    F. Petralia and V. Rao and D. B. Dunson. Repulsive mixtures. Advances in Neural Information Processing Systems (NIPS). 2012

  265. [273]

    Rotnitzky and D

    A. Rotnitzky and D. R. Cox and M. Bottai and J. Robins. Likelihood-based inference with singular information matrix. Bernoulli. 2000

  266. [274]

    Heinrich and J

    P. Heinrich and J. Kahn. Optimal rates for finite mixture estimation. Under review. 2016+

  267. [275]

    L. F. Lee and A. Chesher. Specification testing when score test statistics are identically zero. Journal of Econometrics. 1986

  268. [276]

    J. Chen. Consistency of the MLE under mixture models. arXiv preprint arXiv:1607.01251. 2016

  269. [277]

    Chen and X

    J. Chen and X. Tan and R. Zhang. Inference for normal mixtures in mean and variance. Statistica Sinica. 2008

  270. [278]

    Cox and J

    D. Cox and J. Little and D. O'Shea. Ideals, Varieties, and Algorithms: An Introduction to Computational Algebraic Geometry and Commutative Algebra. 2007

  271. [279]

    B. Stumfel. Solving systems of polynomial equations. 2002

  272. [280]

    Xiao and G

    S. Xiao and G. Zeng. Determination of the limits for multivariate rational functions. Science China Mathematics. 2014

  273. [281]

    N. M. Kiefer. A remark on the parameterization of a model for heterogeneity. 1982

  274. [282]

    Ho and X

    N. Ho and X. Nguyen. Singularity structures and impacts on parameter estimation in finite mixtures of distributions. 2016

  275. [283]

    Toussile and E

    W. Toussile and E. Gassiat. Variable selection in model-based clustering using multilocus genotype data. Advances in Data Analysis and Classification. 2009

  276. [284]

    Gassiat and R

    E. Gassiat and R. V. Handel. The local geometry of finite mixtures. Transaction of the American Mathematical Society. 2014

  277. [285]

    Basu and R

    S. Basu and R. Pollack and M. Roy. Algorithms in real algebraic geometry. 2006

  278. [286]

    R. A. Jacobs and M. I. Jordan and S. J. Nowlan and G. E. Hinton. Adaptive mixtures of local experts. Neural Computation. 1991

  279. [287]

    M. I. Jordan and L. Xu. Convergence results for the EM approach to mixtures of experts architectures. Neural Networks. 1995

  280. [288]

    Heinrich and J

    P. Heinrich and J. Kahn. Strong identifiability and optimal minimax rates for finite mixture estimation. Annals of Statistics. 2017+

  281. [289]

    Balakrishnan and M

    S. Balakrishnan and M. J. Wainwright and B. Yu , Journal =. Statistical guarantees for the

  282. [290]

    Dwivedi and N

    R. Dwivedi and N. Ho and K. Khamaru and M. J. Wainwright and M. I. Jordan and B. Yu , Journal =. Singularity, misspecification, and the convergence rate of

  283. [291]

    Dwivedi and N

    R. Dwivedi and N. Ho and K. Khamaru and M. J. Wainwright and M. I. Jordan and B. Yu , Date-Modified =. AISTATS , Title =

  284. [292]

    Optimal Estimation of

    Wu, Yihong and Yang, Pengkun , year =. Optimal Estimation of. The Annals of Statistics , volume =

  285. [293]

    The Annals of Statistics , volume=

    Strong identifiability and optimal minimax rates for finite mixture estimation , author=. The Annals of Statistics , volume=

  286. [294]

    Jordan , title =

    Nhat Ho and Chiao-Yu Yang and Michael I. Jordan , title =. Journal of Machine Learning Research , year =

  287. [295]

    Mixtures of Experts Unlock Parameter Scaling for Deep

    Johan Samir Obando Ceron and Ghada Sokar and Timon Willi and Clare Lyle and Jesse Farebrother and Jakob Nicolaus Foerster and Gintare Karolina Dziugaite and Doina Precup and Pablo Samuel Castro , booktitle=. Mixtures of Experts Unlock Parameter Scaling for Deep

  288. [296]

    Electronic Journal of Statistics , volume=

    Strong identifiability and parameter learning in regression with heterogeneous response , author=. Electronic Journal of Statistics , volume=. 2025 , publisher=

  289. [297]

    On posterior contraction of parameters and interpretability in

    Guha, Aritra and Ho, Nhat and Nguyen, XuanLong , journal=. On posterior contraction of parameters and interpretability in. 2021 , publisher=

  290. [298]

    Jiang and M

    W. Jiang and M. A. Tanner. On the identifiability of mixtures-of-experts. Neural Networks. 1999

  291. [299]

    M. I. Jordan and R. A. Jacobs. Hierarchical mixtures of experts and the EM algorithm. Neural Computation. 1994

  292. [300]

    Boyd and L

    S. Boyd and L. Vandenberghe. Convex Optimization. 2004

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.