A semiparametric framework clusters high-dimensional elliptical data with heavy tails via cluster-specific centers, a common unknown radial generator, and a shared sparse precision matrix, with GEM algorithm and high-dimensional consistency guarantees.
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume =
8 Pith papers cite this work, alongside 542 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 8roles
background 2representative citing papers
Soft-MSM is a smooth, gradient-enabled version of the context-aware MSM distance for time series alignment that outperforms Soft-DTW alternatives in clustering and nearest-centroid classification.
GPS tracking across theme parks shows visitor movement forms a continuum rather than discrete types, diverges from self-reports, and reverses feature relationships from site to site, requiring local calibration.
A weighted K-means plus decision-tree pipeline learns multi-action policies from observational data and is applied to HCV treatment choices for HIV co-infected patients, finding a high-clearance subgroup and potential cost savings of CAN$3.6-4.9 million.
New hardware-usage-based similarity metrics can identify matching computational kernels between proxy applications and performance suites on both CPU and GPU systems.
A robust sparse clustering method uses spatial medians and automatic feature exclusion to achieve competitive accuracy and better stability than standard K-means on simulated heavy-tailed high-dimensional data.
A scaled base-K positional encoding maps high-dimensional discrete rows injectively into low-dimensional continuous coordinates with approximate Gaussian structure, enabling fast and accurate clustering.
An unsupervised-to-supervised ML pipeline on UK NDNS data discovers four dietary patterns, reproduces them with macro-F1 0.963 using a surrogate classifier, and interprets them via SHAP for potential clinical use.
citing papers explorer
-
Semiparametric Elliptical Mixture Clustering for High-Dimensional Data
A semiparametric framework clusters high-dimensional elliptical data with heavy tails via cluster-specific centers, a common unknown radial generator, and a shared sparse precision matrix, with GEM algorithm and high-dimensional consistency guarantees.
-
Soft-MSM: Differentiable Context-Aware Elastic Alignment for Time Series
Soft-MSM is a smooth, gradient-enabled version of the context-aware MSM distance for time series alignment that outperforms Soft-DTW alternatives in clustering and nearest-centroid classification.
-
Do Waders, Swimmers, and Divers Exist? A GPS-Based Pilot Study of Site-Dependent Visitor Movement in Theme Parks
GPS tracking across theme parks shows visitor movement forms a continuum rather than discrete types, diverges from self-reports, and reverses feature relationships from site to site, requiring local calibration.
-
Policy Learning with Observational Data: The Case of Hepatitis C Treatment for HIV/HCV Co-Infected Patients
A weighted K-means plus decision-tree pipeline learns multi-action policies from observational data and is applied to HCV treatment choices for HIV co-infected patients, finding a high-clearance subgroup and potential cost savings of CAN$3.6-4.9 million.
-
On Similarity of Computational Kernels in our Codes and Proxies
New hardware-usage-based similarity metrics can identify matching computational kernels between proxy applications and performance suites on both CPU and GPU systems.
-
Sparse $K$-spatial-median clustering for high-dimensional data
A robust sparse clustering method uses spatial medians and automatic feature exclusion to achieve competitive accuracy and better stability than standard K-means on simulated heavy-tailed high-dimensional data.
-
Data compression for fast dimension reduction and clustering of high-dimensional discrete data
A scaled base-K positional encoding maps high-dimensional discrete rows injectively into low-dimensional continuous coordinates with approximate Gaussian structure, enabling fast and accurate clustering.
-
An Explainable Unsupervised-to-Supervised Machine Learning Framework for Dietary Pattern Discovery Using UK National Dietary Survey Data
An unsupervised-to-supervised ML pipeline on UK NDNS data discovers four dietary patterns, reproduces them with macro-F1 0.963 using a surrogate classifier, and interprets them via SHAP for potential clinical use.