Pith. sign in

Manifold Hypothesis in Data Analysis: Double Geometrically-Probabilistic Approach to Manifold Dimension Estimation

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Manifold hypothesis states that data points in high-dimensional space actually lie in close vicinity of a manifold of much lower dimension. In many cases this hypothesis was empirically verified and used to enhance unsupervised and semi-supervised learning. Here we present new approach to manifold hypothesis checking and underlying manifold dimension estimation. In order to do it we use two very different methods simultaneously - one geometric, another probabilistic - and check whether they give the same result. Our geometrical method is a modification for sparse data of a well-known box-counting algorithm for Minkowski dimension calculation. The probabilistic method is new. Although it exploits standard nearest neighborhood distance, it is different from methods which were previously used in such situations. This method is robust, fast and includes special preliminary data transformation. Experiments on real datasets show that the suggested approach based on two methods combination is powerful and effective.

citation-role summary

other 1

citation-polarity summary

fields

stat.ML 1

years

2025 1

verdicts

CONDITIONAL 1

roles

other 1

polarities

unclear 1

representative citing papers

A Survey of Dimension Estimation Methods

stat.ML · 2025-07-18 · conditional · novelty 4.0

A broad benchmark of intrinsic dimension estimators shows that no single method or set of hyperparameters works across datasets, and tuned benchmark scores frequently indicate overfitting rather than transferable accuracy.

citing papers explorer

Showing 1 of 1 citing paper.

  • A Survey of Dimension Estimation Methods stat.ML · 2025-07-18 · conditional · none · ref 46 · internal anchor

    A broad benchmark of intrinsic dimension estimators shows that no single method or set of hyperparameters works across datasets, and tuned benchmark scores frequently indicate overfitting rather than transferable accuracy.