REVIEW 3 major objections 5 minor 2 cited by
A Survey of Dimension Estimation Methods
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read No single dimension estimator works across all datasets; most need careful tuning and many overfit benchmarks.
desk verdict A useful, honest benchmark of dimension estimators whose 'overfitting' claim overstates the evidence: it measures hyperparameter sensitivity, not held-out generalization, but the empirical work is solid and deserves peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central organizing framework is the categorization of estimators by the geometric information they use: tangential estimators that detect a local affine tangent structure, parametric estimators that rely on dimension-dependent probability distributions of distances or angles, and estimators based on topological or metric invariants such as minimum spanning trees, k-nearest-neighbour graphs, and magnitude. This categorization is what allows the authors to explain why different estimators fail differently on curvature, noise, boundary effects, and high dimensional throttling, where the neighbourhood size k sets a hard ceiling on the dimension an estimator can return.
What would settle it
A concrete falsifier would be to find a single estimator with a fixed hyperparameter setting that performs accurately on the same benchmark manifolds across a wide range of intrinsic dimensions, curvatures, noise levels, and out-of-distribution datasets, contradicting the paper's claim that no universal estimator or hyperparameter set exists.
Extended reading notes
Core claim
The paper's core claim is stated in its conclusion: "There is no single estimator or set of hyperparameters that can perform well across all settings." This is established empirically by benchmarking a wide cohort of estimators on synthetic manifolds, and by showing that the performance of most estimators depends heavily on the choice of hyperparameters, with the best choice often varying from dataset to dataset. The authors also find that most estimators underestimate intrinsic dimension on high-dimensional datasets, while overestimation can occur on datasets such as SO(n), and that no estimator performs satisfactorily on non-linear manifolds of dimension above six.
Load-bearing premise
The evaluation assumes that the chosen benchmark manifolds and the range of hyperparameters tested represent how dimension estimators will actually be used in practice.
Editorial extensions
If this is right
- Practitioners should not trust a single dimension estimator with default hyperparameters and should instead use a range of estimators and hyperparameters to assess how sensitive the inferred dimension is to modelling assumptions.
- Benchmark results obtained by tuning hyperparameters to the test set should not be interpreted as a predictor of real-world performance, because the paper shows that overfitting to benchmark datasets is frequent.
- On high-dimensional manifolds, tangential estimators face a throttling limit that grows with dimension, and slope-inference estimators can even produce negative estimates, so caution is required.
- For practitioners with limited data, global estimators are generally more suitable, while local estimators require more samples because each neighbourhood must contain enough points.
- Future work on automated or adaptive hyperparameter selection is needed, since the paper finds that no simple rule performs well universally.
Reading between the lines
- The paper's conclusion implies that the apparent accuracy of an estimator on a benchmark should be read as a statement about that benchmark and that tuning process, not about the estimator alone.
- A systematic rule for choosing hyperparameters, for instance based on features of the data before estimation, would be a direct extension of this work and could be tested on the same benchmark suite.
- The finding that slightly positively curved surfaces can be easier for lPCA than flat ones suggests a testable extension: whether this bias generalizes to other tangential estimators or depends on the specific aggregation and thresholding.
- The paper's evidence that estimators differ greatly in variance across random samples implies that reporting a single number for intrinsic dimension without uncertainty is misleading; interval estimates could be derived from running estimators across hyperparameters and subsamples.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper surveys intrinsic dimension estimation methods, organizing them into tangential, parametric, and topological/metric families, and extends the scikit-dimension package with new estimators. The empirical core benchmarks thirteen estimators on synthetic manifolds with intrinsic dimensions 1-70 and ambient dimensions 3-96, using four sample sizes, twenty repeated point sets per condition, and multiple hyperparameter choices. The authors report estimator-specific tables, noise and curvature experiments, a throttling analysis, and qualitative assessments. The central claims are that no single estimator or hyperparameter setting works across all settings and that frequent 'overfitting' in hyperparameter identification suggests limited generalization.
Significance. If the results are read as a sensitivity analysis of the tested estimators within the stated hyperparameter grids, the paper is a substantial and useful contribution. The benchmark is unusually careful: twenty repetitions, standard deviations, four sample sizes, and transparent tables; the authors also disclose in Appendix C that the 'best' hyperparameter is selected with knowledge of the true dimension. The code contribution to scikit-dimension is valuable. The main weakness is interpretive: the abstract and Section 3.3 present the oracle-based sensitivity comparison as evidence of 'overfitting' and poor generalization, which is not what the protocol measures.
major comments (3)
- [Abstract; Section 3.3; Appendix C] The claim that 'overfitting is frequent' and that 'many estimators may not generalise well beyond the datasets on which they have been tested' is not supported by the experimental protocol. In Appendix C, the 'best' hyperparameter for estimator E on manifold M is selected as the minimizer of |d_hat(E,M,H) - d(M)| and the reported performance uses the same 20 point sets used for selection; the 'med abs' and 'med rel' columns use a hyperparameter chosen by minimizing median absolute or relative error against the true dimensions over the same benchmark manifolds. This measures hyperparameter sensitivity and the tension between per-dataset and global choices, not generalization to new samples or new manifolds. To support the abstract's wording, the authors should add a held-out evaluation: tune H on one collection of samples and evaluate on independent samples from the same manifold, and apply H_best(E,M1) to other benchmark manifolds. Without such a test, the evidence supports only 'no single hyperparameter setting is optimal across benchmarks'.
- [Table 3; Appendix C] The qualitative column 'No tailoring of params' in Table 3 is based on the 'med abs' and 'med rel' hyperparameter choices, which are themselves obtained by minimizing the median absolute or relative error against the true dimensions across the benchmark manifolds (Appendix C). This is not an untuned or default setting; it is a globally oracle-tuned fixed parameter. The ticks in Table 3 therefore overstate the 'no tailoring' property. Please either use the scikit-dimension default parameters as the fixed baseline, or rename the column to something like 'insensitivity to a global hyperparameter choice' and adjust the discussion.
- [Section 3.3; Table C1] The statement in Section 3.3 that lPCA 'appears to give the correct dimension every time' is contradicted by Table C1, where the best-hyperparameter estimate for M10d Cubic (true dimension 70) is 55.2 at every sample size. The caveat about hyperparameter tuning does not resolve the contradiction, because the quoted value is already the best over the grid. Please correct the claim or restrict it to the datasets on which lPCA actually succeeds.
minor comments (5)
- [Abstract; Remark 3.2] The abstract's generalizations should carry the qualifier that the empirical conclusions are relative to the hyperparameter grids listed in Appendix D, as the authors themselves acknowledge in Remark 3.2; as written, 'overfitting is frequent' reads as a statement about all hyperparameter ranges.
- [Section 2.3.1] The reference to 'eq. (14)' in the paragraph after the Costa and Hero discussion is a forward reference to Theorem 2.5; please renumber or point to eq. (10), where the Renyi-entropy connection is first visible.
- [Section 3.4.2] The WODCap throttling calculation would benefit from a cleaner derivation: the displayed expression for S(d) mixes Gamma-function notation and the resulting approximation for k is hard to parse; in particular, the notation involving Gamma((d+1)/2, 1/2) should presumably be Gamma((d+1)/2).
- [Section 4; Algorithm 1; Table C10] There are several typographical errors: 'is is' in Section 4, 'Datas set' in Algorithm 1, 'M20 Norm' in the Table C10 caption, and 'Theorem 2.6' in Figure 6 and Section 3.3 where Remark 2.6 is meant.
- [Table 2] The row for M2/M9 ('Affine 3to5 and M9 Affine 3, 20 5, 20') is hard to read; please reformat the table or caption to make the manifold, dimension, and ambient dimension columns unambiguous.
Circularity Check
No circularity: the paper is an empirical benchmark survey whose conclusions are measurements, not derivations from fitted inputs; the 'overfitting' inference is under-supported but not circular.
full rationale
This is a survey and benchmark study, not a derivation chain. The central empirical claims - that no single estimator or hyperparameter set dominates, and that estimators are sensitive to hyperparameter choice - are measured quantities reported from experiments on standard benchmark manifolds, not quantities derived from fitted inputs. The oracle-tuning protocol in Appendix C selects H_best(E,M) = argmin_H |dhat(E,M,H) - d(M)| and compares it with a fixed H_abs chosen across manifolds; this measures per-dataset hyperparameter sensitivity and the tension between per-dataset and global tuning. It is a measurement protocol, and the paper does not present the resulting best-case performance as a prediction about new data. The abstract's wording that 'overfitting is frequent, indicating that many estimators may not generalise well beyond the datasets on which they have been tested' goes beyond what the in-sample protocol can establish - no held-out evaluation is performed - but that is an evidentiary weakness, not circularity: the inference is not forced by the construction, it is an unverified extrapolation. Remark 3.2 itself acknowledges that 'our empirical study is subject to our particular choices of hyperparameter ranges,' explicitly flagging the scope limitation. Theoretical results (Theorems 2.1, 2.2, 2.5, and eqs. (6)-(9)) are imported from external literature (Kleindessner and von Luxburg, Beardwood-Halton-Hammersley, Steele, Schweinhart, Kozma-Lotker-Stupp, Yukich), not from the authors' own prior work; I find no load-bearing self-citations. The survey's taxonomy (tangential, parametric, topological/metric) is presented as an organizing classification, not as a new unification, and it does not rename a known result. The throttling and error-propagation analyses in Sections 3.4.2 and 2.6 are derived algebraically from the estimators' published definitions. Overall, the paper is self-contained as an empirical study; the 'overfitting' claim would need a held-out test to be fully supported, but no step of the paper's argument reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (1)
- Hyperparameter grids and defaults per estimator =
Varied per method; ranges listed in Appendix D
assumptions (3)
- domain assumption The finite sample X is drawn i.i.d. from a distribution supported on a compact smooth submanifold M^d of R^D.
- standard math Classical asymptotic results for Euclidean functionals correctly describe MST and kNN graph scaling used by PH0 and KNN estimators.
- domain assumption The standard scikit-dimension benchmark manifolds with known intrinsic dimension provide a representative testbed for practical dimension estimation.
Cite this review
Pith. "Pith review of A Survey of Dimension Estimation Methods." pith.science (2026). https://pith.science/paper/LUSFVIHX
@misc{pith2026250713887,
author = {Pith},
title = {Pith review of: A Survey of Dimension Estimation Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/LUSFVIHX}},
note = {Machine review of arXiv:2507.13887}
}
read the original abstract
It is a standard assumption that datasets in high dimension have an internal structure which means that they in fact lie on, or near, subsets of a lower dimension. In many instances it is important to understand the real dimension of the data, hence the complexity of the dataset at hand. A great variety of dimension estimators have been developed to find the intrinsic dimension of the data but there is little guidance on how to reliably use these estimators. This survey reviews a wide range of dimension estimation methods, categorising them by the geometric information they exploit: tangential estimators which detect a local affine structure; parametric estimators which rely on dimension-dependent probability distributions; and estimators which use topological or metric invariants. The paper evaluates the performance of these methods, as well as investigating varying responses to curvature and noise. Key issues addressed include robustness to hyperparameter selection, sample size requirements, accuracy in high dimensions, precision, and performance on non-linear geometries. In identifying the best hyperparameters for benchmark datasets, overfitting is frequent, indicating that many estimators may not generalise well beyond the datasets on which they have been tested.
Figures
Figures from the paper (16 more)
Forward citations
Cited by 2 Pith papers
-
HyperShadow: A Benchmark for Detecting 3D Projections of Higher-Dimensional Spatial Objects
Shadows of 4D–6D objects projected to 3D are detectable by learned point-cloud models and by a zero-parameter rigidity residual, but not by intrinsic-dimension estimation.
-
A general framework for adaptive nonparametric dimensionality reduction
Using ABIDE's local neighborhood sizes and intrinsic dimension as plug-in hyperparameters improves LLE, spectral clustering, and UMAP embeddings on benchmark datasets.
Reference graph
Works this paper leans on
-
[1]
A Fractal Dimension for Measures via Persistent Homology
Henry Adams, Elin Farnell, Manuchehr Aminian, Michael Kirby, Joshua Mirth, Rachel Neville, Chris Peterson, and Clayton Shonkwiler. A Fractal Dimension for Measures via Persistent Homology. In Nils A Baas, Gunnar E Carlsson, Gereon Quick, Markus Szymik, and Marius Thaule, editors, Topo- logical Data Analysis, pages 1–31, Cham, 2020. Springer International ...
2020
-
[2]
Estimating the effective dimension of large biological datasets using Fisher separability analysis
Luca Albergante, Jonathan Bac, and Andrei Zinovyev. Estimating the effective dimension of large biological datasets using Fisher separability analysis. In 2019 International Joint Conference on Neural Networks (IJCNN) , pages 1–8, Budapest, 2019. IEEE
2019
-
[3]
Houle, Ken ichi Kawarabayashi, and Michael Nett
Laurent Amsaleg, Oussama Chelly, Teddy Furon, St´ ephane Girard, Michael E. Houle, Ken ichi Kawarabayashi, and Michael Nett. Extreme-value-theoretic estimation of local intrinsic dimension- ality. Data Mining and Knowledge Discovery , 32(6):1768–1805, 11 2018
2018
-
[4]
Intrinsic Dimensionality Estimation within Tight Localities
Laurent Amsaleg, Oussama Chelly, Michael E Houle, Miloˇ s Radovanovi, and Weeris Treeratanajaru. Intrinsic Dimensionality Estimation within Tight Localities. In Proceedings of the 2019 SIAM Inter- national Conference on Data Mining (SDM) , pages 181–189, Calgary, 2019. SIAM
2019
-
[5]
Metric Space Magnitude and Generalisation in Neural Networks
Rayna Andreeva, Katharina Limbeck, Bastian Rieck, and Rik Sarkar. Metric Space Magnitude and Generalisation in Neural Networks. In Timothy Doster, Tegan Emerson, Henry Kvinge, Nina Miolane, Mathilde Papillon, Bastian Rieck, and Sophia Sanborn, editors, Proceedings of 2nd Annual Workshop on Topology, Algebra, and Geometry in Machine Learning (TAG-ML) , vol...
2023
-
[6]
Mirkes, Alexander N
Jonathan Bac, Evgeny M. Mirkes, Alexander N. Gorban, Ivan Tyukin, and Andrei Zinovyev. Scikit- Dimension: A Python Package for Intrinsic Dimension Estimation. Entropy, 23(10):1368, 10 2021
2021
-
[7]
Schwartz
Mukund Balasubramanian and Eric L. Schwartz. The isomap algorithm and topological stability. Science, 295(5552), 2002
2002
-
[8]
Jillian Beardwood, J. H. Halton, and J. M. Hammersley. The shortest path through many points. Mathematical Proceedings of the Cambridge Philosophical Society , 55(4):299–327, 10 1959
1959
Show all 118 references
-
[9]
Er¨ oss, Andr´ as Telcs, and Zolt´ an Somogyv´ ari
Zsigmond Benk˝ o, Marceli Stippinger, Roberta Rehus, Attila Bencze, D´ aniel Fab´ o, Bogl´ arka Hajnal, Lor´ and G. Er¨ oss, Andr´ as Telcs, and Zolt´ an Somogyv´ ari. Manifold-adaptive dimension estimation revisited. PeerJ Computer Science , 8, 2022
2022
-
[10]
Intrinsic Dimension, Persistent Ho- mology and Generalization in Neural Networks
Tolga Birdal, Aaron Lou, Leonidas J Guibas, and Umut Simsekli. Intrinsic Dimension, Persistent Ho- mology and Generalization in Neural Networks. In M Ranzato, A Beygelzimer, Y Dauphin, P S Liang, and J Wortman Vaughan, editors, Advances in Neural Information Processing Systems...
2021
-
[11]
Bayesian PCA
Christopher Bishop. Bayesian PCA. Advances in Neural Information Processing Systems , 11, 1998
1998
-
[12]
Bishop and Yuval Peres
Christopher J. Bishop and Yuval Peres. Fractals in Probability and Analysis . Cambridge University Press, 2016
2016
-
[13]
Intrinsic Dimension Estimation Using Wasserstein Distance
Adam Block, Zeyu Jia, Yury Polyanskiy, and Alexander Rakhlin. Intrinsic Dimension Estimation Using Wasserstein Distance. Journal of Machine Learning Research , 23:1–37, 2022
2022
-
[14]
Remarks on Parallel Analysis
Andreas Buja and Nermin Eyuboglu. Remarks on Parallel Analysis. Multivariate Behavioral Research, 27(4):509–540, 10 1992
1992
-
[15]
Data dimensionality estimation methods: a survey
Francesco Camastra. Data dimensionality estimation methods: a survey. Pattern Recognition , 36(12):2945–2954, 12 2003
2003
-
[16]
Intrinsic dimension estimation: Advances and open prob- lems
Francesco Camastra and Antonino Staiano. Intrinsic dimension estimation: Advances and open prob- lems. Information Sciences, 328:26–41, 1 2016. 39
2016
-
[17]
Intrinsic Dimension Estimation of Data: An Approach Based on Grassberger–Procaccia’s Algorithm
Francesco Camastra and Alessandro Vinciarelli. Intrinsic Dimension Estimation of Data: An Approach Based on Grassberger–Procaccia’s Algorithm. Neural Processing Letters, 14(1):27–34, 8 2001
2001
-
[18]
Campadelli, E
P. Campadelli, E. Casiraghi, C. Ceruti, and A. Rozza. Intrinsic Dimension Estimation: Relevant Techniques and a Benchmark Framework. Mathematical Problems in Engineering , 2015:1–21, 2015
2015
-
[19]
Abanov, Jeffrey Berger, Cameron J
Luca Candelori, Alexander G. Abanov, Jeffrey Berger, Cameron J. Hogan, Vahagn Kirakosyan, Kharen Musaelian, Ryan Samson, James E.T. Smith, Dario Villani, Martin T. Wells, and Mengjia Xu. Robust estimation of the intrinsic dimension of data sets with quantum cognition machine l...
2025
-
[20]
DANCo: An intrinsic dimensionality estimator exploiting angle and norm concentration
Claudio Ceruti, Simone Bassis, Alessandro Rozza, Gabriele Lombardi, Elena Casiraghi, and Paola Campadelli. DANCo: An intrinsic dimensionality estimator exploiting angle and norm concentration. Pattern Recognition, 47(8):2569–2581, 2014
2014
-
[21]
Dimension Detection via Slivers
Siu-Wing Cheng and Man-Kwun Chiu. Dimension Detection via Slivers. In Proceedings of the Twen- tieth Annual ACM-SIAM Symposium on Discrete Algorithms , pages 1001–1010. Society for Industrial and Applied Mathematics, 2009
2009
-
[22]
Costa and Alfred O
Jose A. Costa and Alfred O. Hero. Entropic graphs for manifold learning. In Conference Record of the Asilomar Conference on Signals, Systems and Computers , volume 1, 2003
2003
-
[23]
Costa and Alfred O
Jose A. Costa and Alfred O. Hero. Geodesic entropic graphs for dimension and entropy estimation in manifold learning. IEEE Transactions on Signal Processing , 52(8):2210–2221, 8 2004
2004
-
[24]
Costa and Alfred O
Jose A. Costa and Alfred O. Hero. Learning intrinsic dimension and intrinsic entropy of high- dimensional datasets. In European Signal Processing Conference, volume 06-10-September-2004, 2015
2004
-
[25]
J. M. Craddock and C. R. Flood. Eigenvectors for representing the 500 mb geopotential surface over the Northern Hemisphere. Quarterly Journal of the Royal Meteorological Society , 95(405):576–593, 7 1969
1969
-
[26]
The generalized ratios intrinsic dimension estimator
Francesco Denti, Diego Doimo, Alessandro Laio, and Antonietta Mira. The generalized ratios intrinsic dimension estimator. Scientific Reports, 12(1):20005, 11 2022
2022
-
[27]
Beyond the noise: intrinsic dimension estimation with optimal neighbourhood identification
Antonio Di Noia, Iuri Macocco, Aldo Glielmo, Alessandro Laio, and Antonietta Mira. Beyond the noise: intrinsic dimension estimation with optimal neighbourhood identification. arXiv:2405.15132v2, 5 2024
2024 arXiv
-
[28]
Metric Space Spread, Intrinsic Dimension and the Manifold Hypothesis
Kevin Dunne. Metric Space Spread, Intrinsic Dimension and the Manifold Hypothesis. arXiv:2308.01382, 8 2023
2023 arXiv
-
[29]
Generalization bounds using data- dependent fractal dimensions
Benjamin Dupuis, George Deligiannidis, and Umut Simsekli. Generalization bounds using data- dependent fractal dimensions. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engel- hardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International...
2023
-
[30]
Eckmann and D
J.-P. Eckmann and D. Ruelle. Fundamental limitations for estimating dimensions and Lyapunov exponents in dynamical systems. Physica D: Nonlinear Phenomena , 56(2-3):185–187, 5 1992
1992
-
[31]
Intrinsic dimension estimation for locally under- sampled data
Vittorio Erba, Marco Gherardi, and Pietro Rotondo. Intrinsic dimension estimation for locally under- sampled data. Scientific Reports, 9(1):17133, 11 2019
2019
-
[32]
Estimating the intrinsic dimension of datasets by a minimal neighborhood information
Elena Facco, Maria d’Errico, Alex Rodriguez, and Alessandro Laio. Estimating the intrinsic dimension of datasets by a minimal neighborhood information. Scientific Reports, 7(1):12140, 9 2017
2017
-
[33]
Alternative Definitions of Dimension
Kenneth Falconer. Alternative Definitions of Dimension. In Fractal Geometry, chapter 3, pages 39–58. John Wiley & Sons, Ltd, 2003. 40
2003
-
[34]
Intrinsic dimension estimation of data by principal component analysis
Mingyu Fan, Nannan Gu, Hong Qiao, and Bo Zhang. Intrinsic dimension estimation of data by principal component analysis. arXiv, 2 2010
2010
-
[35]
Manifold-adaptive dimension estimation
Amir massoud Farahmand, Csaba Szepesv´ ari, and Jean-Yves Audibert. Manifold-adaptive dimension estimation. In Proceedings of the 24th international conference on Machine learning , pages 265–272, New York, NY, USA, 6 2007. ACM
2007
-
[36]
Serge Frontier. ´Etude de la d´ ecroissance des valeurs propres dans une analyse en composantes prin- cipales: Comparaison avec le mod` ele du bˆ aton bris´ e.Journal of Experimental Marine Biology and Ecology, 25(1):67–75, 11 1976
1976
-
[37]
Fukunaga and D.R
K. Fukunaga and D.R. Olsen. An Algorithm for Finding Intrinsic Dimensionality of Data. IEEE Transactions on Computers, C-20(2):176–183, 2 1971
1971
-
[38]
Supervised nonlinear dimensionality reduction for visualization and classification
Xin Geng, De Chuan Zhan, and Zhi Hua Zhou. Supervised nonlinear dimensionality reduction for visualization and classification. IEEE Transactions on Systems, Man, and Cybernetics, Part B: Cy- bernetics, 35(6):1098–1107, 12 2005
2005
-
[39]
Elements of Dimensionality Re- duction and Manifold Learning
Benyamin Ghojogh, Mark Crowley, Fakhri Karray, and Ali Ghodsi. Elements of Dimensionality Re- duction and Manifold Learning . Springer International Publishing, Cham, 2023
2023
-
[40]
Persistent magnitude
Dejan Govc and Richard Hepworth. Persistent magnitude. Journal of Pure and Applied Algebra , 225(3):106517, 3 2021
2021
-
[41]
Accurate Estimation of the Intrinsic Dimension Using Graph Distances: Unraveling the Geometric Complexity of Datasets
Daniele Granata and Vincenzo Carnevale. Accurate Estimation of the Intrinsic Dimension Using Graph Distances: Unraveling the Geometric Complexity of Datasets. Scientific Reports, 6, 8 2016
2016
-
[42]
Measuring the strangeness of strange attractors
Peter Grassberger and Itamar Procaccia. Measuring the strangeness of strange attractors. Physica D: Nonlinear Phenomena, 9(1):189–208, 1983
1983
-
[43]
Comparison theorems for the volumes of tubes as generalizations of the Weyl tube formula
Alfred Gray. Comparison theorems for the volumes of tubes as generalizations of the Weyl tube formula. Topology, 21(2):201–228, 1982
1982
-
[44]
Intrinsic Dimensionality Estimation of Submanifolds in $Rˆd$
Hein Matthias and Audibert Jean-Yves. Intrinsic Dimensionality Estimation of Submanifolds in $Rˆd$. In Proceedings of the 22nd International Conference on Machine Learning, pages 289–296, Bonn, 2005. Association for Computing Machinery
2005
-
[45]
Intrinsic dimensionality estimation using Normalizing Flows
Christian Horvat and Jean-Pascal Pfister. Intrinsic dimensionality estimation using Normalizing Flows. In S Koyejo, S Mohamed, A Agarwal, D Belgrave, K Cho, and A Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 12225–12236. Curran Associates, I...
2022
-
[46]
Manifold Hypothesis in Data Analysis: Double Geometrically-Probabilistic Approach to Manifold Dimension Estimation
Alexander Ivanov, Gleb Nosovskiy, Alexey Chekunov, Denis Fedoseev, Vladislav Kibkalo, Mikhail Nikulin, Fedor Popelenskiy, Stepan Komkov, Ivan Mazurenko, and Aleksandr Petiushko. Manifold Hypothesis in Data Analysis: Double Geometrically-Probabilistic Approach to Manifold Dimen...
2021 arXiv
-
[47]
Johnson and Joram Lindenstrauss
William B. Johnson and Joram Lindenstrauss. Extensions of Lipschitz mappings into a Hilbert space. In Richard Beals, Anatole Beck, Alexandra Bellow, and Arshag Hajian, editors, Conference on Modern Analysis and Probability, pages 189–206. American Mathematical Society, Provide...
1984
-
[48]
Low bias local intrinsic dimension estimation from expected simplex skewness
Kerstin Johnsson, Charlotte Soneson, and Magnus Fontes. Low bias local intrinsic dimension estimation from expected simplex skewness. IEEE Transactions on Pattern Analysis and Machine Intelligence , 37(1):196–202, 1 2015
2015
-
[49]
I. T. Jolliffe. Principal Component Analysis. Springer-Verlag, New York, 2002
2002
-
[50]
Manifold Diffusion Geometry: Curvature, Tangent Spaces, and Dimension
Iolo Jones. Manifold Diffusion Geometry: Curvature, Tangent Spaces, and Dimension. arXiv:2411.04100, 11 2024. 41
2024
-
[51]
Bayesian Estimation Approaches for Local Intrinsic Dimensionality
Zaher Joukhadar, Hanxun Huang, and Sarah Monazam Erfani. Bayesian Estimation Approaches for Local Intrinsic Dimensionality. In Edgar Ch´ avez, Benjamin Kimia, Jakub Lokoˇ c, Marco Patella, and Jan Sedmidubsky, editors, Similarity Search and Applications , volume 15268 of Lectu...
2025
-
[52]
The Application of Electronic Computers to Factor Analysis
Henry F Kaiser. The Application of Electronic Computers to Factor Analysis. Educational and Psychological Measurement, 20(1):141–151, 1960
1960
-
[53]
A Geometric View of Data Complexity: Efficient Local Intrinsic Dimension Estimation with Diffusion Models
Hamidreza Kamkari, Brendan Leigh, Ross Rasa Hosseinzadeh, Jesse C Cresswell, and Gabriel Loaiza- Ganem. A Geometric View of Data Complexity: Efficient Local Intrinsic Dimension Estimation with Diffusion Models. Technical report, 2024
2024
-
[54]
Fractal-Based Methods as a Technique for Estimating the Intrinsic Dimensionality of High-Dimensional Data: A Survey
Rasa Karbauskaite and Gintautas Dzemyda. Fractal-Based Methods as a Technique for Estimating the Intrinsic Dimensionality of High-Dimensional Data: A Survey. Informatica. 2016;27(2):257-281, 27(2):257–281, 2016
2016
-
[55]
Additive autoencoder for dimension estimation.Neurocomput- ing, 551, 9 2023
Tommi K¨ arkk¨ ainen and Jan H¨ anninen. Additive autoencoder for dimension estimation.Neurocomput- ing, 551, 9 2023
2023
-
[56]
Is magnitude ’generically continuous’ for finite metric spaces? arXiv:2501.08745, 1 2025
Hirokazu Katsumasa, Emily Roff, and Masahiko Yoshinaga. Is magnitude ’generically continuous’ for finite metric spaces? arXiv:2501.08745, 1 2025
2025 arXiv
-
[57]
Intrinsic Dimension Estimation Using Packing Numbers
Bal´ azs K´ egl. Intrinsic Dimension Estimation Using Packing Numbers. In S Becker, S Thrun, and K Obermayer, editors, Advances in Neural Information Processing Systems 15 , pages 697–704. MIT Press, 2002
2002
-
[58]
The central limit theorem for weighted minimal spanning trees on random points
Harry Kesten and Sungchul Lee. The central limit theorem for weighted minimal spanning trees on random points. The Annals of Applied Probability , 6(2):495–527, 1996
1996
-
[59]
Minimax rates for estimating the dimension of a manifold
Jisu Kim, Alessandro Rinaldo, and Larry Wasserman. Minimax rates for estimating the dimension of a manifold. Journal of Computational Geometry , 10(1):42–95, 2019
2019
-
[60]
Dimensionality estimation without distances
Matth¨ aus Kleindessner and Ulrike Luxburg. Dimensionality estimation without distances. In Guy Lebanon and S. V. N. Vishwanathan, editors, Proceedings of the Eighteenth International Conference on Artificial Intelligence and Statistics , pages 471–479, 2015
2015
-
[61]
The minimal spanning tree and the upper box dimension
Gady Kozma, Zvi Lotker, and Gideon Stupp. The minimal spanning tree and the upper box dimension. Proceedings of the American Mathematical Society, 134(4):1183–1187, 9 2005
2005
-
[62]
On the connectivity threshold for general uniform metric spaces
Gady Kozma, Zvi Lotker, and Gideon Stupp. On the connectivity threshold for general uniform metric spaces. Information Processing Letters, 110(10):356–359, 4 2010
2010
-
[63]
Simple correlation dimension estimator and its use to detect causality
Anna Krakovsk´ a and Martina Chvostekov´ a. Simple correlation dimension estimator and its use to detect causality. Chaos, Solitons and Fractals , 175, 10 2023
2023
-
[64]
J. B. Kruskal. Nonmetric Multidimensional Scaling: A Numerical Method. Psychometrika, 29(2):115– 129, 6 1964
1964
-
[65]
The concentration of measure phenomenon
Michel Ledoux. The concentration of measure phenomenon. Mathematical surveys and monographs, , 89, 2001
2001
-
[66]
On the asymptotic magnitude of subsets of Euclidean space
Tom Leinster and Simon Willerton. On the asymptotic magnitude of subsets of Euclidean space. Geometriae Dedicata, 164(1):287–310, 6 2013
2013
-
[67]
Maximum Likelihood Estimation of Intrinsic Dimension
Elizaveta Levina and Peter J Bickel. Maximum Likelihood Estimation of Intrinsic Dimension. In Advances in Neural Information Processing Systems 17 , 2004
2004
-
[68]
Tangent Space and Dimension Estimation with the Wasserstein Distance
Uzu Lim, Harald Oberhauser, and Vidit Nanda. Tangent Space and Dimension Estimation with the Wasserstein Distance. SIAM Journal on Applied Algebra and Geometry , 8:650–685, 10 2024. 42
2024
-
[69]
Metric space magnitude for evaluating the diversity of latent representations
Katharina Limbeck, Rayna Andreeva, Rik Sarkar, and Bastian Rieck. Metric space magnitude for evaluating the diversity of latent representations. In NIPS ’24: Proceedings of the 38th International Conference on Neural Information Processing Systems, pages 123911–123953, Vancouv...
-
[70]
Riemannian manifold learning
Tong Lin and Hongbin Zha. Riemannian manifold learning. IEEE Transactions on Pattern Analysis and Machine Intelligence , 30(5):796–809, 5 2008
2008
-
[71]
Mini- mum Neighbor Distance Estimators of Intrinsic Dimension
Gabriele Lombardi, Alessandro Rozza, Claudio Ceruti, Elena Casiraghi, and Paola Campadelli. Mini- mum Neighbor Distance Estimators of Intrinsic Dimension. In D Gunopulos, T Hofmann, D Malerba, and M Vazirgiannis, editors, Machine Learning and Knowledge Discovery in Databases. ...
2011
-
[72]
MacKay and Zoubin Ghahramani
David J.C. MacKay and Zoubin Ghahramani. Comments on ’Maximum Likelihood Estimation of Intrinsic Dimension’ by E. Levina and P. Bickel (2004), 2005
2004
-
[73]
Intrinsic Dimension Estimation for Discrete Metrics
Iuri Macocco, Aldo Glielmo, Jacopo Grilli, and Alessandro Laio. Intrinsic Dimension Estimation for Discrete Metrics. Physical Review Letters, 130(6), 2 2023
2023
-
[74]
Geometry of sets and measures in Euclidean spaces
Pertti Mattila. Geometry of sets and measures in Euclidean spaces . Cambridge University Press, Cambridge, 1995
1995
-
[75]
UMAP: Uniform Manifold Approximation and Projection
Leland McInnes, John Healy, Nathaniel Saul, and Lukas Großberger. UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software , 3(29), 2018
2018
-
[76]
Mark W. Meckes. Positive definite metric spaces. Positivity, 17(3):733–757, 9 2013
2013
-
[77]
Mark W. Meckes. Magnitude, Diversity, Capacities, and Dimensions of Metric Spaces. Potential Analysis, 42(2):549–572, 2 2015
2015
-
[78]
Thomas P. Minka. Automatic choice of dimensionality for PCA. In Advances in Neural Information Processing Systems, 2001
2001
-
[79]
Topology (2nd Edition)
James Raymond Munkres. Topology (2nd Edition). Prentice Hall, Inc, 2000
2000
-
[80]
Oganov and Mario Valle
Artem R. Oganov and Mario Valle. How to quantify energy landscapes of solids. Journal of Chemical Physics, 130(10), 2009
2009
-
[81]
Alpha magnitude
Miguel O’Malley, Sara Kalisnik, and Nina Otter. Alpha magnitude. Journal of Pure and Applied Algebra, 227(11):107396, 11 2023
2023
-
[82]
A Novel Approach for Intrinsic Di- mension Estimation
Kadir ¨Oz¸ coban, Murat Manguo˘ glu, and Emrullah Fatih Yetkin. A Novel Approach for Intrinsic Di- mension Estimation. arXiv:2503.09485v1 [cs.LG] 12 Mar 2025 , 3 2025
2025 arXiv
-
[83]
Papaioannou, Ronen Talmon, Ioannis G
Panagiotis G. Papaioannou, Ronen Talmon, Ioannis G. Kevrekidis, and Constantinos Siettos. Time- series forecasting using manifold learning, radial basis function interpolation, and geometric harmonics. Chaos, 32(8), 8 2022
2022
-
[84]
Pettis, Thomas A
Karl W. Pettis, Thomas A. Bailey, Anil K. Jain, and Richard C. Dubes. An Intrinsic Dimensionality Estimator from Near-Neighbor Information. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-1(1):25–37, 1979
1979
-
[85]
Intrinsic dimension estimation based on local adjacency information
Haiquan Qiu, Youlong Yang, and Benchong Li. Intrinsic dimension estimation based on local adjacency information. Information Sciences, 558:21–33, 5 2021
2021
-
[86]
Underestimation modification for intrinsic dimension estimation
Haiquan Qiu, Youlong Yang, and Hua Pan. Underestimation modification for intrinsic dimension estimation. Pattern Recognition, 140, 8 2023
2023
-
[87]
Intrinsic dimension estimation method based on correlation dimension and kNN method
Haiquan Qiu, Youlong Yang, and Saeid Rezakhah. Intrinsic dimension estimation method based on correlation dimension and kNN method. Knowledge-Based Systems, 235, 1 2022. 43
2022
-
[88]
A gen- eral and flexible method for signal extraction from single-cell RNA-seq data
Davide Risso, Fanny Perraudeau, Svetlana Gribkova, Sandrine Dudoit, and Jean Philippe Vert. A gen- eral and flexible method for signal extraction from single-cell RNA-seq data. Nature Communications, 9(1), 12 2018
2018
-
[89]
Rozza, G
A. Rozza, G. Lombardi, C. Ceruti, E. Casiraghi, and P. Campadelli. Novel high intrinsic dimensionality estimators. Machine Learning, 89(1-2):37–65, 10 2012
2012
-
[90]
IDEA: Intrinsic Dimension Estimation Algorithm
Alessandro Rozza, Gabriele Lombardi, Marco Rosa, Elena Casiraghi, and Paola Campadelli. IDEA: Intrinsic Dimension Estimation Algorithm. In Image Analysis and Processing – ICIAP 2011 . Springer, Berlin, Heidelberg., 2011
2011
-
[91]
A Nonlinear Mapping for Data Structure Analysis
John W Sammon. A Nonlinear Mapping for Data Structure Analysis. IEEE Transactions on Com- puters, C-18(5):401–09, 1969
1969
-
[92]
Estimate manifold dimensionality with LID, 12 2020
Aaron Schumacher. Estimate manifold dimensionality with LID, 12 2020
2020
-
[93]
Fractal dimension and the persistent homology of random geometric complexes
Benjamin Schweinhart. Fractal dimension and the persistent homology of random geometric complexes. Advances in Mathematics , 372:107291, 10 2020
2020
-
[94]
Persistent Homology and the Upper Box Dimension
Benjamin Schweinhart. Persistent Homology and the Upper Box Dimension. Discrete & Computational Geometry, 65(2):331–364, 2021
2021
-
[95]
Dimension Estimation Using Random Connection Models
Paulo Serra and Michel Mandjes. Dimension Estimation Using Random Connection Models. Journal of Machine Learning Research, 18(138):1–35, 2017
2017
-
[96]
Hausdorff dimension, heavy tails, and generalization in neural networks
Umut Simsekli, Ozan Sener, George Deligiannidis, and Murat A Erdogdu. Hausdorff dimension, heavy tails, and generalization in neural networks. Advances in Neural Information Processing Systems , 33:5138–5151, 2020
2020
-
[97]
Wasserstein Stability for Persistence Diagrams
Primoz Skraba and Katharine Turner. Wasserstein Stability for Persistence Diagrams. arXiv:2006.16824, 7 2025
2006 arXiv
-
[98]
Diffusion Models Encode the Intrinsic Dimension of Data Manifolds
Jan Pawel Stanczuk, Georgios Batzolis, Teo Deveney, and Carola-Bibiane Sch¨ onlieb. Diffusion Models Encode the Intrinsic Dimension of Data Manifolds. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, ...
2024
-
[99]
Michael Steele
J. Michael Steele. Growth Rates of Euclidean Minimal Spanning Trees with Power Weighted Edges. The Annals of Probability , 16(4), 10 1988
1988
-
[100]
On the Size and Recovery of Submatrices of Ones in a Random Binary Matrix
Xing Sun and Andrew B Nobel. On the Size and Recovery of Submatrices of Ones in a Random Binary Matrix. Journal of Machine Learning Research , 9(80):2431–2453, 2008
2008
-
[101]
Relative Intrinsic Di- mensionality Is Intrinsic to Learning
Oliver J Sutton, Qinghua Zhou, Alexander N Gorban, and Ivan Y Tyukin. Relative Intrinsic Di- mensionality Is Intrinsic to Learning. In Lazaros Iliadis, Antonios Papaleonidas, Plamen Angelov, and Chrisina Jayne, editors, Artificial Neural Networks and Machine Learning – ICANN 2...
2023
-
[102]
LIDL: Local Intrinsic Dimension Estimation Using Approximate Likelihood
Piotr Tempczyk, Rafa l Michaluk, Lukasz Garncarek, Przemys law Spurek, Jacek Tabor, and Adam Golinski. LIDL: Local Intrinsic Dimension Estimation Using Approximate Likelihood. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, edito...
2022
-
[103]
A Global Geometric Framework for Non- linear Dimensionality Reduction
Joshua B Tenenbaum, Vin de Silva, and John C Langford. A Global Geometric Framework for Non- linear Dimensionality Reduction. Science, 290:2319–2323, 2000
2000
-
[104]
G. V. Trunk. Statistical estimation of the intrinsic dimensionality of data collections. Information and Control, 12(5):508–525, 5 1968. 44
1968
-
[105]
Intrinsic dimension estimation for robust detection of ai-generated texts
Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Daniil Cherniavskii, Sergey Nikolenko, Evgeny Burnaev, Serguei Barannikov, and Irina Piontkovskaya. Intrinsic dimension estimation for robust detection of ai-generated texts. Advances in Neural Information Processing Sy...
2023
-
[106]
Ein Fixpunktsatz
Andrey Tychonoff. Ein Fixpunktsatz. Mathematische Annalen, 11:767–776, 1935
1935
-
[107]
Visualizing data using t-SNE
Laurens Van Der Maaten and Geoffrey Hinton. Visualizing data using t-SNE. Journal of Machine Learning Research, 9, 2008
2008
-
[108]
Infinite-Dimensional Topology
Jan van Mill. Infinite-Dimensional Topology. North-Holland, 1st edition, 1988
1988
-
[109]
Verveer and Robert P.W
Peter J. Verveer and Robert P.W. Duin. An Evaluation of Intrinsic Dimensionality Estimators. IEEE Transactions on Pattern Analysis and Machine Intelligence , 17(1):81–86, 1995
1995
-
[110]
Way, Michael Zietz, Vincent Rubinetti, Daniel S
Gregory P. Way, Michael Zietz, Vincent Rubinetti, Daniel S. Himmelstein, and Casey S. Greene. Compressing gene expression data using multiple latent space dimensionalities learns complementary biological representations. Genome Biology, 21(1), 5 2020
2020
-
[111]
Heuristic and computer calculations for the magnitude of metric spaces
Simon Willerton. Heuristic and computer calculations for the magnitude of metric spaces. arXiv:0910.5500, 10 2009
2009 arXiv
-
[112]
Conical dimension as an intrisic dimension estimator and its applications
Xin Yang, Sebastien Michea, and Hongyuan Zha. Conical dimension as an intrisic dimension estimator and its applications. In Chid Apte, David Skillicorn, Bing Liu, and Srinivasan Parthasarathy, editors, Proceedings of the 2007 SIAM International Conference on Data Mining , page...
2007
-
[113]
Adversarial Estimation of Topological Dimension with Harmonic Score Maps
Eric Yeats, Cameron Darwin, Frank Liu, and Hai Li. Adversarial Estimation of Topological Dimension with Harmonic Score Maps. arXiv:2312.06869, 12 2023
2023 arXiv
-
[114]
Joseph E. Yukich. Probability Theory of Classical Euclidean Optimization Problems . Springer Berlin, Heidelberg, 1998
1998
-
[115]
Abstract Latent-Space Construction for Analyzing Large Genomic Data Sets
Wenlan Zang. Abstract Latent-Space Construction for Analyzing Large Genomic Data Sets . PhD thesis, Yale University, 2021. 45 A Comparison of PH and KNN Since PH0 and KNN are derived from the common theory of Euclidean functionals, and are similar in construction, we highlight...
2021
-
[117]
The hyperparameter that minimises the difference between the estimated and intrinsic dimension ˆdbest(E, M) = min H∈HE | ˆd(E, M, H) − d(M )|
-
[118]
While a choice of hyperparameter might be optimal for a certain dataset, the same hyperparameter may lead to a poor performance on another
The performance of the estimator with a fixed choice of hyperparameter; this is either (a) The hyperparameter that minimises the median absolute error across the benchmark manifolds: ˆdabs(E, M) = ˆd(E, M, Habs(E)) Habs(E) = argmin H∈HE median M ∈M | ˆd(E, M, H) − d(M )| (b) T...
-
[2025]
Curran Associates Inc
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.