REVIEW 3 major objections 4 minor 194 references
Unsupervised Machine Learning for Scientific Discovery: Workflow and Best Practices
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Validation-driven workflow yields reliable unsupervised discovery.
desk verdict Useful synthesis of unsupervised best practices, but the case study abandons its own stability-based model selection, so the reliability claim is not demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is Algorithm 1, a stability-driven model-selection loop. For each candidate clustering method and each candidate number of clusters, it repeatedly draws two subsamples of the training data, randomly assigns a different preprocessing pipeline from a grid G to each, clusters both, and scores the agreement of the two clusterings on overlapping observations with the Adjusted Rand Index; the method and cluster count with the highest mean score are chosen. The chosen pipeline is then aggregated across all preprocessing versions by consensus clustering, producing a co-membership matrix that gives per-star local stability, and generalizability is measured by training a random forest on the training cluster labels and computing the agreement between its test-set predictions and test-set cluster labels. Throughout, the grid G encodes the reasonable analytical choices, such as quality-control thresholds, feature subsets, imputation methods, dimension-reduction settings, and clustering hyperparameters, so that the validation explicitly tests sensitivity to the analyst's judgment calls.
What would settle it
The reliability claim would be refuted if the same workflow, run on data from a different spectroscopic survey or on the same stars with a disjoint but equally reasonable set of preprocessing pipelines, produced groupings that barely overlap with the original clusters (low Adjusted Rand Index), or if the random-forest predictor of cluster labels scored near chance on an independent held-out sample.
Extended reading notes
Core claim
The paper's central claim is that unsupervised learning can be made a trustworthy engine for scientific discovery, but only when it is embedded in a validation-driven workflow rather than applied as a one-off algorithm. The workflow consists of translating a research goal into a validatable question, splitting data before preprocessing, exploring and preprocessing with several reasonable options, fitting a range of clustering and dimension-reduction models, selecting the method and cluster count that maximize stability across subsamples and across preprocessing pipelines, and then confirming that the chosen clusters generalize to a held-out test set. Applied to APOGEE globular-cluster stars, the workflow yields eight clusters of which four, namely clusters 1, 2, 6, and 8, are stable under alternative preprocessing choices and generalizable to held-out stars, and whose chemical signatures align with known stellar populations such as iron-rich metal-poor stars and nitrogen-rich stars. On this basis the paper asserts that its workflow has produced reliable and reproducible groupings of stellar abundances despite the abundance space being a continuous spectrum.
Load-bearing premise
The load-bearing premise is that a grouping that stays stable across the particular preprocessing and hyperparameter choices in the grid G, and that generalizes to a held-out test split, counts as scientifically real, an assumption the paper itself flags by acknowledging that the chemical abundance space of Milky Way stars is a continuous spectrum, so the discreteness of clusters is a domain assumption rather than a demonstrated fact.
Editorial extensions
If this is right
- Researchers can adopt a concrete template: split data before preprocessing, enumerate a grid of reasonable analytical choices, select clustering models by stability, and validate generalizability on the held-out split.
- Unsupervised results can be reported with per-cluster trust levels, since the workflow yields local stability and local generalizability scores for each cluster rather than a single global number.
- The case study demonstrates that a solution can be highly stable yet scientifically uninformative: the two-cluster spectral solution recapitulates the known iron-rich versus iron-poor division, so the workflow must be paired with domain knowledge to chase novel groupings.
- Four of the eight K-means clusters, namely clusters 1, 2, 6, and 8, are identified as stable and generalizable, with chemical signatures consistent with known stellar populations, giving astronomers concrete new groupings to investigate.
- Following the workflow can change the conclusion of a field: where prior studies doubted that reliable clustering of the APOGEE abundance space is possible, the workflow produces reproducible groupings, suggesting that the earlier pessimism stemmed partly from unvalidated pipelines.
Reading between the lines
- Because the case study shows that the most stable solution merely recapitulates the known iron-rich versus iron-poor split, a reader can infer that stability alone is insufficient and that the workflow should be paired with a novelty check against established knowledge, a pairing the paper illustrates but does not codify.
- The same validation logic transfers directly to other fields that cluster along continuous spectra, such as cancer cell states or soil types, which the paper mentions only briefly; a concrete extension would be to benchmark the workflow on one such dataset with the same stability and generalizability protocol.
- An implicit cost of the workflow is computational, since the grid of preprocessing, methods, and hyperparameters multiplies the number of fits; a natural extension is an adaptive grid search that concentrates runs in regions where stability is borderline.
- The neighbor-retention criterion used to drop UMAP in favor of t-SNE could, in principle, be reused as a general, dataset-independent check for choosing among dimension-reduction methods before clustering, though the paper only uses it for visualization and input selection.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a structured, model-agnostic workflow for unsupervised machine learning in scientific discovery, covering: formulating validatable questions, data preparation and exploration, using multiple modeling techniques, validation via stability and generalizability, and communication/documentation of results. The recommendations are illustrated with a case study on APOGEE globular-cluster stars, where the authors cluster stars by chemical abundance, select a model by stability (Algorithm 1), validate on held-out data, and identify four clusters (1, 2, 6, 8) as robust and scientifically interpretable. The manuscript also provides an interactive supplement, code, and data.
Significance. If the workflow is taken as a template, this is a timely and practically useful contribution: it addresses a real gap in guidance for unsupervised discovery, and it ships reproducible artifacts (code, data, interactive supplement) and a carefully executed case study. The emphasis on stability across preprocessing choices and on held-out generalizability is a valuable, transferable practice. However, the case study departs from its own model-selection protocol, and the manuscript's central claim that the workflow produced reliable scientific groupings is therefore not fully demonstrated. The paper is a credible candidate for publication after revision, provided the case study is reframed or the protocol is followed strictly.
major comments (3)
- [§3.3 (Modeling and Validation) and Algorithm 1] Algorithm 1 prescribes selecting the clustering method and number of clusters with the highest mean ARI across subsamples and preprocessing pipelines. The text in §3.3 reports that this procedure selects spectral clustering with k=2, and then explicitly sets that result aside ('we instead shift our interest towards the eight clusters generated by K-means') because it recapitulates the known iron-rich/iron-poor split. The K-means k=8 solution is the second-ranked model, chosen after inspecting results, not by the stated stability criterion. Consequently, the subsequent stability and generalizability metrics reported for the k=8 model do not demonstrate that the workflow produces reliable discoveries; they demonstrate that a scientifically interesting solution can be identified after exploratory model hunting. The paper should either present the k=2 spectral solution as the primary validated output of the workflow, or explicitly reframe the k=8 analysis as a secondary, hypothesis-generating step with appropriate caveats about selection bias.
- [§3.3 (Interpretation and Communication) and Figure 4C-D] The designation of clusters 1, 2, 6, and 8 as 'most robust' is made after inspecting the local-stability and local-generalizability results in Figure 4C-D. The same metrics are then used as evidence of reliability for these clusters. Because the clusters are selected on the basis of those very metrics, the reported values are optimistically biased by selection; the paper does not account for this (e.g., no pre-specified rule for 'most robust', no adjustment for multiple clusters, and no reporting of the full distribution of metrics across all clusters). The authors should either pre-register the selection rule, report metrics for all eight clusters without cherry-picking, or explicitly state that the selected clusters' metrics are conditional on post hoc selection and therefore should not be quoted as unbiased reliability estimates.
- [§3.3 (final paragraph)] The paper acknowledges that 'the chemical abundance space of stars in the Milky Way has been established as a continuous spectrum' but then claims that 'our workflow has produced reliable and reproducible groupings of stellar abundances, anchoring data-driven insights into chemical formations of our galaxy.' The stability and generalizability metrics validate that the clustering partitions are reproducible under data perturbations and preprocessing choices; they do not establish that discrete clusters correspond to real stellar populations, especially when the underlying abundance space is acknowledged to be continuous. The paper itself recognizes that 'further research is needed to scientifically validate these clusters,' but this caveat appears only in the discussion of the case study, not in the abstract or the concluding claim. The central claim should be tempered to distinguish statistical reproducibility from scientific validity, or the paper should explicitly state a domain assumption that discrete clusters are meaningful despite the continuous spectrum.
minor comments (4)
- [§3.3 (Modeling and Validation)] The phrase 'we shortlist two promising methods' is vague; Figure 3A presumably shows a full ranking, but the reader cannot determine why precisely these two are shortlisted or whether the choice is based on a threshold. A brief explanation of the shortlisting criterion would improve transparency.
- [§3.2.1, Table 1] The table lists UMAP as a dimension-reduction option, but §3.3 reports that UMAP is dropped from all subsequent analyses after the neighborhood-retention exploration. This is a reasonable decision, but it should be flagged as a deviation from the planned grid so that the set G used in Algorithm 1 is precisely defined (the text later says 'we drop UMAP from consideration in all subsequent analyses').
- [§2.1.4 (Generalizability)] The generalizability measure uses a random forest trained on training cluster labels to predict test cluster labels, but the paper does not discuss the choice of the random forest's hyperparameters or whether the predictions are sensitive to them. A one-sentence justification or sensitivity note would strengthen the validation description.
- [§3.3 (Figure 3B caption)] The caption for Figure 3B states that the consensus matrix is computed 'across every run of the full clustering pipeline: imputation methods, feature sets, and DR methods.' Clarify whether this includes all preprocessing pipelines in G or only those used in the final model; if only a subset is used, the caption should say so.
Circularity Check
Partial circularity in the case study: the clusters called 'most robust' are selected using the same stability/generalizability metrics that are then reported as evidence of reliability.
-
fitted input called prediction
[Section 3.3, 'Interpretation and Communication' (Figure 4C-D), and Section 3.2.2(a) / Algorithm 1]
"Due to the strong generalizability and stability performance of clusters 1, 2, 6, and 8 in 4C and D, we determine these four groupings as the most robust, and hence suitable for scientific interpretation and communication to collaborators."
Clusters 1, 2, 6, and 8 are chosen as the interpretable/final groupings because they score highest on the local stability and generalizability metrics. Those same metrics are then cited as proof that the workflow yields reliable and reproducible groupings. The validation measure is therefore used both to select the outcome and to certify it, so the reported robustness is partially manufactured by the selection rule. The held-out test set supplies some independence for a pre-specified model, but here the test-set metrics were inspected before designating the robust clusters, so they cannot serve as an unbiased confirmation.
full rationale
The workflow recommendations themselves are not circular: they are grounded in standard data-science practice and external frameworks such as PCS [187], and the self-citations [2, 63] are motivational rather than load-bearing. The circularity is confined to the case-study evidence. Algorithm 1 defines the best model as the most stable (m*, k*), and the paper reports that this is spectral clustering with k=2; the authors then abandon that result ('we instead shift our interest towards the eight clusters generated by K-means') and later designate clusters 1, 2, 6, and 8 as most robust after inspecting their stability/generalizability metrics. Presenting those same metrics as validation of the selected groupings is a select-on-the-metric-then-certify-with-the-same-metric loop. This is partial rather than total circularity because the metrics include held-out test generalizability and alternative preprocessing sweeps, and the authors explicitly call for future external astronomical validation ('we emphasize further research is needed to scientifically validate these clusters'). The case study's central reliability claim is therefore weakened, but the paper's general workflow retains independent content; a score of 5 reflects this partial circularity.
Assumptions & free parameters
free parameters (5)
- Number of clusters k =
8
- t-SNE perplexity =
100
- Spectral clustering n_neighbors =
60
- S/N quality control threshold =
70 (baseline)
- Log surface gravity threshold =
3.6 (baseline)
assumptions (6)
- domain assumption Stability and generalizability across reasonable analytical choices are valid criteria for trusting unsupervised findings.
- standard math Adjusted Rand Index suitably measures clustering agreement.
- domain assumption Random forest predictions of cluster labels provide a valid measure of cluster generalizability.
- domain assumption The considered preprocessing pipelines G are representative of the space of reasonable choices.
- domain assumption Chemical abundance features are informative for grouping stars by origin.
- domain assumption Discrete clusters in abundance space correspond to scientifically meaningful stellar populations, despite acknowledged continuity.
Cite this review
Pith. "Pith review of Unsupervised Machine Learning for Scientific Discovery: Workflow and Best Practices." pith.science (2026). https://pith.science/paper/J4HBP2QM
@misc{pith2026250604553,
author = {Pith},
title = {Pith review of: Unsupervised Machine Learning for Scientific Discovery: Workflow and Best Practices},
year = {2026},
howpublished = {\url{https://pith.science/paper/J4HBP2QM}},
note = {Machine review of arXiv:2506.04553}
}
read the original abstract
Unsupervised machine learning is widely used to mine large, unlabeled datasets to make data-driven discoveries in critical domains such as climate science, biomedicine, astronomy, chemistry, and more. However, despite its widespread utilization, there is a lack of standardization in unsupervised learning workflows for making reliable and reproducible scientific discoveries. In this paper, we present a structured workflow for using unsupervised learning techniques in science. We highlight and discuss best practices starting with formulating validatable scientific questions, conducting robust data preparation and exploration, using a range of modeling techniques, performing rigorous validation by evaluating the stability and generalizability of unsupervised learning conclusions, and promoting effective communication and documentation of results to ensure reproducible scientific discoveries. To illustrate our proposed workflow, we present a case study from astronomy, seeking to refine globular clusters of Milky Way stars based upon their chemical composition. Our case study highlights the importance of validation and illustrates how the benefits of a carefully-designed workflow for unsupervised learning can advance scientific discovery.
Figures
Reference graph
Works this paper leans on
-
[1]
Alam and N
S. Alam and N. Yao. The impact of preprocessing steps on the accuracy of machine learning algorithms in sentiment analysis.Computational and Mathematical Organization Theory, 25:319–335, 2019
2019
-
[2]
Allen, L
G. Allen, L. Gan, and L. Zheng. Interpretable machine learning for discovery: Statistical challenges and opportunities.Annual Review of Statistics and Its Application, 11, 2023
2023
-
[3]
C. J. Alpert, A. B. Kahng, and S.-Z. Yao. Spectral partitioning with multiple eigenvectors. Discrete Applied Mathematics, 90(1-3):3–26, 1999
1999
-
[4]
Anders, C
F. Anders, C. Chiappini, B. X. Santiago, G. Matijevič, A. B. Queiroz, M. Steinmetz, and G. Guiglion. Dissecting stellar chemical abundance space with t-sne.Astronomy & Astro- physics, 619:A125, 2018
2018
-
[5]
Anders, P
F. Anders, P. Gispert, B. Ratcliffe, C. Chiappini, I. Minchev, S. Nepal, A. B. d. A. Queiroz, J. A. Amarante, T. Antoja, G. Casali, et al. Spectroscopic age estimates for apogee red- giant stars: Precise spatial and kinematic trends with age in the galactic disc.Astronomy & Astrophysics, 678:A158, 2023
2023
-
[6]
L. K. Andersen and B. J. Reading. A supervised machine learning workflow for the reduction of highly dimensional biological data.Artificial Intelligence in the Life Sciences, 5:100090, 2024
2024
-
[7]
Armstrong, G
G. Armstrong, G. Rahman, C. Martino, D. McDonald, A. Gonzalez, G. Mishne, and R. Knight. Applications and comparison of dimensionality reduction methods for microbiome data.Frontiers in Bioinformatics, 2:821861, 2022
2022
-
[8]
Arnold, L
C. Arnold, L. Biedebach, A. Küpfer, and M. Neunhoeffer. The role of hyperparameters in machine learning models and how to tune them.Political Science Research and Methods, 12(4):841–848, 2024
2024
Show all 194 references
-
[9]
Arthur and S
D. Arthur and S. Vassilvitskii. k-means++: The advantages of careful seeding. Technical report, Stanford, 2006
2006
-
[10]
Ashtari, R
N. Ashtari, R. Mullins, C. Qian, J. Wexler, I. Tenney, and M. Pushkarna. From discovery to adoption: Understanding the ml practitioners’ interpretability journey. InProceedings of the 2023 ACM Designing Interactive Systems Conference, pages 2304–2325, July 2023
2023
-
[11]
Baeza, J
I. Baeza, J. G. Fernández-Trincado, S. Villanova, D. Geisler, D. Minniti, E. R. Garro, B. Bar- buy, T. C. Beers, and R. R. Lane. Apogee-2s mg–al anti-correlation of the metal-poor globular cluster ngc 2298.Astronomy & Astrophysics, 662:A47, 2022
2022
-
[12]
Baron and D
D. Baron and D. Poznanski. The weirdest sdss galaxies: results from an outlier detection algorithm.Monthly Notices of the Royal Astronomy Society, Nov. 2016. arXiv:1611.07526 [astro-ph]
2016 arXiv
-
[13]
S. Becker. Unsupervised learning procedures for neural networks.International Journal of Neural Systems, 2(01n02):17–33, 1991. 24
1991
-
[14]
Belokurov and A
V. Belokurov and A. Kravtsov. Nitrogen enrichment and clustered star formation at the dawn of the galaxy.Monthly Notices of the Royal Astronomical Society, 525(3):4456–4473, 2023
2023
-
[15]
Belouafa, F
S. Belouafa, F. Habti, S. Benhar, B. Belafkih, S. Tayane, S. Hamdouch, A. Bennamara, and A. Abourriche. Statistical tools and approaches to validate analytical methods: methodology and practical examples.International Journal of Metrology and Quality Engineering, 8:9, 2017
2017
-
[16]
Ben-Hur, A
A. Ben-Hur, A. Elisseeff, and I. Guyon. A stability based method for discovering structure in clustered data. InBiocomputing 2002, pages 6–17. World Scientific, 2001
2002
-
[17]
L. M. Bennett and H. Gadlin. Collaboration and team science: from theory to practice, 2012
2012
-
[18]
L. Berni. Searching for chemo-kinematic structures in the milky way halo with deep clustering algorithms.arXiv preprint arXiv:2409.11429, 2024
2024 arXiv
-
[19]
Bialopetravičius and D
J. Bialopetravičius and D. Narbutis. Deriving star cluster parameters with convolutional neu- ral networks-ii. extinction and cluster-background classification.Astronomy & Astrophysics, 633:A148, 2020
2020
-
[20]
Biswas, M
S. Biswas, M. Wardat, and H. Rajan. The art and practice of data science pipelines: A comprehensive study of data science pipelines in theory, in-the-small, and in-the-large. In Proceedings of the 44th International Conference on Software Engineering, pages 2091–2103, May 2022
2022
-
[21]
Bousquet and A
O. Bousquet and A. Elisseeff. Stability and generalization.Journal of Machine Learning Research, 2(Mar):499–526, 2002
2002
-
[22]
J. Bovy. The chemical homogeneity of open clusters.The Astrophysical Journal, 817(1):49, 2016
2016
-
[23]
Bovy, H.-W
J. Bovy, H.-W. Rix, and D. W. Hogg. The milky way has no distinct thick disk.The Astrophysical Journal, 751(2):131, 2012
2012
-
[24]
Bravo-Merodio, J
L. Bravo-Merodio, J. A. Williams, G. V. Gkoutos, and A. Acharjee. -omics biomarker identi- fication pipeline for translational medicine.Journal of translational medicine, 17:1–10, 2019
2019
-
[25]
Bresciani and M
S. Bresciani and M. Eppler. The pitfalls of visual representations: A review and classifi- cation of common errors made while designing and interpreting visualizations.Sage Open, 5(4):2158244015611451, 2015
2015
-
[26]
Buder, J
S. Buder, J. Kos, E. Wang, M. McKenzie, M. Howell, S. Martell, M. Hayden, D. Zucker, T. Nordlander, B. Montet, et al. The galah survey: Data release 4.arXiv preprint arXiv:2409.19858, 2024
2024 arXiv
-
[27]
Bullmore and D
E. Bullmore and D. Bassett. Brain graphs: Graphical models of the human brain connectome. Annual Review of Clinical Psychology, 7(1):113–140, 2011
2011
-
[28]
Carpenter, J
J. Carpenter, J. Bartlett, T. Morris, A. Wood, M. Quartagno, and M. Kenward.Multiple Imputation and Its Application. John Wiley & Sons, 2023
2023
-
[29]
Casamiquela, A
L. Casamiquela, A. Castro-Ginard, F. Anders, and C. Soubiran. The (im) possibility of strong chemical tagging.Astronomy & Astrophysics, 654:A151, 2021. 25
2021
-
[30]
Cavallo, L
L. Cavallo, L. Spina, G. Carraro, L. Magrini, E. Poggio, T. Cantat-Gaudin, M. Pasquato, S. Lucatello, S. Ortolani, and J. Schiappacasse-Ulloa. Parameter estimation for open clusters using an artificial neural network with a quadtree-based feature extractor.The Astronomical Jou...
2023
-
[31]
Chandola, A
V. Chandola, A. Banerjee, and V. Kumar. Anomaly detection: A survey.ACM Computing Surveys (CSUR), 41(3):1–58, 2009
2009
-
[32]
B. Chen, E. D’Onghia, S. A. Pardy, A. Pasquali, C. B. Motta, B. Hanlon, and E. K. Grebel. Chemodynamical clustering applied to apogee data: Rediscovering globular clusters.The Astrophysical Journal, 860(1):70, 2018
2018
-
[33]
M. Chiao. Young and rich stars.Nature Physics, 11(5):377–377, 2015
2015
-
[34]
M. Chiao. A gauge of stellar age.Nature Astronomy, 3(8):687–687, 2019
2019
-
[35]
Chiappini, F
C. Chiappini, F. Anders, T. d. S. Rodrigues, A. Miglio, J. Montalbán, B. Mosser, L. Girardi, M. Valentini, A. Noels, T. Morel, et al. Young [α/fe]-enhanced stars discovered by corot and apogee: What is their origin?Astronomy & Astrophysics, 576:L12, 2015
2015
-
[36]
Y. L. Chow, S. Singh, A. E. Carpenter, and G. P. Way. Predicting drug polypharmacology from cell morphology readouts using variational autoencoder latent space arithmetic.PLoS computational biology, 18(2):e1009888, 2022
2022
-
[37]
A. E. Chua, L. D. Pfeifer, E. R. Sekera, A. B. Hummon, and H. Desaire. Workflow for evaluating normalization tools for omics data using supervised and unsupervised machine learning.Journal of the American Society for Mass Spectrometry, 34(12):2775–2784, 2023
2023
-
[38]
Clare.Communicating Clearly about Science and Medicine: Making Data Presentations as Simple as Possible
J. Clare.Communicating Clearly about Science and Medicine: Making Data Presentations as Simple as Possible... But No Simpler. Routledge, New York, 2017
2017
-
[39]
A. J. Clarke, V. P. Debattista, D. L. Nidever, S. R. Loebman, R. C. Simons, S. Kassin, M. Du, M. Ness, D. B. Fisher, T. R. Quinn, et al. The imprint of clump formation at high redshift–i. a discα-abundance dichotomy.Monthly Notices of the Royal Astronomical Society, 484(3):347...
2019
-
[40]
Collaboration et al
G. Collaboration et al. Gaia data release 3: Summary of the content and survey properties. Astronomy & Astrophysics, 674:A1, 2023
2023
-
[41]
Crone, S
S. Crone, S. Lessmann, and R. Stahlbock. The impact of preprocessing on data mining: An evaluation of classifier sensitivity in direct marketing.European Journal of Operational Research, 173(3):781–800, 2006
2006
-
[42]
De Silva, K
G. De Silva, K. Freeman, J. Bland-Hawthorn, S. Martell, E. W. De Boer, M. Asplund, S. Keller, S. Sharma, D. Zucker, T. Zwitter, et al. The galah survey: scientific motivation. Monthly Notices of the Royal Astronomical Society, 449(3):2604–2617, 2015
2015
-
[43]
V. P. Debattista, D. J. Liddicott, O. A. Gonzalez, L. B. e Silva, J. A. Amarante, I. Lazar, M.Zoccali, E.Valenti, D.B.Fisher, T.Khachaturyants, etal. Theimprintofclumpformation at high redshift. ii. the chemistry of the bulge.The Astrophysical Journal, 946(2):118, 2023
2023
-
[44]
Denny and A
M. Denny and A. Spirling. Text preprocessing for unsupervised learning: Why it matters, when it misleads, and what to do about it.Political Analysis, 26(2):168–189, 2018. 26
2018
-
[45]
T. G. Dietterich. Ensemble methods in machine learning. InInternational Workshop on Multiple Classifier Systems, pages 1–15, Berlin, Heidelberg, June 2000. Springer Berlin Hei- delberg
2000
-
[46]
H. Dike, Y. Zhou, K. Deveerasetty, and Q. Wu. Unsupervised learning based on artificial neural network: A review. In2018 IEEE International Conference on Cyborg and Bionic Systems (CBS), pages 322–327. IEEE, October 2018
2018
-
[47]
J. H. Do and D. K. Choi. Clustering approaches to identifying gene expression patterns from dna microarray data.Molecules and Cells, 25(2):279–288, 2008
2008
-
[48]
Donor, P
J. Donor, P. M. Frinchaboy, K. Cunha, J. E. O’Connell, C. A. Prieto, A. Almeida, F. Anders, R. Beaton, D. Bizyaev, J. R. Brownstein, R. Carrera, C. Chiappini, R. Cohen, D. A. García- Hernández, D. Geisler, S. Hasselquist, H. Jönsson, R. R. Lane, S. R. Majewski, D. Minniti, C. ...
2020
-
[49]
Drton and M
M. Drton and M. D. Perlman. Multiple testing and error control in gaussian graphical model selection.Statistical Science, 22(3):430–449, 2007
2007
-
[50]
Ebert-Uphoff and Y
I. Ebert-Uphoff and Y. Deng. Causal discovery for climate research using graphical models. Journal of Climate, 25(17):5648–5665, 2012
2012
-
[51]
El Naqa, D
I. El Naqa, D. Ruan, G. Valdes, A. Dekker, T. McNutt, Y. Ge, Q. Wu, J. Oh, M. Thor, W. Smith, and A. Rao. Machine learning and modeling: Data, validation, communication challenges.Medical Physics, 45(10):e834–e840, 2018
2018
-
[52]
Evergreen
S. Evergreen. Effective data visualization: The right chart for the right data.SAGE Publi- cations, 2019
2019
-
[53]
Farahani, W
F. Farahani, W. Karwowski, and N. Lighthall. Application of graph theory for identifying con- nectivity patterns in human brain networks: A systematic review.Frontiers in Neuroscience, 13:585, 2019
2019
-
[54]
J.G.Fernández-Trincado, T.C.Beers, D.Minniti, B.Tang, S.Villanova, D.Geisler, A.Pérez- Villegas, and K. Vieira. Aluminium-enriched metal-poor stars buried in the inner galaxy. Astronomy & Astrophysics, 643:L4, 2020
2020
-
[55]
J. G. Fernández-Trincado, T. C. Beers, V. M. Placco, E. Moreno, A. Alves-Brito, D. Minniti, B. Tang, A. Pérez-Villegas, C. Reylé, A. C. Robin, et al. Discovery of a new stellar subpopu- lation residing in the (inner) stellar halo of the milky way.The Astrophysical Journal Lett...
2019
-
[56]
J. G. Fernandez-Trincado, T. C. Beers, B. Tang, E. Moreno, A. Pérez-Villegas, and M. Ortigoza-Urdaneta. Chemodynamics of newly identified giants with a globular cluster like abundance patterns in the bulge, disc, and halo of the milky way.Monthly Notices of the Royal Astronomi...
2019
-
[57]
J. G. Fernández-Trincado, L. Chaves-Velasquez, A. Pérez-Villegas, K. Vieira, E. Moreno, M. Ortigoza-Urdaneta, and L. Vega-Neme. Dynamical orbital classification of selected n-rich stars with gaia data release 2 astrometry.Monthly Notices of the Royal Astronomical Society, 495(...
2020
-
[58]
reproducibility and replicability in science
H. Fineberg, V. Stodden, and X. Meng. Highlights of the us national academies report on “reproducibility and replicability in science”.Harvard Data Science Review, 2(4), 2020
2020
-
[59]
Fornito, A
A. Fornito, A. Zalesky, and M. Breakspear. Graph analysis of the human connectome: Promise, progress, and pitfalls.NeuroImage, 80:426–444, 2013
2013
-
[60]
Fraix-Burnet, C
D. Fraix-Burnet, C. Bouveyron, and J. Moultaka. Unsupervised classification of sdss galaxy spectra | astronomy & astrophysics (a&a).astronomy & Astrophysics, 649:A53, 2021
2021
-
[61]
Frank and I
H. Frank and I. Hatak. Doing a research literature review. InHow to Get Published in the Best Entrepreneurship Journals, pages 94–117. Elgar, 2014
2014
-
[62]
A. J. Gamez, C. S. Zhou, A. Timmermann, and J. Kurths. Nonlinear dimensionality reduction in climate data.Nonlinear Processes in Geophysics, 11(3):393–398, 2004
2004
-
[63]
L. Gan, T. M. Zikry, and G. I. Allen. Are machine learning interpretations reliable? a stability study on global interpretations.arXiv preprint arXiv:2505.15728, 2025
2025 arXiv
-
[64]
Garcia-Dias, C
R. Garcia-Dias, C. A. Prieto, J. S. Almeida, and I. Ordovás-Pascual. Machine learning in apogee-unsupervised spectral classification with k-means.Astronomy & Astrophysics, 612:A98, 2018
2018
-
[65]
Gastner and G
M. Gastner and G. Ódor. The topology of large open connectome networks for the human brain.Scientific Reports, 6(1):27249, 2016
2016
-
[66]
Ghojogh, M
B. Ghojogh, M. Crowley, F. Karray, and A. Ghodsi.Elements of Dimensionality Reduction and Manifold Learning. Springer, 2023
2023
-
[67]
Gilmore, S
G. Gilmore, S. Randich, M. Asplund, J. Binney, P. Bonifacio, J. Drew, S. Feltzing, A. Fergu- son, R. Jeffries, G. Micela, et al. The gaia-eso public spectroscopic survey.The Messenger, 147:25–31, 2012
2012
-
[68]
Gleicher, D
M. Gleicher, D. Albers, R. Walker, I. Jusufi, C. Hansen, and J. Roberts. Visual comparison for information visualization.Information Visualization, 10(4):289–309, 2011
2011
-
[69]
G. Gobo. Sampling, representativeness and generalizability. InQualitative Research Practice, pages 405–426. Sage, 2004
2004
-
[70]
Goder and V
A. Goder and V. Filkov. Consensus clustering algorithms: Comparison and refinement. In Proceedings of the Tenth Workshop on Algorithm Engineering and Experiments (ALENEX), pages 109–117. Society for Industrial and Applied Mathematics, Jan. 2008
2008
-
[71]
Gratton, C
R. Gratton, C. Sneden, and E. Carretta. Abundance variations within globular clusters. Annu. Rev. Astron. Astrophys., 42(1):385–440, 2004
2004
-
[72]
Greene, P
D. Greene, P. Cunningham, and R. Mayer. Unsupervised learning and clustering. InMachine Learning Techniques for Multimedia: Case Studies on Organization and Retrieval, pages 51–
-
[73]
Handl, J
J. Handl, J. Knowles, and D. Kell. Computational cluster validation in post-genomic data analysis.Bioinformatics, 21(15):3201–3212, 2005
2005
-
[74]
Hastie, R
T. Hastie, R. Tibshirani, and J. Friedman.Undirected Graphical Models, pages 625–648. Springer, 2009. 28
2009
-
[75]
Hawkins, P
K. Hawkins, P. Jofré, T. Masseron, and G. Gilmore. Using chemical tagging to redefine the interface of the galactic disc and halo.Monthly Notices of the Royal Astronomical Society, 453:758–774, Oct. 2015. ADS Bibcode: 2015MNRAS.453..758H
2015
-
[76]
M. R. Hayden, J. Bovy, J. A. Holtzman, D. L. Nidever, J. C. Bird, D. H. Weinberg, B. H. Andrews, S.R.Majewski, C.AllendePrieto, F.Anders, T.C.Beers, D.Bizyaev, C.Chiappini, K.Cunha, P.Frinchaboy, D.A.García-Herńandez, A.E.GarcíaPérez, L.Girardi, P.Harding, F. R. Hearty, J. A. ...
2015
-
[77]
Healy.Data Visualization: A Practical Introduction
K. Healy.Data Visualization: A Practical Introduction. Princeton University Press, 2024
2024
-
[78]
J. Heil, V. Häring, B. Marschner, and B. Stumpe. Advantages of fuzzy k-means over k- means clustering in the classification of diffuse reflectance soil spectra: A case study with west african soils.Geoderma, 337:11–21, 2019
2019
-
[79]
M. Hong, S. Tao, L. Zhang, L. Diao, X. Huang, S. Huang, S. Xie, Z. Xiao, and H. Zhang. Rna sequencing: New technologies and applications in cancer research.Journal of Hematology & Oncology, 13:1–16, 2020
2020
-
[80]
Hotelling
H. Hotelling. Analysis of a complex of statistical variables into principal components.Journal of educational psychology, 24(6):417, 1933
1933
-
[81]
F. Hu, Z. Lu, H. Wong, and T. P. Yuen. Analysis of air quality time series of hong kong with graphical modeling.Environmetrics, 27(3):169–181, 2016
2016
-
[82]
Huang, Y
H. Huang, Y. Wang, C. Rudin, and E. P. Browne. Towards a comprehensive evaluation of dimension reduction methods for transcriptomic data visualization.Communications biology, 5(1):719, 2022
2022
-
[83]
Islam and S
M. Islam and S. Jin. An overview of data visualization. In2019 International Conference on Information Science and Communications Technologies (ICISCT), pages 1–7. IEEE, Novem- ber 2019
2019
-
[84]
K. J. Jager, C. Zoccali, A. Macleod, and F. W. Dekker. Confounding: what it is and how to deal with it.Kidney International, 73(3):256–260, 2008
2008
-
[85]
Janvrin, R
D. Janvrin, R. Raschke, and W. Dilla. Making sense of complex data using interactive data visualization.Journal of Accounting Education, 32(4):31–48, 2014
2014
-
[86]
Josse and F
J. Josse and F. Husson. Selecting the number of components in principal component analysis usingcross-validationapproximations.Computational Statistics & Data Analysis, 56(6):1869– 1879, 2012
2012
-
[87]
S. G. Kane, V. Belokurov, M. Cranmer, S. Monty, H. Zhang, and A. Ardern-Arentsen. The ones that got away: chemical tagging of globular cluster-origin stars with gaia bp/rp spectra. Monthly Notices of the Royal Astronomical Society, 536(3):2507–2524, 2025
2025
-
[88]
Kaufman and P
L. Kaufman and P. Rousseeuw.Finding Groups in Data: An Introduction to Cluster Analysis. John Wiley & Sons, 2009. 29
2009
-
[89]
Kim and Y
Y. Kim and Y. Cho. Predicting drug–gene–disease associations by tensor decomposition for network-based computational drug repositioning.Biomedicines, 11(7):1998, 2023
1998
-
[90]
Springer Berlin Heidelberg, 2008
2008
-
[91]
Kleijnen
J. Kleijnen. Validation of models: statistical techniques and data availability. InProceedings of the 31st conference on Winter simulation: Simulation—a bridge to the future - Volume 1, pages 647–654, Dec. 1999
1999
-
[92]
Kobak and G
D. Kobak and G. C. Linderman. Initialization is critical for preserving global data structure in both t-sne and umap.Nature Biotechnology, 39(2):156–157, 2021
2021
-
[93]
Koivisto
T. Koivisto. Efficient data analysis pipeline. InData Science for Natural Sciences Seminar, pages 1–4, 2019
2019
-
[94]
M. R. Krumholz, M. R. Bate, H. G. Arce, J. E. Dale, R. Gutermuth, R. I. Klein, Z.- Y. Li, F. Nakamura, and Q. Zhang. Star cluster formation and feedback.arXiv preprint arXiv:1401.2473, 2014
2014 arXiv
-
[95]
Kuzilek and M
J. Kuzilek and M. Cavus. Rashomon effect in educational research: Why more is better than one for measuring the importance of the variables?arXiv preprint arXiv:2412.12115, 2024
2024 arXiv
-
[96]
König, J
I. König, J. Malley, C. Weimar, H. Diener, and A. Ziegler. Practical experiences on the necessity of external validation.Statistics in Medicine, 26(30):5499–5511, 2007
2007
-
[97]
Lange, V
T. Lange, V. Roth, M. Braun, and J. Buhmann. Stability-based validation of clustering solutions.Neural Computation, 16(6):1299–1323, 2004
2004
-
[98]
S. L. Lauritzen and N. A. Sheehan. Graphical models for genetic analyses.Statistical Science, 18(4):489–514, 2003
2003
-
[99]
Legnardi, A
M. Legnardi, A. Milone, L. Armillotta, A. Marino, G. Cordoni, A. Renzini, E. Vesperini, F. D’Antona, M. McKenzie, D. Yong, et al. Constraining the original composition of the gas forming first-generation stars in globular clusters.Monthly Notices of the Royal Astronomical Soci...
2022
-
[100]
P. Li, H. Luo, B. Ji, and J. Nielsen. Machine learning for data integration in human gut microbiome.Microbial Cell Factories, 21(1):241, 2022
2022
-
[101]
L. Liao, H. Li, W. Shang, and L. Ma. An empirical study of the impact of hyperparameter tuning and model optimization on the performance properties of deep neural networks.ACM Transactions on Software Engineering and Methodology (TOSEM), 31(3):1–40, 2022
2022
-
[102]
H. Liu, K. Roeder, and L. Wasserman. Stability approach to regularization selection (stars) for high dimensional graphical models.Advances in Neural Information Processing Systems, 23, 2010
2010
-
[103]
Z. Liu, R. Ma, and Y. Zhong. Assessing and improving reliability of neighbor embedding methods: a map-continuity perspective.arXiv preprint arXiv:2410.16608, 2024
2024 arXiv
-
[104]
Ma and A
T. Ma and A. Zhang. Omics informatics: from scattered individual software tools to inte- grated workflow management systems.IEEE/ACM Transactions on Computational Biology and Bioinformatics, 14(4):926–946, 2016. 30
2016
-
[105]
J. T. Mackereth, R. A. Crain, R. P. Schiavon, J. Schaye, T. Theuns, and M. Schaller. The origin of diverseα-element abundances in galaxy discs.Monthly Notices of the Royal Astro- nomical Society, 477(4):5072–5089, 2018
2018
-
[106]
Marino, A
A. Marino, A. Milone, E. Dondoglio, A. Renzini, G. Cordoni, H. Jerjen, A. Karakas, E. La- gioia, M. Legnardi, M. McKenzie, et al. The metallicity variations along the chromosome maps: The globular cluster 47 tucanae.The Astrophysical Journal, 958(1):31, 2023
2023
-
[107]
S. L. Martell, J. P. Smolinski, T. C. Beers, and E. K. Grebel. Building the galactic halo from globular clusters: evidence from chemically unusual red giants.Astronomy & Astrophysics, 534:A136, 2011
2011
-
[108]
D. G. Mayer and D. G. Butler. Statistical validation.Ecological Modelling, 68(1-2):21–32, 1993
1993
-
[109]
McInnes, J
L. McInnes, J. Healy, and J. Melville. Umap: Uniform manifold approximation and projection for dimension reduction.arXiv preprint arXiv:1802.03426, 2018
2018 arXiv
-
[110]
X. Meng. Reproducibility, replicability, and reliability.Harvard Data Science Review, 2(4):10, 2020
2020
-
[111]
A. P. Milone and A. F. Marino. Multiple populations in star clusters.Universe, 8(7):359, 2022
2022
-
[112]
R. Mojena. Hierarchical grouping methods and stopping rules: an evaluation.The Computer Journal, 20(4):359–363, 1977
1977
-
[113]
Molnar.Interpretable Machine Learning
C. Molnar.Interpretable Machine Learning. Lulu.com, 2020
2020
-
[114]
Monti, P
S. Monti, P. Tamayo, J. Mesirov, and T. Golub. Consensus clustering: A resampling-based method for class discovery and visualization of gene expression microarray data.Machine Learning, 52:91–118, 2003
2003
-
[115]
Monti, P
S. Monti, P. Tamayo, J. Mesirov, and T. Golub. Consensus clustering: a resampling-based method for class discovery and visualization of gene expression microarray data.Machine Learning, 52:91–118, 2003
2003
-
[116]
Morrison-Smith, C
S. Morrison-Smith, C. Boucher, A. Sarcevic, N. Noyes, C. O’Brien, N. Cuadros, and J. Ruiz. Challenges in large-scale bioinformatics projects.Humanities and Social Sciences Communi- cations, 9(1):1–9, 2022
2022
-
[117]
Munappy, J
A. Munappy, J. Bosch, and H. Olsson. Data pipeline management in practice: Challenges and opportunities. InProduct-Focused Software Process Improvement: 21st International Conference, PROFES 2020, Turin, Italy, November 25–27, 2020, Proceedings 21, pages 168–
2020
-
[118]
D. L. Nidever, J. A. Holtzman, C. A. Prieto, S. Beland, C. Bender, D. Bizyaev, A. Burton, R.Desphande, S.W.Fleming, A.E.G.Pérez, etal. Thedatareductionpipelinefortheapache point observatory galactic evolution experiment.The Astronomical Journal, 150(6):173, 2015. 31
2015
-
[119]
National Academies Press, Washington, DC, 2019
National Academies of Sciences, Engineering, and Medicine.Reproducibility and Replicability in Science. National Academies Press, Washington, DC, 2019
2019
-
[120]
Pagnini, P
G. Pagnini, P. Di Matteo, M. Haywood, A. Mastrobuono-Battisti, F. Renaud, M. Mondelin, O. Agertz, P. Bianchini, L. Casamiquela, S. Khoperskov, et al. Abundance ties: Nephele and the globular cluster population accreted withωcen-based on apogee dr17 and gaia edr3. Astronomy & A...
2025
-
[121]
Olaode, G
A. Olaode, G. Naghdy, and C. Todd. Unsupervised classification of images: A review.Inter- national Journal of Image Processing, 8(5):325–342, 2014
2014
-
[122]
F. Pat, S. Juneau, V. Böhm, R. Pucha, A. G. Kim, A. Bolton, C. Lepart, D. Green, and A. D. Myers. Reconstructing and classifying sdss dr16 galaxy spectra with machine-learning and dimensionality reduction algorithms.arXiv preprint arXiv:2211.11783, 2022
2022 arXiv
-
[123]
Parkavi, A
A. Parkavi, A. Jawaid, S. Dev, and M. Vinutha. The patterns that don’t exist: Study on the effects of psychological human biases in data analysis and decision making. In2018 3rd Inter- national Conference on Computational Systems and Information Technology for Sustainable Solu...
2018
-
[124]
J. Pick, C. Kasper, H. Allegue, N. Dingemanse, N. Dochtermann, K. Laskowski, M. Lima, H. Schielzeth, D. Westneat, J. Wright, and Y. Araya-Ajoy. Describing posterior distributions of variance components: Problems and the use of null distributions to aid interpretation. Methods ...
2023
-
[125]
Perini, C
L. Perini, C. Galvin, and V. Vercruyssen. A ranking stability measure for quantifying the robustness of anomaly detection methods. InECML PKDD 2020 Workshops: Workshops of the European Conference on Machine Learning and Knowledge Discovery in Databases, pages 397–408. Springer...
2020
-
[126]
C. A. Prieto, S. Majewski, R. Schiavon, K. Cunha, P. Frinchaboy, J. Holtzman, K. Johnston, M. Shetrone, M. Skrutskie, V. Smith, et al. Apogee: the apache point observatory galactic evolutionexperiment.Astronomische Nachrichten: Astronomical Notes, 329(9-10):1018–1021, 2008
2008
-
[127]
Price-Jones, J
N. Price-Jones, J. Bovy, J. J. Webb, C. Allende Prieto, R. Beaton, J. R. Brownstein, R. E. Cohen, K. Cunha, J. Donor, P. M. Frinchaboy, et al. Strong chemical tagging with apogee: 21 candidate star clusters that have dissolved across the milky way disc.Monthly Notices of the R...
2020
-
[128]
Purkayastha, I
S. Purkayastha, I. Mondal, S. Sarkar, P. Goyal, and J. Pillai. Drug-drug interactions pre- diction based on drug embedding and graph auto-encoder. In2019 IEEE 19th International Conference on Bioinformatics and Bioengineering (BIBE), pages 547–552. IEEE, 2019
2019
-
[129]
Probst, A.-L
P. Probst, A.-L. Boulesteix, and B. Bischl. Tunability: Importance of hyperparameters of machine learning algorithms.Journal of Machine Learning Research, 20(53):1–32, 2019
2019
-
[130]
B. L. Ratcliffe, M. K. Ness, K. V. Johnston, and B. Sen. Tracing the assembly of the milky way’s disk through abundance clustering.The Astrophysical Journal, 900(2):165, 2020
2020
-
[131]
W. M. Rand. Objective criteria for the evaluation of clustering methods.Journal of the American Statistical association, 66(336):846–850, 1971
1971
-
[132]
Reggiani, K
H. Reggiani, K. C. Schlaufman, and A. R. Casey. Iron-rich metal-poor stars and the astro- physics of thermonuclear events observationally classified as type ia supernovae. i. establishing the connection.The Astronomical Journal, 166(3):128, 2023
2023
-
[133]
Raza and N
K. Raza and N. Singh. A tour of unsupervised deep learning for medical image analysis. Current Medical Imaging, 17(9):1059–1077, 2021. 32
2021
-
[134]
J. Roberts. Multiple view and multiform visualization. InVisual Data Exploration and Analysis VII, volume 3960, pages 176–185. SPIE, February 2000
2000
-
[135]
J. Roberts. On encouraging multiple views for visualization. InProceedings. 1998 IEEE Conference on Information Visualization. An International Conference on Computer Visual- ization and Graphics (Cat. No. 98TB100246), pages 8–14. IEEE, July 1998
1998
-
[136]
Roscher, B
R. Roscher, B. Bohn, M. F. Duarte, and J. Garcke. Explainable machine learning for scientific insights and discoveries.IEEE Access, 8:42200–42216, 2020
2020
-
[137]
Roohi, K
A. Roohi, K. Faust, U. Djuric, and P. Diamandis. Unsupervised machine learning in pathol- ogy: The next frontier.Surgical Pathology Clinics, 13(2):349–358, 2020
2020
-
[138]
V. Roth, T. Lange, M. Braun, and J. Buhmann. A resampling approach to cluster validation. Compstat: Proceedings in Computational Statistics, pages 123–128, 2002
2002
-
[139]
M. Rostami. Increasing model generalizability for unsupervised visual domain adaptation. In Conference on Lifelong Learning Agents, pages 281–293. PMLR, Nov. 2022
2022
-
[140]
Rudin, C
C. Rudin, C. Zhong, L. Semenova, M. Seltzer, R. Parr, J. Liu, S. Katta, J. Donnelly, H. Chen, and Z. Boner. Amazing things come from having many good models.arXiv preprint arXiv:2407.04846, 2024
2024 arXiv
-
[141]
D. B. Rubin. Multiple imputation. InFlexible Imputation of Missing Data, Second Edition, pages 29–62. Chapman and Hall/CRC, 2018
2018
-
[142]
Sadoddin and A
R. Sadoddin and A. A. Ghorbani. A comparative study of unsupervised machine learning and data mining techniques for intrusion detection. InInternational Workshop on Machine Learning and Data Mining in Pattern Recognition, pages 404–418, Berlin, Heidelberg, July
-
[143]
Rusta, S
E. Rusta, S. Salvadori, V. Gelli, I. Koutsouridou, and A. Marconi. Linking high-z and low-z: Are we observing the progenitors of the milky way with jwst?The Astrophysical Journal Letters, 974(2):L35, 2024
2024
-
[144]
Samuel, F
S. Samuel, F. Löffler, and B. König-Ries. Machine learning pipelines: Provenance, repro- ducibility and fair data principles. InInternational Provenance and Annotation Workshop, pages 226–230, Cham, June 2020. Springer International Publishing
2020
-
[145]
J. L. Sanders and P. Das. Isochrone ages for 3 million stars with the second gaia data release. Monthly Notices of the Royal Astronomical Society, 481(3):4093–4110, 2018
2018
-
[146]
Sainburg, L
T. Sainburg, L. McInnes, and T. Gentner. Parametric umap embeddings for representation and semisupervised learning.Neural Computation, 33(11):2881–2907, 2021
2021
-
[147]
R. P. Schiavon, S. G. Phillips, N. Myers, D. Horta, D. Minniti, C. Allende Prieto, B. An- guiano, R. L. Beaton, T. C. Beers, J. R. Brownstein, et al. The apogee value-added cata- logue of galactic globular cluster stars.Monthly Notices of the Royal Astronomical Society, 528(2)...
2024
-
[148]
R. P. Schiavon, O. Zamora, R. Carrera, S. Lucatello, A. Robin, M. Ness, S. L. Martell, V. V. Smith, D. García-Hernández, A. Manchado, et al. Chemical tagging with apogee: discovery of a large population of n-rich stars in the inner galaxy.Monthly Notices of the Royal Astronomi...
2017
-
[149]
Sarhadi, D
A. Sarhadi, D. H. Burn, G. Yang, and A. Ghodsi. Advances in projection of climate change impacts using supervised nonlinear dimensionality reduction techniques.Climate Dynamics, 48:1329–1351, 2017. 33
2017
-
[150]
Schweinsberg, M
M. Schweinsberg, M. Feldman, N. Staub, O. R. van den Akker, R. C. van Aert, M. A. Van As- sen, Y. Liu, T. Althoff, J. Heer, A. Kale, and Z. Mohamed. Same data, different conclusions: Radical dispersion in empirical results when independent analysts operationalize and test the ...
2021
-
[151]
Sengupta, D
E. Sengupta, D. Garg, T. Choudhury, and A. Aggarwal. Techniques to eliminate human bias in machine learning. In2018 International Conference on System Modeling & Advancement in Research Trends (SMART), pages 226–230. IEEE, 2018
2018
-
[152]
Schratz, J
P. Schratz, J. Muenchow, E. Iturritxa, J. Richter, and A. Brenning. Hyperparameter tuning and performance assessment of statistical and machine-learning algorithms using spatial data. Ecological Modelling, 406:109–120, 2019
2019
-
[153]
Silburt and I
J. Silburt and I. Aubert. Morphious: an unsupervised machine learning workflow to detect the activation of microglia and astrocytes.Journal of Neuroinflammation, 19(1):24, 2022
2022
-
[154]
Simmhan, C
Y. Simmhan, C. Van Ingen, A. Szalay, R. Barga, and J. Heasley. Building reliable data pipelines for managing community data using scientific workflows. In2009 Fifth IEEE In- ternational Conference on e-Science, pages 321–328. IEEE, December 2009
2009
-
[155]
L. Shi, P. He, B. Liu, K. Fu, and Q. Wu. A robust generalization of isomap for new data. In2005 International Conference on Machine Learning and Cybernetics, volume 3, pages 1707–1712. IEEE, August 2005
2005
-
[156]
Slade, A
E. Slade, A. M. Brearley, A. Coles, M. J. Hayat, P. M. Kulkarni, A. S. Nowacki, R. A. Oster, M. A. Posner, G. Samsa, H. Spratt, et al. Essential team science skills for biostatisticians on collaborative research teams.Journal of Clinical and Translational Science, 7(1):e243, 2023
2023
-
[157]
Soenen, E
J. Soenen, E. Van Wolputte, L. Perini, V. Vercruyssen, W. Meert, J. Davis, and H. Blockeel. The effect of hyperparameter tuning on the comparative evaluation of unsupervised anomaly detection methods. InProceedings of the KDD’21 Workshop on Outlier Detection and De- scription,...
2021
-
[158]
Sinoquet.Probabilistic Graphical Models for Genetics, Genomics, and Postgenomics
C. Sinoquet.Probabilistic Graphical Models for Genetics, Genomics, and Postgenomics. OUP Oxford, 2014
2014
-
[159]
D. J. Stekhoven and P. Bühlmann. Missforest—non-parametric missing value imputation for mixed-type data.Bioinformatics, 28(1):112–118, 2012. 34
2012
-
[160]
V. Stodden. Theme editor’s introduction to reproducibility and replicability in science.Har- vard Data Science Review, 2(4), 2020
2020
-
[161]
Solan, D
Z. Solan, D. Horn, E. Ruppin, and S. Edelman. Unsupervised learning of natural languages. Proceedings of the National Academy of Sciences, 102(33):11629–11634, 2005
2005
-
[162]
E. D. Sun, R. Ma, and J. Zou. Dynamic visualization of high-dimensional data.Nature Computational Science, 3(1):86–100, 2023
2023
-
[163]
W. Sun, O. Nasraoui, and P. Shafto. Evolution and impact of bias in human and machine learning algorithm interaction.PLOS ONE, 15(8):e0235502, 2020
2020
-
[164]
Sucar.Probabilistic Graphical Models
L. Sucar.Probabilistic Graphical Models. Advances in Computer Vision and Pattern Recog- nition. Springer London, 2015
2015
-
[165]
Tibau, C
X.-A. Tibau, C. Reimers, C. Requena-Mesa, and J. Runge. Spatio-temporal autoencoders in weather and climate research.Deep Learning for the Earth Sciences: A Comprehensive Approach to Remote Sensing, Climate Science, and Geosciences, pages 186–203, 2021
2021
-
[166]
Tibshirani and G
R. Tibshirani and G. Walther. Cluster validation by prediction strength.Journal of Compu- tational and Graphical Statistics, 14(3):511–528, 2005
2005
-
[167]
Springer Nature, 2019
N.Suri, M.Murty, andG.Athithan.Outlier Detection: Techniques and Applications. Springer Nature, 2019
2019
-
[168]
Tiwari and A
A. Tiwari and A. K. Sekhar. Workflow based framework for life science informatics.Compu- tational biology and chemistry, 31(5-6):305–319, 2007
2007
-
[169]
S. D. Tremaine, J. Ostriker, and L. Spitzer Jr. The formation of the nuclei of galaxies. i-m31. Astrophysical Journal, vol. 196, Mar. 1, 1975, pt. 1, p. 407-411., 196:407–411, 1975
1975
-
[170]
Ting and D
Y.-S. Ting and D. H. Weinberg. How many elements matter?The Astrophysical Journal, 927(2):209, 2022
2022
-
[171]
A. Unwin. Why is data visualization important? what is important in data visualization. Harvard Data Science Review, 2(1):1, 2020
2020
-
[172]
Van der Maaten and G
L. Van der Maaten and G. Hinton. Visualizing data using t-sne.Journal of machine learning research, 9(11), 2008
2008
-
[173]
Turner, A
C. Turner, A. Fuggetta, L. Lavazza, and A. Wolf. A conceptual basis for feature engineering. Journal of Systems and Software, 49(1):3–15, 1999
1999
-
[174]
H. Wang, J. Wang, C. Dong, Y. Lian, D. Liu, and Z. Yan. A novel approach for drug-target interactions prediction based on multimodal deep autoencoder.Frontiers in Pharmacology, 10:1592, 2020
2020
-
[175]
M. Ward, G. Grinstein, and D. Keim.Interactive Data Visualization: Foundations, Tech- niques, and Applications. AK Peters/CRC Press, 2010
2010
-
[176]
Van Der Maaten, E
L. Van Der Maaten, E. O. Postma, and H. J. Van Den Herik. Dimensionality reduction: A comparative review.Journal of Machine Learning Research, 10(66–71):13, 2009
2009
-
[177]
N. M. Webb and R. J. Shavelson. Generalizability theory: Overview. InEncyclopedia of Statistics in Behavioral Science, volume 2, pages 717–719. Wiley, 2005
2005
-
[178]
H. J. Weerts, A. C. Mueller, and J. Vanschoren. Importance of tuning hyperparameters of machine learning algorithms.arXiv, 2020. arXiv preprint arXiv:2007.07588
2020 arXiv
-
[179]
D. Watson. On the philosophy of unsupervised learning.Philosophy & Technology, 36(2):28, 2023. 35
2023
-
[180]
Wickham and G
H. Wickham and G. Grolemund.R for Data Science, volume 2. O’Reilly Media, Sebastopol, CA, 2017
2017
-
[181]
Willis and V
C. Willis and V. Stodden. Trust but verify: How to leverage policies, workflows, and in- frastructure to ensure computational reproducibility in publication.Harvard Data Science Review, 2(4), 2021
2021
-
[182]
M. Weis, S. Papadopoulos, L. Hansel, T. Lüddecke, B. Celii, P. Fahey, E. Wang, J. Bae, A. Bodor, D. Brittain, and J. Buchanan. An unsupervised map of excitatory neurons’ den- dritic morphology in the mouse visual cortex.bioRxiv, pages 2022–12, 2022
2022
-
[183]
L. Xia, C. Lee, and J. J. Li. Statistical method scdeed for detecting dubious 2d single- cell embeddings and optimizing t-sne and umap hyperparameters.Nature Communications, 15(1):1753, 2024
2024
-
[184]
Springer International Publishing, 2020
2020
-
[186]
B. Yu. Stability.Bernoulli, 19(4):1484 – 1500, 2013
2013
-
[187]
Xu and D
R. Xu and D. Wunsch.Clustering. John Wiley & Sons, 2008
2008
-
[188]
T. Yarkoni. The generalizability crisis.Behavioral and Brain Sciences, 45:e1, 2022
2022
-
[189]
Zahid, T
H. Zahid, T. Mahmood, and N. Ikram. Enhancing dependability in big data analytics en- terprise pipelines. InSecurity, Privacy, and Anonymity in Computation, Communication, and Storage: 11th International Conference and Satellite Workshops, SpaCCS 2018, Mel- bourne, NSW, Austra...
2018
-
[190]
Yu and K
B. Yu and K. Kumbier. Veridical data science.Proceedings of the National Academy of Sciences, 117(8):3920–3929, 2020
2020
-
[191]
Yu and H
T. Yu and H. Zhu. Hyper-parameter optimization: A review of algorithms and applications. arXiv preprint arXiv:2003.05689, 2020
2003 arXiv
-
[192]
T. M. Zikry, S. C. Wolff, J. S. Ranek, H. M. Davis, A. Naugle, N. Luthra, A. A. Whitman, K. M. Kedziora, W. Stallaert, M. R. Kosorok, et al. Cell cycle plasticity underlies fractional resistance to palbociclib in er+/her2- breast tumor cells.Proceedings of the National Academy...
2024
-
[193]
C. Zelaya. Towards explaining the effects of data preprocessing on machine learning. In2019 IEEE 35th International Conference on Data Engineering (ICDE), pages 2086–2090. IEEE, April 2019
2019
-
[194]
L. Zhu, C. Lee, D. Margolis, and L. Najafizadeh. Decoding cortical brain states from widefield calcium imaging data using visibility graph.Biomedical Optics Express, 9(7):3017–3036, 2018
2018
-
[2007]
Springer Berlin Heidelberg
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.