Pith. sign in

REVIEW 2 cited by

Assessing and improving reliability of neighbor embedding methods: a map-continuity perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.16608 v2 pith:4LPX3ZNE submitted 2024-10-22 stat.ME cs.LGstat.COstat.ML

classification stat.MEcs.LGstat.COstat.ML
keywords embeddingdatahigh-dimensionalintroducelearningmapsmethodsneighbor
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visualizing high-dimensional data is essential for understanding biomedical data and deep learning models. Neighbor embedding methods, such as t-SNE and UMAP, are widely used but can introduce misleading visual artifacts. We find that the manifold learning interpretations from many prior works are inaccurate and that the misuse stems from a lack of data-independent notions of embedding maps, which project high-dimensional data into a lower-dimensional space. Leveraging the leave-one-out principle, we introduce LOO-map, a framework that extends embedding maps beyond discrete points to the entire input space. We identify two forms of map discontinuity that distort visualizations: one exaggerates cluster separation and the other creates spurious local structures. As a remedy, we develop two types of point-wise diagnostic scores to detect unreliable embedding points and improve hyperparameter selection, which are validated on datasets from computer vision and single-cell omics.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Uncovering smooth structures in single-cell data with PCS-guided neighbor embeddings

    stat.ML 2025-06 conditional novelty 6.0 of 10

    NESS uses stability across random initializations of neighbor embeddings to improve and assess smooth single-cell representations.

  2. Unsupervised Machine Learning for Scientific Discovery: Workflow and Best Practices

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A best-practices workflow for unsupervised scientific discovery, illustrated by a stability- and generalizability-driven clustering case study of Milky Way globular clusters using APOGEE data.

Pith tools