Pith. sign in

REVIEW

Wide Area VISTA Extra-galactic Survey (WAVES): Unsupervised star-galaxy separation on the WAVES-Wide photometric input catalogue using UMAP and {rm{scriptsize HDBSCAN}}

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11611 v2 pith:F6IOFO2O submitted 2024-06-17 astro-ph.GA

Wide Area VISTA Extra-galactic Survey (WAVES): Unsupervised star-galaxy separation on the WAVES-Wide photometric input catalogue using UMAP and {rm{scriptsize HDBSCAN}}

classification astro-ph.GA
keywords galaxiesmethodbaselinescriptsizecatalogueclassifiersurveywide
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Star-galaxy separation is a crucial step in creating target catalogues for extragalactic spectroscopic surveys. A classifier biased towards inclusivity risks including spurious stars, wasting fibre hours, while a more conservative classifier might overlook galaxies, compromising completeness and hence survey objectives. To avoid bias introduced by a training set in supervised methods, we employ an unsupervised machine learning approach. Using photometry from the Wide Area VISTA Extragalactic Survey (WAVES)-Wide catalogue comprising 9-band $u-K_s$ data, we create a feature space with colours, fluxes, and apparent size information extracted by ${\rm P{\scriptsize RO} F{\scriptsize OUND}}$. We apply the non-linear dimensionality reduction method UMAP (Uniform Manifold Approximation and Projection) combined with the classifier ${\rm{\scriptsize HDBSCAN}}$ to classify stars and galaxies. Our method is verified against a baseline colour and morphological method using a truth catalogue from Gaia, SDSS, GAMA, and DESI. We correctly identify 99.72% of galaxies within the AB magnitude limit of $Z = 21.2$, with an F1 score of $0.9971 \pm 0.0018$ across the entire ground truth sample, compared to $0.9879 \pm 0.0088$ from the baseline method. Our method's higher purity ($0.9967 \pm 0.0021$) compared to the baseline ($0.9795 \pm 0.0172$) increases efficiency, identifying 11% fewer galaxy or ambiguous sources, saving approximately 70,000 fibre hours on the 4MOST instrument. We achieve reliable classification statistics for challenging sources including quasars, compact galaxies, and low surface brightness galaxies, retrieving 95.1%, 84.6%, and 99.5% of them respectively. Angular clustering analysis validates our classifications, showing consistency with expected galaxy clustering, regardless of the baseline classification.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.