REVIEW 2 major objections 5 minor 30 references
Identifying Ring Galaxies in DESI Legacy Imaging Surveys Using Machine Learning Methods
T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims a two-stage Swin Transformer pipeline finds 8,052 new ring galaxies in DESI Legacy Imaging Surveys with 64.87 percent precision.
desk verdict A useful new ring-galaxy catalog with measured precision, but the central labels rest on an undocumented single-observer visual inspection and a subjective training cut. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the two-stage classification design built on the Swin Transformer, a vision model whose shifted-window self-attention captures both local and long-range image structure. Stage one (Swin T1) is a binary classifier trained on 4,113 verified ring-galaxy images versus 35,000 non-ring images, with data augmentation; stage two (Swin T2) is a second binary classifier trained to distinguish rings from spirals and barred spirals, which are the main contaminants. The second stage is what converts a high-recall first pass into a usable candidate list: it raises application precision from about 9 percent on the extra candidates to roughly 65 percent on the retained set. The paper also compares the Swin Transformer against ResNet18 and VGG16 on the same data and selects the Swin architecture for its higher F1 score.
What would settle it
Visually classify a random sample of DR9 galaxies with spectroscopic redshift between 0.01 and 0.20 and r-band magnitude below 17.5 that were not used in training, including faint and ambiguous ring cases, and compare with the model's predictions; a recall much lower on faint rings than on prominent rings would show the catalog misses a substantial population.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that a two-stage Swin Transformer classifier can identify ring galaxies in a wide-area imaging survey at a precision competitive with or better than previous machine-learning efforts, without relying on simulated training data. The first stage separates ring galaxies from all other galaxies; the second stage removes spiral and barred spiral galaxies that dominate the false positives. Combining three second-stage models and removing duplicates yielded 18,802 unique candidates, of which 12,196 were visually confirmed as true rings, an overall precision of 64.87 percent. After removing 4,144 objects overlapping with earlier catalogs, 8,052 objects remain as new discoveries. The paper also reports that ring galaxies show smaller color changes with redshift than non-ring galaxies, consistent with a more homogeneous population.
Load-bearing premise
The load-bearing premise is that the 4,113 visually verified ring-galaxy images used for training represent all ring galaxies in the survey; if faint or ambiguous rings are common, the model will systematically miss them and the 8,052-object catalog will be incomplete and biased.
Editorial extensions
If this is right
- The published catalog gives astronomers 8,052 new ring galaxies with positions, spectroscopic redshifts, and g/r/z fluxes, a substantial expansion of the known sample.
- The two-stage scheme shows a practical way to hunt rare morphologies in large imaging surveys when positive examples are scarce, since the second stage is explicitly built to remove the dominant false-positive class.
- The reported 64.87 percent precision implies that roughly 6,606 of the 18,802 unique candidates are still non-rings, so statistical studies using the full candidate union should account for contamination.
- At 64.87 percent, the visual-inspection precision exceeds the 58.9 percent reported for a prior machine-learning ring search that relied on simulated training data, and the method avoids simulated data altogether.
Reading between the lines
- Because the positive training set dropped 4,774 images whose rings were not prominent or were hard to identify, the catalog is likely biased toward prominent rings; faint rings in DR9 are probably underrepresented even if the reported precision is accurate.
- The redshift and color distributions presented in the paper therefore describe the detectable prominent-ring population rather than the intrinsic ring-galaxy population, since the training selection and the survey magnitude cut shape them.
- The same two-stage architecture could transfer to other rare morphological classes, such as polar-ring galaxies or tidal dwarf candidates, by keeping the first stage broad and retraining the second stage on the dominant contaminant class.
- A testable extension is to run the trained models on a deeper or bluer survey and measure whether the precision remains stable outside the spectroscopic redshift and magnitude cuts used here.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a two-stage machine-learning pipeline, based on a Swin Transformer, to identify ring galaxies in DESI Legacy Imaging Surveys DR9. Stage 1 separates ring galaxies from all other galaxies; Stage 2 removes spiral and barred-spiral contaminants. The authors train on 4,113 visually vetted ring galaxy images and apply the model to 573,668 galaxies with z_spec=0.01-0.20 and mag_r<17.5. Stage 1 yields 49,264 candidates; Stage 2, using three balanced classifiers, yields 18,802 unique candidates. Visual inspection of all 18,802 candidates confirms 12,196 true ring galaxies (overall precision 64.87%), and after removing 4,144 overlaps with training samples and prior catalogs, the paper reports 8,052 newly discovered ring galaxies. A machine-readable catalog is provided at DOI 10.5281/zenodo.15545272.
Significance. If the catalog is reliable, it is a useful addition to the relatively small set of confirmed ring galaxies and extends the search to the DESI Legacy Surveys footprint. The paper has several strengths: it visually inspects every final candidate rather than a sample, it compares three architectures and several class-imbalance configurations, and it makes the resulting catalog publicly available. The two-stage design is sensible for reducing contamination from spiral and barred-spiral galaxies. However, the central quantitative claims—the 64.87% precision and the 8,052-object count—depend entirely on a visual inspection procedure that is described only as 'systematic visual inspection', with no stated criteria, no number of inspectors, and no inter-rater agreement measure. For a catalog paper, this is a load-bearing reproducibility gap that must be addressed before the central claims can be fully accepted.
major comments (2)
- [Section 5.1, Table 3] The reported precision of 64.87% (12,196/18,802) and the final catalog membership of 8,052 objects rest entirely on visual inspection of 18,802 candidate images, yet the manuscript provides no inspection protocol: it does not state how many inspectors were involved, whether they were blinded to the model prediction, what explicit criteria defined a 'true ring galaxy', or how disagreements were resolved. Because the training positives were also selected by subjective visual judgement (Section 2.2), the ground-truth labels for both training and final validation are single-observer judgements with no demonstrated reproducibility. Please provide a detailed protocol and, ideally, an independent re-inspection of a random subset with an inter-rater agreement statistic (e.g., Cohen's kappa), so that both the precision and the final count are anchored by reproducible measurements.
- [Section 2.2] The positive training set was constructed by visually removing 4,774 of 8,887 images whose rings were 'not prominent or difficult to clearly identify'. No quantitative or operational criteria are given for this removal. This biases the classifier toward prominent, cleanly resolved rings and means that the resulting catalog is likely incomplete for faint, edge-on, or poorly resolved ring galaxies. Consequently, the redshift and color distributions in Figures 7 and 8 inherit this selection bias, and the paper should explicitly state that the 8,052 objects are a subset selected by these criteria rather than a complete census. At minimum, the criteria for excluding training images should be specified and the expected impact on completeness discussed.
minor comments (5)
- [Section 2.2, Table 1] The cross-matched counts in the text (2,598 Nair & Abraham; 185 Timmis & Shamir; 443 Shamir; 1,151 Krishnakumar & Kalmbach; 3,657 Galaxy Zoo 2; 853 GALAXY CRUISE) differ from the 'used ring galaxies' counts in Table 1 (1,087; 71; 252; 774; 1,726; 203). I assume Table 1 lists post-filter counts, but this should be stated explicitly, and the treatment of galaxies appearing in multiple input catalogs should be clarified.
- [Figures 4 and 5] The captions state that 'AUC values are identical across all three models', while the text in Section 4.2.1 says the AUC values 'differ by only 0.001'. Please make the reported values consistent and give the actual AUC numbers.
- [Section 2.3] The list of augmentation operations implies ten or more transformed versions per image, but the text and Figure 2 mention '8 samples'. Please clarify how many augmented images are generated per original image and which operations are randomly applied versus always applied.
- [Section 4.2.2] The construction of the Stage 2 negative samples is not fully specified: the text says spiral and barred-spiral galaxies were balanced using 'a random sampling method', but it does not state how these morphological types were identified, from which catalog, or what the resulting class balance was after sampling. Please provide this information for reproducibility.
- [Section 5.1] When reporting the 9% precision on the 1,000 randomly selected images from the difference set between Swin T1 8-8 and Swin T1 3-8, the sample size and the implied binomial uncertainty should be given, since 9% is based on 1,000 samples and the uncertainty is nontrivial.
Circularity Check
No significant circularity: the catalog and precision are produced by an independently trained model plus external visual verification, with no fitted target or self-citation chain.
full rationale
The derivation chain is self-contained against the paper's claimed inputs. The positive training set (4113 images) is assembled from six independent external catalogs (Nair & Abraham 2010; Hart et al. 2016; Timmis & Shamir 2017; Shamir 2020; Krishnakumar & Bryce Kalmbach 2022; Tanaka et al. 2023) after cross-matching to DR9, and the negative set is drawn from Galaxy Zoo 2 non-ring classifications (Section 2.2). The two-stage Swin models are trained on these labels and then applied to 573,668 previously unseen DR9 galaxy images (Section 5.1); no parameter is fitted to the final '8052 new ring galaxies' count. The claimed precision 64.87% = 12,196/18,802 is a directly measured quantity obtained by systematic visual inspection of the model's candidate set, and the final catalog is defined by removing overlaps with the training samples and existing catalogs (Hart et al. 2016; Shimakawa et al. 2024; Krishnakumar & Kalmbach 2024; Abraham et al. 2024). The paper does not rely on a self-citation chain or uniqueness theorem; the external catalogs are independent prior results, not outputs of this paper. The absence of a documented inter-rater agreement protocol for the visual inspection is a reproducibility and validation limitation, but it does not make the derivation circular: the human labels are not obtained by solving the model's equations and are not equivalent to the model's outputs by construction. Therefore no circular step can be exhibited.
Assumptions & free parameters
free parameters (4)
- stage-1 training augmentation factors =
3x and 8x
- stage-2 positive-to-negative ratios =
1:2, 1:3, 1:5
- classification threshold =
0.5
- visually excluded positive images =
4774
assumptions (5)
- domain assumption A ring galaxy can be reliably recognized from a single 256x256 three-band Legacy Surveys cutout.
- domain assumption Galaxy Zoo 2 galaxies with ring count<1 are a reliable non-ring population.
- domain assumption The different positive-label sources correspond to the same morphological class of ring galaxies.
- domain assumption The authors' visual classification is an acceptable ground truth for the final catalog.
- domain assumption The parent sample defined by z_spec=0.01-0.20 and mag_r<17.5 is a meaningful target population.
Cite this review
Pith. "Pith review of Identifying Ring Galaxies in DESI Legacy Imaging Surveys Using Machine Learning Methods." pith.science (2026). https://pith.science/paper/VU4SVXFF
@misc{pith2026250616090,
author = {Pith},
title = {Pith review of: Identifying Ring Galaxies in DESI Legacy Imaging Surveys Using Machine Learning Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/VU4SVXFF}},
note = {Machine review of arXiv:2506.16090}
}
read the original abstract
The formation and evolution of ring structures in galaxies are crucial for understanding the nature and distribution of dark matter, galactic interactions, and the internal secular evolution of galaxies. However, the limited number of existing ring galaxy catalogs has constrained deeper exploration in this field. To address this gap, we introduce a two-stage binary classification model based on the Swin Transformer architecture to identify ring galaxies from the DESI Legacy Imaging Surveys. This model first selects potential candidates and then refines them in a second stage to improve classification accuracy. During model training, we investigated the impact of imbalanced datasets on the performance of the two-stage model. We experimented with various model combinations applied to the datasets of the DESI Legacy Imaging Surveys DR9, processing a total of 573,668 images with redshifts ranging from z_spec = 0.01-0.20 and magr <17.5. After applying the two-stage filtering and conducting visual inspections, the overall Precision of the models exceeded 64.87%, successfully identifying a total of 8052 newly discovered ring galaxies. With our catalog, the forthcoming spectroscopic data from DESI will facilitate a more comprehensive investigation into the formation and evolution of ring galaxies.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Abazajian, K. N., Adelman-McCarthy, J. K., Ag¨ ueros, M. A., et al. 2009, ApJS, 182, 543, doi: 10.1088/0067-0049/182/2/543
-
[2]
Abolfathi, B., Aguado, D. S., Aguilar, G., et al. 2018, ApJS, 235, 42, doi: 10.3847/1538-4365/aa9e8a
-
[3]
Automated Detection of Galactic Rings from SDSS Images
Abraham, L., Abraham, S., Kembhavi, A. K., et al. 2024, arXiv e-prints, arXiv:2404.04484, doi: 10.48550/arXiv.2404.04484
work page Pith review arXiv doi:10.48550/arxiv.2404.04484 2024
-
[4]
Abraham, S., Aniyan, A. K., Kembhavi, A. K., Philip, N. S., & Vaghmare, K. 2018, MNRAS, 477, 894, doi: 10.1093/mnras/sty627
-
[5]
F., Argudo-Fern´ andez, M., et al
Almeida, A., Anderson, S. F., Argudo-Fern´ andez, M., et al. 2023, ApJS, 267, 44, doi: 10.3847/1538-4365/acda98
-
[6]
1995, The Catalog of Southern Ringed Galaxies:
Buta, R. 1995, The Catalog of Southern Ringed Galaxies:
work page 1995
-
[7]
Erratum, IOP, doi: 10.1086/192176
-
[8]
1996, FCPh, 17, 95
Buta, R., & Combes, F. 1996, FCPh, 17, 95
1996
Show all 30 references
-
[9]
C., & Pan-STARRS Team
Chambers, K. C., & Pan-STARRS Team. 2017, in American Astronomical Society Meeting Abstracts, Vol. 229, American Astronomical Society Meeting Abstracts #229, 223.03
2017
-
[10]
J., Lang, D., et al
Dey, A., Schlegel, D. J., Lang, D., et al. 2019, AJ, 157, 168, doi: 10.3847/1538-3881/ab089d D’Onghia, E., Mapelli, M., & Moore, B. 2008, MNRAS, 389, 1275, doi: 10.1111/j.1365-2966.2008.13625.x
2019
-
[11]
2024, A&A, 683, A32, doi: 10.1051/0004-6361/202245215
Fernandez, J., Alonso, S., Mesa, V., & Duplancic, F. 2024, A&A, 683, A32, doi: 10.1051/0004-6361/202245215
2024 doi
-
[12]
Garcia-Ribera, E., P´ erez-Montero, E., Garc ´ ıa-Benito, R., & V ´ ılchez, J. M. 2015, in Highlights of Spanish Astrophysics VIII, ed. A. J. Cenarro, F. Figueras, C. Hern´ andez-Monteagudo, J. Trujillo Bueno, & L. Valdivielso, 372–372
2015
-
[13]
2020, ApJS, 251, 28, doi: 10.3847/1538-4365/abc0ed
Goddard, H., & Shamir, L. 2020, ApJS, 251, 28, doi: 10.3847/1538-4365/abc0ed
2020 doi
-
[14]
E., Bamford, S
Hart, R. E., Bamford, S. P., Willett, K. W., et al. 2016, MNRAS, 461, 3663, doi: 10.1093/mnras/stw1588
2016 doi
-
[15]
1979, ApJ, 227, 714, doi: 10.1086/156782
Kormendy, J. 1979, ApJ, 227, 714, doi: 10.1086/156782
1979 doi
- [16]
-
[17]
Krishnakumar, H., & Kalmbach, J. B. 2024, The Astronomical Journal, 168, 191, doi: 10.3847/1538-3881/ad7132
2024 doi
-
[18]
Krizhevsky, A., Sutskever, I., & Hinton, G. E. 2012, Advances in neural information processing systems, 25 15
2012
-
[19]
2021, in Proceedings of the IEEE/CVF international conference on computer vision, 10012–10022
Liu, Z., Lin, Y., Cao, Y., et al. 2021, in Proceedings of the IEEE/CVF international conference on computer vision, 10012–10022
2021
-
[20]
1976, ApJ, 209, 382, doi: 10.1086/154730
Lynds, R., & Toomre, A. 1976, ApJ, 209, 382, doi: 10.1086/154730
1976 doi
-
[21]
F., Nelson, E., & Petrillo, K
Madore, B. F., Nelson, E., & Petrillo, K. 2009, ApJS, 181, 572, doi: 10.1088/0067-0049/181/2/572
2009 doi
-
[22]
B., & Abraham, R
Nair, P. B., & Abraham, R. G. 2010, ApJS, 186, 427, doi: 10.1088/0067-0049/186/2/427
2010 doi
-
[23]
1987, ApJ, 320, 454, doi: 10.1086/165562
Giovanelli, R. 1987, ApJ, 320, 454, doi: 10.1086/165562
1987 doi
-
[24]
2018, MNRAS, 476, 5365, doi: 10.1093/mnras/sty613
Sedaghat, N., & Mahabal, A. 2018, MNRAS, 476, 5365, doi: 10.1093/mnras/sty613
2018 doi
-
[25]
2020, MNRAS, 491, 3767, doi: 10.1093/mnras/stz3297
Shamir, L. 2020, MNRAS, 491, 3767, doi: 10.1093/mnras/stz3297
2020 doi
-
[26]
2024, PASJ, 76, 191, doi: 10.1093/pasj/psae002
Shimakawa, R., Tanaka, M., Ito, K., & Ando, M. 2024, PASJ, 76, 191, doi: 10.1093/pasj/psae002
2024 doi
-
[27]
V., & Reshetnikov, V
Smirnov, D. V., & Reshetnikov, V. P. 2022, MNRAS, 516, 3692, doi: 10.1093/mnras/stac2549
2022 doi
-
[28]
2023, PASJ, 75, 986, doi: 10.1093/pasj/psad055
Tanaka, M., Koike, M., Naito, S., et al. 2023, PASJ, 75, 986, doi: 10.1093/pasj/psad055
2023 doi
-
[29]
2017, ApJS, 231, 2, doi: 10.3847/1538-4365/aa78a3
Timmis, I., & Shamir, L. 2017, ApJS, 231, 2, doi: 10.3847/1538-4365/aa78a3
2017 doi
-
[30]
2023, Journal of Cosmology and Astroparticle Physics, 2023, 097
Zhou, R., Ferraro, S., White, M., et al. 2023, Journal of Cosmology and Astroparticle Physics, 2023, 097
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.