{"id":"888dda02-2b8f-42fc-bff4-7dcf4355ee07","arxiv_id":"2506.16090","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A two-stage Swin Transformer pipeline identified 8,052 ring galaxy candidates in DESI Legacy Surveys DR9 with 64.87% visual-inspection precision.","lead":"Two machine learning classifiers were applied to 573,668 galaxy images from the DESI Legacy Surveys and, after manual review, produced a catalog of 8,052 newly identified ring galaxies. Larger ring galaxy samples are needed to study galaxy interactions and dark matter, so this catalog gives the field a substantially bigger statistical base.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The final catalog and precision rest entirely on an undocumented visual inspection; without inter-rater agreement or stated criteria, the 8,052 count is not independently anchored.","rationale":"The reader's conditional verdict is correct, but my load-bearing concern is different from the reader's weakest assumption. The reader focused on training-set representativeness (completeness bias). I focus on the reliability of the visual ground truth used to validate the final catalog and to measure precision. Both concerns involve human visual classification, hence 'partial' agreement. The 8,052 count and the 64.87% precision are both produced by an undocumented visual inspection. If that inspection is noisy, the catalog membership is uncertain and the precision estimate is not reproducible. This is more central than completeness because it directly attacks the correctness of the headline number, not just its completeness. The paper's public catalog and reproducible workflow are real strengths, and the authors did perform full visual inspection, so rejection is not warranted. However, the missing protocol means the central measurement is not yet independently anchored. The reader's conditional verdict already includes requests for honest discussion of selection bias and code/weights; I would add a requirement to report inter-rater agreement or an independent re-inspection sample. That does not change the verdict category, so I recommend UNCHANGED.","tokens_in":12605,"tokens_out":10176,"duration_ms":119671,"concrete_test":"Select a random sample of ~500 images from the 18,802 candidates (or from the 8,052 catalog), have at least two independent annotators, blinded to model outputs and to each other, classify each as ring or non-ring using a pre-specified definition, and compute Cohen's kappa and the resulting precision. If kappa is below ~0.8 or the precision differs from 64.87% by more than ~3 percentage points, the reported numbers are not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims—8,052 newly discovered ring galaxies and an overall precision of 64.87% (12,196/18,802)—are both derived from the authors' visual inspection of 18,802 candidates (Section 5.1). Yet the manuscript gives no protocol for this inspection: no number of inspectors, no blinding, no explicit ring-definition criteria, and no inter-rater agreement metric. Earlier, in Section 2.2, the training positives were also pruned by subjective visual criteria ('not prominent or difficult to clearly identify'), removing 4,774 of 8,887 images. This means the ground-truth labels that define both the precision and the catalog membership are single-observer judgments with no demonstrated reproducibility. For faint, edge-on, or poorly resolved rings, different inspectors could easily disagree, shifting both the 64.87% precision and the final 8,052 count. This is an omitted validation of the central measurement, not an internal contradiction.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a two-stage machine-learning pipeline, based on a Swin Transformer, to identify ring galaxies in DESI Legacy Imaging Surveys DR9. Stage 1 separates ring galaxies from all other galaxies; Stage 2 removes spiral and barred-spiral contaminants. The authors train on 4,113 visually vetted ring galaxy images and apply the model to 573,668 galaxies with z_spec=0.01-0.20 and mag_r<17.5. Stage 1 yields 49,264 candidates; Stage 2, using three balanced classifiers, yields 18,802 unique candidates. Visual inspection of all 18,802 candidates confirms 12,196 true ring galaxies (overall precision 64.87%), and after removing 4,144 overlaps with training samples and prior catalogs, the paper reports 8,052 newly discovered ring galaxies. A machine-readable catalog is provided at DOI 10.5281/zenodo.15545272.","tokens_in":12697,"tokens_out":4931,"duration_ms":54339,"significance":"If the catalog is reliable, it is a useful addition to the relatively small set of confirmed ring galaxies and extends the search to the DESI Legacy Surveys footprint. The paper has several strengths: it visually inspects every final candidate rather than a sample, it compares three architectures and several class-imbalance configurations, and it makes the resulting catalog publicly available. The two-stage design is sensible for reducing contamination from spiral and barred-spiral galaxies. However, the central quantitative claims—the 64.87% precision and the 8,052-object count—depend entirely on a visual inspection procedure that is described only as 'systematic visual inspection', with no stated criteria, no number of inspectors, and no inter-rater agreement measure. For a catalog paper, this is a load-bearing reproducibility gap that must be addressed before the central claims can be fully accepted.","major_comments":[{"comment":"The reported precision of 64.87% (12,196/18,802) and the final catalog membership of 8,052 objects rest entirely on visual inspection of 18,802 candidate images, yet the manuscript provides no inspection protocol: it does not state how many inspectors were involved, whether they were blinded to the model prediction, what explicit criteria defined a 'true ring galaxy', or how disagreements were resolved. Because the training positives were also selected by subjective visual judgement (Section 2.2), the ground-truth labels for both training and final validation are single-observer judgements with no demonstrated reproducibility. Please provide a detailed protocol and, ideally, an independent re-inspection of a random subset with an inter-rater agreement statistic (e.g., Cohen's kappa), so that both the precision and the final count are anchored by reproducible measurements.","section":"Section 5.1, Table 3"},{"comment":"The positive training set was constructed by visually removing 4,774 of 8,887 images whose rings were 'not prominent or difficult to clearly identify'. No quantitative or operational criteria are given for this removal. This biases the classifier toward prominent, cleanly resolved rings and means that the resulting catalog is likely incomplete for faint, edge-on, or poorly resolved ring galaxies. Consequently, the redshift and color distributions in Figures 7 and 8 inherit this selection bias, and the paper should explicitly state that the 8,052 objects are a subset selected by these criteria rather than a complete census. At minimum, the criteria for excluding training images should be specified and the expected impact on completeness discussed.","section":"Section 2.2"}],"minor_comments":[{"comment":"The cross-matched counts in the text (2,598 Nair & Abraham; 185 Timmis & Shamir; 443 Shamir; 1,151 Krishnakumar & Kalmbach; 3,657 Galaxy Zoo 2; 853 GALAXY CRUISE) differ from the 'used ring galaxies' counts in Table 1 (1,087; 71; 252; 774; 1,726; 203). I assume Table 1 lists post-filter counts, but this should be stated explicitly, and the treatment of galaxies appearing in multiple input catalogs should be clarified.","section":"Section 2.2, Table 1"},{"comment":"The captions state that 'AUC values are identical across all three models', while the text in Section 4.2.1 says the AUC values 'differ by only 0.001'. Please make the reported values consistent and give the actual AUC numbers.","section":"Figures 4 and 5"},{"comment":"The list of augmentation operations implies ten or more transformed versions per image, but the text and Figure 2 mention '8 samples'. Please clarify how many augmented images are generated per original image and which operations are randomly applied versus always applied.","section":"Section 2.3"},{"comment":"The construction of the Stage 2 negative samples is not fully specified: the text says spiral and barred-spiral galaxies were balanced using 'a random sampling method', but it does not state how these morphological types were identified, from which catalog, or what the resulting class balance was after sampling. Please provide this information for reproducibility.","section":"Section 4.2.2"},{"comment":"When reporting the 9% precision on the 1,000 randomly selected images from the difference set between Swin T1 8-8 and Swin T1 3-8, the sample size and the implied binomial uncertainty should be given, since 9% is based on 1,000 samples and the uncertainty is nontrivial.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a catalog paper whose central deliverable is a set of 8,052 objects. The lack of a documented visual inspection protocol is a reproducibility issue that goes to the validity of the headline number, and I would not be comfortable accepting the catalog claim until that is fixed. The Section 2.2/Table 1 count discrepancy, while probably resolvable, should also be cleaned up. The topic is within the scope of the journal, and the public catalog is a positive feature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper delivers a genuinely new data product: 8,052 ring galaxy candidates in DESI Legacy Surveys DR9, with a measured precision of 64.87% from visual inspection of all 18,802 final candidates. The catalog is public. If the count holds up, it more than doubles the bright low-z sample, which is useful for statistical work and DESI spectroscopy follow-up. That alone makes it worth reading.\n\nThe ML pipeline is standard: a two-stage Swin Transformer, first to separate rings from everything else, second to remove spiral/barred spirals. Nothing novel algorithmically, but the careful comparison of training ratios and the decision to verify a 1000-image subset of the difference between two stage-1 outputs shows they thought about precision tradeoffs. They also cross-matched against prior catalogs and removed overlaps, so the 'new' claim seems real.\n\nWhere it gets soft: the training sample. They started with 8,887 positive images and removed 4,774 as 'not prominent or difficult to clearly identify.' That's a huge subjective cut. If weak rings are common in the survey, the model will systematically miss them, and the redshift/color distributions in Figures 7 and 8 inherit that selection. The paper never addresses this. Also, Section 2.2 and Table 1 give different numbers for the final positive set—the text says 4,113 after removal, Table 1 sums to 4,113 but the source breakdown doesn't match the cross-matches described. That needs fixing.\n\nThe bigger concern, and the one the stress-test notes, is the visual inspection that anchors everything. There's no protocol: no number of inspectors, no independent agreement check, no stated criteria for 'ring.' For the precision of 64.87% to be meaningful, someone else needs to be able to reproduce those labels. As written, the central count rests on single-observer judgment. I'm not saying they didn't do it carefully, just that the paper doesn't let a reader check.\n\nNo code or model weights are released, which is a problem for a machine-learning paper but not fatal for a catalog paper if someone wants to use the catalog itself.\n\nBottom line: this is a plausible, useful catalog with an honest measured precision, but the subjective pruning and lack of inspection protocol mean the completeness and the precision numbers are not independently anchored. A good referee would ask for documentation of the inspection, a code/weight release, and an honest bias discussion. I'd send it to review, because the catalog deserves scrutiny and the issues are fixable.","headline":"A useful new ring-galaxy catalog with measured precision, but the central labels rest on an undocumented single-observer visual inspection and a subjective training cut.","tokens_in":13364,"tokens_out":2130,"would_cite":true,"duration_ms":21250,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims a two-stage Swin Transformer pipeline finds 8,052 new ring galaxies in DESI Legacy Imaging Surveys with 64.87 percent precision.","keywords":["ring galaxies","Swin Transformer","DESI Legacy Imaging Surveys","galaxy morphology classification","machine learning","galaxy catalog","data augmentation","galaxy interactions"],"falsifier":"Visually classify a random sample of DR9 galaxies with spectroscopic redshift between 0.01 and 0.20 and r-band magnitude below 17.5 that were not used in training, including faint and ambiguous ring cases, and compare with the model's predictions; a recall much lower on faint rings than on prominent rings would show the catalog misses a substantial population.","tokens_in":12323,"feed_emoji":"🔭","tokens_out":6682,"duration_ms":61618,"temperature":0.7,"pith_summary":"This paper claims that a two-stage binary classifier built on the Swin Transformer can find ring galaxies in the DESI Legacy Imaging Surveys more reliably than earlier single-pass machine-learning searches. Applied to 573,668 galaxy images with spectroscopic redshifts 0.01–0.20 and r-band magnitude below 17.5, the pipeline produced candidates that, after visual inspection, reached an overall precision of 64.87 percent. The result is a catalog of 8,052 newly discovered ring galaxies with positions, redshifts, and fluxes. If the claim holds, the catalog substantially expands the known ring-galaxy sample available for studying dark matter, galaxy interactions, and galaxy evolution.","feed_headline":"Machine learning finds 8,052 new ring galaxies","feed_subtitle":"A two-stage Swin Transformer screened 573,668 DESI images at 64.87 percent precision.","key_machinery":"The load-bearing mechanism is the two-stage classification design built on the Swin Transformer, a vision model whose shifted-window self-attention captures both local and long-range image structure. Stage one (Swin T1) is a binary classifier trained on 4,113 verified ring-galaxy images versus 35,000 non-ring images, with data augmentation; stage two (Swin T2) is a second binary classifier trained to distinguish rings from spirals and barred spirals, which are the main contaminants. The second stage is what converts a high-recall first pass into a usable candidate list: it raises application precision from about 9 percent on the extra candidates to roughly 65 percent on the retained set. The paper also compares the Swin Transformer against ResNet18 and VGG16 on the same data and selects the Swin architecture for its higher F1 score.","core_discovery":"On the paper's own terms, the central discovery is that a two-stage Swin Transformer classifier can identify ring galaxies in a wide-area imaging survey at a precision competitive with or better than previous machine-learning efforts, without relying on simulated training data. The first stage separates ring galaxies from all other galaxies; the second stage removes spiral and barred spiral galaxies that dominate the false positives. Combining three second-stage models and removing duplicates yielded 18,802 unique candidates, of which 12,196 were visually confirmed as true rings, an overall precision of 64.87 percent. After removing 4,144 objects overlapping with earlier catalogs, 8,052 objects remain as new discoveries. The paper also reports that ring galaxies show smaller color changes with redshift than non-ring galaxies, consistent with a more homogeneous population.","pith_inferences":["Because the positive training set dropped 4,774 images whose rings were not prominent or were hard to identify, the catalog is likely biased toward prominent rings; faint rings in DR9 are probably underrepresented even if the reported precision is accurate.","The redshift and color distributions presented in the paper therefore describe the detectable prominent-ring population rather than the intrinsic ring-galaxy population, since the training selection and the survey magnitude cut shape them.","The same two-stage architecture could transfer to other rare morphological classes, such as polar-ring galaxies or tidal dwarf candidates, by keeping the first stage broad and retraining the second stage on the dominant contaminant class.","A testable extension is to run the trained models on a deeper or bluer survey and measure whether the precision remains stable outside the spectroscopic redshift and magnitude cuts used here."],"forward_implications":["The published catalog gives astronomers 8,052 new ring galaxies with positions, spectroscopic redshifts, and g/r/z fluxes, a substantial expansion of the known sample.","The two-stage scheme shows a practical way to hunt rare morphologies in large imaging surveys when positive examples are scarce, since the second stage is explicitly built to remove the dominant false-positive class.","The reported 64.87 percent precision implies that roughly 6,606 of the 18,802 unique candidates are still non-rings, so statistical studies using the full candidate union should account for contamination.","At 64.87 percent, the visual-inspection precision exceeds the 58.9 percent reported for a prior machine-learning ring search that relied on simulated training data, and the method avoids simulated data altogether."],"supporting_citations":[{"why":"Supplies the DESI Legacy Imaging Surveys DR9 images that are the search and training data for the whole study.","marker":"Dey et al. (2019)"},{"why":"Introduces the Swin Transformer architecture on which both binary classifiers are built.","marker":"Liu et al. (2021)"},{"why":"Contributes the largest single source of verified ring galaxies used as positive training samples.","marker":"Nair & Abraham (2010)"},{"why":"Provides Galaxy Zoo 2 ring and non-ring classifications used to build both positive and negative samples.","marker":"Hart et al. (2016)"},{"why":"Contributes ring galaxy candidates from a flood-fill search used as additional positive training samples.","marker":"Timmis & Shamir (2017)"},{"why":"Contributes ring galaxy candidates from SDSS and earlier redshift and color trends that the paper compares against.","marker":"Shamir (2020)"},{"why":"Adds ring galaxy samples to the positive training set.","marker":"Krishnakumar & Bryce Kalmbach (2022)"},{"why":"Adds GALAXY CRUISE ring galaxy identifications to the positive training set.","marker":"Tanaka et al. (2023)"},{"why":"Provides the prior machine-learning ring search with 58.9 percent precision that this paper uses as a performance baseline.","marker":"Krishnakumar & Kalmbach (2024)"},{"why":"Provides a comparison catalog of ring-like candidates used for cross-checking and for the low-redshift distribution discussion.","marker":"Abraham et al. (2024)"}],"fun_headline_variants":["AI IDs 8,052 ring galaxies in DESI survey","Swin Transformer finds 8,052 new ring galaxies","Two-stage AI discovers 8,052 ring galaxies","ML nets 8,052 ring galaxies in DESI imaging"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 4,113 visually verified ring-galaxy images used for training represent all ring galaxies in the survey; if faint or ambiguous rings are common, the model will systematically miss them and the 8,052-object catalog will be incomplete and biased.","fun_headline_variants_meta":{"raw":{"variants":["AI IDs 8,052 ring galaxies in DESI survey","Swin Transformer finds 8,052 new ring galaxies","Two-stage AI discovers 8,052 ring galaxies","ML nets 8,052 ring galaxies in DESI imaging"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000832,"raw_usage":{"total_tokens":3615,"prompt_tokens":911,"completion_tokens":2704,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":2636}},"tokens_in":527,"tokens_out":2704,"duration_ms":19531,"temperature":1.0,"reasoning_tokens":2636,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:44:46.204364+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Visually classify a random sample of DR9 galaxies with spectroscopic redshift between 0.01 and 0.20 and r-band magnitude below 17.5 that were not used in training, including faint and ambiguous ring cases, and compare with the model's predictions; a recall much lower on faint rings than on prominent rings would show the catalog misses a substantial population.","supporting_citations":[{"cited_title":"2017, ApJS, 231, 2, doi: 10.3847/1538-4365/aa78a3","cited_arxiv_id":null,"evidence_quote":"Contributes ring galaxy candidates from a flood-fill search used as additional positive training samples."},{"cited_title":"2020, MNRAS, 491, 3767, doi: 10.1093/mnras/stz3297","cited_arxiv_id":null,"evidence_quote":"Contributes ring galaxy candidates from SDSS and earlier redshift and color trends that the paper compares against."},{"cited_title":"Analysis of Ring Galaxies Detected Using Deep Learning with Real and Simulated Data","cited_arxiv_id":"2210.11428","evidence_quote":"Adds ring galaxy samples to the positive training set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior machine-learning ring search with 58.9 percent precision that this paper uses as a performance baseline."},{"cited_title":"Automated Detection of Galactic Rings from SDSS Images","cited_arxiv_id":"2404.04484","evidence_quote":"Provides a comparison catalog of ring-like candidates used for cross-checking and for the low-redshift distribution discussion."}],"review_version":1}