{"id":"f03cf2b7-9fd0-4c2a-b9dd-e23ad6361f7b","arxiv_id":"2502.03142","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"DES-trained transformer ensembles, applied without fine-tuning to two-magnitudes-deeper HSC imaging, recovered most known and newly found LSBGs in Abell 194, yielding a 171-object catalogue.","lead":"Astronomers trained transformer AI models on shallower Dark Energy Survey images and applied them without retraining to deeper Hyper Suprime-Cam images of the Abell 194 cluster, identifying 171 faint low-surface-brightness galaxies, 87 of them new. The test asks whether survey-specific machine-learning classifiers can be reused on future deeper surveys such as LSST and Euclid.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 93% transfer TPR is inflated because the HSC test sample largely overlaps the DES training catalogue: no Abell 194 exclusion is stated, so the headline generalization claim is not independently measured.","rationale":"The reader's weakest_assumption identifies the same core issue: the 93% TPR is computed on a final sample that substantially overlaps the DES training set, with no stated exclusion of Abell 194 sources, and the ground truth is visual rather than spectroscopic. I agree with that reading. The concern is load-bearing because the abstract's quantitative evidence for cross-survey transfer is precisely that TPR. Still, the 87 new HSC-only discoveries, the re-detection of known DES objects in independent deeper imaging, and the comparison with Zaritsky et al. (2023) UDGs provide partial support for the method and the sample, so the appropriate outcome is a conditional acceptance with a concrete revision, not rejection. Since the reader already returned CONDITIONAL, my stress-test does not change the verdict; it reinforces the requested revision to report performance on objects absent from training and to release the full catalogue with membership caveats.","tokens_in":33893,"tokens_out":5534,"duration_ms":53559,"concrete_test":"Retrain the DES ensembles after explicitly removing all Abell 194 objects from the training/validation/test split (and state how many were removed), then apply the models to the same 977 preselected HSC candidates; report TPR/FPR separately for the 84 DES-known and 87 HSC-only LSBGs. If the TPR on the 87 unseen objects drops materially below 93%, the headline transfer claim should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 trains on the Tanoglidis et al. (2021b) and Thuruthipilly et al. (2024b) DES LSBG catalogue, which contains the Abell 194 LSBGs-DES sample; no removal of Abell 194 objects from the training set is stated anywhere. Section 5.1 then reports that 95 of the 96 DES-sample LSBGs were re-identified by the model in HSC and that 84 of them remain in the final 171-object sample on which TPR = 93% is computed in Sec. 4.1. Because 18,532 of 27,873 LSBGs were randomly drawn for training, roughly two-thirds of those 84 objects are expected to have been in the training set. The HSC test cutouts are indeed new images, but the model is being scored on galaxies it was trained to recognise from DES; this measures within-sample consistency across surveys, not generalisation to unseen LSBGs. The 87 HSC-only galaxies are visually confirmed but not spectroscopically verified, and no TPR is quoted for them separately. Therefore the central claim of successful no-fine-tuning transfer is not established by the reported 93%.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains LSBG DETR and LSBG ViT ensembles on DES DR1 data from the Tanoglidis et al. (2021b) and Thuruthipilly et al. (2024b) catalogues, standardizes DES and HSC images to surface-brightness units, applies the ensembles without fine-tuning to SExtractor-selected candidates in deep HSC observations of Abell 194, and then uses GALFIT Sérsic fits and visual inspection to produce 171 LSBGs, including 87 new discoveries and 28 UDGs. The paper reports a 93% true positive rate on HSC without fine-tuning, compares the new sample with previous DES and Zaritsky catalogues, and uses the UDG counts and radial densities to argue for a cluster-mass scaling and for UDGs as an extended dwarf-galaxy population.","tokens_in":34107,"tokens_out":11102,"duration_ms":96803,"significance":"If the transfer claim survives a clean evaluation, the paper would be a useful demonstration that surface-brightness-standardized transformer models can be applied across surveys without fine-tuning, with direct implications for LSST and Euclid. The compiled sample of 171 LSBGs (28 UDGs) with GALFIT parameters and GALEX photometry is a useful resource, and the comparison with previous DES catalogues highlights concrete improvements in masking and local sky subtraction. The code and training data are public, which makes the evaluation issue checkable and fixable. The astrophysical conclusions about UDG abundance and radial density are plausible but are contingent on cluster membership and on the transfer evaluation being clean.","major_comments":[{"comment":"The training set in §2.1 is built from the combined Tanoglidis et al. (2021b) and Thuruthipilly et al. (2024b) DES LSBG catalogue without any stated exclusion of Abell 194 sources, and §5.1 reports that 84 objects in the final HSC sample come from exactly that catalogue. Since 18,532 of 27,873 LSBGs were randomly drawn for training, roughly two-thirds of those 84 objects are expected to have been in the training set. The 93% TPR quoted in §4.1 is therefore largely a measure of cross-survey consistency for galaxies the model has already seen in DES, not of generalization to unseen LSBGs. I request that the authors retrain with all Abell 194 sources excluded, or otherwise demonstrate that the overlap is negligible, and that they report the TPR separately for the 87 HSC-only sources.","section":"§2.1, §4.1, §5.1"},{"comment":"The TPR of 159/171 is evaluated against a 'true' sample whose 87 new objects are labelled by the authors' own visual inspection, and no FPR or confusion matrix for the HSC field is reported. From the numbers in §4.1, roughly 113 of the 272 ML-selected candidates were rejected as non-LSBGs, which implies a non-negligible false-positive rate that should be stated explicitly; a high TPR alone is not sufficient to establish successful transfer because it can be achieved by classifying everything as an LSBG. I request the full 2x2 confusion matrix on the HSC candidate set, with TPR and FPR, or an independent validation set for the HSC-only sources.","section":"§4.1, §4.2"},{"comment":"Section 4.1 treats all 171 LSBGs as Abell 194 members even though 12 lie beyond R200 and none of the 87 new sources has a spectroscopic redshift, and this assumption is carried into the UDG abundance comparison with Karunakaran & Zaritsky (2023) in §5.2.1 and the mass-normalized radial density profile in §5.2.2. Background or foreground interlopers would inflate the UDG count and alter the radial distribution, so the support for a linear N_UDG-M200 relation and for the normalized density profile is not established by these data. I request a membership-restricted analysis, using available redshifts and a statistical background correction, or at least an explicit quantitative assessment of how plausible interloper fractions change the conclusions.","section":"§4.1, §5.2.1, §5.2.2"}],"minor_comments":[{"comment":"The x-axis labels in Figure 10 ('10 1 100') and the axis labels in Figure 12 ('102 103 104') appear to be missing superscript formatting and should be typeset as powers of ten.","section":"Figures 10 and 12"},{"comment":"The labels and running-median legend use 'LSBS' in several places; this should be 'LSBG'.","section":"Figure 14 and §5.3.1"},{"comment":"The caption cites 'Zaritsky +21' while the text and reference list use 2023; please make the citation consistent.","section":"Figure 13"},{"comment":"There is a typo 'Fig, 16' and an odd spacing in the section header 'T rends in color'; these should be corrected in proofreading.","section":"§5.3.2"},{"comment":"The catalogue table would be easier to use if it included a flag for sources beyond R200 and for the 12 DES-sample objects reclassified as non-LSBGs, since the text discusses these subsets.","section":"Table B.1"},{"comment":"The visual inspection is described as independent and performed by two authors, but no inter-rater agreement statistic or explicit rule for resolving disagreements is given; reporting the disagreement rate would improve reproducibility.","section":"§3.6"}],"recommendation":"major_revision","confidential_remarks":"The training/test overlap is the main obstacle; it is a fixable evaluation issue rather than a reason to reject, because the HSC-only subset appears to be detected at a plausible rate even if the exact value is not yet established. I would encourage the editor to ask for a retrained evaluation with Abell 194 excluded and for the full HSC confusion matrix."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper is the first real no-fine-tuning cross-survey test of a transformer LSBG detector, and it comes with a useful new catalogue. The weakness is that the headline TPR is inflated by training set overlap, and the astrophysical conclusions rest on unverified sample purity.\n\nThe genuinely new bit: they train transformer ensembles on DES, standardize pixel values to surface brightness, and apply them to deeper HSC images of Abell 194 without any fine-tuning. They recover 95 of 96 DES-known LSBGs and add 87 new ones, with careful comparison against earlier catalogues. The methodology is clearly described, the masking and Sérsic fitting are thoughtful, and the authors are honest about why the model misses the faintest sources. The code is public, which is a plus.\n\nThe soft spot is the 93% TPR claim. The training sample is drawn from the same DES LSBG catalogue that contains the Abell 194 sources; no exclusion is stated. With ~18.5k of ~27.9k LSBGs used for training, roughly two-thirds of the 84 DES-sample objects in the final catalogue were likely in the training set. So the model is being scored on galaxies it has already seen, just in new images. That makes 93% a within-sample consistency measure rather than a generalization test. Recomputing from their numbers, the 87 HSC-only galaxies have a TPR of roughly 86% (75 of 87, assuming the model detected 159 total, of which 84 are DES sample). That is still respectable, but it should be reported separately and not buried in an aggregate number. The ground truth is only visual inspection; there are no spectroscopic confirmations, and all 171 objects are treated as cluster members even though 12 lie beyond R200. So the UDG scaling and dwarf–galaxy continuity conclusions rest on an unverified purity.\n\nNone of this kills the transfer-learning idea, but the paper would be stronger if the authors explicitly separated known-sample from new-sample TPR and acknowledged the overlap. I would send it to peer review, but I'd ask for that revision before acceptance.","headline":"New cross-survey transfer test for LSBG detection, but the 93% TPR is partly a training-set artifact; still worth refereeing.","tokens_in":34707,"tokens_out":7494,"would_cite":false,"duration_ms":63506,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Transformer models trained on DES imaging find 171 low-surface-brightness galaxies in deeper HSC data, with 93% recall and no fine-tuning.","keywords":["low surface brightness galaxies","transfer learning","transformers","Abell 194","ultra-diffuse galaxies","Hyper Suprime-Cam","Dark Energy Survey","galaxy clusters"],"falsifier":"A blind, independent census of the same HSC field—carried out by human annotators or a separately trained model without access to the DES catalogue—that counts a substantially different number of LSBGs would falsify the 93% claim. Spectroscopic redshifts for all 171 candidates, showing many are not at the cluster redshift, would falsify the cluster-member assumption and the UDG scaling conclusion.","tokens_in":33675,"feed_emoji":"🌌","tokens_out":8024,"duration_ms":68540,"temperature":0.7,"pith_summary":"The paper asks whether a machine-learning model trained on a shallower survey can be trusted to find low-surface-brightness galaxies in a deeper survey. It takes transformer models trained on Dark Energy Survey images and runs them, without fine-tuning, on Hyper Suprime-Cam images of the Abell 194 cluster, which go about two magnitudes deeper. After converting both surveys' images to common surface-brightness units, the transformer ensembles recover 159 of 171 LSBGs, a 93% true positive rate, with 12 more found by visual inspection. The resulting catalogue, the largest for Abell 194 to date, contains 28 ultra-diffuse galaxies, and its UDG counts and radial distribution support a near-linear scaling with cluster mass. The paper concludes that transfer learning across survey depth is feasible with proper normalization, opening the door to applying existing trained models to upcoming deep surveys.","feed_headline":"93% recall: transfer learning finds faint cluster galaxies","feed_subtitle":"Transformer models trained on DES images detect low-surface-brightness galaxies in deeper HSC data without fine-tuning.","key_machinery":"The load-bearing mechanism is the pixel-level surface-brightness normalization: DES and HSC images are both converted to µJy arcsec$^{-2}$ before being resized into 64$\\times$64 pixel cutouts, so the models see equivalent pixel statistics despite different zero points, pixel scales, and PSFs. The second piece is the transformer ensemble itself: four LSBG Detection Transformers and four LSBG Vision Transformers, each tuned with different hyperparameters, whose averaged probabilities classify each cutout as LSBG or contaminant at a 0.5 threshold. The pipeline closes with GALFIT single-component Sérsic fits on masked, locally sky-subtracted images, which re-derive $r_{\\rm eff,g}$ and $\\mu_{\\rm eff,g}$ to enforce the sample definition and reject poor fits and contaminants.","core_discovery":"The central discovery is that transformer models trained exclusively on DES imaging can be transferred to HSC data of a different depth and resolution, achieving a true positive rate of 93% in identifying LSBGs without any fine-tuning, provided the inputs are standardized to pixel-level surface brightness. The paper further claims that this process yields a sample of 171 LSBGs in the Abell 194 cluster (87 new), of which 28 meet the UDG criteria. The UDG abundance and mass-normalized radial density profile agree with a near-linear log-scale relation between UDG count and cluster halo mass, and the UDGs lie at the extended diffuse end of the dwarf-galaxy size-luminosity plane, suggesting they are part of a continuous dwarf population rather than a distinct class.","pith_inferences":["A natural next test is a fully blind independent search of the same HSC field; if the human or algorithmic census differs by more than the reported 12 missed objects, the 93% TPR would need revising.","The paper does not recalibrate the UDG–halo-mass relation despite its Abell 194 point lying above the Karunakaran & Zaritsky prediction; a homogeneous reanalysis across many clusters with matched UDG selection would determine whether the offset is real.","Surface-brightness normalization alone does not correct for PSF or depth differences; applying this transfer approach to surveys with substantially different seeing (e.g., space vs ground) may require an additional PSF-homogenisation step.","The central FUV$-$NUV colour gradient is based on only 15$-$20 detections; stacking or deeper far-UV imaging would test whether the quenching trend is a physical signal or a selection artefact."],"forward_implications":["If the transfer result holds, models trained on existing DES data can be applied directly to deeper surveys such as LSST and Euclid, avoiding the need to assemble new large training sets for each survey.","The 171-object catalogue roughly doubles the known LSBG population of Abell 194 and provides a larger statistical base for studying how cluster environment shapes faint galaxies.","The measured UDG abundance in Abell 194 sits on the literature $N_{\\rm UDG}$–$M_{200}$ relation, suggesting that UDG counts can serve as a halo-mass tracer once selection effects are controlled.","The overlap of UDGs with dwarf galaxies in the size-luminosity plane argues against UDGs being a discrete population, which should guide simulations of dwarf galaxy evolution.","The identified failure modes—underrepresentation of faint LSBGs in training and confusion near bright galaxies—point to specific data-augmentation and sample-balancing strategies for future models."],"supporting_citations":[{"why":"Supplies the DES LSBG and contaminant catalogue that forms the training data and the LSBG definition used in this work.","marker":"Tanoglidis et al. (2021b)"},{"why":"Contributed the transformer models and the extended DES LSBG sample; the paper applies these models to HSC without retraining.","marker":"Thuruthipilly et al. (2024b)"},{"why":"Provided the masking and local sky-subtraction recipe that improves Sérsic fits and re-measures of LSBG parameters.","marker":"Bautista et al. (2023)"},{"why":"Gives the UDG number–halo mass scaling relation against which the Abell 194 UDG count is compared.","marker":"Karunakaran & Zaritsky (2023)"},{"why":"Establishes the UDG definition (reff > 1.5 kpc, µ0 > 24 mag arcsec−2) used for classification.","marker":"van Dokkum et al. (2015a)"},{"why":"Supplies the Abell 194 virial radius and mass estimates used to define cluster membership and compare densities.","marker":"Rines et al. (2003)"},{"why":"Provides an earlier UDG catalogue for Abell 194 that the new sample is cross-checked against.","marker":"Zaritsky et al. (2023)"}],"fun_headline_variants":["Transfer learning nets 93% hit rate on faint galaxies","DES-trained AI finds faint galaxies in deeper HSC data","93% recall without fine-tuning in faint galaxy search","87 new faint galaxies in Abell 194 via transfer learning","Shallow survey training, deep survey finds: 93% TPR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported 93% true positive rate is measured against a final 171-galaxy sample that the authors themselves assembled using the same fitting and inspection pipeline, and which they treat as complete without confirming membership for every galaxy.","fun_headline_variants_meta":{"raw":{"variants":["Transfer learning nets 93% hit rate on faint galaxies","DES-trained AI finds faint galaxies in deeper HSC data","93% recall without fine-tuning in faint galaxy search","87 new faint galaxies in Abell 194 via transfer learning","Shallow survey training, deep survey finds: 93% TPR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1716,"prompt_tokens":1069,"completion_tokens":647,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":685,"completion_tokens_details":{"reasoning_tokens":563}},"tokens_in":685,"tokens_out":647,"duration_ms":5781,"temperature":1.0,"reasoning_tokens":563,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T05:45:34.113630+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A blind, independent census of the same HSC field—carried out by human annotators or a separately trained model without access to the DES catalogue—that counts a substantially different number of LSBGs would falsify the 93% claim. Spectroscopic redshifts for all 171 candidates, showing many are not at the cluster redshift, would falsify the cluster-member assumption and the UDG scaling conclusion.","supporting_citations":[{"cited_title":"& Zaritsky, D","cited_arxiv_id":null,"evidence_quote":"Gives the UDG number–halo mass scaling relation against which the Abell 194 UDG count is compared."},{"cited_title":"J., Kurtz, M","cited_arxiv_id":null,"evidence_quote":"Supplies the Abell 194 virial radius and mass estimates used to define cluster membership and compare densities."}],"review_version":1}