{"id":"967fb97c-5b49-471a-a702-2e42b641f11d","arxiv_id":"1908.03610","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A systematic comparison of ten machine learning classifiers on Dark Energy Survey galaxy images shows that convolutional neural networks outperform classical methods, reaching roughly 99 percent accuracy after the authors relabel galaxies they believe Galaxy Zoo 1 misclassified.","lead":"This paper compares ten machine learning methods for sorting galaxy images into ellipticals and spirals, using Dark Energy Survey images labeled by Galaxy Zoo 1 volunteers. It finds that convolutional neural networks win, and it argues that many remaining errors are actually wrong labels in the training catalog.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline accuracy relies on circular relabeling and excludes uncertain galaxies; needs blind independent validation before the >0.99 claim is accepted.","rationale":"The central claim combines a method comparison and a headline accuracy. The method comparison appears internally consistent and supports CNN as the best of the ten tested approaches. However, the abstract's accuracy of about 0.99, and the improved average over 0.99, depend on two fragile decisions: changing GZ1 test labels on the basis of the model's own repeated failures, and reporting accuracy only on the classifiable p>=0.8 subset. The reader's weakest assumption correctly identifies the circularity in the relabeling procedure. My stress test adds that the conditional accuracy on classifiable galaxies is a separate but related inflation of the headline number. The recommended test is a blind expert re-labeling of the disputed objects plus a recomputation of accuracy on all 1,000 test galaxies with original labels. If the blind labels confirm the CNN's relabeling, the concern is resolved; if not, the paper should present the lower, unconditional accuracy. This does not overturn the ranking result, but it means the paper should be accepted only conditional on satisfying this validation.","tokens_in":25074,"tokens_out":4515,"duration_ms":53498,"concrete_test":"Build a blind gold set: give two independent expert classifiers (or a fresh GZ-style voting run on DES cutouts) the DES stamps for all galaxies whose labels were changed in Section 5.2.4, plus the 22 high-probability failures from Section 5.2.1, without showing CNN or GZ1 labels, and ask them to classify each as elliptical, spiral, or lenticular. Then recompute the CNN accuracy two ways: (i) on all 1,000 test galaxies using original GZ1 labels, counting uncertain p<0.8 cases as errors, and (ii) with the new gold labels replacing the changed labels. If the experts agree with GZ1 on a substantial fraction of the relabeled objects, or if the full-sample accuracy on original labels falls materially below 0.99, the headline accuracy claim should be revised or reported conditionally.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The method ranking itself is supported by consistent comparisons on the original labels, but the headline accuracy numbers are not unbiased measurements on the original test set. In Section 5.2.4 and Table 8, 'confirmed' GZ1 misclassifications are selected using frequency criteria based on CNN failures, with no independent expert relabeling described for the full list. The same CNN's predictions then define the new labels, and those labels are changed in both the training and test sets before reporting accuracies of 0.991/0.994 in Table 9. This is circular: the model's own disagreements define the ground truth used to evaluate the model. Additionally, Table 5 and Table 9 report accuracy only for the Nclassifiable subset with p >= 0.8, excluding 'uncertain' galaxies from the denominator, so the abstract's '~0.99' is conditional on discarding the hardest cases. Counting uncertain galaxies as errors would put the full-sample accuracy well below 0.99. The CNN-best ranking is less affected because it is established before relabeling, but the paper's most public-facing claims require external validation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a systematic comparison of ten supervised machine learning methods (including CNN, KNN, LR, SVM, RF, MLPC, and their variants with RBMs) for binary morphological classification of galaxies into ellipticals and spirals using DES imaging with Galaxy Zoo 1 (GZ1) visual labels. The authors evaluate pixel inputs, HOG features, and a CNN-specific combination input, and consistently find CNN to be the best method. The paper further investigates misclassifications, identifies lenticular galaxies as low-confidence objects, and claims that after purifying the training and test labels of ~2.5% of galaxies misclassified by GZ1, CNN reaches an average accuracy of 0.991 (best 0.994). The central claims are the CNN ranking and the >0.99 accuracy figure.","tokens_in":25312,"tokens_out":2737,"duration_ms":27618,"significance":"If the >0.99 accuracy claim were valid, this would be a strong benchmark for automated morphological classification on DES-like imaging and a useful guide for method selection. The paper deserves credit for a systematic, consistent comparison: a fixed 1,000-galaxy test set, three to five reruns, and ROC curves with uncertainty bands are used throughout. The rediscovery of lenticulars as uncertain objects is an interesting empirical result. However, the headline accuracy is not an unbiased measurement as presented; the method comparison and the accuracy claim must be evaluated separately, and the latter currently rests on a circular relabeling procedure and on discarding uncertain galaxies.","major_comments":[{"comment":"The reported accuracies of 0.991 and 0.994 after 'correcting' GZ1 labels are not independent measurements. The confirmed misclassifications are defined by the CNN's own repeated failures: a galaxy is 'confirmed' if it appears at least four times in total failures and at least once among high-probability failures (Table 8). These same labels are then changed in the test set and used to compute the accuracy. Because the model's disagreements define the ground truth, the evaluation is circular, and the abstract's claim of approximately 0.99 accuracy is not supported as an unbiased estimate. I recommend either independently validating the relabeled objects (e.g., by expert visual classification of the full confirmed list, not only the three unanimous-disagreement galaxies in Section 4.5) or clearly reframing the abstract and conclusions to present the corrected-label accuracy as a post-hoc consistency check rather than the model's predictive accuracy.","section":"Section 5.2.4, Table 8, Table 9"},{"comment":"The headline accuracy is computed only for the N_classifiable subset (p >= 0.8) and excludes 'uncertain' galaxies. For example, Table 5 reports accuracy 0.974 on 912 classifiable galaxies, while 88 of the 1,000 test galaxies are excluded; Table 9 similarly reports 0.991 on 976 galaxies with 16 uncertain. The abstract's '~0.99' is therefore conditional on discarding the hardest cases. If the uncertain galaxies are counted as errors, the full-sample accuracy is substantially lower (e.g., in Table 5, dataset 2: 912 * 0.974 / 1000 is about 0.889). The abstract and conclusions should state explicitly that the accuracy applies to the high-confidence subsample, not to the full test sample.","section":"Section 5.1, Table 5, Table 9"},{"comment":"The claimed ~2.5% misclassification rate in GZ1 is derived from the CNN-based frequency criteria, but the paper does not describe an independent verification procedure for the full 'confirmed' list (Fig. 14). The authors visually confirm three unanimous-disagreement galaxies in Section 4.5, but no such external verification is reported for the 25 or so confirmed objects. Without independent labels or a described expert inspection protocol, the misclassification rate is also a product of the CNN's behavior and is not independently established.","section":"Section 5.2.4"}],"minor_comments":[{"comment":"The phrase 'or a investigation' contains a grammar error; it should be 'or an investigation'.","section":"Abstract"},{"comment":"The column header 'Nuncetain' appears to be a typo for 'Nuncertain'.","section":"Table 5 and Table 9"},{"comment":"The sentence describing the added Gaussian noise is garbled: 'it is big enough to make a detectable but change of pixel values' should be rewritten for clarity.","section":"Section 2.1.1"},{"comment":"The description of the purification procedure would benefit from a clearer account of the number of iterations: the text states 'After carrying out this purification twice,' but the preceding description of rerunning five times on each new training set is ambiguous about how the two iterations relate to the five reruns.","section":"Section 5.2.4"},{"comment":"The table footnote refers to 'the sixth method' when describing 'CNN (GPU)', but CNN (GPU) is not the sixth row; this should be corrected.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The method-comparison portion of the paper is sound and reasonably thorough, and the CNN-best conclusion is consistent across ROC curves, accuracy, and multiple reruns. My main concern is the circularity of the relabeling procedure in Section 5.2.4 and the conditional nature of the accuracy metric. These issues are fixable with independent validation or substantial reframing, so I recommend major revision rather than rejection. The authors should be encouraged to consult the stress-test note, which correctly identifies the same load-bearing problem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has two distinct parts: a careful comparison of ten classifiers on DES imaging, and an attempt to 'correct' Galaxy Zoo labels. The first part is solid and worth citing; the second part's headline numbers don't survive contact with the methods.\n\nWhat's actually new: a controlled, equal-footing comparison of ten machine-learning methods on the same DES pixel data, with a held-out test set and repeated reruns. CNN wins, which is not surprising given the literature, but the systematic comparison is useful for survey-scale work. The combined raw+HOG input as separate CNN channels is a small, legitimate twist. The failure analysis also does something nice: objects with low probability in both classes are often visually lenticular, which is a sensible sanity check.\n\nWhere it falls apart: Section 5.2.4. The 'confirmed misclassified' galaxies are selected by how often the CNN's own predictions disagree with Galaxy Zoo (frequency criteria in Table 8). Those same CNN predictions then define the new labels, and the labels are changed in the test set before reporting accuracy of 0.991/0.994. That is circular in the plainest sense: the model defines the truth that evaluates the model. The claim that ~2.5% of GZ1 labels are wrong rests on the same logic. To make this work you'd need blind human inspection of the contested galaxies, or an independent source of truth. The paper shows images of a few examples, which are convincing, but that's not a systematic validation of roughly 70 relabeled objects.\n\nThere's a second, quieter problem. The abstract's '~0.99' comes from Tables 5 and 9, which report accuracy only for objects with predicted probability p >= 0.8, the 'classifiable' subset. Uncertain galaxies are left out of the denominator. With 16-19 uncertain objects per run out of 1000, even treating those as errors drops the full-sample accuracy below 0.99. The paper is not upfront about that conditionality.\n\nThe method ranking itself is less affected, because it is established on the original labels before any relabeling. I believe the CNN-beats-the-rest conclusion would survive independent scrutiny. The headline accuracy numbers, in contrast, are not measurements; they are the result of a self-fulfilling procedure.\n\nWho's this for? Anyone building a morphology pipeline for DES or LSST, especially those deciding whether to bother with pixels versus parameters. It deserves a serious referee, because the benchmark is useful, but the referee's main job should be to force either independent validation of the relabeling or honest reporting of accuracy on the unchanged test set with uncertain galaxies counted. I'd send it to review, but I would not accept the >0.99 claim without fix.","headline":"The method comparison is useful and likely correct; the headline accuracy after 'correcting' Galaxy Zoo labels is circular and should not be trusted.","tokens_in":25900,"tokens_out":1789,"would_cite":false,"duration_ms":21242,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"With pixel input alone, a convolutional neural network outperforms nine other machine-learning methods at separating ellipticals from spirals on Dark Energy Survey images, reaching an average accuracy above 0.99 once roughly 2.5 per cent…","keywords":["galaxy morphology","machine learning","convolutional neural networks","Dark Energy Survey","Galaxy Zoo 1","pixel input","lenticular galaxies","classification accuracy"],"falsifier":"Have at least three independent experienced classifiers, blind to both the CNN probabilities and the Galaxy Zoo labels, inspect the DES images of the 22 high-confidence mismatches and the 8 suspected misclassifications shown in the paper's Figures 11 and 15. If a substantial fraction of the 'confirmed' relabelled galaxies do not show unambiguous structures agreeing with the CNN, the corrected-label accuracy of 0.994 overstates the method.","tokens_in":24878,"feed_emoji":"🔭","tokens_out":6517,"duration_ms":66735,"temperature":0.7,"pith_summary":"At stake is which supervised machine-learning method should be trusted to classify millions of galaxy images automatically. Using about 2,800 Dark Energy Survey galaxies with visual labels from Galaxy Zoo 1, the paper pits ten methods against one another on a two-way problem: elliptical versus spiral. It argues that the convolutional neural network is the most successful of these when fed raw image pixels, and that its mistakes are diagnostic: high-confidence mismatches are usually cases where the visual label itself is wrong, while low-confidence objects are mostly lenticular (S0) galaxies. With the flagged label errors corrected, the CNN reaches an average test accuracy above 0.99, with 0.994 as the best single run. The practical payoff would be a pre-trained model able to sort DES-level images of millions of galaxies without new human labelling.","feed_headline":"CNN tops nine rivals on galaxy shape sorting","feed_subtitle":"Trained on DES pixels and Galaxy Zoo labels, it reaches ~99 percent accuracy after correcting bad labels.","key_machinery":"The load-bearing object is a convolutional neural network with three convolutional layers (32, 64, and 128 filters), each followed by max pooling, two fully connected hidden layers of 1024 units with dropout, and a two-class softmax output. A distinctive input mode, called the combination input, stacks the raw linearly-scaled $50 \\times 50$ stamp together with its Histogram of Oriented Gradients (HOG) feature image so the CNN reads both at once; this outperforms either input alone. Training uses rotated copies with added Gaussian noise, balanced elliptical/spiral counts, and a classification criterion $p \\geq 0.8$ that separates confident galaxies from uncertain ones. The uncertainty class is where the paper's discovery of lenticulars emerges.","core_discovery":"The paper's central claim is that, for binary elliptical/spiral classification from image pixels alone on Dark Energy Survey stamps, a convolutional neural network outperforms K-nearest neighbours, logistic regression, support vector machines, random forests, multi-layer perceptrons, and their restricted-Boltzmann-machine variants. With about 2,800 Galaxy Zoo 1 labelled galaxies rotated into roughly 100,000 training samples, the CNN reaches an accuracy near 0.95 with balanced data and combined raw-plus-HOG input. Raising the classification threshold to $p \\geq 0.8$ lifts accuracy to about 0.987 by marking low-confidence objects as uncertain, and most of those uncertain objects look like lenticulars on DES images. The paper also claims that about 2.5 per cent of the Galaxy Zoo 1 labels in this sample are wrong, visible when DES's sharper, deeper images expose structure that SDSS lacked; after correcting those labels and retraining, the average accuracy over five runs exceeds 0.99.","pith_inferences":["Because the same CNN defines which labels are 'confirmed misclassifications' and then measures its own accuracy on the corrected labels, the 0.994 figure is best read as an upper bound until independent visual confirmation of those roughly 2.5 per cent of objects.","If the lenticular finding generalizes, the probability gap between the two softmax outputs can be used as a rough morphological axis: S0 candidates, mergers, and edge-on discs may all collect at intermediate probabilities, giving three effective classes from a binary trainer.","A direct extension would ask whether the corrected labels also improve the non-CNN methods, since the paper retrains only the CNN after purification; the ranking of the ten methods could shift if all of them received the cleaner training set."],"forward_implications":["A pre-trained CNN on DES-style pixel stamps can be applied directly to millions of galaxies, giving binary morphology without additional human labels.","Balancing the training set by class is necessary when using pixel input; unbalanced sets systematically depress elliptical recall in most methods.","The uncertainty channel, $p < 0.8$, is an inexpensive way to flag objects that need human inspection, and in practice those objects are dominated by lenticulars.","Galaxy Zoo 1 labels carry measurable contamination, about 2.5 per cent in this sample, and machine-learning triage can flag suspects for reclassification in future surveys.","HOG features help most pixel-based methods but not K-nearest neighbours, and their benefit largely disappears when a restricted Boltzmann machine already compresses the features."],"supporting_citations":[{"why":"Supplies the Galaxy Zoo 1 volunteer classifications used as ground-truth labels for training and testing.","marker":"Lintott et al. 2008"},{"why":"Provides the later Galaxy Zoo 1 data release that defines the selection and agreement thresholds for the visual labels.","marker":"Lintott et al. 2011"},{"why":"Provides the bias-corrected elliptical/spiral probabilities used to assign the GZ1 labels.","marker":"Bamford et al. 2009"},{"why":"Basis for the CNN architecture and for rotation-based data augmentation in galaxy image classification.","marker":"Dieleman et al. 2015"},{"why":"Supplies the Gaussian-noise augmentation precedent and CNN application that the preprocessing follows.","marker":"Huertas-Company et al. 2015"},{"why":"Introduces the Histogram of Oriented Gradients descriptor that the paper applies as an image feature input.","marker":"Dalal & Triggs 2005"},{"why":"Describes the DES Year 1 GOLD catalogue from which the stripe-82 sample is selected.","marker":"Drlica-Wagner et al. 2018"},{"why":"Documents DES depth and resolution, which the paper uses to argue that some Galaxy Zoo labels are visibly wrong.","marker":"Abbott et al. 2018"}],"fun_headline_variants":["CNN outlearns nine rivals on galaxy shapes","Deep learning wins galaxy classification shootout","AI sorts galaxies with 99% accuracy after fixing labels","Convolutional nets beat classic ML on galaxy types","Galaxy Zoo labels corrected by sharper DES images"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the galaxies the CNN repeatedly flags as 'confirmed misclassified' truly have wrong Galaxy Zoo 1 labels, so that relabelling them and then reporting accuracy is a fair test rather than circular self-confirmation; this enters where the model's own repeated failures decide the corrected truth (Section 5.2.4).","fun_headline_variants_meta":{"raw":{"variants":["CNN outlearns nine rivals on galaxy shapes","Deep learning wins galaxy classification shootout","AI sorts galaxies with 99% accuracy after fixing labels","Convolutional nets beat classic ML on galaxy types","Galaxy Zoo labels corrected by sharper DES images"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000575,"raw_usage":{"total_tokens":2766,"prompt_tokens":1048,"completion_tokens":1718,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":1646}},"tokens_in":664,"tokens_out":1718,"duration_ms":11558,"temperature":1.0,"reasoning_tokens":1646,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:07:39.954965+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have at least three independent experienced classifiers, blind to both the CNN probabilities and the Galaxy Zoo labels, inspect the DES images of the 22 high-confidence mismatches and the 8 suspected misclassifications shown in the paper's Figures 11 and 15. If a substantial fraction of the 'confirmed' relabelled galaxies do not show unambiguous structures agreeing with the CNN, the corrected-label accuracy of 0.994 overstates the method.","supporting_citations":[],"review_version":1}