{"id":"f8e82aaa-4e6d-49e1-9a46-e9cacc726f80","arxiv_id":"2412.03778","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"Using redshift-bin-dedicated hybrid unsupervised-supervised classification of CANDELS galaxies, the authors report a roughly constant disk fraction (~60%) and spheroid fraction (~30%) across 0.2<z<2.4, challenging visual-classification-based literature.","lead":"This paper classifies about 14,000 CANDELS galaxies into disks, spheroids, and irregulars using a hybrid machine-learning method, and finds that the disk fraction stays near 60% from z=0.2 to 2.4. The authors argue that earlier reports of a declining disk fraction were biased by visual classification, a claim that would reshape galaxy evolution if confirmed.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The FERENGI degradation test in §4.2 is circular: bin-dedicated CNNs were trained on labels produced by the same metric/SOM pipeline they are meant to validate, so high retention rates only show self-consistency, not that the constant disk fraction is unbiased.","rationale":"The paper's central claim is that disk, spheroid, and irregular fractions are nearly constant across 0.2<z<2.4, contrary to earlier visual-classification studies. The only way to establish this is to show the classification labels are unbiased at each redshift. The reader correctly identifies the absence of ground-truth validation as the weakest assumption. I add a sharper point: the paper's own degradation-based validation in Section 4.2 is logically circular, making the claim not merely under-validated but presented with an invalid confidence estimate. The bin-dedicated CNN models trained on the same pipeline's labels cannot serve as an independent check of those labels. The internal consistency checks (ensemble F1 scores, agreement with Lee et al. 2024) are useful but cannot substitute for an external yardstick; both methods may share the same metric-related sensitivity to rest-frame wavelength and resolution. The proposed mock test breaks the circularity by providing known intrinsic morphologies at every redshift. Because this flaw is addressable with existing simulation tools and the paper otherwise contains a coherent, reproducible methodology, a CONDITIONAL verdict remains appropriate, but the condition must include external validation against known morphologies, not just agreement with another unsupervised method. No ad hominem is intended; the issue is an argumentative gap in the validation chain.","tokens_in":22508,"tokens_out":6216,"duration_ms":67794,"concrete_test":"Generate mock CANDELS-like F814W images from a hydrodynamic simulation (e.g., TNG50/IllustrisTNG or EAGLE) for galaxies with known intrinsic morphology labels, using stellar kinematics and rest-frame optical light to define true disks/spheroids/irregulars. Redshift these galaxies to the 11 bins from z=0.2 to 2.4 using the same FERENGI degradation, then run the full pipeline (MEGG metrics, SOM/IsoData labeling, CNN training, 0.1/0.2/0.8/0.9 thresholds) and compare the recovered fractions per bin to the intrinsic fractions. If the recovered disk fraction is ~60% while the intrinsic disk fraction drops measurably beyond z~1, the central claim fails; if recovered matches intrinsic, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing vulnerability is in the validation of the labels, not in the ML implementation. Section 3.2 builds disk/spheroid labels from 'prominent clusters' of MEGG metrics measured in F814W, with IsoData thresholds, and Section 3.5 defines irregulars purely as CNN probability between 0.2 and 0.8. No independent morphological ground truth is used at any step. Section 4.2 then claims to validate the method by degrading local disks/spheroids with FERENGI and finding ~92-95% retention (Figure 9). This test is circular: the bin-dedicated CNN models used for inference were trained on real high-z galaxies whose labels were produced by the very same metric/SOM pipeline. If that pipeline systematically mislabels degraded galaxies at high z, the CNN learns those mislabels as 'correct', so high retention only demonstrates self-consistency. The flat ~60% disk and ~30% spheroid fractions could then arise because the confidence thresholds (0.1/0.9, 0.2/0.8) shunt ambiguous high-z systems into the irregular class instead of allowing disk fraction to decline, exactly the trend the paper attributes to visual bias. Agreement with Lee et al. (2024) does not break this circle, since that work also uses unsupervised non-parametric metric classification without a known physical ground truth.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a hybrid unsupervised-supervised classification of ~14,000 CANDELS galaxies at 0.2<z<2.4, using MEGG non-parametric metrics, SOM-based labeling, and Xception CNNs with hundred-model ensembles per redshift bin. The authors define disk, spheroid, and irregular classes via CNN probability thresholds and report a nearly constant disk fraction (~60%) and spheroid fraction (~30%) over the full redshift range, attributing the decline seen in earlier work to visual classification bias. They also use FERENGI to simulate redshifted local galaxies and claim that up to 18% of disks can be misclassified by degradation, and they interpret mass-dependent trends as evidence for merger-driven spheroid formation.","tokens_in":22791,"tokens_out":4834,"duration_ms":47057,"significance":"If the constant-fraction result is correct, it would substantially revise the picture of morphological evolution at z<2.4 and strengthen the case against visual inspection as a ground truth for high-redshift morphology. The paper has real strengths: bin-dedicated training with ensembles, explicit investigation of model transfer across redshift, and a quantitative degradation analysis. However, the central claim rests on a labeling chain in which the same metric distributions used to define labels also define the classes, and the degradation test is circular; without an independent ground truth, the reported fractions cannot be distinguished from a re-expression of the metric thresholds.","major_comments":[{"comment":"The disk/spheroid labels are generated by selecting 'prominent clusters' of the MEGG metrics (M20, E, Gini, G2) with SOMbrero and IsoData, and the irregular class is defined entirely by the CNN probability interval 0.2-0.8. Consequently the reported fractions are not an independent measurement of morphology but a classification of the same metric distributions used to build the labels. The central claim of a constant ~60% disk fraction requires external validation, e.g., against simulations with known intrinsic morphology or against a separate high-resolution data set, before it can support the conclusion that visual classifications are biased.","section":"Sections 3.2 and 3.5"},{"comment":"The FERENGI degradation test is circular. The bin-dedicated CNNs used to classify the artificially redshifted low-z disks/spheroids were trained on real high-z galaxies whose labels were produced by the same SOM/metric pipeline. High retention rates (92-95%) therefore only show that the pipeline reproduces its own labels under degradation; they do not demonstrate that the constant disk fraction is unbiased. I would like to see the degradation test repeated with models trained on labels from independent sources, such as hydrodynamical simulations, or at least a demonstration that the SOM labels at high z are not dominated by resolution effects.","section":"Section 4.2 and Figure 9"},{"comment":"The luminosity evolution term is fixed to Mz' = Mz0 - (1 x z') without justification, and the paper states that establishing a more accurate form is beyond scope. Since the claimed 18% disk misclassification depends on this term, the analysis should include a sensitivity test over a plausible range of evolution parameters; otherwise the degradation correction is itself an uncontrolled free parameter.","section":"Section 4.2, Eq. (8) and surrounding text"},{"comment":"There is an internal inconsistency in the degradation statistics. The text reports mean retention fractions of 92% (spheroids) and 95% (disks), yet also states 'observed decrease of about 15% on average'; a 92-95% retention corresponds to a 5-8% decrease, not 15%. The abstract's 'up to 18%' also needs to be reconciled with the mean values in Figure 9 and with the actual maximum per redshift bin, and the summary bullet in Section 6 says 'up to 17%' rather than 18%.","section":"Section 5, Figure 9 and Section 6"}],"minor_comments":[{"comment":"The quantity called 'F1 score' is defined as the fraction of galaxies that do not change class; this is not the standard F1 metric, which is the harmonic mean of precision and recall. Please rename the quantity to 'agreement fraction' or compute the actual F1 score.","section":"Section 4.1 and Figure 6"},{"comment":"The notation M2O and M20 is used interchangeably; the text introduces M2O and Eq. (2) defines M20, and Table 3 uses M20. Please unify the notation throughout.","section":"Section 2.2 and Table 3"},{"comment":"There are several typos and grammatical slips, including 'we adopt, we adopt' in Section 1, 'It is notable that we we tend to agree' in the Figure 10 caption, 'Irregural Fraction' in Figure 9, and 'Figure 4 exhibit' in Section 3.3. A careful proofread is needed.","section":"Throughout"},{"comment":"The paper states that 14,736 galaxies remain after cleaning in Section 2.1 but Section 6 reports classifying 13,988 galaxies; please clarify the relation between these numbers, e.g., removal of 'Unclassifiable' objects and any other cuts.","section":"Section 6 and Section 2.1"},{"comment":"The data availability statement says the data are available from the corresponding author upon reasonable request; for reproducibility of the central claim, please release the trained models, the morphological labels, and the code used to generate the fractions.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The paper's main scientific claim is potentially important, but I am not fully convinced that the validation is independent. The editor may wish to solicit a second opinion on the circularity concern and on whether the comparison with Lee et al. (2024) provides sufficient external support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the headline: the claim of an almost constant ~60% disk fraction out to z~2.4 is exactly the kind of result that would matter if true. But I don't think this paper demonstrates it yet. The labels come from the same metric/SOM pipeline that trains the CNN, and the irregular class is defined by a hand-set probability window (0.2–0.8). The FERENGI 'validation' in §4.2 doesn't break the loop: the bin-dedicated models were trained on labels produced by that same pipeline, so high retention rates mainly show self-consistency.\n\nWhat's genuinely new and good: dividing the classification into Δz=0.2 bins and comparing bin-dedicated vs single-wide-range models is a real improvement, and the ~25% disagreement between them is a useful caution for the community. The FERENGI degradation experiment is a sensible idea, and the paper is transparent about its limitations—it states that the luminosity evolution term is fixed at -1.0 for illustrative purposes and that irregular classifications should be treated cautiously. That honesty counts for something.\n\nThe soft spots are the ones that matter more. First, there is no independent ground truth anywhere in the chain: the SOM 'prominent clusters' are selected from MEGG metrics measured in a single band, the CNN inherits those labels, and the irregular fraction is set by thresholds. If the metrics fail to separate disks from spheroids at high z, the constant disk fraction is an artifact, not a discovery. Second, the quoted uncertainties are only ensemble scatter (Qσ); they don't include the systematic uncertainty from threshold choices or the evolution-term assumption, so they understate the error budget. The agreement with Lee et al. (2024) is encouraging but doesn't resolve the issue, since that work also uses unsupervised metric classification without known ground truth.\n\nWho gets value: methodologically oriented readers, especially those building automated morphology pipelines and worrying about redshift-dependent biases. The paper is a good cautionary example of why visual-classification-based training sets may be biased. But as a measurement of the true disk/spheroid evolution, I'd want to see the pipeline validated on simulations with known morphologies, or on a low-z sample with secure visual labels, before believing the headline.\n\nRecommendation: send it out. The question is important and the method is not obviously wrong. A good referee should push for external validation, error bars that cover threshold choices, and release of code and catalogs.","headline":"A headline claim that would reshape the disk-fraction narrative, but the labeling loop and hand-set thresholds leave it under-supported; still worth a referee's time.","tokens_in":23370,"tokens_out":2743,"would_cite":false,"duration_ms":27177,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hybrid metric-and-CNN classifier finds a flat disk fraction from z=0.2 to z=2.4.","keywords":["galaxy morphology","CANDELS","hybrid classification","self-organizing maps","convolutional neural networks","visual classification bias","disk fraction","high-redshift galaxies"],"falsifier":"Take the same CANDELS galaxies and measure the same MEGG metrics on rest-frame near-infrared JWST images; if at z>1 the F814W-based disks are classified as spheroids or irregulars by the rest-frame metrics in large numbers, or if the rest-frame disk fraction falls with redshift, the constant-fraction claim collapses. A simpler check: degrade local disks with FERENGI and have experienced astronomers classify them by eye; if disks remain recognizable, the claim that visual inspection is systematically biased is weakened.","tokens_in":22272,"feed_emoji":"🌀","tokens_out":7617,"duration_ms":72575,"temperature":0.7,"pith_summary":"The paper sets out to classify roughly 14,000 CANDELS galaxies between redshift 0.2 and 2.4 into disks, spheroids, and irregulars without any step of visual inspection. Using redshift-bin-specific models trained on labels produced by an unsupervised clustering of non-parametric morphological metrics, it finds that the fraction of disks stays near 60 percent and spheroids near 30 percent across 8.5 billion years of cosmic time. This directly contradicts earlier work reporting that disk galaxies become rare beyond z~1. The paper argues that those earlier results were biased by visual classification, which loses disk features to surface-brightness dimming and shrinking angular size at high redshift, and it reports close agreement with a JWST-based unsupervised study. If the claim holds, the epoch of disk formation is earlier and the morphological evolution of galaxies over this interval is much weaker than previously thought.","feed_headline":"Disk fraction holds near 60 percent out to z=2.4","feed_subtitle":"A no-human-eyes method blames visual bias for earlier claims that disks disappeared at high redshift.","key_machinery":"The argument is carried by a hybrid pipeline. First, four non-parametric metrics - the second moment of light $M_{20}$, entropy, Gini coefficient, and gradient pattern asymmetry $G_2$ - are measured in F814W cutouts within a Petrosian ellipse. A Self-Organizing Map clusters these metric vectors, and an IsoData threshold per metric selects 'prominent clusters' that become the disk and spheroid training labels, with no human labeling. For each redshift bin of width 0.2, an ensemble of one hundred convolutional networks is trained on these labels, and the averaged output probability assigns final classes: below 0.1 disk, above 0.9 spheroid, between 0.2 and 0.8 irregular. The FERENGI code is then used to artificially redshift low-redshift galaxies to test how flux dimming and angular-size changes alter the same metrics and classifications.","core_discovery":"The central discovery is that, once human visual judgment is removed from labeling, the global morphology mix of massive galaxies is remarkably stable: disks constitute roughly 60 percent, spheroids roughly 30 percent, and irregulars roughly 10 percent over 0.2<z<2.4, with the fitted slope for disk fraction $m_{\\rm disk}=0.08\\pm0.03$ consistent with a flat trend. The paper further finds that a single classifier trained on the full redshift range disagrees with bin-dedicated classifiers for about 25 percent of galaxies, mostly above z~1, and that simulated cosmological degradation converts up to about 18 percent of disk galaxies into apparent spheroids or irregulars. It attributes the declining disk fractions reported by visual-classification studies to these effects, and points to the close match with an unsupervised JWST analysis as independent support. In the mass-resolved sample, the fraction of massive spheroids ($M_{\\rm stellar}\\geq10^{10.5}\\,M_\\odot$) rises by about 40 percent toward lower redshift while the massive disk fraction falls by about 20 percent, a complementarity the paper reads as evidence that merging massive disks builds spheroids.","pith_inferences":["If the flat global fractions are real, then the widely reported 'disk decline' is largely a selection and methodology effect; a direct consequence is that disk assembly was already complete by z~2 for galaxies above $10^9\\,M_\\odot$, which is a stronger statement than the paper explicitly makes.","The irregular class, defined as CNN probability between 0.2 and 0.8, is a method-internal category; part of its mild increase at z>1 could be degraded disks rather than genuinely disturbed systems, as the FERENGI experiment hints.","A testable extension would be to apply the same metric-plus-SOM pipeline to rest-frame near-infrared JWST images of the same CANDELS fields; agreement would strengthen the claim, while disagreement would reveal a band-dependent bias in the F814W labels.","Since the method uses only one band, the agreement with JWST results suggests wavelength-dependent metric variation is modest; this could be quantified by measuring MEGG metrics in multiple bands for the same galaxies at fixed rest-frame wavelength."],"forward_implications":["Earlier visual-based catalogs that report a decreasing disk fraction beyond z~1 should be re-examined; under this method the disk fraction is flat to z=2.4.","Supervised classifiers trained on visually labeled high-redshift samples inherit a bias that this method avoids by deriving labels from metrics alone.","Morphology models that cover wide redshift ranges in one shot mislabel about a quarter of galaxies compared to redshift-bin-specific models; future surveys should bin in redshift.","Corrections for surface-brightness dimming and angular-size degradation are needed before comparing fractions across redshift; up to 18 percent of disks can be lost to apparent spheroids or irregulars.","Massive disk and spheroid fractions change in opposite directions, implying mergers transform massive disks into spheroids, while the overall mix stays constant because low-mass galaxies remain disk-dominated."],"supporting_citations":[{"why":"Supplies the underlying EGG-based hybrid method, the expected ~90 percent accuracy, and the parameter definitions that this work updates.","marker":"K24"},{"why":"Objective unsupervised JWST classification whose redshift and mass-resolved fractions match the paper's results; the key external comparison.","marker":"Lee et al. (2024)"},{"why":"Provides the FERENGI code that simulates flux dimming, angular-size changes, and luminosity evolution for degraded galaxies.","marker":"Barden et al. (2008)"},{"why":"Earlier morphological evolution study used as the comparison for the z~1 transition and for the cross-bin classifier mismatch trend.","marker":"Conselice et al. (2011)"},{"why":"Visual-classification CANDELS catalog that the paper contrasts against, representing the visually biased high-redshift labels.","marker":"Kartaltepe et al. (2015)"},{"why":"Visual classification at z>2 reporting high spheroid and low disk fractions; one of the earlier claims the paper challenges.","marker":"Mortlock et al. (2013)"}],"fun_headline_variants":["Disk fraction holds at 60% out to z=2.4","No human eyes: disk fraction stable to z=2.4","Visual bias blamed for false disk decline to z=2.4","Cosmological dimming misclassifies 18% of disk galaxies","Massive disk mergers may build spheroids"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the prominent SOM clusters, built from metrics measured in the single F814W band, are the true disk and spheroid populations at every redshift, and that CNN probabilities between 0.2 and 0.8 faithfully mark irregulars.","fun_headline_variants_meta":{"raw":{"variants":["Disk fraction holds at 60% out to z=2.4","No human eyes: disk fraction stable to z=2.4","Visual bias blamed for false disk decline to z=2.4","Cosmological dimming misclassifies 18% of disk galaxies","Massive disk mergers may build spheroids"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001072,"raw_usage":{"total_tokens":4603,"prompt_tokens":1173,"completion_tokens":3430,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":789,"completion_tokens_details":{"reasoning_tokens":3341}},"tokens_in":789,"tokens_out":3430,"duration_ms":23816,"temperature":1.0,"reasoning_tokens":3341,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:05:56.731351+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same CANDELS galaxies and measure the same MEGG metrics on rest-frame near-infrared JWST images; if at z>1 the F814W-based disks are classified as spheroids or irregulars by the rest-frame metrics in large numbers, or if the rest-frame disk fraction falls with redshift, the constant-fraction claim collapses. A simpler check: degrade local disks with FERENGI and have experienced astronomers classify them by eye; if disks remain recognizable, the claim that visual inspection is systematically biased is weakened.","supporting_citations":[{"cited_title":"H., Park C., Hwang H","cited_arxiv_id":null,"evidence_quote":"Objective unsupervised JWST classification whose redshift and mass-resolved fractions match the paper's results; the key external comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier morphological evolution study used as the comparison for the z~1 transition and for the cross-bin classifier mismatch trend."},{"cited_title":"S., et al., 2015, The Astrophysical Journal Supplement Series, 221, 11","cited_arxiv_id":null,"evidence_quote":"Visual-classification CANDELS catalog that the paper contrasts against, representing the visually biased high-redshift labels."}],"review_version":1}