{"id":"db936199-4b49-4159-9399-d520f4f533da","arxiv_id":"1908.07046","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A random forest trained on BPT-classified SDSS galaxies can classify z=0.3-0.8 emission line galaxies into four subtypes using only optical features, with the best performance among four machine learning methods tested.","lead":"Astronomers trained machine learning classifiers on low-redshift galaxies to sort intermediate-redshift galaxies into star-forming, composite, AGN, and LINER types using only optical spectra and colors. The random forest classifier reaches 93% accuracy for star-forming galaxies and roughly 66 to 72% for the rarer classes, which is useful for upcoming surveys such as DESI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper quotes low-redshift test-set accuracies as if they apply at 0.32<z<0.8, but the intermediate-z mapping is untested and the only evidence, stacked spectra, is acknowledged to be affected by selection and contamination.","rationale":"The reader's weakest_assumption identifies the same core issue: the trained mapping between the eight features and BPT-derived class is assumed to be redshift-invariant, but this is not directly validated for 0.32<z<0.8. I agree, and this is genuinely load-bearing because the paper's stated purpose is to classify intermediate-z galaxies, while all reported accuracies come from a low-z test set under a different S/N selection. The paper is transparent about the limitation in Sec. 5, and the stacked-spectrum comparison is honestly described as affected by selection effects and contamination. There is no internal inconsistency in the ML methodology; the cross-validation, code release, and feature analysis are sound. My proposed internal redshift-split test can be run immediately with the public code and would either support or weaken the transfer assumption, so the condition the reader placed on acceptance is appropriate. I therefore do not change the reader's CONDITIONAL verdict.","tokens_in":18755,"tokens_out":3862,"duration_ms":39909,"concrete_test":"Retrain the RF on the lower-redshift half of the z<0.32 sample (e.g., z<0.16), apply the intermediate-z selection cuts from Sec. 2.3 (S/N>3 on [OII], Hbeta, [OIII], no [NII]/Halpha/[SII] requirement) to the upper-redshift half (0.16<=z<0.32), and compute per-class accuracy against BPT labels. If accuracies drop by more than the quoted 1-sigma uncertainties for any subtype, the redshift transfer is not established; the definitive follow-up would be NIR spectroscopy (e.g., MOSDEF, LEGA-C) of a representative intermediate-z subsample to compute a true confusion matrix.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 2 accuracies (93.4% SFG, 69.4% composite, 71.8% AGN, 65.7% LINER) are computed on the z<0.32 test set with the same S/N selection as the training set (Sec. 2.1), not on the 0.32<z<0.8 sample described in Sec. 2.3. The intermediate-z sample requires S/N>3 on [OII], Hbeta, and [OIII] only, omits the [NII]/Halpha/[SII] cuts, and covers a redshift range where the physical mapping between the eight features and BPT class is assumed, not demonstrated, to be identical. The most important feature, [OIII]/Hbeta (Fig. 8), is sensitive to ionization parameter and can evolve with redshift, shifting high-z star-forming galaxies toward the AGN locus; the paper presents no direct validation of this transfer. The stacked-spectrum comparison (Fig. 14) is indirect: the authors themselves note that the intermediate-z composites and LINERs have significantly higher equivalent widths than their low-z counterparts and are contaminated by AGNs and SFGs (Sec. 5). Thus the central claim that the RF classifies 0.32<z<0.8 ELGs with the quoted accuracies is not supported by the measurements in this paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains four supervised machine learning classifiers (KNN, SVC, random forest, MLP) on 28,869 SDSS/eBOSS galaxies at z<0.32, labeled into star-forming, composite, AGN, and LINER types using the BPT diagnostic diagram. Input features are [OIII]/Hb, [OII]/Hb, [OIII] line width, stellar velocity dispersion, and four k-corrected colors. The random forest achieves the best AUC scores and per-subtype accuracies of 93.4%, 69.4%, 71.8%, and 65.7%. The authors apply the trained random forest to 49,272 galaxies at 0.32<z<0.8 and compare stacked spectra with BPT-classified low-redshift stacks, concluding that the intermediate-redshift classifications are correct. The public code and trained models are released.","tokens_in":19063,"tokens_out":6564,"duration_ms":60279,"significance":"If valid, the method would fill a real gap: enabling four-type emission-line galaxy classification from optical spectra alone at 0.3<z<0.8, which is currently prevented by the shift of [NII], Ha, and [SII] out of the optical window. This would be useful for DESI, PFS, and 4MOST. The paper's strengths are the systematic comparison of four algorithms with k-fold cross-validation and error bars, the feature-importance analysis, and the release of code and trained models. The main weakness is that the central intermediate-redshift claim rests on an unvalidated transfer of a low-redshift decision boundary, not on direct measurement.","major_comments":[{"comment":"The accuracies quoted in the abstract (93.4%, 69.4%, 71.8%, 65.7%) and in Section 4.6 are measured on the z<0.32 test sample described in Section 2.1. That sample is selected with S/N>3 on [NII], Halpha, and [SII], and its labels come from the BPT diagram. These numbers are not measured on the 0.32<z<0.8 sample of Section 2.3, which requires S/N>3 only on [OII], Hb, and [OIII] and has no BPT labels. The only evidence for intermediate-z performance is the stacked-spectrum comparison in Figure 14, which the authors themselves state (Section 5) is affected by the stronger-line selection and by AGN/SFG contamination of the composite and LINER stacks. Therefore the central claim that the RF classifier classifies 0.32<z<0.8 ELGs with the quoted accuracies is not supported by the measurements. The paper should either provide direct validation (e.g., near-IR spectroscopy or a simulated transfer test) or be reframed as a low-redshift classifier with an unvalidated application.","section":"Section 5, Table 2"},{"comment":"The statement that the RF classifier 'gives as consistent a classification as the BPT diagram' is partly circular: the training labels for all z<0.32 galaxies are defined by the same Kauffmann et al. (2003) and Kewley et al. (2006) demarcation lines used to construct Figure 12. Reproducing those lines on the held-out test set is a sanity check, not an independent validation. The comparison with the KEx diagram in Figure 13 is also expected because two of the eight RF features, [OIII]/Hb and sigma([OIII]), are exactly the axes of that diagram; therefore the consistency does not constitute external confirmation of the intermediate-z classification.","section":"Section 4.8, Figure 12"},{"comment":"The application to 0.32<z<0.8 assumes that the mapping between the eight features and the BPT-defined physical class is identical to the mapping at z<0.32, but this assumption is not tested. In particular, [OIII]/Hb, which has the highest feature importance (Figure 8), is sensitive to ionization parameter and may evolve with redshift, potentially shifting high-z star-forming galaxies toward the AGN locus. The paper presents no check of this redshift invariance, such as a comparison to the existing small samples with NIR spectroscopy (e.g., MOSDEF), so the transfer of the decision boundary remains an unsupported extrapolation.","section":"Section 5"}],"minor_comments":[{"comment":"The text reports average AUC scores of 0.892, 0.895, 0.931, and 0.892 for KNN, SVC, RF, and MLP, but Table 1 lists the MLP average as 0.906; the text value should be corrected.","section":"Section 4.6, Table 1"},{"comment":"The column headers in Tables 1 and 2 list LINERs before AGNs, while the text consistently gives the order SFGs, composites, AGNs, LINERs; reorder the columns for consistency.","section":"Table 1, Table 2"},{"comment":"The sentence 'the accuracies of the RF classification for SFGs are 93.4%, 69.4%, 71.8%' should read 'for SFGs, composites, and AGNs, respectively,' because three values are listed.","section":"Section 4.8"},{"comment":"The citation 'Baldwin, Philips, & Terlevich 1981' contains a typo: the second author is Phillips, as in the reference list.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The low-redshift benchmark is solid and the public code is a useful contribution. The main issue is the mismatch between the paper's framing and what is actually measured: the quoted accuracies apply to the z<0.32 test set, not to the intermediate-redshift sample, and the stacked-spectrum evidence is indirect and partly circular. I recommend major revision, with a clear statement that the intermediate-redshift accuracies are unverified and with a quantitative test (NIR data or simulations) or an explicit reframing of the contribution as a low-z classifier with a speculative high-z application."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read: the valuable part is the low-z benchmark, not the intermediate-z claim. The paper builds a four-way BPT classifier from optical-only features, compares KNN/SVC/RF/MLP properly, reports ROC/AUC and confusion matrices, and ships code and trained models. I don't know of an earlier optical-only diagram that separates SF, composite, AGN, and LINER into four classes at once, so that part is a real contribution for DESI/PFS/4MOST planning. The low-redshift evaluation is honest: six-fold CV, class-balanced training, 4-feature versus 8-feature comparison, and a useful feature-importance analysis. The confusion matrix makes clear that LINERs and composites bleed into each other, which is the kind of information observers need. The prior literature is cited sensibly, and the comparison to KEx and MEx diagrams is fair. That part deserves to be used and cited.\n\nThe stress-test concern lands. Table 2's accuracies—93.4, 69.4, 71.8, 65.7—are from the z<0.32 test set with the heavier S/N cuts, not from the 0.32<z<0.8 sample. The intermediate-z sample has different cuts and no BPT ground truth, so those numbers should not be quoted as intermediate-z accuracies. The stacked-spectrum comparison (Fig. 14) is indirect and, as the authors acknowledge, the composite and LINER stacks are contaminated and selection-biased. Since [OIII]/Hbeta is the most important feature and is ionization-parameter sensitive, the redshift-invariance assumption deserves at least a prominent caveat. The agreement with KEx and MEx diagrams is suggestive but those diagrams are not ground truth either.\n\nI want to be clear about what is not wrong. The training labels and low-z test labels both use BPT, so Fig. 12 is partly a sanity check rather than independent confirmation—but the low-z classification problem is well posed and the performance numbers are reliable where they actually apply. The circularity concern is real but minor; the extrapolation concern is the main issue. The authors state the ideal test is NIR spectra, which is the right caveat.\n\nBottom line: this is a solid, useful paper for observers and ML practitioners in galaxy evolution. It deserves a serious referee, but the abstract and conclusions should be revised so intermediate-z accuracy is not implied. If the referee asks for that reframing plus a look at a NIR-confirmed sample when available, the paper will be stronger for it.","headline":"A solid low-redshift classifier benchmark with public code; the intermediate-redshift application is plausible but not validated, and the accuracy claims should be reframed accordingly.","tokens_in":19652,"tokens_out":3123,"would_cite":true,"duration_ms":33261,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A random forest trained on low-redshift galaxies can classify intermediate-redshift emission-line galaxies into four excitation classes using only optical data.","keywords":["emission line galaxies","galaxy classification","random forest","active galactic nuclei","LINERs","intermediate redshift","optical diagnostics","machine learning"],"falsifier":"Take a few hundred $0.32<z<0.8$ galaxies with published rest-frame near-infrared spectra covering $[\\mathrm{N\\,II}]/\\mathrm{H}\\alpha$ and $[\\mathrm{S\\,II}]/\\mathrm{H}\\alpha$, classify them by the standard BPT criteria, and compare with the random forest predictions; if the per-class agreement lies well below the reported 93.4%, 69.4%, 71.8%, and 65.7% accuracies—particularly because composite and LINER predictions drift systematically with redshift—the transfer assumption is falsified.","tokens_in":18605,"feed_emoji":"🔭","tokens_out":11429,"duration_ms":97998,"temperature":0.7,"pith_summary":"At redshifts between 0.3 and 0.8, the emission-line ratios used in the standard optical diagnostic diagrams—$[\\mathrm{N\\,II}]/\\mathrm{H}\\alpha$ and $[\\mathrm{S\\,II}]/\\mathrm{H}\\alpha$—are shifted out of the optical window, so galaxies could not be assigned to the conventional four excitation classes (star-forming, composite, AGN, LINER) without expensive near-infrared spectra. The paper claims that this four-way classification can instead be learned from eight features that remain measurable: $[\\mathrm{O\\,III}]/\\mathrm{H}\\beta$, $[\\mathrm{O\\,II}]/\\mathrm{H}\\beta$, the $[\\mathrm{O\\,III}]$ line width, the stellar velocity dispersion, and four rest-frame colors. A random forest trained on low-redshift galaxies labeled by the standard diagrams reaches accuracies of 93.4% for star-forming galaxies, 69.4% for composites, 71.8% for AGNs, and 65.7% for LINERs, and the stacked spectra of intermediate-redshift galaxies so classified match the low-redshift stacks. If right, this gives upcoming wide-field optical surveys a practical way to do emission-line-galaxy science at intermediate redshift without near-infrared follow-up.","feed_headline":"Four galaxy types now classifiable at z=0.3–0.8 with optical data","feed_subtitle":"A random forest keeps 60–93% accuracy on four galaxy classes, letting optical-only surveys do the job.","key_machinery":"The load-bearing object is a random forest classifier—an ensemble of decision trees whose votes assign each galaxy a class—trained on low-redshift BPT labels with eight features: four spectroscopic ($[\\mathrm{O\\,III}]/\\mathrm{H}\\beta$, $[\\mathrm{O\\,II}]/\\mathrm{H}\\beta$, $[\\mathrm{O\\,III}]$ line width, stellar velocity dispersion $\\sigma_*$) and four rest-frame colors ($u-g$, $g-r$, $r-i$, $i-z$) k-corrected to $z=0.1$. The random forest was selected after comparing k-nearest neighbors, support vector classifier, and a multi-layer perceptron; it had the highest average area-under-the-ROC-curve score (0.931) and the best balance of per-class accuracies. Feature importance in the trained forest shows that $[\\mathrm{O\\,III}]/\\mathrm{H}\\beta$, the $[\\mathrm{O\\,III}]$ line width, and $g-r$ do most of the work, which is consistent with earlier two-dimensional diagnostics built from the same physics. The same machinery, with only the four spectroscopic features, retains most of the performance, which is what makes the method usable when imaging is absent.","core_discovery":"The central claim is that a random forest classifier, trained on 28,869 low-redshift ($z<0.32$) galaxies whose labels come from the BPT emission-line diagnostic diagram (the standard $[\\mathrm{O\\,III}]/\\mathrm{H}\\beta$ versus $[\\mathrm{N\\,II}]/\\mathrm{H}\\alpha$ classification), correctly transfers the four-way classification to 49,272 galaxies at $0.32<z<0.8$. The transfer works because the classifier learns the boundary between classes in a feature space made only of quantities available from optical spectra and broad-band photometry at those redshifts. On a held-out low-redshift test sample the reported accuracies are 93.4% (star-forming), 69.4% (composite), 71.8% (AGN), and 65.7% (LINER), with the four-feature spectroscopic-only version only a few percent lower. The paper's direct evidence for the high-redshift transfer is that the stacked rest-frame 3400–5050 Å spectra of the intermediate-$z$ classes closely match the stacked spectra of the same classes defined by BPT at low $z$, and that the intermediate-$z$ class counts fall where the kinematic–excitation and mass–excitation diagrams would put them. This establishes, conditionally on that consistency, that a four-subtype physical classification of intermediate-redshift galaxies is achievable from optical data alone.","pith_inferences":["The high importance of $[\\mathrm{O\\,III}]$ line width and $[\\mathrm{O\\,III}]/\\mathrm{H}\\beta$ suggests the random forest is effectively relearning the kinematic–excitation diagram, with $g-r$ substituting for stellar mass; a testable prediction is that the learned boundary in the line-width-versus-ratio plane should track that demarcation out to higher redshift.","Because the labels come from low-redshift BPT classifications, the classifier can only be as good as the assumption that the same line-ratio physics separates the classes at higher redshift; a targeted near-infrared sample of a few hundred $0.32<z<0.8$ galaxies would directly calibrate the transfer, something the paper lists as the ideal test but does not perform.","The method is survey-agnostic in the sense that the four spectroscopic features can be measured by any optical spectrograph, so retraining on a different instrument's line-flux system is a straightforward extension; the reported accuracies are tied to this particular survey's noise properties and should be re-measured per survey.","Composite galaxies, the hardest class at 69.4% accuracy, are exactly the transition population between star formation and AGN activity; preserving them as a separate class lets one map where that transition happens across cosmic time, provided the confusion matrix is folded into the analysis."],"forward_implications":["Surveys at $0.32<z<0.8$ can obtain four-type classifications for each galaxy in real time from spectra plus photometry, without waiting for near-infrared follow-up; the paper classifies all 49,272 intermediate-redshift galaxies this way.","A spectra-only random forest retains most of the accuracy (92.3%, 63.7%, 67.3%, 60.8%), so the method survives in fields without multi-band imaging.","The classifier reproduces BPT classifications at low redshift and its intermediate-redshift outputs line up with the kinematic–excitation and mass–excitation boundaries, so it can serve as a star-forming/AGN selection tool that also preserves the composite and LINER distinction.","The confusion matrix gives practical error budgets: composites leak into star-forming galaxies at 23.8%, AGNs leak into LINERs at 18.8%, and LINERs leak into composites at 28.4%, so class fractions in a survey sample can be corrected."],"supporting_citations":[{"why":"Introduces the BPT optical diagnostic diagram that defines the four excitation classes used as training labels.","marker":"Baldwin et al. 1981"},{"why":"Provides the alternative diagnostic diagrams whose demarcations also enter the four-way labeling.","marker":"Veilleux & Osterbrock 1987"},{"why":"Supplies the star-forming/composite/AGN demarcation line applied to the low-redshift training sample.","marker":"Kauffmann et al. 2003"},{"why":"Supplies the AGN/LINER boundary needed to split the active population into AGNs and LINERs.","marker":"Kewley et al. 2006"},{"why":"Provides the k-correction method used to bring the u, g, r, i, z photometry to rest-frame z=0.1 colors for both samples.","marker":"Blanton & Roweis 2007"},{"why":"Demonstrates the [OII]/H-beta versus [OIII]/H-beta diagnostic that justifies two of the spectroscopic input features.","marker":"Lamareille 2010"},{"why":"Shows color versus [OIII]/H-beta separation of star-forming galaxies and AGNs, supporting the use of rest-frame colors as features.","marker":"Yan et al. 2011"},{"why":"Defines the kinematic–excitation diagram with [OIII] line width and [OIII]/H-beta, which motivates features and provides a consistency check.","marker":"Zhang & Hao 2018"},{"why":"Defines the mass–excitation diagram used to check that intermediate-redshift random forest classes fall in the expected stellar-mass regions.","marker":"Juneau et al. 2011"},{"why":"Provides the spectral stacking code used to build the stacked spectra that validate the intermediate-redshift classifications.","marker":"Comparat et al. 2016"}],"fun_headline_variants":["Random forest best for four-way galaxy classification at z=0.3–0.8","Optical-only machine learning identifies AGN, LINER, composite, star-forming at z<0.8","ML classifies intermediate-redshift galaxies into BPT types from optical data","Four galaxy types from optical spectra and colors alone via ML"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole transfer rests on the assumption that the relation between the eight measurable features and the four physically defined classes is the same at $0.32<z<0.8$ as it is at $z<0.32$; at the higher redshifts there is no direct $[\\mathrm{N\\,II}]/\\mathrm{H}\\alpha$ or $[\\mathrm{S\\,II}]/\\mathrm{H}\\alpha$ ground truth, only stacked-spectrum consistency.","fun_headline_variants_meta":{"raw":{"variants":["Random forest best for four-way galaxy classification at z=0.3–0.8","Optical-only machine learning identifies AGN, LINER, composite, star-forming at z<0.8","ML classifies intermediate-redshift galaxies into BPT types from optical data","Four galaxy types from optical spectra and colors alone via ML"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000931,"raw_usage":{"total_tokens":4137,"prompt_tokens":1246,"completion_tokens":2891,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":862,"completion_tokens_details":{"reasoning_tokens":2803}},"tokens_in":862,"tokens_out":2891,"duration_ms":22037,"temperature":1.0,"reasoning_tokens":2803,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:27:50.941047+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a few hundred $0.32<z<0.8$ galaxies with published rest-frame near-infrared spectra covering $[\\mathrm{N\\,II}]/\\mathrm{H}\\alpha$ and $[\\mathrm{S\\,II}]/\\mathrm{H}\\alpha$, classify them by the standard BPT criteria, and compare with the random forest predictions; if the per-class agreement lies well below the reported 93.4%, 69.4%, 71.8%, and 65.7% accuracies—particularly because composite and LINER predictions drift systematically with redshift—the transfer assumption is falsified.","supporting_citations":[{"cited_title":"M., Tremonti, C., et al.\\ 2003, , 346, 1055","cited_arxiv_id":null,"evidence_quote":"Supplies the star-forming/composite/AGN demarcation line applied to the low-redshift training sample."},{"cited_title":"R., & Roweis, S.\\ 2007, , 133, 734","cited_arxiv_id":null,"evidence_quote":"Provides the k-correction method used to bring the u, g, r, i, z photometry to rest-frame z=0.1 colors for both samples."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates the [OII]/H-beta versus [OIII]/H-beta diagnostic that justifies two of the spectroscopic input features."},{"cited_title":"C., Newman, J","cited_arxiv_id":null,"evidence_quote":"Shows color versus [OIII]/H-beta separation of star-forming galaxies and AGNs, supporting the use of rest-frame colors as features."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the kinematic–excitation diagram with [OIII] line width and [OIII]/H-beta, which motivates features and provides a consistency check."},{"cited_title":"M., & Salim, S.\\ 2011, , 736, 104","cited_arxiv_id":null,"evidence_quote":"Defines the mass–excitation diagram used to check that intermediate-redshift random forest classes fall in the expected stellar-mass regions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the spectral stacking code used to build the stacked spectra that validate the intermediate-redshift classifications."}],"review_version":1}