{"id":"e19889f1-f699-4013-b944-44104b7f9f54","arxiv_id":"2412.14676","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A neural network using six SED-derived galaxy properties predicts Ly-alpha emitters with 77% true positive and 14% false positive rates, validated against JWST spectroscopic samples.","lead":"A neural network trained on six galaxy properties (star formation rate, stellar mass, UV brightness, age, UV slope, and dust) predicts whether a galaxy emits Ly-alpha radiation, with 77% true positive and 14% false positive rates. The model could let astronomers identify large numbers of Ly-alpha-emitting galaxies from photometric data alone, helping map the epoch of cosmic reionization.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training-set selection against low-mass non-LAEs makes the reported FPR and the JWST low-mass LAE fractions depend on an unvalidated extrapolation; the model's main application to M*~1e8.5 galaxies is not independently supported.","rationale":"The reader's weakest-assumption analysis identifies exactly the same load-bearing issue: the absence of low-mass non-LAEs in the training set biases predictions in the mass range of the JWST application. The paper is transparent about this, which is a point in its favor, but transparency does not remove the extrapolation. The internal VANDELS-only AUC check shows the moderate-mass classification is stable, and the COSMOS2020/SC4K comparison is qualitatively sensible, so the central method is not without support. However, the 'independent' JWST validation lacks a non-LAE control in the low-mass bin; the reported 91% success rate is inflated by the model's prior toward low-mass LAEs, and the paper's own M* > 1e9 sub-sample gives 67%. The Ly-alpha fraction and the z>7 bubble-size argument inherit this same weakness because they identify low-mass blue galaxies without detected Ly-alpha as intrinsic LAEs. The reader's CONDITIONAL verdict remains appropriate; my read does not move it.","tokens_in":20654,"tokens_out":4594,"duration_ms":42978,"concrete_test":"Assemble spectroscopically confirmed non-LAEs at 3<z<6 with M*<1e9 Msun from public JWST/NIRSpec surveys in the same fields (e.g., JADES or CEERS), run the same CIGALE SED fitting and the same trained network on them, and compare the P(LAE) distribution with confirmed LAEs in the same mass bin. If the median P(LAE) for low-mass non-LAEs exceeds 0.7, or if the false positive rate at P(LAE)>0.7 is well above the claimed 14%, then the model does not transfer to the JWST mass regime and the low-mass LAE fractions and bubble-size inference are overestimates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the six SED-derived properties cover the target population. It fails in the mass regime where the model is actually applied. Non-LAEs come only from VANDELS, a magnitude-limited survey with median M* near 1e9 Msun; all MUSE galaxies are LAEs by construction. The paper itself states (Sec 3.3, Sec 4.2) that almost every galaxy with MUV > -19 or M* < 1e9 Msun is assigned P(LAE)>0.9 because no low-mass non-LAEs are in the training set. The JWST sample has median log M* = 8.5, below the training median of 9.1, so the headline '77% TPR / 14% FPR' and '91% of JWST LAEs have P>0.7' are not representative of the target mass range. The 91% validation is largely a restatement of the bias: for M* > 1e9 Msun the success rate drops to 67% (12/18). The inferred Ly-alpha fraction and the z=7.1/7.5 bubble-size interpretation both rely on classifying low-mass blue galaxies as intrinsically LAE (e.g., a 1e8.2 Msun galaxy with beta=-2.4 and P>0.9 is assumed to be an intrinsic LAE despite no Ly-alpha detection). If low-mass non-LAEs with blue slopes exist, predicted LAE fractions and the bubble-size conclusion are systematically biased. The limitation is acknowledged in Sec 4.2 but not corrected or quantified in the JWST application.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains a neural network to classify high-redshift galaxies as Lyα emitters (LAEs) or non-LAEs from six SED-derived properties (SFR, stellar mass, MUV, age, UV slope β, and E(B−V)). The training set combines 926 VANDELS galaxies (520 LAEs and 406 non-LAEs by visual inspection) and 507 MUSE galaxies (all LAEs, selected by line detection). At a probability threshold P(LAE)>0.7, the model reports 77% true positive rate and 14% false positive rate on a held-out test set, with AUC=0.88. The authors validate the model on independent samples: SC4K narrow-band LAEs and spectroscopically confirmed JWST LAEs (91% of the latter have P(LAE)>0.7). They apply the model to a JWST photometric sample to derive the Lyα fraction at 3<z<6, compare it with previous measurements after an EW>25 Å correction, and use predicted P(LAE) for spectroscopically observed galaxies at z≈7.1 and 7.5 to argue for moderate-size ionized bubbles (R≲1 pMpc) rather than a single large bubble.","tokens_in":21103,"tokens_out":4522,"duration_ms":40741,"significance":"If the method were robust across the mass range of interest, it would provide a practical way to construct large LAE samples from photometric data alone and to statistically correct Lyα fraction measurements for IGM attenuation during reionization. The paper has several strengths: it uses multiple independent validation samples (SC4K and JWST LAEs), it explicitly compares with a random-forest classifier and finds comparable performance, and it reports permutation feature importance with uncertainty. The 5-fold cross-validation and Monte Carlo noise injection are reasonable. However, the central application to the JWST sample is compromised by a known and acknowledged training-set bias: the absence of low-mass non-LAEs makes the model's output nearly deterministic in the very mass range where the JWST sample (median log M*=8.5) lies. The 91% JWST-LAE recovery is largely a restatement of this bias, and the mass-restricted success rate drops to 67%. The Lyα fraction and bubble-size conclusions therefore rest on extrapolations that are not independently supported.","major_comments":[{"comment":"The training set contains no low-mass non-LAEs: all MUSE galaxies are LAEs by construction and VANDELS is a magnitude-limited sample. The paper itself states (Sec. 3.3) that 'almost all of the galaxies with MUV > −19 or M* < 10^9 Msun are classified as LAEs with P(LAE)>0.9'. The JWST sample has median log(M*/Msun)=8.5, below the training median of 9.1 (Sec. 2.5). Consequently, the reported TPR/FPR and the model's predictions for the JWST sample are not validated in the regime where the model is actually applied. The 91% recovery of spectroscopically confirmed JWST LAEs is dominated by low-mass galaxies; for M*>10^9 Msun the success rate is 67% (12/18, Sec. 4.4.2). This is a load-bearing issue for the central application, and the authors should quantify the mass-dependent performance and either restrict the scientific conclusions to the validated mass range or construct a training set with low-mass non-LAEs (e.g., from deeper surveys) to demonstrate that the model's behavior is not an artifact of the missing class.","section":"Sec. 3.3 and Sec. 2.5"},{"comment":"The Lyα fraction comparison with previous studies applies a correction factor for the fraction of LAEs with EW0>25 Å measured from the training sample, with values 0.22 and 0.48 in the two MUV bins. This correction is applied to the model-predicted LAEs in the JWST sample, whose mass distribution differs markedly from the training sample. If the true EW>25 Å fraction depends on mass (as the paper's own MUV dependence suggests it may), then applying the training-sample ratio to a lower-mass population is not justified. The claim that 'the expected Lyα fraction X25_Lyα are consistent with the previous results' is therefore only as strong as the assumption that the correction is mass-independent within each MUV bin. The authors should derive the correction in mass-matched bins or explicitly test its sensitivity to the mass distribution of the predicted LAEs.","section":"Sec. 4.4.3"},{"comment":"The constraint on ionized bubble size at z≈7.18 and 7.49 interprets spectroscopically non-detected, low-mass galaxies with P(LAE)>0.9 as intrinsic LAEs whose Lyα is attenuated by the IGM, and then uses this to argue for moderate-size bubbles (R<1 pMpc) rather than a single large bubble. This interpretation assumes the model's predictions remain valid in the reionization-era galaxy population and in the low-mass regime. The paper acknowledges (Sec. 3.3) that the model assigns P(LAE)>0.9 to essentially all galaxies with MUV>−19 or M*<10^9 Msun, and a specific example used in the argument has M*=10^8.2 Msun and β=−2.4. This is precisely the regime where the model is unvalidated; the mass-restricted JWST validation (67% for M*>10^9) does not cover it. The bubble-size conclusion is load-bearing and should be re-evaluated, for example by showing that the conclusion is unchanged when using a conservative threshold that is calibrated on the mass range where the model has demonstrable predictive power, or by explicitly modeling the expected false-positive rate for low-mass blue non-LAEs.","section":"Sec. 4.4.5"}],"minor_comments":[{"comment":"The caption defines 'TP, FN, TP, and TN' in the FPR/TPR formulas, which appears to be a typo for 'TP, FN, FP, and TN'.","section":"Sec. 3.3 / Fig. 6 caption"},{"comment":"The flux calibration between VANDELS and MUSE uses only 25 objects and yields a median ratio of 0.55 with large scatter (16th–84th percentile 0.24–1.11). While this correction does not affect the binary LAE/non-LAE labels (which are based on visual inspection), it enters the EW>25 Å fraction used in Sec. 4.4.3; the large scatter should be propagated into that correction or discussed as a systematic uncertainty.","section":"Sec. 2.1"},{"comment":"The abstract states '91% of LAEs spectroscopically confirmed by JWST have a probability of LAE higher than 70%' but does not mention the mass-dependence of this success rate; the paper should qualify the statement to reflect the 67% rate for M*>10^9 Msun.","section":"Abstract and Sec. 4.4.2"},{"comment":"The permutation feature importance shows that SFR, E(B−V), and age are consistent with zero, but the paper correctly notes that correlated features can mask importance. This caveat is well placed, but the interpretation in Sec. 4.1 that 'age does not impact the model output' is stronger than the method supports given the large age uncertainties shown in Fig. 3.","section":"Sec. 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid technical contribution in the sense that it clearly documents the training set, the network architecture, and the validation procedure, and it does not oversell the raw classifier performance. However, the gap between the validated mass range and the mass range of the main application is substantial, and the paper's own text (Sec. 3.3, Sec. 4.2) admits the limiting bias. This is a correctable issue: the authors could re-do the JWST application in the validated mass range, provide a mass-dependent calibration of P(LAE), and downgrade the bubble-size claim to a tentative suggestion. Given that the manuscript is otherwise well-structured and the methodology is reproducible in principle, major revision is appropriate. I would encourage the editor to request the trained model or code as a condition of acceptance, as this would substantially increase the paper's utility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWhat you should know up front: this is a real, reasonably careful application of a neural network to LAE classification from SED-derived properties, and it is not circular—the classifier is tested on independent data (SC4K, JWST-confirmed LAEs). But the headline numbers carry a selection bias that the paper itself describes, and the JWST application and the HII bubble constraint lean on that bias more than the abstract lets on.\n\nWhat is genuinely new: the NN architecture for this task (prior work used kNN and random forests), the permutation feature importance ranking (β, MUV, M*), and the practical claim that you can build large LAE samples from broadband photometry alone. The training data construction is careful: Lyα fluxes measured from spectra, SED fits with CIGALE, MC dropout-style ensemble to propagate parameter errors. For moderate-mass galaxies (M* > 1e9 Msun) the classifier performs about as well as Napolitano's random forest, which the authors acknowledge. That is a fair, credible result.\n\nThe soft spot is exactly where the stress-test lands. The MUSE sample is all LAEs by construction; VANDELS supplies the non-LAEs but does not reach below roughly 1e9 Msun. So the training set has no low-mass non-LAEs, and the paper openly says (Sec 3.3, 4.2) that almost every galaxy with MUV > -19 or M* < 1e9 gets P(LAE) > 0.9. The JWST sample they apply the model to has median log M* = 8.5, below the training median of 9.1. That is an extrapolation, and the reported 77% TPR / 14% FPR are not representative of that regime. The 91% JWST success rate is dominated by the low-mass bias; for M* > 1e9 it falls to 67% (12/18). So the validation numbers, while honestly reported with this breakdown, do not support the low-mass inferences. The Lyα fraction comparison is also partly calibrated with EW>25 fractions measured from the training set, which makes it a consistency check. The z~7.1/7.5 bubble-size argument rests on a handful of galaxies and assumes the six-property → Lyα relation does not evolve with redshift; calling that a \"strong constraint\" is an overstatement.\n\nNone of this sinks the paper. The authors are transparent about the bias and do not hide the M*>1e9 drop. The core classifier is a reasonable, useful tool for moderate-mass samples, and the feature importance analysis is informative. What needs to change is the framing of the JWST and reionization applications: either restrict the claims to the mass range where the model is calibrated, or re-train/calibrate with deeper non-LAE samples, and soften the bubble-size wording to \"suggestive\" rather than \"strong constraint.\"\n\nI would send this to a serious referee. It is a legitimate contribution that will improve with revision. The referee should push on the mass extrapolation and the bubble-size statistics, but the paper is not a desk reject.","headline":"Useful NN classifier with honest limitations, but the low-mass extrapolation and the bubble-size claim are softer than the abstract suggests.","tokens_in":21558,"tokens_out":3120,"would_cite":true,"duration_ms":25951,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Photometry alone finds 91% of JWST-confirmed Lyα galaxies","keywords":["Lyα emitters","neural network","SED fitting","galaxy physical properties","cosmic reionization","JWST","Lyα fraction","ionized bubbles"],"falsifier":"Take a spectroscopically observed sample of galaxies at 3<z<6 with stellar masses below $10^{9}$ Msun that demonstrably lack Lyα emission (for example from JWST/NIRSpec follow-up of photometrically selected star-forming galaxies) and run them through the network; if a substantial fraction receive P(LAE)>0.7, the claimed 14% false positive rate is violated for the low-mass regime and the Lyα fraction predictions for JWST galaxies are biased high.","tokens_in":20452,"feed_emoji":"🔭","tokens_out":5488,"duration_ms":36775,"temperature":0.7,"pith_summary":"This paper claims that a neural network trained on spectroscopic surveys can predict whether a distant galaxy emits Lyα radiation using only six physical properties estimated from broadband photometry: star formation rate, stellar mass, UV magnitude, age, UV slope, and dust attenuation. The model achieves a 77% true positive rate and a 14% false positive rate when a threshold of P(LAE)>0.7 is used, and it flags 91% of spectroscopically confirmed JWST Lyα emitters as LAEs. If this works, astronomers could build large, continuous-redshift samples of Lyα emitters from photometric data alone, bypassing expensive spectroscopy and narrowband filters. The model also predicts a Lyα fraction that rises from z=3 to z=6 in line with previous work, and it suggests that ionized regions around z~7 galaxies are moderate-sized bubbles rather than a single large bubble.","feed_headline":"Photometry alone finds 91% of JWST-confirmed Lyα galaxies","feed_subtitle":"A six-property classifier trained on VANDELS and MUSE promises continuous-redshift LAE surveys and new bubble-size constraints.","key_machinery":"The central object is the neural network classifier itself: five hidden layers of 64 nodes each with eLU activation, 25% dropout, and a sigmoid output, trained with Adam on binary cross-entropy and ensembled by Monte-Carlo noise injection over the six input parameters plus 5-fold cross-validation. The six inputs—SFR, stellar mass, $M_{\\mathrm{UV}}$, age, UV slope $\\beta$, and $E(B-V)$—are derived by CIGALE SED fitting, with redshift fixed to spectroscopic or photometric values. The network's role is to replace single-parameter cuts (like a $\\beta$ threshold) with a multivariate boundary; permutation feature importance then identifies $\\beta$, $M_{\\mathrm{UV}}$, and $M_*$ as the decisive drivers. For the reionization application, the same network is used to predict intrinsic Lyα emission before IGM attenuation, and the mismatch between predicted LAEs and observed Lyα detections is read as a signature of neutral gas outside the bubbles.","core_discovery":"The paper's central claim is that a feedforward neural network with five hidden layers, trained on the VANDELS and MUSE spectroscopic samples, captures the nonlinear mapping between six SED-derived galaxy properties and the presence of Lyα emission well enough to serve as a photometric LAE classifier. At the operating point P(LAE)>0.7, the classifier reaches 77% completeness and 14% contamination in held-out test data, and the area under the ROC curve is 0.88. When applied to public JWST photometric catalogs, 91% of spectroscopically confirmed LAEs receive P(LAE)>0.7, and the model predicts an EW0>25 Å Lyα fraction whose redshift evolution matches published measurements. Applying the same classifier to reionization-era CEERS galaxies, the paper argues that comparing predicted intrinsic LAEs with observed Lyα detections favours separate moderate-sized ionized bubbles (R≲1 pMpc) over a single large bubble at z≈7.18 and z≈7.49.","pith_inferences":["If the training bias toward low-mass LAEs is corrected with deeper surveys, the same architecture could yield accurate predictions for the faintest galaxies, where JWST now provides the needed non-LAE spectra.","The model's explicit finding that P(LAE) does not correlate with Lyα flux or EW suggests the network is learning a binarized escape condition rather than a quantitative transfer function; a regression head trained on the measured fluxes might recover EW information the current classifier throws away.","The 91% JWST match rate is partly a consequence of the JWST sample being low-mass and blue-$\\beta$, exactly the regime where the training set has no non-LAEs; extending the test to mass-complete samples would sharpen the bubble-size inference.","Because the network inputs are SED-derived, systematic errors in SED fitting (e.g., photometric-redshift outliers) will propagate directly into P(LAE); tying the Monte-Carlo noise injection to full SED posterior draws rather than Gaussian parameter errors would make the probability output more reliable."],"forward_implications":["Photometric surveys can be mined for Lyα emitters over a continuous redshift window without spectroscopy or narrowband filters, enabling large LAE samples at 3<z<6.","The predicted Lyα fraction, corrected to EW0>25 Å, reproduces the rise from z=3 to z=6 seen in previous work, validating the model for population statistics.","Applying the model to reionization-era galaxies and comparing with observed Lyα detections distinguishes intrinsic LAEs from IGM-attenuated ones, yielding constraints on ionized bubble size; the paper finds R≲1 pMpc bubbles at z≈7.18 and z≈7.49.","Spatial maps of predicted LAEs can be compared with non-LAE overdensities to test whether LAEs trace the underlying matter distribution in different redshift slices."],"supporting_citations":[{"why":"Supplies the VANDELS spectroscopic catalogue and broadband photometry used to construct the training sample of LAEs and non-LAEs.","marker":"Garilli et al. (2021)"},{"why":"Provides the MUSE spectroscopic catalogue of galaxies, which contributes the low-mass LAE population to the training set.","marker":"Schmidt et al. (2021)"},{"why":"Supplies the Lyα flux and equivalent-width measurements for MUSE galaxies that define the LAE labels in training.","marker":"Kerutt et al. (2022)"},{"why":"CIGALE SED fitting code that produces the six input physical properties from broadband photometry for all galaxies.","marker":"Boquien et al. (2019)"},{"why":"Permutation feature importance method used to rank the six inputs and identify β, MUV, and M* as most influential.","marker":"Altmann et al. (2010)"},{"why":"Random-forest LAE classifier used as the performance baseline for the neural network comparison.","marker":"Napolitano et al. (2023)"},{"why":"JWST/NIRSpec spectroscopically confirmed LAEs that independently validate the model predictions.","marker":"Jones et al. (2024)"},{"why":"Spectroscopic observations of z≈7.1 galaxies in CEERS used to infer the size of ionized bubbles from predicted versus detected Lyα.","marker":"Chen et al. (2024)"}],"fun_headline_variants":["Photometry-only neural net finds 91% of JWST Lyα galaxies","Neural net recovers 9/10 Lyα emitters from photometry alone","Six galaxy traits predict Lyα, 91% match with JWST","No spectra needed: model catches most Lyα galaxies","Lyα finder: 77% true positive, 14% false positive"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The mapping between the six physical properties and Lyα emission learned from VANDELS and MUSE remains valid for the fainter, lower-mass JWST galaxies and for galaxies at higher redshift, even though the training sample contains almost no low-mass non-LAEs.","fun_headline_variants_meta":{"raw":{"variants":["Photometry-only neural net finds 91% of JWST Lyα galaxies","Neural net recovers 9/10 Lyα emitters from photometry alone","Six galaxy traits predict Lyα, 91% match with JWST","No spectra needed: model catches most Lyα galaxies","Lyα finder: 77% true positive, 14% false positive"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000463,"raw_usage":{"total_tokens":2384,"prompt_tokens":1082,"completion_tokens":1302,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":698,"completion_tokens_details":{"reasoning_tokens":1204}},"tokens_in":698,"tokens_out":1302,"duration_ms":10176,"temperature":1.0,"reasoning_tokens":1204,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:00:38.928066+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a spectroscopically observed sample of galaxies at 3<z<6 with stellar masses below $10^{9}$ Msun that demonstrably lack Lyα emission (for example from JWST/NIRSpec follow-up of photometrically selected star-forming galaxies) and run them through the network; if a substantial fraction receive P(LAE)>0.7, the claimed 14% false positive rate is violated for the low-mass regime and the Lyα fraction predictions for JWST galaxies are biased high.","supporting_citations":[],"review_version":1}