{"id":"a113852c-b9e9-4b9c-8329-548a00a6346f","arxiv_id":"2412.12060","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":11,"one_line_summary":"For star-forming galaxies with rest-frame EW(Hα) ≳ 500 Å (or about 8% nebular continuum), spectral fits that ignore nebular emission systematically overestimate stellar masses by ~0.6-1 dex.","lead":"This study compares two spectral fitting codes on 500 star-forming galaxies to quantify when the glow from ionized gas, the nebular continuum, must be included when measuring galaxy properties. It finds a clear threshold, at about 500 Å rest-frame Hα equivalent width or an 8% nebular contribution, above which ignoring that glow biases stellar mass estimates by 0.6 dex or more.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central threshold rests on FADO's nebular model as ground truth; no external photoionization benchmark is provided, so a systematic bias in FADO's X_neb would shift the 8%/500 Å threshold and the reported 0.6 dex mass bias.","rationale":"I find the central inference—that omission of nebular continuum causes systematic errors that grow with EW(Hα)—to be well supported by the two paired comparisons (FADO vs STARLIGHT, and FADO FC vs PS), with the internal control plausibly isolating the nebular component from code-specific differences. The remaining weak point is the calibration of the nebular contribution itself: the paper never benchmarks FADO's Krüger et al. (1995) nebular continuum against an independent photoionization code or against an observable like the Balmer jump. This matters because the reported threshold and mass-error amplitudes are expressed in FADO's X_neb units; the reader's weakest_assumption identifies this same point. I do not see an internal inconsistency that would overturn the result, and the CONDITIONAL verdict already captures the need for external validation. The concrete Cloudy-mock test would settle whether the concern lands; if it passes, the headline threshold can be trusted at the stated precision, and if it fails, the threshold and mass-bias numbers would need recalibration rather than abandonment of the qualitative conclusion. I therefore recommend no change to the reader's CONDITIONAL verdict.","tokens_in":26228,"tokens_out":7512,"duration_ms":73117,"concrete_test":"Build a set of mock SDSS-like spectra with known stellar populations (BC03 SSPs, Chabrier IMF) and nebular continua generated independently of FADO—e.g., with Cloudy photoionization models—spanning EW(Hα) from ~50 to ~1500 Å, with known true X_neb and true stellar mass. Run FADO in full-consistency and pure-stellar modes on these mocks, and compare recovered X_neb and M* to the truth. If the recovered X_neb is biased by more than ~20–30% at fixed EW(Hα), or if the recovered stellar-mass offset between FC and PS differs from the true bias by more than ~0.2 dex near the threshold, the paper's quantitative threshold and mass-error calibration need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The quantitative headline—X_neb≈8% and EW(Hα)≈500 Å for a significant impact, with stellar mass overestimated by ~0.6 dex on average and up to 2 dex—is calibrated entirely on FADO's full-consistency output. FADO computes the nebular continuum from standard photoionisation prescriptions (Krüger et al. 1995: two-photon, free-free, free-bound; Sect. 3) within a self-consistent loop, but the paper does not validate that component against an independent photoionization calculation or against a direct measurement of the nebular continuum. The FADO FC-vs-PS comparison (Sect. 5, Fig. 11) is a good control for code-specific fitting differences and strengthens the claim that nebular continuum matters, but it cannot detect a systematic error in FADO's nebular model itself: both modes share the same prescription. If FADO overproduces (or underproduces) nebular continuum for a given stellar population, then X_neb, the EW→X_neb conversion in Eq. 1, the 8% threshold, and the magnitude of the inferred mass error are all expressed on a biased scale. This is the load-bearing assumption because every quantitative conclusion in Sect. 6, including the z≈2–6 high-redshift extrapolation, is downstream of it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper seeks to establish a quantitative threshold for when neglecting the nebular continuum in optical spectral fitting leads to significant biases in derived galaxy properties. Using 500 star-forming SDSS-DR7 galaxies deliberately spread across more than two decades in EW(Halpha), the authors fit each spectrum with FADO (full-consistency stellar+nebular modelling) and STARLIGHT (stellar-only), and additionally run FADO in pure-stellar mode as a control. They find that the rest-frame EW(Halpha) and EW(Hbeta) are tight tracers of the median optical nebular contribution X_neb, and claim that for X_neb of about 8% (EW(Halpha) about 500 A, or about 375 A in a pseudo-continuum convention) the neglect of nebular emission becomes significant. Above this threshold, STARLIGHT overestimates stellar masses by 0.6 dex on average and up to 2 dex in the most extreme cases, with additional effects on age and metallicity. The paper then extrapolates the threshold to high redshift, arguing that galaxies with M* = 10^7-10^11 Msun cross it on average at z about 2-6.","tokens_in":26695,"tokens_out":9960,"duration_ms":91904,"significance":"If the threshold result holds, it provides a practical, citable benchmark for observers and modellers: a rest-frame EW(Halpha) cut that separates regimes where nebular continuum modelling is or is not required. This is especially timely for JWST and upcoming MOONS spectroscopy of high-redshift star-forming galaxies. The study has genuine methodological strengths: both codes are applied with the same spectral basis and extinction law, a second SSP library is used as a robustness check, and the FADO pure-stellar-mode comparison (Fig. 11) helps separate the effect of including nebular emission from code-specific fitting differences. The tracer relations in Eqs. (1)-(5) are simple and usable. The main weaknesses are statistical: the headline threshold is tied to manually chosen bin edges rather than a formal breakpoint, the paper's own 'significant difference' criterion is applied inconsistently, and the absolute X_neb scale is calibrated entirely on FADO's nebular prescription without external validation. These issues are addressable in revision, and I do not see grounds for rejection.","major_comments":[{"comment":"The threshold criterion is not consistently applied. Under the paper's own definition of a significant difference (difference larger than 0.2 dex in either direction), the 100<=EW(Halpha)<500 A bin already contains 55% significantly different galaxies for M_curr (21% FD, 34% ST) and 57% for M_ever (22% FD, 35% ST). Thus, a statistically significant difference between the two codes is identifiable well below EW(Halpha)=500 A; what actually changes at 500 A is that the differences become one-sided, with STARLIGHT overestimating mass for more than 70% of galaxies. The statement that 'we can identify the value EW(Halpha)=500 A as the threshold for which a statistically significant difference between the two codes is identifiable' conflates the presence of significant differences with a directional bias. A formal change-point or sliding-window analysis of the continuous FADO-STARLIGHT difference as a function of EW(Halpha), with an uncertainty on the breakpoint, is needed to support the threshold claim.","section":"Section 5.1, Table 1"},{"comment":"The 500 A value is not fitted; it is one of the manually chosen bin edges (100, 500, 1000 A), and no uncertainty is quoted for the threshold or for its X_neb equivalent. Because Eq. (1) is steep in log-log space, the mapping to X_neb about 8% inherits a nontrivial uncertainty, and based on Table 1 the transition could plausibly lie anywhere within the 100-500 A interval. Please derive the threshold and its confidence interval from the continuous relation between the FADO-STARLIGHT difference and EW(Halpha) or X_neb, and quote the corresponding uncertainty on X_neb.","section":"Section 5.1, Figs. 7-10; Eq. (1)"},{"comment":"The absolute calibration of X_neb, and therefore Eq. (1), the 8% threshold, and the 0.6-2 dex mass-bias estimates, rests entirely on FADO's nebular continuum prescription (Krueger et al. 1995; two-photon, free-free, and free-bound emission). The pure-stellar-mode control in Fig. 11 is a good check that including a nebular component matters, but it cannot detect a systematic error in FADO's nebular model itself because both modes share the same prescription. If FADO systematically over- or under-predicts the nebular continuum for a given stellar population, the reported X_neb scale and threshold would shift. This is a correctness-risk concern rather than a claim of internal circularity: the internal comparison is valid, but the absolute scale is model-dependent. I recommend validating X_neb against an independent photoionisation calculation (e.g., Cloudy run with the same ionising SEDs) or against direct nebular-continuum measurements, and at minimum stating this limitation explicitly and assessing how plausible variations in the nebular prescription would move the threshold.","section":"Section 3 and Section 6.1"},{"comment":"The paper's own significance criterion does not support the claim that neglecting nebular emission significantly affects stellar metallicity. For light- and mass-weighted metallicity, the mean FADO-STARLIGHT differences are 0.03 and 0.003 dex, and even in the EW(Halpha)>=1000 A bin the light-weighted mean difference reaches only 0.09 dex. The fractions of galaxies with significant differences in the highest bins never exceed 51% for Z_L and 49% for Z_M, and the distributions retain a strong peak near zero. The abstract's statement that stellar mass, age, and metallicity are all significantly impacted is therefore an overclaim. The age-metallicity degeneracy argument in Sect. 6.1 is an inference rather than a direct measurement; please either soften the conclusions or provide direct evidence that the fitted stellar populations change above the threshold in a way that propagates systematically to the metallicity estimates.","section":"Section 6.1, Table 1, and Fig. 9"},{"comment":"The conversion of the threshold to a pseudo-continuum EW appears arithmetically inconsistent. The text states that FADO's EW estimates are on average 0.1 dex (about 25%) higher than pseudo-continuum-based estimates. Starting from 500 A, dividing by 10^0.1 = 1.26 gives about 397 A, whereas applying a 25% reduction directly gives 375 A; the paper quotes 375 A. Please clarify which operation is intended and re-derive the scaled thresholds for Hbeta, [OIII] lambda5007, and HeI lambda5876 quoted in the summary, since these values will be used by observers who adopt pseudo-continuum conventions.","section":"Section 6.1 and Section 3"}],"minor_comments":[{"comment":"There is a grammatical issue in the abstract ('it has not been established a clear threshold'), and the text contains the typo 'redshfit' in Section 1 and again in Section 6.2.","section":"Abstract and Section 1"},{"comment":"The quoted thresholds in terms of EW(Hbeta), sSFR, EW([OIII] lambda5007), and EW(HeI lambda5876) are derived by plugging X_neb=8% into the fitted relations, but no uncertainties are propagated from the fit parameters. Please provide confidence intervals for these derived threshold values.","section":"Equations (1)-(5) and Section 6.1"},{"comment":"The construction of Fig. 12 is described only verbally: the Faisst et al. (2016) redshift evolution is normalized using mean EW(Halpha) values from Cardoso et al. (2022) in stellar-mass bins. Please specify the bin definitions, the number of galaxies per bin, and how the scatter in the EW(Halpha)-redshift relation is treated, so the claimed z about 2-6 crossing can be assessed quantitatively.","section":"Section 6.2, Fig. 12"},{"comment":"The choice of 0.2 dex as the 'significant difference' threshold is reasonable and well referenced, but it is introduced after the analysis rather than as a pre-registered criterion; a sentence noting that the qualitative conclusions are not sensitive to using 0.15 or 0.25 dex would strengthen the presentation.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the manuscript is within the scope of A&A and addresses a timely question for high-redshift spectroscopy. My main concerns are the statistical rigour of the headline threshold (manual bin edges, inconsistent use of the 'significant difference' criterion), the overclaim regarding metallicity, and the absence of external validation of FADO's nebular prescription. All of these are addressable in a revision, and I would be willing to review a revised version. I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful paper. The authors establish a simple criterion—rest-frame EW(Hα)≳500 Å (or X_neb≳8%)—above which neglecting the nebular continuum biases stellar masses by ~0.6 dex on average, up to 2 dex. That criterion is exactly what the high-z community needs for JWST-era fitting.\n\nWhat's new: prior work showed nebular emission matters for EELGs, but nobody had calibrated the threshold or provided EW-based tracers with fitted relations. The linear relations between X_neb and EW(Hα)/EW(Hβ)/sSFR are new and immediately usable, and the pseudo-continuum scaling (500→375 Å) is a practical service.\n\nThe strongest part is the double comparison. FADO vs STARLIGHT could be criticized as comparing two codes with different algorithms, but the FADO pure-stellar mode control removes that: within the same code, turning on the nebular component produces the same EW(Hα)=500 Å threshold and ~0.75 dex mass differences at the highest EWs. That is a clean demonstration that the effect is driven by nebular modelling, not by code quirks. The robustness check with a different SSP basis also helps.\n\nWhere it's soft: the threshold is defined by bin edges, not a formal breakpoint fit, and the 'significant' criterion of 0.2 dex is adopted from earlier accuracy tests, not derived here. The highest EW bin has only 55 galaxies. The abstract includes metallicity among significantly affected properties, but the body says there is no clear metallicity threshold—that needs fixing. The stress-test worry about FADO being ground truth is fair for the X_neb=8% calibration (which inherits FADO's Krüger et al. photoionisation prescriptions), but it does not sink the EW=500 Å threshold: that is anchored by the differential comparisons, which remain meaningful even if FADO's absolute nebular scale shifts. A cross-check with Prospector or BAGPIPES would stiffen the calibration. The z=2–6 extrapolation depends on Faisst et al.'s EW evolution and is more speculative, though clearly labelled as average expectations.\n\nWho it's for: anyone fitting optical spectra of star-forming galaxies, especially EELG and high-z JWST work. I'd send it to review; the core result is solid, and the revisions are about honest presentation rather than redoing the analysis.","headline":"A practical, usable threshold for when nebular continuum matters; the mass effect is real and well-controlled, but the abstract overstates the metallicity result and the X_neb scale rests on FADO.","tokens_in":27214,"tokens_out":2811,"would_cite":true,"duration_ms":25708,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that once the median optical nebular contribution passes about 8% (rest-frame EW(H$\\alpha$) $\\simeq$ 500 Å), spectral fits that model only starlight systematically overestimate stellar masses by $\\sim$0.6 dex on average…","keywords":["nebular continuum","spectral fitting","stellar mass","equivalent width","star-forming galaxies","FADO","STARLIGHT","high-redshift galaxies"],"falsifier":"Refit the same SDSS spectra with a second, independently calibrated photoionisation model for the nebular continuum and compare the resulting $X_{neb}$ and stellar masses with FADO and STARLIGHT. If the independent model does not yield comparable nebular fractions and does not reproduce the $\\sim$0.6 dex lower masses relative to pure-stellar fits above EW(H$\\alpha$)$\\simeq$500 Å, the threshold as stated is not robust to the choice of nebular model.","tokens_in":26061,"feed_emoji":"🌌","tokens_out":11712,"duration_ms":90902,"temperature":0.7,"pith_summary":"The paper aims to put a number on when the nebular continuum becomes impossible to ignore in galaxy spectral fitting. Fitting 500 star-forming SDSS-DR7 galaxies with FADO (stellar plus nebular, self-consistent) and STARLIGHT (pure stellar), it finds that for a median optical nebular contribution $X_{neb}\\gtrsim8\\%$, corresponding to rest-frame EW(H$\\alpha$)$\\simeq$500 Å, the two codes diverge: pure-stellar fits return stellar masses higher by $\\sim$0.6 dex on average and up to 2 dex in the most extreme cases. It also establishes that EW(H$\\alpha$) and EW(H$\\beta$) are tight linear tracers of $X_{neb}$, giving observers a simple screen for when nebular modelling matters. If the threshold holds, most galaxies with $M_*\\sim10^7$–$10^{11}\\,M_\\odot$ will cross it on average at $z\\sim2$–6, making nebular continuum modelling a standard requirement for high-redshift spectroscopy.","feed_headline":"Nebular glow can inflate galaxy masses by 0.6 dex","feed_subtitle":"Above EW(Hα) ~ 500 Å, spectral fits that skip the nebular continuum overestimate galaxy stellar mass.","key_machinery":"The central object is the median optical nebular contribution $X_{neb}$, defined as the median over 3000–9000 Å of the ratio of FADO's nebular continuum to its total continuum. FADO, the reference code, self-consistently fits stellar and nebular emission using standard photoionisation prescriptions (two-photon, free-free and free-bound emission) and ensures the best-fit stellar population reproduces the observed nebular features. STARLIGHT, applied with a purely stellar base, is the contrast case. The argument is carried by binning the sample by rest-frame EW(H$\\alpha$) and by the FADO-versus-STARLIGHT difference in derived properties, with differences above 0.2 dex treated as significant.","core_discovery":"On its own terms, the paper's discovery is that the optical nebular continuum is a first-order ingredient, not a small correction, once it contributes a median $\\sim$8% of the 3000–9000 Å continuum. Comparing FADO and STARLIGHT on identical spectra, the derived stellar masses agree for EW(H$\\alpha$)$<$100 Å, then diverge systematically above the threshold: for EW(H$\\alpha$)$\\geq$500 Å, more than 70% of galaxies get pure-stellar masses higher by an average of $\\sim$0.6 dex, reaching up to 2 dex for EW(H$\\alpha$)$\\geq$1000 Å, with mass-weighted stellar ages also affected in the most extreme bin. The mechanism is that a flat nebular continuum makes stellar-only codes compensate with older stellar populations, and splitting the observed continuum into stellar plus nebular components lowers the stellar light and its mass-to-light ratio, hence the lower masses. A control run with FADO in pure-stellar mode reproduces the same threshold, indicating the difference is driven by the nebular component rather than only by code-specific behaviour.","pith_inferences":["The paper does not test whether its threshold transfers to photometric SED fitting, but if it does, galaxies at $z>2$ whose stellar masses are derived from pure-stellar templates would be systematically overestimated whenever their rest-optical EWs are high.","A practical extension the authors leave implicit: the EW(H$\\alpha$)$\\simeq$500 Å (or 375 Å) cut could serve as a trigger in large surveys, sending only galaxies above it to self-consistent nebular fits and keeping stellar-only fits for the rest.","Because the control compares FADO with itself, the absolute location of the threshold is partly contingent on FADO's nebular prescription; an independent multi-code comparison would show how much of the 8% value is physical rather than code-specific.","If future JWST rest-optical spectra confirm the same $X_{neb}$–EW relation at $z>3$, the $\\sim$0.6 dex mass correction would propagate into derived SFRs and specific star formation rates, potentially shifting high-redshift scaling relations."],"forward_implications":["For galaxies with rest-frame EW(H$\\alpha$)$\\geq$500 Å (or $\\sim$375 Å under pseudo-continuum definitions), stellar masses from pure-stellar fits are overestimated by $\\sim$0.6 dex on average and up to 2 dex.","EW(H$\\alpha$) and EW(H$\\beta$) can be used to estimate $X_{neb}$ through the paper's linear relations, so a single emission-line measurement can flag when nebular modelling is required.","On average, galaxies with stellar masses between $10^7$ and $10^{11}\\,M_\\odot$ reach the threshold at $z\\sim$2–6, implying that high-redshift surveys should routinely include nebular continuum in their spectral models.","Mass-weighted stellar ages diverge significantly for EW(H$\\alpha$)$\\geq$1000 Å, while light-weighted ages and metallicities show weaker or mixed differences; stellar mass is the cleanest property affected.","Below the threshold, differences between the codes are within the 0.2 dex significance level, but the paper notes that subtle effects may still exist and are hard to assess."],"supporting_citations":[{"why":"Defines FADO and its self-consistent stellar plus nebular continuum modelling; supplies the reference method for estimating the nebular contribution.","marker":"Gomes & Papaderos 2017"},{"why":"Provides STARLIGHT, the pure-stellar spectral fitting code whose differences from FADO define the impact threshold.","marker":"Cid Fernandes et al. 2005"},{"why":"Supplies the SDSS-DR7 FADO fits from which the 500-galaxy sample is drawn, plus earlier FADO-versus-STARLIGHT comparisons.","marker":"Cardoso et al. 2022"},{"why":"Supplies the photoionisation prescription for two-photon, free-free and free-bound emission used to compute the nebular continuum in FADO.","marker":"Krüger et al. 1995"},{"why":"Provides the empirical rest-frame EW(H$\\alpha$) redshift evolution used to predict when average galaxies cross the threshold at $z\\sim$2–6.","marker":"Faisst et al. 2016"},{"why":"Documents the offset between FADO and pseudo-continuum EW definitions, used to scale the 500 Å threshold to 375 Å.","marker":"Miranda et al. 2023"},{"why":"Shows that nebular modelling changes inferred properties for SDSS extreme emission-line galaxies, the population the threshold is designed to flag.","marker":"Breda et al. 2022"}],"fun_headline_variants":["Nebular continuum omission overstates galaxy mass by 0.6 dex","Above EW(Hα) 500 Å, galaxy masses inflated without nebular fit","Ignoring nebular light adds 0.6 dex to galaxy mass estimates","Nebular glow matters: mass bias up to 2 dex at high EW(Hα)"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that FADO's nebular continuum model is the true one; if its prescription for how ionised gas radiates is biased, the measured nebular fractions, the 8% threshold, and the claimed $\\sim$0.6 dex mass overestimates would all shift.","fun_headline_variants_meta":{"raw":{"variants":["Nebular continuum omission overstates galaxy mass by 0.6 dex","Above EW(Hα) 500 Å, galaxy masses inflated without nebular fit","Ignoring nebular light adds 0.6 dex to galaxy mass estimates","Nebular glow matters: mass bias up to 2 dex at high EW(Hα)"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000927,"raw_usage":{"total_tokens":4054,"prompt_tokens":1111,"completion_tokens":2943,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":727,"completion_tokens_details":{"reasoning_tokens":2856}},"tokens_in":727,"tokens_out":2943,"duration_ms":18852,"temperature":1.0,"reasoning_tokens":2856,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:18:59.088077+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Refit the same SDSS spectra with a second, independently calibrated photoionisation model for the nebular continuum and compare the resulting $X_{neb}$ and stellar masses with FADO and STARLIGHT. If the independent model does not yield comparable nebular fractions and does not reproduce the $\\sim$0.6 dex lower masses relative to pure-stellar fits above EW(H$\\alpha$)$\\simeq$500 Å, the threshold as stated is not robust to the choice of nebular model.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines FADO and its self-consistent stellar plus nebular continuum modelling; supplies the reference method for estimating the nebular contribution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides STARLIGHT, the pure-stellar spectral fitting code whose differences from FADO define the impact threshold."},{"cited_title":"L., Capak, P., Hsieh, B","cited_arxiv_id":null,"evidence_quote":"Provides the empirical rest-frame EW(H$\\alpha$) redshift evolution used to predict when average galaxies cross the threshold at $z\\sim$2–6."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the offset between FADO and pseudo-continuum EW definitions, used to scale the 500 Å threshold to 375 Å."}],"review_version":1}