{"id":"2495b95e-84bc-4d1e-8130-1b2cb46fe136","arxiv_id":"1908.05626","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using average rather than instantaneous star formation rates reproduces [OII] luminosities within 5% for typical simulated galaxies, and large-scale clustering is robust to the luminosity estimation method.","lead":"This paper tests whether the [OII] emission line luminosity of simulated galaxies can be estimated from average star formation rates instead of instantaneous ones. It finds the approximation works within 5% for typical galaxies, and that large-scale clustering is insensitive to the estimation method.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The <5% average-vs-instantaneous SFR accuracy claim is validated only inside SAG; SAGE and Galacticus use coarser 10-step time averaging, so the central transferability assumption is untested.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing concern: the 5% agreement is measured only in SAG, whose 25-step time resolution is finer than the 10-step resolution of SAGE and Galacticus, and the paper extends the conclusion to those models without a direct test. I agree that this is the most exposed point in the central claim. The concern is concrete and testable: if coarser time averaging changes the average-vs-instantaneous L[OII] ratio beyond 5%, then the abstract's general statement is not supported, and the SAGE/Galacticus luminosity functions in Sec. 3.4 carry an unquantified systematic error. The clustering conclusion is more robust because it is based on ratios between proxies and the same get_emlines reference within each model, so even a failure of the transferability test would not necessarily overturn the large-scale clustering result. I therefore keep the reader's CONDITIONAL verdict: the paper is valuable and carefully executed internally, but the headline accuracy claim needs a boundary condition or an explicit test in a second model. No ad hominem and no accusation of inconsistency is intended; this is an extrapolation risk that the authors could settle with a modest computational check.","tokens_in":39531,"tokens_out":4721,"duration_ms":47782,"concrete_test":"Run SAG with coarser time sampling matching SAGE/Galacticus: modify SAG to output the SFR averaged over 10 equally spaced substeps instead of 25, recompute get_emlines L[OII], and compare the instantaneous-to-average luminosity function ratios in Fig. 7. If the <5% region for attenuated L[OII] <= 10^42.2 erg/s shifts, shrinks, or disappears, the abstract should restrict the accuracy claim to SAG and the SAGE/Galacticus results should be re-interpreted. A complementary check is to instrument SAGE or Galacticus for a subset of snapshots to output the last-substep SFR, then compare L[OII] derived from the last-substep SFR versus the full-average SFR; this directly tests the transferability assumption without relying on SAG.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim that post-processing L[OII] from average SFRs is accurate to <5% for dust-attenuated L[OII] <= 10^42.2 erg/s rests entirely on an internal test within SAG. In Sec. 3.3, the instantaneous and average SFRs of SAG are fed to get_emlines, and the resulting luminosity functions agree within 5% over the stated range. However, Sec. 2.1.4 states that SAG subdivides the time between snapshots into 25 steps (each ~10-25 Myr at z~1), while SAGE and Galacticus use only 10 steps and output only average SFRs. The average SFR over coarser steps is a poorer proxy for the recent star formation that powers [OII] emission; the paper itself notes at the start of Sec. 3 that a time-averaged SFR 'can include contributions from stellar populations older than those responsible for generating the nebular emission.' No test shows that SAG's 5% margin survives the factor-2.5 coarser time sampling or the different star formation histories of SAGE and Galacticus. Nevertheless, Sec. 3.4 applies average SFRs to SAGE and Galacticus, and the abstract and conclusions phrase the 5% result as a general statement. Additionally, the dust parameters cos(theta)=0.60 and omega=0.80 are tuned to match the same DEEP2+VVDS luminosity functions used for comparison (Sec. 3.2), so the LF agreement in Fig. 9 is not an independent validation. The clustering conclusion in Sec. 4.4.2 is less exposed because it compares proxies against get_emlines within each model, but the get_emlines reference itself inherits the untested average-SFR substitution. The load-bearing concern is therefore an extrapolation risk in the central accuracy claim, not an internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses three semi-analytic models (SAG, SAGE, Galacticus) run on the same MultiDark Planck 2 simulation to predict the properties of [O II] emitters at 0.6 < z < 1.2, comparing with DEEP2-Firefly and DEEP2+VVDS observations. The authors compute L[O II] with the get_emlines code, which ideally requires instantaneous SFRs; only SAG provides these, while SAGE and Galacticus output only average SFRs. Using SAG, they test average versus instantaneous SFR as inputs and report <5% differences in the resulting luminosity functions for a limited luminosity range. They also derive simple proxies for L[O II] from SFR and u/g magnitudes, and test these proxies against luminosity functions, halo occupation distributions, and the projected two-point correlation function, finding that the clustering amplitude above about 1 h^-1 Mpc is robust to the choice of proxy.","tokens_in":39907,"tokens_out":6905,"duration_ms":65620,"significance":"If the stated accuracy of using average SFRs holds beyond SAG, the paper provides useful practical guidance for adding emission-line luminosities to semi-analytic mock catalogues for DESI/Euclid-type surveys. The work has clear strengths: the three SAMs share the same dark-matter simulation and halo catalogues, the average-versus-instantaneous SFR test in Sec. 3.3 is a controlled internal comparison, the public data release includes DEEP2-Firefly quantities and model emission-line luminosities, and the dust implementation is made available on GitHub. The clustering conclusion is well supported by the proxy comparisons in Sec. 4.4.2. However, the central 5% accuracy claim is currently stated more generally than what is actually tested, and part of the luminosity-function agreement with observations is calibrated rather than predicted. These issues are fixable but require changes to the presentation and, ideally, an additional test.","major_comments":[{"comment":"The 5% average-versus-instantaneous SFR result is established only within SAG, which subdivides the time between snapshots into 25 steps. SAGE and Galacticus split time into 10 steps and output only average SFRs, as stated in Sec. 2.1.4, and the paper itself notes at the beginning of Sec. 3 that a time-averaged SFR can include contributions from stellar populations older than those responsible for the nebular emission. No test shows that SAG's 5% margin survives the factor-of-2.5 coarser time sampling or the different star formation histories of SAGE and Galacticus, yet Sec. 3.4 applies average SFRs to these models and the abstract and Sec. 5 state the 5% result as a general finding. Please either add a direct test using a coarser SFR output from a model that also tracks a recent burst component, or explicitly restrict the claim to SAG and to the tested SFR time resolution.","section":"Secs. 2.1.4, 3.3, 3.4 and abstract"},{"comment":"The dust attenuation parameters cos(theta)=0.60 and omega_lambda=0.80 are chosen to obtain the best agreement with the DEEP2+VVDS luminosity functions, as stated in Sec. 3.2. The agreement shown in Fig. 9 is therefore partly a calibration product rather than an independent validation of the get_emlines L[O II] prescription. The paper should explicitly acknowledge this in the conclusions where the luminosity-function agreement is highlighted, and ideally show the sensitivity of the comparison to these two parameters. The internal average-versus-instantaneous ratio analysis in Sec. 3.3 is less affected by this issue because both inputs share the same attenuation model.","section":"Sec. 3.2 and Fig. 9"},{"comment":"The abstract states that post-processing L[O II] from average SFRs is accurate to <5% for dust-attenuated L[O II] <~ 10^42.2 erg/s, but Sec. 3.3 reports <5% agreement only in the narrower range 10^41-10^42.2 erg/s and states that the discrepancy grows to about 20% at fainter luminosities. The lower bound is essential and should not be dropped. The paper also gives conflicting lower bounds for the same range (10^41 in Sec. 3.3 and the summary versus 10^40.9 in the final paragraph of Sec. 3.3) and conflicting SFR ranges (10^-0.2 to 10^1.6 in Sec. 3.3 versus 1 to 10^1.5 in Sec. 5). These numbers need to be reconciled because the applicability range is part of the central claim.","section":"Abstract, Sec. 3.3, Sec. 5"}],"minor_comments":[{"comment":"The caption says the relations use 'the instantaneous SFR for sage and average SFR for sage and galacticus'; the first occurrence should be SAG rather than sage.","section":"Table 2 caption"},{"comment":"The SAGE panels are labelled 'SAGE inst z=0.94' although SAGE provides only average SFRs; this is inconsistent with the main text and should be corrected to avoid confusion.","section":"Fig. C1 and Table C1 caption"},{"comment":"The text refers to the 'BTP diagram' when describing the Baldwin-Phillips-Terlevich diagram; this should read 'BPT diagram'.","section":"Sec. 3.1"},{"comment":"The DEEP2-Firefly comparison is described with the redshift range '0.9 < z < 11' in the text accompanying Fig. 8; this should be '0.9 < z < 1.1'.","section":"Sec. 3.3, Fig. 8"},{"comment":"When describing the proxy luminosity functions, the shading is said to come from 100 Gaussian realisations but the covariance between the proxies and L[O II] is not discussed; a brief statement that the scatter is treated as independent would clarify the uncertainty estimate.","section":"Sec. 4.4.1"}],"recommendation":"major_revision","confidential_remarks":"The internal SAG test in Sec. 3.3 is credible and the clustering robustness result is well supported. The main revision should focus on limiting the 5% claim to what is actually tested, or adding a test for coarser time averaging, and on presenting the Fig. 9 luminosity-function agreement as partly calibrated rather than as an independent validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read if you build or use emission-line mock catalogs. The new quantitative result is that feeding a SAM's average SFR into get_emlines recovers the [OII] LF from instantaneous SFR within 5% for attenuated L[OII] ≤ 10^42.2, at least in SAG. The clustering robustness across methods is a useful practical conclusion.\n\nWhat the paper does well: it is honest about the components. The average-vs-instantaneous test in Sec 3.3 is clean, with the SFR function and LF ratios, and the 5% band is clearly defined. They show the redshift evolution of that agreement in Appendix B. The comparison across three SAMs on the same MDPL2 halo catalog is valuable. They ship the derived catalogs and DEEP2-FF data. The proxy analysis (SFR, u/g magnitudes) is systematic and the conclusion that simple L[OII] estimates suffice for large-scale clustering but not for LFs is sensible and well-hedged.\n\nSoft spots. The 5% claim is established only in SAG, which uses 25 substeps per snapshot. SAGE and Galacticus use 10 steps and output only average SFRs; the paper applies the 5% finding to them in Sec 3.4 without testing that the coarser time sampling preserves the agreement. The paper even notes that average SFR can include older stellar populations. So the abstract's general phrasing slightly overstates what is demonstrated. The dust attenuation parameters (cosθ=0.60, ω=0.80) are tuned to match the same DEEP2+VVDS LFs used for validation, so the absolute LF agreement in Fig 9 is partly fitted. That doesn't break the average-vs-instantaneous comparison, which uses the same dust prescription for both inputs, but it does weaken the claim of independent LF validation. The clustering conclusion is on safer ground because it compares proxies against each model's own get_emlines reference, though that reference itself inherits the untested average-SFR substitution in SAGE and Galacticus.\n\nWho it's for: anyone constructing synthetic ELG catalogs for DESI/Euclid/PFS, or interpreting HOD/clustering of [OII] emitters. Not a breakthrough, but a solid calibration paper with genuinely useful numbers.\n\nRecommendation: send to a competent referee. The main caveat should be addressed—either by validating the average-SFR substitution in a second SAM or by softening the general claim. I'd also like to see the dust parameters marginalized or at least acknowledged as fitted. None of this is disqualifying; the paper deserves peer review.","headline":"Solid methods paper with a genuinely useful new result on average vs instantaneous SFR for [OII] mocks, and one untested extrapolation that should be fixed before the 5% claim goes general.","tokens_in":40608,"tokens_out":2620,"would_cite":true,"duration_ms":24265,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Estimating [O II] luminosity from a galaxy's average star formation rate agrees with the full photoionisation calculation to within five percent for typical emitters, and the large-scale clustering of these galaxies is independent of the…","keywords":["emission line galaxies","[O II] luminosity","semi-analytic galaxy formation models","star formation rate","galaxy clustering","luminosity function","halo occupation distribution","DEEP2 survey"],"falsifier":"Re-run the two models that lack instantaneous star formation output with a time-substepping scheme that resolves the last ~10–25 Myr of star formation, recompute [O II] luminosities with get_emlines using that resolved SFR, and compare the resulting dust-attenuated luminosity function at L[OII] < $10^{42}$.2 erg/s with the one obtained from their average SFR; a discrepancy larger than 5% would show the claim is model-specific.","tokens_in":39350,"feed_emoji":"🔭","tokens_out":13070,"duration_ms":102610,"temperature":0.7,"pith_summary":"This paper asks whether [O II] emission-line luminosities of model galaxies can be computed after the fact from time-averaged star formation rates, rather than the instantaneous rates that a standard nebular-emission code prefers. Using the one galaxy formation model that outputs both quantities, the authors show that substituting the average rate changes dust-attenuated [O II] luminosities by less than five percent for galaxies with L[OII] below about $10^{42}$.2 erg/s. They then test simple linear proxies for [O II] luminosity—star formation rate and rest-frame UV u- and g-band magnitudes—across three semi-analytic models of galaxy formation, and find that the clustering amplitude of [O II] emitters on scales above 1 $h^{-1}$ Mpc is unchanged no matter which proxy is used. The practical consequence is that fast, cheap luminosity estimates are adequate for building mock catalogues and forecasting large-scale clustering for upcoming emission-line surveys, even though they are not accurate enough to predict luminosity functions.","feed_headline":"Post-processing [O II] from average SFRs is accurate to 5 percent","feed_subtitle":"Clustering on large scales is unaffected by how [O II] luminosity is computed, so quick proxies can build survey mocks.","key_machinery":"The central object is the photoionisation-based mapping from star formation rate and cold-gas metallicity to [O II] luminosity, implemented in the public get_emlines code. In this prescription, $L(\\lambda_j) = 1.37\\times 10^{-12}\\, Q_{H^0}\\, F(\\lambda_j,q,Z_{\\rm cold})/F({\\rm H}\\alpha,q,Z_{\\rm cold})$, where $Q_{H^0} \\propto \\mathrm{SFR}$ is the hydrogen-ionising photon rate and the flux ratio comes from a pre-computed grid of H II region models with ionisation parameter $q(Z) = q_0 (Z_{\\rm cold}/Z_0)^{-\\gamma}$. The argument proceeds by comparing this mapping fed with the instantaneous SFR (available in SAG only) against the same mapping fed with the average SFR, and then by testing linear proxies built from SFR and observed-frame u/g magnitudes as substitutes for the full calculation.","core_discovery":"The central claim is that the post-processing computation of [O II] luminosity from average star formation rates is accurate for model galaxies with dust-attenuated L[OII] ≲ $10^{42}$.2 erg/s (less than 5% discrepancy), and that in all three models considered the amplitude of the clustering at scales above 1 $h^{-1}$ Mpc remains unchanged independently of the method used to derive L[OII]. This is established by comparing, within the SAG model, the [O II] luminosity functions obtained by feeding the get_emlines photoionisation prescription with instantaneous versus average SFRs; the two agree within 5% over the range $10^{41}$–$10^{42}$.2 erg/s for attenuated luminosities. The clustering robustness is demonstrated by computing projected two-point correlation functions for galaxies selected at L[OII] > $10^{40}$.4 erg/s, using direct code output, three linear proxies (SFR, u-band and g-band absolute magnitudes), and an observational SFR–metallicity conversion; across SAG, SAGE, and Galacticus, the methods agree within roughly 5–12% on scales above 1 $h^{-1}$ Mpc.","pith_inferences":["Because the 5% agreement was only established in one model, models with coarser output steps or burstier star formation histories (the other two considered here) could exceed that threshold; re-running them with shorter substeps would test this directly.","The clustering robustness means that BAO and redshift-space-distortion forecasts for upcoming emission-line surveys can safely use simple SFR or magnitude proxies, even though any measurement relying on accurate luminosity functions cannot.","The u/g magnitude proxies, if calibrated on real photometric surveys, could provide a way to estimate [O II] luminosity for galaxies without spectra, assuming the model correlations hold in the real Universe.","The finding that the 5% agreement region widens with redshift suggests the average-SFR approximation becomes more reliable at higher z, possibly because star formation is steadier; this could be checked by extending the SAG comparison beyond z = 1.2."],"forward_implications":["Semi-analytic galaxy formation models that output only average star formation rates can be post-processed with get_emlines to obtain [O II] luminosities within 5% of the instantaneous-rate result for typical emitters, so no re-simulation with finer time steps is needed for those galaxies.","Simple linear L[OII] proxies based on SFR or rest-frame UV u/g magnitudes yield luminosity functions whose shapes depend on the model, so they are not reliable for predicting number densities of bright [O II] emitters.","The projected two-point clustering of [O II] emitters above 1 h^-1 Mpc is robust to the luminosity-estimation method, with agreement within roughly 5–12% across models and estimators, supporting the use of cheap proxies in survey forecasts.","The halo occupation distribution of [O II] emitters changes little when different luminosity proxies are used, which justifies using simple proxies when building mock catalogues that need HODs."],"supporting_citations":[{"why":"Supplies the get_emlines photoionisation code and calibration that converts star formation rate and metallicity into [O II] luminosity; it is the method whose inputs are being tested.","marker":"Orsi et al. 2014"},{"why":"The SAG model, the only one of the three that outputs instantaneous star formation rates, providing the test bed for the average-versus-instantaneous comparison.","marker":"Cora et al. 2018"},{"why":"The SAGE model, one of the two models that only provide average SFRs, used to test whether the post-processing result transfers.","marker":"Croton et al. 2016"},{"why":"The Galacticus model, the other average-SFR-only model, used to test transferability of the approximation.","marker":"Benson 2012"},{"why":"Presents the public catalogues that run the three models on the same dark-matter simulation, supplying the galaxy populations analysed.","marker":"Knebe et al. 2018"},{"why":"The DEEP2 survey data that provide the observed [O II] emitters and spectra used for comparison with model predictions.","marker":"Newman et al. 2013"},{"why":"One of the observational SFR-plus-metallicity conversions used as an L[OII] proxy and as a baseline in clustering and luminosity-function comparisons.","marker":"Kewley et al. 2004"},{"why":"The Firefly spectral fitting code that produces the DEEP2-Firefly physical properties (SFR, stellar mass, metallicity) compared with the models.","marker":"Wilkinson et al. 2017"}],"fun_headline_variants":["Average SFRs yield [O II] luminosities within 5%","[O II] from average SFRs: 5% accurate, clustering robust","Quick [O II] proxies preserve large-scale clustering","For [O II] luminosity, average SFRs suffice within 5%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The five-percent agreement was measured inside only one of the three models, where 'instantaneous' star formation means the mass of stars formed over the last ~10–25 Myr; the paper assumes the same insensitivity to SFR time resolution holds for the other two models, which subdivide time into ten steps and have different star formation histories, and this transferability is asserted rather than tested.","fun_headline_variants_meta":{"raw":{"variants":["Average SFRs yield [O II] luminosities within 5%","[O II] from average SFRs: 5% accurate, clustering robust","Quick [O II] proxies preserve large-scale clustering","For [O II] luminosity, average SFRs suffice within 5%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001229,"raw_usage":{"total_tokens":5160,"prompt_tokens":1167,"completion_tokens":3993,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":783,"completion_tokens_details":{"reasoning_tokens":3912}},"tokens_in":783,"tokens_out":3993,"duration_ms":26432,"temperature":1.0,"reasoning_tokens":3912,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:07:44.710847+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the two models that lack instantaneous star formation output with a time-substepping scheme that resolves the last ~10–25 Myr of star formation, recompute [O II] luminosities with get_emlines using that resolved SFR, and compare the resulting dust-attenuated luminosity function at L[OII] < $10^{42}$.2 erg/s with the one obtained from their average SFR; a discrepancy larger than 5% would show the claim is model-specific.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the get_emlines photoionisation code and calibration that converts star formation rate and metallicity into [O II] luminosity; it is the method whose inputs are being tested."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Presents the public catalogues that run the three models on the same dark-matter simulation, supplying the galaxy populations analysed."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The DEEP2 survey data that provide the observed [O II] emitters and spectra used for comparison with model predictions."},{"cited_title":"J., Geller M","cited_arxiv_id":null,"evidence_quote":"One of the observational SFR-plus-metallicity conversions used as an L[OII] proxy and as a baseline in clustering and luminosity-function comparisons."},{"cited_title":"M., Maraston C., Goddard D., Thomas D., Parikh T., 2017, , 472, 4297","cited_arxiv_id":null,"evidence_quote":"The Firefly spectral fitting code that produces the DEEP2-Firefly physical properties (SFR, stellar mass, metallicity) compared with the models."}],"review_version":1}