{"id":"21e7e5bc-bc80-41c8-9452-79db1526735b","arxiv_id":"2509.02741","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Fit-level differences among BPASS spectral libraries, IMFs, and metallicity treatments shift inferred stellar mass, age, and SFR by up to 0.27, 0.19, and 1.4 dex and change reconstructed mass assembly histories by up to 12 percent.","lead":"This paper measures how much the choice of stellar population model changes the galaxy properties astronomers infer from light. Fitting identical mock galaxies with different spectral libraries, IMFs, and metallicity assumptions shifts inferred mass, age, and star formation rate by more than the quoted observational errors, and can flip a galaxy from star-forming to quiescent.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline SFR uncertainty of 1.4 dex and the star-forming-to-quiescent flip stem from the BaSeL library, which the paper itself deems unsuitable for these data.","rationale":"The central claim is quantitative: SED modelling choices introduce scatter of 0.27 dex in mass, 0.19 dex in age, and 1.4 dex in SFR, exceeding survey uncertainties. The paper is a careful controlled experiment with useful validation in Appendices B through D, and the qualitative message that model uncertainties are non-negligible is credible. However, the specific numbers are dominated by one or two outlier models. The 1.4 dex SFR offset is BaSeL-only; the age offset of 0.19 plus or minus 0.11 is marginal; the mass offset of 0.27 dex is AP-only, with C3K and CKC within 0.04 dex of the reference. The most load-bearing weakness is therefore not the v2.2.1-as-truth assumption flagged by the Reader (a standard limitation of mock experiments), but the paper's own admission that BaSeL is unsuitable for the data quality it claims to model. Including BaSeL in the headline uncertainty budget without a model-selection step mixes a known-bad model with genuine alternatives. A secondary internal inconsistency supports the need for revision: Conclusion (iii) says fixing Z=0.014 causes roughly 0.4 and 0.08 dex offsets from the true values, yet Table 5 shows the fixed and variable-metallicity runs agree to about 0.01 dex, and the offset from EAGLE is common to all metallicity prescriptions. These issues argue for conditional acceptance with a request to recompute headline statistics on a plausibility-filtered model set and to correct the metallicity wording; they do not invalidate the paper's core finding.","tokens_in":36340,"tokens_out":13419,"duration_ms":112746,"concrete_test":"Recompute the Table 3 offsets and the star-forming/quiescent classification fractions excluding the BaSeL library, and also compute the standard deviation of inferred parameters across the remaining libraries (v2.2.1, AP, C3K, CKC) rather than the maximum offset from v2.2.1. If the SFR spread drops from 1.4 dex to roughly 0.5 dex and the classification flip disappears, the Abstract's headline numbers and the 'transforming a galaxy' claim should be rephrased as conditional on including a library the paper itself deems unsuitable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 9(i) and the Abstract report SFR variations of 1.4±1.0 dex and state that spectral library choice can flip a galaxy between star-forming and quiescent. Table 3 shows the 1.4 dex SFR offset belongs entirely to the BaSeL library (-1.4±1.0 dex); the next-largest SFR offset is AP at 0.52±0.19 dex. Section 4.2 nonetheless concludes that \"the low resolution of the BaSeL spectra renders them unsuitable for fitting to the high-quality simulated galaxies used in this study, as they fail to capture key spectral features necessary for robust parameter inference.\" The paper provides no goodness-of-fit or Bayesian model comparison allowing an observer to reject BaSeL for SDSS+VISTA-quality data, so the quoted absolute uncertainties conflate a plausible alternative library (AP) with a template grid that cannot resolve the data's features. Excluding BaSeL, mass variation is about 0.27 dex (AP) with C3K and CKC below 0.05 dex, and SFR variation falls to about 0.5 dex; the qualitative conclusion that model uncertainties exceed observational errors survives, but the specific Abstract numbers and the classification-flip claim are anchored to a model the authors themselves reject.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Using 18 z=0 galaxies from the EAGLE simulation, the authors generate mock SDSS spectra and VISTA YJH photometry with the BPASS v2.2.1 framework, then fit the mock observations with bagpipes using BPASS variants that vary the stellar spectral library (AP, BaSeL, C3K, CKC), the IMF shape, and the metallicity treatment (fixed, variable, evolving). All other fitting priors are identical. The paper reports offsets in derived stellar mass, age, SFR and extinction relative to the v2.2.1 reference fits, and shows that spectral-library choice can change mass, age and SFR by up to 0.27, 0.19 and 1.4 dex respectively, alter galaxy classifications between star-forming and quiescent, and shift reconstructed mass assembly histories by up to ~12 per cent. It concludes that SED modelling uncertainties exceed typical survey observational uncertainties and should be propagated in future survey analyses.","tokens_in":36565,"tokens_out":9645,"duration_ms":84475,"significance":"The controlled mock-fitting design is a genuine strength: identical priors, known input properties, a no-dust control, a spectrum-generation validation, and an IMF direction-reversal test make the sensitivity analysis easy to follow. The qualitative message—that hard-coded SPS assumptions introduce systematic offsets larger than quoted observational errors—is plausible and important for survey-based galaxy evolution studies, and the paper explicitly does not claim to identify the correct model. However, the quantitative headline is anchored to the BaSeL library, which the authors themselves conclude is unsuitable for these data, and the reported \"absolute uncertainties\" are mean offsets from one reference model rather than a model-to-model dispersion. These issues affect the abstract, the conclusions, and the survey-error-budget message, so they need to be addressed before the specific numbers can be used.","major_comments":[{"comment":"The headline SFR uncertainty of 1.4±1.0 dex and the star-forming-to-quiescent flip are driven entirely by the BaSeL grid, which the paper itself rejects. Table 3 assigns the 1.4 dex SFR offset to BaSeL, and Section 4.2 concludes that the low resolution of the BaSeL spectra renders them unsuitable for fitting to the high-quality simulated galaxies used in this study. If BaSeL is excluded, the largest library-induced offsets are 0.27 dex in mass and 0.52 dex in SFR (AP), with C3K and CKC below 0.05 dex; the qualitative conclusion survives, but the specific Abstract numbers and the classification-flip claim do not. Because the fits are framed as blind (Section 3.3), the paper needs either a quantitative model-comparison step (Bayesian evidence, chi-square, or posterior predictive checks) that justifies including or excluding BaSeL, or a restatement of the headline statistics based on libraries appropriate to SDSS-resolution data.","section":"§4.2, Table 3, Abstract, §9(i)"},{"comment":"The quantities quoted as \"absolute uncertainties\" are mean offsets relative to the v2.2.1 reference, with the quoted scatter being the standard error of that mean, not a dispersion across model choices. For example, the mass offsets for AP and BaSeL (0.27 and 0.23 dex) are both positive, so the 0.27 dex value is a systematic bias of one model rather than the width of the model-induced scatter; similarly the 1.4 dex SFR number is the offset of BaSeL rather than a spread. The paper should separate systematic offsets from random scatter and report a model-to-model dispersion (e.g., standard deviation or full range across libraries) if these numbers are to be used as \"uncertainties\" in survey error budgets.","section":"§9(i), Table 3"},{"comment":"Because the mock observations are generated with BPASS v2.2.1 and every fit uses a BPASS variant with the same evolutionary tracks, binary treatment, dust prescription and SFH parameterisation, the quoted numbers measure internal sensitivity within the BPASS framework, not the total uncertainty of SED modelling. The paper is transparent about the v2.2.1 assumption in Section 3.2, but the Abstract and the survey-implications discussion (Section 8) phrase the results as uncertainties in \"SED modelling\" more generally. Appendix A provides relevant perspective: switching to single-star-only tracks changes mass by 0.18±0.11 dex and SFR by 0.48±0.18 dex (Table A1), comparable to several library differences. The headline numbers should therefore be described as lower bounds conditional on the BPASS framework rather than as absolute modelling uncertainties.","section":"§3.2, Appendix A, §8"}],"minor_comments":[{"comment":"In Table 1, the C03 row lists three numeric entries ('-2.3', '1.0', '300') under the four numeric columns; since the text says the C03 IMF has an exponential low-mass cutoff at 1 Msun and an upper slope of -2.3, the table formatting should be corrected or the row removed from the power-law parameter columns.","section":"Table 1"},{"comment":"The Figure 3 caption contains a duplicated word ('fromfrom') that should be removed.","section":"Figure 3 caption"},{"comment":"The phrase 'systemic uncertainties' near the end of Section 5 should be 'systematic uncertainties'.","section":"Section 5"},{"comment":"The statement that fixing metallicity at Z=0.012 or Z=0.022 produces SFR and mass offsets of ~0.2-0.4 dex is inferred from the trend in Fig. 11 rather than from fits run with those fixed values; this should be stated explicitly to avoid the impression that those offsets were directly measured.","section":"Section 6.2"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of the journal and the core idea is sound, but the Abstract's headline numbers should not appear without reconciling the BaSeL exclusion issue. Once the statistics are reworked or explicitly qualified, this could become a solid methods contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline numbers in the abstract are the thing to know. The 1.4 dex SFR variation and the star-forming-to-quiescent flips come entirely from the BaSeL library, which Section 4.2 explicitly concludes is unsuitable for fitting these simulated SDSS/VISTA-quality data. Excluding BaSeL, the SFR spread drops to about 0.5 dex. That does not kill the paper's central point, but it makes the abstract misleading as written.\n\nWhat the paper does well is build a genuinely useful controlled experiment. It takes 18 EAGLE galaxies, generates mock SDSS spectra and VISTA photometry with one SPS framework (BPASS v2.2.1), and then fits with the same framework but four different spectral libraries, five IMFs, and three metallicity treatments, holding all fitting priors fixed. That sort of end-to-end sensitivity test within a single SPS code is not something the prior literature had. The appendices are a real strength: Appendix B validates the mock-spectrum generator against EAGLE photometry, Appendix C runs a no-dust control, and Appendix D reverses the direction of the IMF test to check the systematics behave sensibly. The paper is also honest about its own circularity: it is measuring sensitivity to model choice, not claiming the v2.2.1 truth is the real universe.\n\nThe stress-test note is correct, and it matters. The abstract presents the full spread across all libraries as if there were a four-way tie, when in fact the largest term is a library the paper itself says cannot resolve the data's features. There is no formal goodness-of-fit or Bayesian evidence to justify keeping BaSeL in the headline range, or to let a reader reject it cleanly. A serious referee should require the authors to separate a \"plausible alternative\" set (AP, C3K, CKC) from BaSeL and present both, with the conclusion drawn from the former and BaSeL treated as a bounded, known-pathology case. The qualitative point — that model uncertainty exceeds the <0.2 dex observational errors usually quoted — survives even for AP alone.\n\nA smaller but real issue is reproducibility. The preprint does not ship the pipeline code or the derived fit catalogs with a commit hash, and the headline numbers depend on hand-set choices like E(B-V)=0.3, fixed Z=0.014, and the specific 18 galaxies. That is addressable, but it currently makes independent checking harder than it should be.\n\nBottom line: this deserves peer review. The experimental design is sound, the validation is exemplary, and the core result is useful for anyone working with bagpipes, BPASS, or SED fitting generally. The revisions needed are around presentation and transparency, not the physics. I would take it to reading group and would likely cite the AP and IMF systematics in my own work.","headline":"A careful controlled SED-fitting uncertainty study whose headline numbers are inflated by a template library the authors themselves reject; the qualitative conclusion survives, but the abstract needs reframing.","tokens_in":37123,"tokens_out":2217,"would_cite":true,"duration_ms":22340,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The choice of stellar spectral library, not observational precision, sets the floor on how accurately galaxy masses, ages, and star formation rates can be inferred from SED fitting.","keywords":["SED fitting","stellar population synthesis","BPASS","stellar spectral libraries","initial mass function","metallicity assumptions","EAGLE simulation","galaxy stellar mass"],"falsifier":"Fitting a sample of nearby galaxies with independently known stellar masses and star formation histories, from resolved stellar populations or dynamical measurements, using the same library variants; if the recovered scatter around the known values is much smaller than the quoted $0.27$ dex for mass and $1.4$ dex for star formation rate, the claim that spectral library choice sets the survey error floor would be refuted.","tokens_in":36119,"feed_emoji":"🔭","tokens_out":8080,"duration_ms":65476,"temperature":0.7,"pith_summary":"The paper tries to establish that the underexamined modelling ingredients of spectral energy distribution (SED) fitting introduce systematic errors in galaxy properties that exceed the observational uncertainties reported by surveys. Using 18 galaxies from the EAGLE simulation, the authors generate mock SDSS spectra and VISTA photometry with one stellar population synthesis framework, BPASS v2.2.1, then refit them with variants that change only the stellar spectral library, initial mass function, or metallicity assumption. They find that switching spectral libraries shifts inferred stellar mass by $0.27\\pm0.09$ dex, age by $0.19\\pm0.11$ dex, and star formation rate by $1.4\\pm1.0$ dex, enough to flip a galaxy's classification from star-forming to quiescent. The paper concludes that as survey data improve, modelling uncertainty, not measurement noise, will limit what can be said about galaxy evolution, so these choices must be part of any error budget.","feed_headline":"SED model choices shift galaxy masses and ages beyond survey noise","feed_subtitle":"Inferred masses shift by 0.27 dex and star formation rates by 1.4 dex purely from spectral library choice.","key_machinery":"The central object is a controlled mock-observation loop: a set of 18 EAGLE galaxies at $z\\simeq0$ whose particle data are turned into mock SDSS spectra and VISTA photometry using the default BPASS v2.2.1 population-synthesis models; these mock observations are then refit with the bagpipes SED-fitting code using BPASS variants in which only one ingredient changes, whether the stellar spectral library, the IMF slope, or the metallicity prescription. The loop converts hidden assumptions into a measured systematic scatter, because the deviation of each refit from the v2.2.1 result isolates the effect of that single choice while the true galaxy properties are known.","core_discovery":"The central claim is that the specific stellar spectral library embedded in a stellar population synthesis model is a dominant, hard-coded source of systematic uncertainty in SED fitting, larger than the observational noise quoted by large surveys. Working in a controlled setting where the truth is known, with mock spectra built from EAGLE galaxies through the default BPASS v2.2.1 framework, the paper shows that fitting with the AP or BaSeL libraries systematically biases inferred stellar masses higher by roughly 0.23 to 0.27 dex and ages higher by 0.17 to 0.19 dex, while BaSeL suppresses inferred star formation rates by up to 1.4 dex, sometimes reclassifying galaxies as quiescent. IMF slope changes shift masses and star formation rates at the 0.1 to 0.2 dex level, and fixing an incorrect metallicity biases mass and star formation rate even when the average effect over a stacked sample is small. The reconstructed cosmic mass assembly history changes by up to about 12 percent depending on the choices, and the authors contend that with upcoming wide surveys the field is entering a model-limited rather than observation-limited regime.","pith_inferences":["The offsets' direction is sample dependent: the paper works with old, $z\\sim0$ galaxies where red stellar populations dominate, so shallower IMFs imply higher inferred masses; a high-redshift young-population sample would plausibly show the opposite sign, as the paper's cited literature already suggests.","Because the mock observations are generated and fit within one BPASS family, the quoted scatter measures the internal flexibility of that framework; combining BPASS with structurally independent stellar population codes based on single-star isochrones would either bracket a wider systematic range or reveal that BPASS variants already span most of it.","A testable extension would reverse the loop: generate mock spectra from one of the alternative libraries as truth; if the recovered offsets do not mirror the v2.2.1-based ones, the modelling error budget cannot be summarized by a single number and must be reported per model pair.","The paper's warning about photometry-only surveys could be sharpened by running the same fitting with photometry alone; if the scatter does not shrink when spectroscopy is removed, the claim that spectroscopy increases sensitivity to library-induced systematics would need revision.",""],"forward_implications":["Survey error bars on stellar mass and star formation rate that quote only observational noise (typically reported as less than 0.2 dex) understate the true uncertainty by roughly a factor of two or more.","Analyses of galaxy demographics should treat quiescent and star-forming fractions as partly model-dependent, since the BaSeL library alone can move galaxies across the classification boundary.","Reconstructed cosmic star formation and mass assembly histories carry up to roughly 10 to 12 percent systematic uncertainty at early times, dominated by spectral library and metallicity prescription choice.","For stacked samples of local massive galaxies above $10^9\\,M_\\odot$, fixing metallicity is acceptable, but only if the fixed value matches the sample's mass range and redshift.","Photometry-only surveys such as Euclid and Roman will not escape these systematics with better data; the paper suggests the bottleneck shifts to model assumptions.",""],"supporting_citations":[{"why":"Supplies the default BPASS v2.2.1 stellar population models used both to generate the mock truth spectra and as the reference fitting template.","marker":"Stanway & Eldridge 2018"},{"why":"Introduces the BPASS v2.3.1 framework and documents the alternative stellar spectral libraries whose templates are tested.","marker":"Byrne & Stanway 2023"},{"why":"Provides the bagpipes SED-fitting code used to infer all galaxy parameters from the mock observations.","marker":"Carnall et al. 2018"},{"why":"Describes the EAGLE simulation from which the 18 mock galaxies and their known star formation and metallicity histories are drawn.","marker":"Schaye et al. 2015"},{"why":"One of the alternative stellar spectral libraries; fitting with it drives the +0.27 dex mass and +0.52 dex star formation rate offsets.","marker":"Allende Prieto et al. 2018"},{"why":"The BaSeL library; its low resolution drives the largest star formation rate suppression, up to -1.4 dex, and the star-forming to quiescent reclassification.","marker":"Westera et al. 2002"},{"why":"The CKC library, the near-identical predecessor of the v2.2.1 default, used as a consistency check on the spectral library comparison.","marker":"Conroy et al. 2014"},{"why":"The C3K library, the updated CKC successor, whose narrower wavelength grid introduces a small but systematic negative mass offset.","marker":"Choi et al. 2016"},{"why":"Provides the dust attenuation curves applied both when generating the mock observations and when fitting them.","marker":"Salim et al. 2018"},{"why":"Supplies the non-parametric star formation history binning and Student's-t priors used in all bagpipes fits.","marker":"Leja et al. 2019a"}],"fun_headline_variants":["Spectral library choice alone shifts galaxy masses by 0.27 dex","SED model choice flips some galaxies from star-forming to quiescent","SED options bias masses by 0.27 dex and star formation by 1.4 dex","Model choice, not measurement noise, dominates galaxy property errors","SED library uncertainty exceeds survey noise in galaxy property estimates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison rests on the assumption that a real galaxy's emission is well approximated by the BPASS v2.2.1 models, since the mock spectra are generated from that framework and then used as the reference for measuring offsets.","fun_headline_variants_meta":{"raw":{"variants":["Spectral library choice alone shifts galaxy masses by 0.27 dex","SED model choice flips some galaxies from star-forming to quiescent","SED options bias masses by 0.27 dex and star formation by 1.4 dex","Model choice, not measurement noise, dominates galaxy property errors","SED library uncertainty exceeds survey noise in galaxy property estimates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000496,"raw_usage":{"total_tokens":2484,"prompt_tokens":1050,"completion_tokens":1434,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":1337}},"tokens_in":666,"tokens_out":1434,"duration_ms":12569,"temperature":1.0,"reasoning_tokens":1337,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:34:18.571581+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fitting a sample of nearby galaxies with independently known stellar masses and star formation histories, from resolved stellar populations or dynamical measurements, using the same library variants; if the recovered scatter around the known values is much smaller than the quoted $0.27$ dex for mass and $1.4$ dex for star formation rate, the claim that spectral library choice sets the survey error floor would be refuted.","supporting_citations":[],"review_version":2}