{"id":"43526e13-339d-4b78-a942-0c4177337754","arxiv_id":"2508.16525","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For the same galaxy spectra, codes that assume a single stellar population retrieve a flat, slightly bottom-heavy IMF, while flexible codes find a second, metal-poor component, with simulations showing the SFH assumption drives the difference.","lead":"This paper applies four popular stellar-population fitting codes to the same two galaxies and shows that the recovered initial mass function depends strongly on whether a code assumes one simple stellar population or allows a flexible mix of ages and metallicities. The simulations suggest that precise answers from single-population codes can be confidently wrong when a galaxy has two components, which matters for claims of IMF variations in elliptical galaxies.","discovery_kind":"extension","skeptic_critique":null,"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares four full-spectral-fitting codes (ALF and PyStaff, which assume a single stellar population, versus Starlight and pPXF, which fit non-parametric SFHs) on the same Magellan/IMACS spectra of NGC1399 and NGC1404, and on mock spectra of increasing star-formation-history complexity. The headline claims are: (i) SSP-assuming codes are precise and accurate only for pure SSP inputs, returning biased ages, metallicities, and IMF slopes for composite populations; (ii) flexible codes detect secondary components but with larger scatter and some biases; (iii) for the two observed galaxies, ALF and PyStaff find a single old, metal-rich population with a flat super-Salpeter IMF, while Starlight and pPXF suggest two stellar components with different metallicities and IMF slopes; and (iv) the assumption on allowed SFH is crucial for all retrieved stellar properties, including the IMF. The simulations and real-data fits are mutually consistent under the paper's interpretation.","tokens_in":29448,"tokens_out":4645,"duration_ms":60929,"significance":"The paper addresses a timely and important question: how much of the IMF slope and other stellar-population measurements depend on the fitting code's SFH assumption. Its main strengths are the controlled comparison on the same data with the same Conroy et al. (2018) models, the use of both MCMC and Monte-Carlo error estimation, and the detailed appendices on elemental abundances, residuals, and parameter degeneracies. If the central claims hold, the paper would give useful practical guidance for future stellar-population work. However, the simulation evidence is essentially a self-recovery test within the same model grid, and the real-data two-component decomposition is made by subjective inspection of weight distributions. These limitations affect the strength of the conclusions about the two galaxies but do not negate the value of the code comparison itself.","major_comments":[{"comment":"The mock spectra are generated from the Conroy et al. (2018) template grid and are then fit with the same Conroy et al. (2018) templates by all four codes. This measures each code's ability to invert its own forward model, not the fidelity of the model grid to real stellar populations. The paper uses these simulations to conclude that Starlight is reliable on real data ('we are inclined to rely only on Starlight results', Section 5) and that SSP codes 'return erroneous results' for composite SFHs. That extrapolation is load-bearing but not justified without an independent forward model (e.g., EMILES or another spectral library with independent response functions). The ranking of codes could change if the true stellar populations are not within the Conroy et al. grid. Please either add such a test or explicitly reframe the simulation claims as statements about internal model consistency.","section":"Section 3.5; Table 5"},{"comment":"The detection of two stellar components in NGC1399 and NGC1404 and the component fractions in Figure 9 are obtained by 'isolating' components in the three-dimensional weights matrix 'with the help of contour plots'. This is not a formal decomposition: there is no model comparison, no criterion for when a secondary peak is significant, and no stated uncertainty on the component fractions. Given that pPXF produces spurious sub-solar metallicity peaks even for SSP mock inputs (Section 5, point 1; Figure 21), visual inspection of weight distributions is insufficient to support the claim that 'it is likely that both galaxies do have a double stellar population' (Section 5). A quantitative significance test, or a substantially weakened conclusion, is needed.","section":"Section 4; Figures 5-9"},{"comment":"The CSP3 simulation was explicitly 'tuned to resemble' the real-data results for NGC1399 and NGC1404. Its successful recovery is then used as evidence that Starlight and pPXF can recover the two components in the real galaxies. This is partially circular: the simulation confirms that the codes can recover the assumed input, not that the assumed input is actually present in the observations. The paper should clearly distinguish between a simulation designed to match the proposed hypothesis and an independent validation of that hypothesis.","section":"Section 5; Table 5"},{"comment":"The statement 'we have found that the majority of stars have top-heavy IMF' is not supported by the quantitative component fractions presented in Figure 9. That figure indicates that the main (top-heavy) component is dominant only in the inner half of the galaxies, while the secondary (bottom-heavy) component reaches roughly 50% near R/Re ~ 0.2. Without an aperture- or mass-weighted integration of the radial component fractions, the majority-of-stars claim is not established. The more cautious wording in Section 5 ('the super-solar metallicity component associated with a TH IMF') is consistent with the data.","section":"Section 6; Figure 9"}],"minor_comments":[{"comment":"ALF normally permits a second younger component; the paper states this was disabled for the main comparison but enabled for some simulations. It would help to state explicitly in the simulation section which runs allowed the second component in ALF and whether this choice affects the interpretation of Figure 19.","section":"Section 3.2"},{"comment":"The number of realizations (10) and the noise level used for the mock spectra are mentioned only later in the appendix. Please state them in Section 3.5, since they are relevant to the precision claims in Section 5.","section":"Section 3.5"},{"comment":"The footnote says 'we have not tested for multiple populations', which sits oddly next to the strong language about 'a double stellar population'. Please align the wording so the exploratory nature of the two-component interpretation is consistent throughout.","section":"Section 5; footnote 3"},{"comment":"The text says the middle point 'deviates by ~2σ' from the original value, but the plotted error bars and the method for computing σ are not described for this comparison figure. Please clarify what is plotted and how σ is calculated.","section":"Figures 12-13"},{"comment":"The statement 'a threshold ... is unfortunately impossible to uniquely define' is a useful caveat, but the subsequent recommendation to first use Starlight-like codes and then switch to SSP codes would be more actionable if the criterion for 'confirmed to be consistent with a SSP' were specified in terms of the weight distributions shown in Figures 5 and 6.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"This is a borderline major-revision / reject case. The circular self-recovery nature of the simulations is a serious limitation, but the paper's core comparative experiment is still valuable and can be made sound by reframing the claims and adding an independent forward-model test. The real-data two-component detection needs to be treated as a hypothesis, not a robust measurement. I would not reject at this stage; the needed changes are within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first comparison I know that focuses on IMF-slope retrieval with ALF, PyStaff, Starlight, and pPXF on the same spectra and the same model grid. That is a real gap in the literature, and the paper fills it. The core methodological result, that SSP-based codes are precise but biased for composite populations while non-parametric codes find the extra component at the cost of precision, comes through clearly. The simulation suite with increasing SFH complexity (SSP through CSP3) is the right tool for this job. The authors also do the field a service by being unusually transparent about systematics: they test two wavelength ranges, discuss residuals at length, quantify cross-correlations, and partially address the non-solar abundance asymmetry in Appendix A. The citation pattern is fair, including their own prior work where it is actually relevant.\n\nThe soft spots are, in order of severity. First, the simulations are closed loop: mocks come from the Conroy et al. (2018) grid and all four codes fit against that same grid. At best this measures each code's ability to invert its own forward model, not absolute astrophysical accuracy. The paper never states this limitation in so many words, and readers could walk away thinking the simulations validate real-galaxy IMF slopes. Second, the two-component decomposition in NGC 1399/1404 is subjective, done by eye on weight distributions and contour plots. There is no automated, reproducible criterion for when a secondary component is declared real. And their own simulations show pPXF is biased for exactly these cases, so the double-population claim rests largely on Starlight alone. The abstract and summary say 'it is likely that both galaxies do have a double stellar population' and 'the majority of stars have top-heavy IMF'; the discussion is more hedged, and that gap matters. Third, the abundance-fitting asymmetry between codes is a confounder, partially addressed but not fully resolved.\n\nNone of this is fatal. The lesson, the SFH assumption can silently bias age, metallicity, and IMF retrieval, is important and well demonstrated. If the authors add an explicit circularity statement and an automated component-detection rule, this becomes a solid paper. As it stands, it deserves a serious referee, but with requested revisions before acceptance.","headline":"A genuinely useful first head-to-head of IMF retrieval across four full spectral fitting codes, but the two-component astrophysical claims for NGC 1399/1404 are softer than the abstract implies.","tokens_in":29988,"tokens_out":2964,"would_cite":true,"duration_ms":38115,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Which stellar initial mass function a galaxy appears to have depends on the star-formation history the fitting code assumes — and two Fornax ellipticals most likely host two stellar populations with different IMFs.","keywords":["initial mass function","full spectral fitting","stellar populations","star formation history","elliptical galaxies","Fornax cluster","NGC1399","NGC1404"],"falsifier":"Rebuild the simulation ladder with an independent forward model: mock spectra generated from a different stellar population synthesis code and spectral library, with known two-component star-formation histories and observed-like noise, then fit with the same four codes. If ALF and PyStaff still return precise-but-biased single-population answers and Starlight/pPXF still recover both components, the central claim stands; if the ranking flips or the components blur, the result is an artifact of the shared template grid. A complementary observational check: measure the radial mass-to-light gradie","tokens_in":29368,"feed_emoji":"🌌","tokens_out":15072,"duration_ms":137415,"temperature":0.7,"pith_summary":"This paper asks how robustly full spectral fitting codes recover the stellar initial mass function (IMF) and other stellar population properties, and answers that the code's star-formation assumption largely decides the answer. Four public codes fit the same high-signal optical and near-infrared spectra of the Fornax ellipticals NGC1399 and NGC1404, using identical stellar population models. ALF and PyStaff, which assume a single simple stellar population (SSP), return one old population with a metallicity declining with radius and a flat super-Salpeter IMF, while Starlight and pPXF, which fit mixtures of populations, find two components of similar old age with different metallicities and IMF slopes. Mock-spectra tests with increasing star-formation complexity show why: SSP-assuming codes are the most precise and accurate only when the true population is a pure SSP, and become biased, with misleadingly small error bars, as soon as a second component is present. The paper concludes that the star-formation assumption is crucial for all retrieved stellar properties, and that both galaxies most likely host a double stellar population in which the majority of stars have a top-heavy IMF.","feed_headline":"Four fit codes, one dataset, two IMF answers","feed_subtitle":"Codes assuming one stellar population see a flat IMF; flexible codes uncover a hidden second component.","key_machinery":"Four public full spectral fitting codes, paired by design into two classes: ALF and PyStaff, which assume a single simple stellar population (SSP), and Starlight and pPXF, which solve for the best-fit weighted composition of template SSPs. All four fit the same optical+NIR spectra of NGC1399 and NGC1404 from the same Conroy et al. (2018) grid, with the low-mass IMF slope parametrized as one power-law index over 0.08–1 solar masses. The discriminating experiment is a ladder of mock spectra of increasing star-formation complexity — pure SSP; a secondary younger, metal-poor component; and two components with distinct IMF slopes — which isolates how the SFH assumption changes accuracy, precision","core_discovery":"The star-formation-history assumption made during the fit controls every retrieved stellar property, including the IMF. On the same Magellan spectra of NGC1399 and NGC1404, the SSP-assuming codes ALF and PyStaff return one old population with a flat super-Salpeter IMF, while Starlight and pPXF, which fit template mixtures, reveal two populations — a dominant metal-rich, top-heavy-IMF component and a metal-poor, bottom-heavy one whose fraction grows with radius to ~50%. Mock spectra show the SSP codes win on a pure SSP, but any second component makes them biased, with misleadingly narrow errors. The flat IMF is a trade-off between the two real components; both galaxies likely host a double st","pith_inferences":["Editorial inference: because the simulations generate mock spectra from the same Conroy et al. (2018) grid the codes fit, they measure each code's ability to invert its own forward model; if that grid misrepresents real stellar populations (abundance response functions, IMF parametrization), the code ranking and the two-component recovery could change. An independent forward model would be the dec","Editorial inference: the claim that the majority of stars have a top-heavy IMF depends on the light/mass fractions Starlight and pPXF assign to each component; an independent check would compare the radial mass-to-light gradient implied by the IMF slopes with dynamical mass estimates from the same velocity dispersion profiles.","Editorial inference: the proposed workflow — flexible-SFH code first, SSP code second — is testable on a sample of galaxies with a priori known inputs; a natural extension is applying it to galaxies currently reported to have flat IMF gradients to see whether they split into two components as these two do.","Editorial inference: the double-population interpretation is consistent with, but does not prove, the two-phase formation scenario; spatially resolved stellar-population maps of these galaxies (e.g., from IFU data) could confirm whether the secondary component is spatially associated with an accreted envelope."],"forward_implications":["Any IMF slope, age, or metallicity measured from these galaxies under an SSP assumption — including this paper's own PyStaff-based flat super-Salpeter IMF — should be read as a mass-weighted trade-off between two real components, not as a property of a single population.","The widely reported radial pattern of bottom-heavy centers and Milky-Way-like outskirts may be an artifact of averaging two components whose mass fractions change with radius.","The local IMF–metallicity relation is not necessarily positive within a galaxy: here the metal-rich component carries the top-heavy IMF and the metal-poor component a bottom-heavy one, opposite to the SSP-based relation found in other studies.","Despite different masses and cluster positions, both galaxies share a similar two-component, all-old star-formation history, consistent with early accretion events during the formation of the Fornax cluster.","Future full spectral fitting of passive galaxies should first search for secondary components with a flexible-SFH code before applying high-precision SSP codes, ideally with a code that fits elemental abundances as free parameters within a non-parametric SFH."],"supporting_citations":[{"why":"Supplies the stellar population synthesis models and elemental-abundance response functions that all four codes fit; also the basis of ALF.","marker":"Conroy et al. 2018"},{"why":"Supplies pPXF, the penalized pixel-fitting engine used by one of the two non-parametric codes.","marker":"Cappellari 2017"},{"why":"Supplies Starlight, the base-spectrum superposition code that anchors the flexible-SFH side of the comparison.","marker":"Cid Fernandes et al. 2005"},{"why":"Supplies PyStaff, the SSP-assuming code whose setup (emcee sampling, wavelength sections) is used here.","marker":"Vaughan et al. 2018a"},{"why":"Earlier PyStaff analysis of NGC1399 that provides an external baseline for the IMF, age, and metallicity trends.","marker":"Vaughan et al. 2018b"},{"why":"Prior PyStaff analysis of NGC1404 on the same data; supplies the comparison baseline and the abundance-extrapolation discussion.","marker":"Feldmeier-Krause et al. 2021"},{"why":"Prior work quantifying IMF retrieval biases; motivates the fitting setup and the single power-law IMF parametrization.","marker":"Lonoce et al. 2021"},{"why":"Two-phase galaxy formation scenario used to interpret the main and secondary components as in-situ and accreted.","marker":"Naab et al. 2009"},{"why":"Two-stage IMF-variation model the paper weighs its predominantly top-heavy finding against.","marker":"Weidner et al. 2013"}],"fun_headline_variants":["Fitting code assumptions flip the stellar IMF result","Star-formation history assumption decides the IMF you get","Flexible codes uncover a second stellar population in ellipticals","SSP assumptions bias the IMF: flexible codes show two populations","On the same data, stellar fitting codes disagree on the IMF"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The simulations that anchor the interpretation assume that mock spectra drawn from the Conroy et al. (2018) template grid, with noise taken from the observed data, behave like real galaxies; because the fitting codes use that same grid, the tests measure each code's capacity to invert its own forward model, not how faithfully the grid reproduces real stellar populations.","fun_headline_variants_meta":{"raw":{"variants":["Fitting code assumptions flip the stellar IMF result","Star-formation history assumption decides the IMF you get","Flexible codes uncover a second stellar population in ellipticals","SSP assumptions bias the IMF: flexible codes show two populations","On the same data, stellar fitting codes disagree on the IMF"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000255,"raw_usage":{"total_tokens":1463,"prompt_tokens":853,"completion_tokens":610,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":540}},"tokens_in":597,"tokens_out":610,"duration_ms":7786,"temperature":1.0,"reasoning_tokens":540,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:14:36.897374+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rebuild the simulation ladder with an independent forward model: mock spectra generated from a different stellar population synthesis code and spectral library, with known two-component star-formation histories and observed-like noise, then fit with the same four codes. If ALF and PyStaff still return precise-but-biased single-population answers and Starlight/pPXF still recover both components, the central claim stands; if the ranking flips or the components blur, the result is an artifact of the shared template grid. A complementary observational check: measure the radial mass-to-light gradie","supporting_citations":[],"review_version":1}