{"id":"bc3d42e1-cc74-4857-b50d-af3c2dd7f1d7","arxiv_id":"2411.08945","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"EAGLE and TNG100 simulations largely reproduce the spectral covariance of SDSS galaxies, but AGN feedback differences, traced to black hole seeding, cause subtle mismatches in quiescent and AGN populations.","lead":"This paper compares the absorption-line spectra of real galaxies from SDSS with synthetic spectra from the EAGLE and TNG100 cosmological simulations, using a principal component analysis of continuum-subtracted light. It finds that both simulations reproduce the overall spectral variance of real galaxies, but that differences in black hole seeding and feedback create subtle mismatches in quiescent and AGN populations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The AGN/BH-seeding conclusion rests on unvalidated classification cuts that differ between EAGLE and TNG100; a photoionization-based BPT mock could settle whether the mismatch is physical or a selection artifact.","rationale":"The reader identified the classification equivalence as the weakest assumption, and this is indeed the most load-bearing point. The paper's strongest claim has two parts: first, that real and simulated spectra are consistent in spectral covariance; second, that the residual mismatch is driven by AGN feedback, specifically central BH seeding. The first part is supported mainly by visual inspection of latent-space contours without a quantitative consistency metric, which is a separate concern. However, the second part, which makes the paper's distinctive physical claim, depends directly on the mapping between BPT classes and the sSFR/lambda_Edd cuts. The thresholds in Table A1 and Figure 1 are calibrated to match global fractions, not to reproduce individual BPT classifications, and the thresholds differ between EAGLE and TNG100. Because the two simulations use different AGN definitions, the comparison of AGN latent-space positions and SSP-fit age/metallicity trends is not apples-to-apples. The small post-homogenisation AGN sample sizes further weaken the inference: splitting into 33rd and 67th percentile stacks leaves only dozens of galaxies per stack, so the apparent AGN divergence could easily be statistical noise. A concrete test using a physically motivated BPT mock, or even a uniform lambda_Edd threshold across both simulations, would directly settle whether the reported AGN mismatch is physical or an artifact of the classification scheme. For these reasons I agree with the reader's conditional verdict: the paper is promising and the methodology is interesting, but the central causal attribution is not yet securely established. I recommend no change to the reader's verdict, as the proposed test is exactly the kind of condition that would need to be met for the conclusion to be accepted.","tokens_in":36060,"tokens_out":5583,"duration_ms":58670,"concrete_test":"Recompute the AGN and Q subsamples of EAGLE and TNG100 by post-processing the simulated gas with a photoionization-based BPT classifier (for example, the Hirschmann et al. 2023 pipeline or a CLOUDY-based line-ratio method) rather than the sSFR/lambda_Edd cuts, preserving the same mass homogenisation, and repeat the latent-space projection and SSP stack fitting of Section 6. If the EAGLE-vs-TNG100 AGN divergence and the PC1 age trend survive this physically grounded classification, the BH-seeding attribution is supported; if they weaken or disappear, the current conclusion is a selection artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim — that the latent-space mismatch in AGN (and quiescent) spectra traces AGN subgrid physics, specifically BH seeding — rests on the equivalence between the SDSS BPT classes and the simulation cuts in sSFR and lambda_Edd. Section 3.2 and Table A1 show the thresholds are chosen to reproduce global BPT fractions, not validated against a physical mapping. Critically, the thresholds are not the same for the two simulations: EAGLE defines AGN as lambda_Edd > -2, while TNG100 uses lambda_Edd > -0.6 (Fig. 1). Thus 'AGN' denotes a different region of accretion parameter space in each simulation, and the comparison of AGN latent-space locations or SSP-fit age/metallicity trends may reflect this selection difference rather than a genuine difference in stellar-population variance. The problem is compounded by the tiny homogenised AGN samples (104 EAGLE, 244 TNG, versus 917/263 SDSS in Table 1): the 33rd/67th percentile stacks used for the PC1 age inference contain only roughly 35-80 galaxies, so the claimed AGN divergence could be noise. If the classification mapping is not faithful, the disagreement attributed to BH seeding is an artifact, and the headline that spectral covariance is a valid model-independent test of simulations is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the PCA-based spectral covariance method of Sharbaf et al. (2023) to synthetic spectra constructed from EAGLE and TNG100 galaxies, matching the SDSS instrumental setup, noise properties, and stellar-mass distributions. It projects the continuum-subtracted synthetic spectra onto SDSS-derived eigenvectors for star-forming, AGN, and quiescent subsamples, compares the resulting latent-space locations, fits SSP ages and metallicities to percentile stacks, and examines the simulated star formation histories. The central claim is that the optical absorption-line covariance of real galaxies is broadly reproduced by both simulations, with star-forming populations in good agreement, while AGN and, downstream, quiescent populations show discrepancies that the authors attribute to differences in AGN subgrid prescriptions, specifically central black hole seeding.","tokens_in":36374,"tokens_out":6621,"duration_ms":59602,"significance":"If the conclusions hold, the paper offers a genuinely model-independent benchmark for galaxy-formation simulations that goes beyond scaling relations, and it identifies a concrete, falsifiable point of divergence among EAGLE, TNG100, and SDSS in the AGN population. The construction of the synthetic spectra is careful: realistic noise is drawn from SDSS inverse variances, velocity dispersions are matched, stellar-mass distributions are homogenised before comparison, and the analysis code and synthetic data are made publicly available. These strengths are substantial. However, the headline comparison currently lacks any quantitative statistical test, and the simulated-galaxy classification underlying the AGN conclusion is calibrated to global fractions rather than validated against a physical mapping. These issues must be addressed before the causal attribution to black-hole seeding can be accepted.","major_comments":[{"comment":"The equivalence between the SDSS BPT classification and the simulation cuts in (sSFR, lambda_Edd) is assumed rather than demonstrated. The thresholds are chosen to reproduce the observed SF/AGN/Q fractions, as stated explicitly in Section 3.2 and Table A1, and they differ between the two simulations: EAGLE defines AGN as lambda_Edd > -2 while TNG100 uses lambda_Edd > -0.6 (Fig. 1). Because the 'AGN' samples therefore occupy different regions of accretion parameter space, the latent-space and SSP-fit differences attributed to AGN feedback (Figs. 6-8 and Section 7) may partly be a selection artifact rather than a physical difference in stellar populations. A concrete test would be to generate synthetic BPT classifications for the same simulated galaxies, for example via photoionization modelling, and check whether the adopted cuts recover those classes; at minimum, the authors should show that the conclusions are robust to varying the thresholds within reasonable ranges.","section":"Section 3.2, Table A1, Appendix A"},{"comment":"The statement that 'real and simulated spectra are consistent regarding spectral covariance' is not backed by any statistical test. The text describes overlapping contours and qualitative separation, but no two-sample test (e.g., KS, AD, or a bootstrap overlap fraction) is computed between the PC distributions, and no error bars account for the finite sizes of the simulated subsamples (Table 1 lists 104 EAGLE AGN, 244 TNG100 AGN, 740-1045 quiescent, and 2094-2460 star-forming galaxies after homogenisation). Without such quantification, 'consistency' and 'discrepancy' cannot be distinguished from noise, particularly for the small AGN samples. Please add a distributional test on the PC projections and report effect sizes or confidence intervals.","section":"Section 6, Figures 5 and B1"},{"comment":"The SSP-fit comparison uses stacks of the lowest and highest 33rd/67th percentile projections. For the EAGLE AGN sample this means each stack contains only about 35 galaxies (104 total after homogenisation; Table 1), and for TNG100 AGN about 81 galaxies. The MCMC contours show the fitting uncertainty of the stacked spectrum but not the galaxy-to-galaxy sampling variance within each stack or the uncertainty due to the stack construction. The claim that 'AGN galaxies in EAGLE show different metallicities in the opposite direction to the age-metallicity degeneracy' (Section 6.1) may therefore be driven by a small number of objects. Please include bootstrap resampling of the stacks, or an equivalent procedure, and report the number of galaxies per stack so that the statistical weight of the AGN divergence can be assessed.","section":"Section 6.1, Figures 6-8, Table 1"},{"comment":"The attribution of the latent-space mismatch to black-hole seeding is only circumstantial. The paper demonstrates that EAGLE and TNG100 differ in the redshift distribution of BH seeding and in the distribution of lambda_Edd, but those quantities also depend on the accretion and feedback prescriptions, and no control test (e.g., varying only the seeding prescription within one simulation) is presented. The abstract wording 'could lead to the mismatch' is appropriately cautious, but the Discussion later states 'We ascribe this difference to the quenching mechanisms adopted by the simulations', which overstates the evidence. Please either soften this causal language or add a test that directly varies the seeding prescription.","section":"Section 7, Figure 13"}],"minor_comments":[{"comment":"The text and figure axes contain several typos: '67rd' should be '67th', 'Lock back time' should be 'Lookback time', and Section 7 contains 'intringuingspike' rather than 'intriguing spike'.","section":"Section 6.1, Figures 9-11"},{"comment":"The caption contains 'desccribed' and the appendix text 'thhe'; please correct these typographical errors.","section":"Figure B1 caption, Appendix B"},{"comment":"The caption states that the SDSS sample corresponds to the one homogenised with EAGLE, while Figure B1 shows SDSS sets for both simulations; please clarify in each panel which homogenised SDSS sample is being displayed.","section":"Section 6, Figure 5 caption"},{"comment":"The projection formula is written explicitly only for PC1, although the analysis uses the first three principal components; please state that the equation is illustrative and give the general form, or define the projections for PC2 and PC3 explicitly.","section":"Section 5, Equation (2)"},{"comment":"The noiseless comparison is shown only for EAGLE; repeating the exercise for TNG100 would make the claim that noise modelling is not the dominant systematic more convincing.","section":"Appendix D, Figure D1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well-structured and the synthetic spectral construction is a genuine strength. The main methodological gap is the absence of a quantitative statistical comparison in latent space, combined with an unvalidated mapping between BPT classes and simulation cuts in sSFR and lambda_Edd. Both issues are fixable within the scope of the paper, but until they are addressed the AGN/black-hole-seeding conclusion should not be considered established. I would be supportive of publication after a major revision that adds the requested tests and softens the causal claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a careful application of the PCA-SDSS covariance diagnostic to synthetic spectra from EAGLE and TNG100. The headline—that real and simulated spectra are consistent in spectral covariance, with the main differences in the AGN subset—is plausible but not firmly established. The BH-seeding explanation is explicitly framed as a candidate, and that level of caution is appropriate.\n\nWhat is new: projecting synthetic spectra from two major simulations onto eigenvectors derived from SDSS data, with realistic noise and mass matching, then tracing differences back to star formation histories and subgrid physics. The SF populations match well in both simulations; AGN and quiescent populations show discrepancies. The paper is methodical: they use 33rd/67th percentile stacks because of small samples, test the role of noise in an appendix, and the data and code are public. All of that is real credit.\n\nThe soft spots are real and land where the stress-test note points. First, the simulated galaxies are classified into SF/AGN/Q using cuts in sSFR and lambda_Edd that are calibrated to reproduce SDSS BPT fractions, but the thresholds differ between EAGLE and TNG100 (lambda_Edd > -2 versus > -0.6). So 'AGN' means different accretion regimes in the two simulations, and the AGN latent-space comparison may partly reflect that selection difference. Second, the homogenised AGN samples are small (104 EAGLE, 244 TNG, versus 917 SDSS), so the claimed AGN divergence could be noise. Third, the consistency of the latent-space contours is asserted visually; no quantitative test is offered. A photoionization-based BPT mock would settle whether the classification mapping is physical.\n\nNone of this sinks the paper. The authors do not overclaim; they say the seeding difference 'could lead' to the mismatch. But the central diagnostic would be stronger with a formal consistency measure and a validation of the simulated classification. The citation pattern is fine, and the self-reference to PCA-SDSS is legitimate since this is a direct extension.\n\nThe paper is for people comparing simulations to observations and for those working on subgrid feedback. It deserves a serious referee and I would engage with it.","headline":"Careful application of spectral covariance to simulations; the AGN mismatch is suggestive, but unvalidated classification cuts and small samples leave the BH-seeding claim unproven.","tokens_in":36899,"tokens_out":2059,"would_cite":true,"duration_ms":20810,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Spectral covariance tests reveal that EAGLE and TNG100 simulations capture the variance of real galaxy spectra, with subtle AGN feedback mismatches tied to black hole seeding.","keywords":["spectral covariance","principal component analysis","galaxy formation simulations","EAGLE","IllustrisTNG","AGN feedback","galaxy quenching","SDSS spectra"],"falsifier":"If a non-AGN physical process (e.g., stellar feedback, environment) were found to be the dominant cause of the spectral variance differences between EAGLE and TNG100, the central claim would fail. A concrete test: run a simulation with identical subgrid prescriptions to TNG100 but with EAGLE-style BH seeding (earlier and lower mass seeds), and check whether the latent-space distribution of AGN and quiescent galaxies shifts to match the SDSS data and EAGLE's PC1 gradient.","tokens_in":35923,"feed_emoji":"🔭","tokens_out":1588,"duration_ms":18063,"temperature":0.7,"pith_summary":"The paper asks whether cosmological hydrodynamical simulations reproduce not just the average properties of galaxies but the full pattern of spectral variation seen in real galaxies. It applies principal component analysis to continuum-subtracted SDSS spectra and projects synthetic spectra from the EAGLE and Illustris TNG100 simulations onto the same eigenbasis. The central claim is that the simulated spectra are largely consistent with the observed spectral covariance, but subtle differences—especially in the AGN and quiescent subsets—can be traced to differences in how the simulations seed and grow central black holes. If this is right, spectral covariance offers a model-independent benchmark for galaxy formation simulations, and AGN feedback seeding is a key lever for improving them.","feed_headline":"Spectral variance reveals AGN seeding as key simulation mismatch","feed_subtitle":"EAGLE and TNG100 spectra match real galaxies' covariance, but black hole seeding times and AGN feedback leave a detectable imprint.","key_machinery":"The central machinery is principal component analysis on the covariance matrix of continuum-subtracted, emission-line-masked optical spectra (3800–4200 Å and 5000–5400 Å windows). Spectra from SDSS define the eigenbasis; synthetic spectra from EAGLE and TNG100 are projected onto the same eigenvectors to locate each simulated galaxy in the latent space of PC1–PC3. The argument then uses two supporting tools: SSP spectral fitting of stacked spectra split by PC projection percentiles to read off stellar age and metallicity, and star formation histories constructed directly from simulated stellar particles to interpret the latent-space positions physically.","core_discovery":"The paper's central discovery is that the covariance structure of real galaxy spectra, as encoded by the first three principal components of SDSS spectra in two optical windows, is largely reproduced by both EAGLE and TNG100 synthetic spectra, but with identifiable deviations. The first principal component is predominantly driven by stellar age, and the age sequence SF→AGN→Q is preserved in both simulations. However, the AGN and quiescent subsets show differences in the distribution of spectral variance that the authors trace to the subgrid implementation of AGN feedback, specifically the black hole seeding time: TNG100 seeds black holes later (sharp peak at z≈2) and with a higher seed mass and halo mass threshold, whereas EAGLE seeds earlier and more evenly in redshift. The resulting differences in the Eddington-ratio distribution lead to noticeably different quenching patterns and star formation histories for quiescent galaxies. The paper argues that spectral covariance is a powerful, model-independent test of simulations, complementing scaling relations.","pith_inferences":["The paper's methodology could be extended to other simulations (e.g., TNG50, SIMBA, ASTRID) to see whether the AGN-mismatch signature is generic or specific to the EAGLE/TNG flavor of subgrid physics.","The latent-space comparison could be applied to emission-line windows as well, as a direct test of the mapping from sSFR/λEdd cuts to BPT classes.","The authors implicitly suggest that the later, more massive black hole seeding in TNG100 leads to a stronger but more abrupt quenching, which might also produce observable imprints in the halo occupation statistics or the scatter in the star-forming main sequence at higher redshift.","One could test the age-metallicity degeneracy interpretation directly by constructing mock spectra with known age-metallicity combinations and projecting them onto the SDSS eigenbasis, thus calibrating the PC axes in terms of physical parameters."],"forward_implications":["Spectral covariance provides a purely data-driven benchmark for galaxy formation simulations, independent of physical parameter choices, so simulation comparisons can move beyond bulk scaling relations.","The AGN feedback subgrid prescriptions in simulations need to reproduce not just the quenched fraction but the detailed shape of the spectral variance, meaning black hole seeding and accretion prescriptions are directly testable against observed spectra.","The explicit SF→AGN→Q evolutionary sequence in latent space gives a physical ordering of galaxy populations that simulations should reproduce; the overlap between SF and AGN in simulations indicates a specific deficiency in how AGN activity is coupled to star formation.","Differences in the star formation histories of quiescent galaxies between EAGLE and TNG100 directly reflect the different quenching mechanisms, so latent-space projections can be used to diagnose which physical process is responsible for quenching in a given simulation.","Using both blue and red spectral windows cross-checks the results and shows that the age sensitivity of PC1 is robust across wavelength ranges, strengthening the interpretation that PC1 tracks stellar age."],"supporting_citations":[{"why":"The PCA-SDSS study that established the eigenbasis and the blue/red spectral windows, providing the covariance analysis that this paper applies to simulations.","marker":"Sharbaf et al. (2023)"},{"why":"The reference for the EAGLE simulation, defining the subgrid physics, box size, and calibration used to generate the synthetic spectra.","marker":"Schaye et al. (2015)"},{"why":"The reference for the IllustrisTNG simulation, defining the TNG100 run, its cosmological parameters, and the subgrid model.","marker":"Pillepich et al. (2018b)"},{"why":"The prior work that used the same sSFR and λEdd classification scheme for simulated galaxies and the stellar mass homogenisation, which this paper follows.","marker":"Angthopo et al. (2021)"},{"why":"The paper describing the AGN feedback model in IllustrisTNG, including the BH seeding mass and the kinetic/thermal mode switch, which is central to the discrepancy claim.","marker":"Weinberger et al. (2017)"}],"fun_headline_variants":["Spectral covariance exposes AGN seeding mismatch","AGN seeding time splits simulation galaxy spectra","EAGLE and TNG100 differ on black hole seeding","Covariance test flags AGN feedback in simulations","Quenching mismatch traced to AGN seeding in models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The classification of simulated galaxies into star-forming, AGN, and quiescent using direct cuts in sSFR and λEdd is assumed to be equivalent to the BPT emission-line classification of SDSS galaxies, even though the thresholds are tuned to match observed fractions.","fun_headline_variants_meta":{"raw":{"variants":["Spectral covariance exposes AGN seeding mismatch","AGN seeding time splits simulation galaxy spectra","EAGLE and TNG100 differ on black hole seeding","Covariance test flags AGN feedback in simulations","Quenching mismatch traced to AGN seeding in models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000269,"raw_usage":{"total_tokens":1663,"prompt_tokens":1028,"completion_tokens":635,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":561}},"tokens_in":644,"tokens_out":635,"duration_ms":6439,"temperature":1.0,"reasoning_tokens":561,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:13:09.280868+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a non-AGN physical process (e.g., stellar feedback, environment) were found to be the dominant cause of the spectral variance differences between EAGLE and TNG100, the central claim would fail. A concrete test: run a simulation with identical subgrid prescriptions to TNG100 but with EAGLE-style BH seeding (earlier and lower mass seeds), and check whether the latent-space distribution of AGN and quiescent galaxies shifts to match the SDSS data and EAGLE's PC1 gradient.","supporting_citations":[{"cited_title":"G., Dalla Vecchia C., Pillepich A., 2021, @doi [ ] 10.1093/mnras/staa3294 , https://ui.adsabs.harvard.edu/abs/2021MNRAS.502.3685A 502, 3685","cited_arxiv_id":null,"evidence_quote":"The prior work that used the same sSFR and λEdd classification scheme for simulated galaxies and the stellar mass homogenisation, which this paper follows."}],"review_version":1}