{"id":"2dc501a5-eedd-408d-a8fd-797de91eef09","arxiv_id":"2508.21152","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A unified empirical equation, fitted separately to three simulation suites, links galaxy star formation history shape to halo mass, baryon fraction, black hole mass, and feedback strength.","lead":"This paper uses the CAMELS suite of simulations to test how changing the strengths of supernova and black hole feedback alters the star formation histories of galaxies. It finds that a single mathematical form, with model-specific coefficients, can describe the average galaxy star formation history across three different simulation models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eqn 18 is fit to normalizing-flow-sampled SFHs with only qualitative 1P validation; a held-out direct-simulation test is needed to support the unified-equations claim.","rationale":"The reader's CONDITIONAL verdict is appropriate. The qualitative results—stellar feedback effects, AGN model-dependence, interactions, baryon cycling—are plausible and supported by direct LH measurements (Figs 9–11) and prior literature. The load-bearing risk is specifically the predictive general-equation claim. The paper itself flags the main caveats (mass range, box size, resolution, ASTRID tails) and the DPL caveat, which is to its credit. The missing piece is a quantitative out-of-sample check against simulations, not against the emulator on which the equations were built. I therefore agree with the reader's weakest assumption, and keep CONDITIONAL. This is not an objection to the physical interpretation; it is a request for the specific validation that would convert the empirical equations from in-sample description to a testable relation.","tokens_in":48099,"tokens_out":4289,"duration_ms":46699,"concrete_test":"Hold out a random subset of CAMELS LH boxes (e.g., 20 per model) before any emulator training or equation construction. For each held-out box, compute the average SFH of 100 galaxies directly from star particles, fit the double power-law (Eqn 13), and compare the resulting (α, β, τ, ϕ, η) with Eqn 18/Table 2 predictions. Report per-model RMS residuals and the fraction of direct SFHs within the emulator's 1σ credible interval. Also repeat the ASTRID comparison separately for boxes whose direct average SFHs have sustained late-time tails; if those residuals dominate or the DPL fits fail, restrict the claim or add a tail parameter. If the held-out residuals are comparable to the training residuals and within a pre-defined tolerance (e.g., 10% of the parameter dynamic range), the unified-equations claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that one functional form (Eqn 18) with Table 2 coefficients relates SFH shape parameters to Mhalo, fb, MBH/Mhalo and feedback parameters across TNG/SIMBA/ASTRID—rests on a chain: emulator → double power-law fit → GAM/symbolic-regression term selection → coefficient fitting. The weakest link is the emulator. All SFHs used to build Eqn 18 are sampled from the normalizing flow trained on LH boxes (Sec 3.1; 'Unless otherwise mentioned, all the SFHs in the following sections are generated by sampling the trained normalizing flows'). The only validation shown is qualitative: Appendix A/Fig 19 compares median flow samples with the 1P runs and states 'qualitative agreement', but no quantitative coverage, calibration, or held-out error statistic is reported. If the emulator systematically smooths or shifts average SFHs in under-sampled regions of the 6D parameter space, the DPL parameters—and therefore all coefficients and interaction terms in Eqn 18—encode emulator bias rather than the simulation's feedback response. The acknowledged DPL failure for a subset of ASTRID SFHs with sustained late-time tails (Sec 3.2) compounds this: those SFHs have poorly constrained α and τ, yet they are included in the ASTRID fits that set Table 2 coefficients. Because the equations are selected and fit with no reported coefficient uncertainties or held-out test, the 'single set of equations' claim is currently supported only by in-sample agreement on emulator-generated data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how variations in stellar feedback, AGN feedback, and cosmology affect the average star formation histories (SFHs) of galaxies at z~0 in the CAMELS IllustrisTNG, SIMBA, and ASTRID simulations. The authors train normalizing-flow emulators on the CAMELS Latin Hypercube (LH) sets, use the emulators to sample average SFHs, and fit these with a double power-law (Eq. 13) to obtain shape parameters {α, β, τ, φ, η}. They then use random forests, generalized additive models, and symbolic regression to construct a common set of equations (Eq. 18) with model-specific coefficients (Table 2) that relate SFH shape parameters to halo mass, baryon fraction, black hole-to-halo mass ratio, and CAMELS feedback/cosmology parameters. The paper also presents interaction analyses, a reparameterization of supernova feedback in terms of mass/energy loading, an SBI-based proof-of-concept inference of CAMELS parameters from SFHs, and several supplementary observational diagnostics.","tokens_in":48449,"tokens_out":3033,"duration_ms":37665,"significance":"If the central claim holds—that one functional form with model-dependent coefficients describes SFH shape across three different hydrodynamical codes—this would be a useful empirical framework for connecting SFH observations to feedback physics and for interpreting CAMELS parameter-space studies. The paper's strengths include its use of the public CAMELS suite, explicit treatment of galaxy selection and resolution limits, the combination of multiple ML methods, and the extensive diagnostic figures. However, the equations in Eq. 18 are fitted to emulator-sampled SFHs, the emulator validation is only qualitative, and the coefficients carry no reported uncertainties or out-of-sample tests. These gaps are load-bearing for the 'single set of equations' claim, so the result is currently a promising but not fully supported empirical summary.","major_comments":[{"comment":"All SFHs used to build Eqn. 18 are generated by sampling the trained normalizing flows ('Unless otherwise mentioned, all the SFHs in the following sections are generated by sampling the trained normalizing flows', §3.1). The only comparison to actual simulation outputs is the qualitative 1P validation in Fig. 19, which states 'qualitative agreement' but reports no quantitative coverage, calibration, or held-out error statistic. Since Eqn. 16 minimizes loss on emulator-sampled SFHs, any emulator bias in under-sampled regions of the 6D parameter space propagates directly into the DPL parameters and hence into every coefficient in Table 2. I request a quantitative validation of the emulator against held-out LH boxes and/or the 1P runs, with per-parameter residuals and coverage diagnostics, and a demonstration that the derived Eqn. 18 coefficients are stable when the emulator is retrained or","section":"§3.1, Appendix A (Fig. 19)"},{"comment":"The central equations are presented without any measure of predictive accuracy. The coefficients in Table 2 have no uncertainties, and the loss defined in Eq. (16) is minimized on the same emulator-generated data used to select the terms. No residual plots, R² values, or held-out predictions are shown for the DPL parameters. The paper even uses the GAM loss as a 'proxy of the Bayes risk' (§3.3.2), but does not compare Eq. 18 against that benchmark. Without out-of-sample validation, the claim that 'a single set of equations ... can describe the SFHs across all three CAMELS models' is supported only by in-sample agreement. Please provide bootstrap/subsample coefficient uncertainties and a held-out evaluation (e.g., fitting coefficients on a training subset and evaluating on a withheld LH subset, or predicting one model's coefficients from another).","section":"§6.1.1, Eq. (18), Table 2"},{"comment":"The paper acknowledges that the double power-law 'does not describe a subset of ASTRID SFHs with sustained late-time tails' (§3.2). These ASTRID SFHs are nevertheless included in the fits that determine the ASTRID coefficients in Table 2, where α and τ are the least constrained parameters. Because Eqn. 18 is claimed to hold across all three models, the fraction of average SFHs that are poorly represented by the DPL form must be quantified, and the sensitivity of the derived coefficients to excluding or reparameterizing these cases should be shown. If the DPL failure is non-negligible in ASTRID, the corresponding rows of Table 2 may encode an artifact of the fitting form rather than the feedback response.","section":"§3.2, §4.2, §6.1.1"},{"comment":"The claim of a 'single set of equations' is weakened by the fact that the coefficients are free to vary per model in Eq. (16). If the functional form is the same but every coefficient differs, the statement reduces to 'each model can be fit by a member of a parametric family'. The paper needs to demonstrate what is shared beyond the functional form—for example, that the same terms remain important across models, that coefficients can be predicted from model properties, or that the equations generalize to held-out models/parameters. As written, the evidence for universality is largely the symbolic-regression term frequencies in Fig. 12, which are qualitative and in-sample.","section":"§6.1.1, Eq. (16), Eq. (18)"}],"minor_comments":[{"comment":"The validation text in Appendix A repeats 'qualitative agreement' twice. Please state explicitly which quantitative metrics were computed (e.g., coverage, calibration, chi-square) and whether any failed.","section":"§3.1, Appendix A"},{"comment":"The text defines α as the falling slope and β as the rising slope, but Figure 1 and Eq. (13) can be misread because the two power-law terms are symmetric. Consider adding a sentence explicitly defining the relation between α, β and the t<τ versus t>τ behavior.","section":"§3.2, Eq. (13)"},{"comment":"The χ² metric in Eq. (17) has unusual units (SFR² over SFR² integrated over time). Please clarify whether the integrand is intended to be a dimensionless ratio or whether the normalization is meant to produce a time-averaged statistic.","section":"§4.3, Eq. (17)"},{"comment":"The random forest feature importances in Fig. 13 are shown without error bars or sensitivity checks. Given that they are used as a 'sanity check' for Eq. 18, a bootstrap estimate would strengthen the comparison.","section":"§3.3.1"},{"comment":"The equation for η is written separately with no coefficients in Table 2. For completeness, state explicitly that η = 12.5Ωm − 3 is used for all three models, and whether a coefficient uncertainty was estimated.","section":"§6.1.1, Eq. (18)"},{"comment":"There are several typographical and formatting issues, including 'early rimes' in §6.1.1, inconsistent use of 'Mhalo' (log Mhalo vs Mhalo) in Eq. (18) and Table 2, and missing figure cross-references in the text. A careful proofread is recommended.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid CAMELS-based phenomenological study with a large amount of useful diagnostic material, but the central 'general equations' claim currently rests on emulator-generated, in-sample data with only qualitative validation. I would not reject the paper: the qualitative SFH trends and the mass-versus-energy-loading reparameterization are interesting and likely robust. The revision should focus on (1) quantitative emulator validation against held-out simulations, (2) uncertainties and out-of-sample metrics for Eq. 18, and (3) an explicit treatment of the DPL failure mode for ASTRID tails. If those can be supplied, the paper could be a valuable contribution to the CAMELS literature and to SFH-based feedback inference."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you work on SFH shape statistics or the CAMELS suite; the qualitative results will survive. The thing to be careful about is the central equation set, which is a summary of emulator output rather than a tested law.\n\nWhat the paper does well: it gives the first systematic comparison of how average galaxy SFHs respond to stellar and AGN feedback variations across TNG, SIMBA, and ASTRID within CAMELS. The double power-law parametrization is sensible for average SFHs, and the authors are honest about where it fails, notably the ASTRID late-time tails mentioned in Section 3.2. The qualitative findings, stellar feedback dominates low-mass regulation, AGN feedback is model-dependent, and stellar/AGN feedback couple through gas supply and BH growth, are plausible and consistent with prior literature. I also found the reparametrization of stellar feedback into mass loading versus energy per SFR (Section 6.2) genuinely useful; it explains why seemingly similar parameters produce opposite effects across codes and gives observers a more physical knob to think about. The paper also ships the machinery as public tools (LtU-ILI flows, CAMELS data), which is real value.\n\nWhere I agree with the reader's reservations: the unified equations in Eqn 18 are fitted to SFHs sampled from the normalizing-flow emulator, and the only emulator validation shown is qualitative agreement with the 1P runs (Appendix A). There are no coefficient uncertainties on Table 2, no held-out test against directly simulated SFHs, and the DPL-inadequate ASTRID tails are still fed into the fits. The paper's own text admits this: 'all the SFHs in the following sections are generated by sampling the trained normalizing flows.' So the 'single set of equations across all three models' claim is currently descriptive, not predictive. That does not sink the paper, because the equation system is presented as an empirical framework and the authors list several caveats, but it does mean the headline claim should be softened or backed with a real validation run.\n\nWho gets value: anyone comparing simulations to observations of SFHs, or calibrating sub-grid feedback. It deserves a serious referee. My recommendation is to send it out, and to ask the authors to add one out-of-sample test using CV or 1P boxes directly, report coefficient uncertainties, and quantify how many average SFHs the double power law fails to describe.","headline":"A careful and useful survey of how feedback shapes SFHs across three CAMELS models, but the headline unified equations are in-sample fits to an emulator and should not be treated as validated predictions.","tokens_in":49037,"tokens_out":1234,"would_cite":true,"duration_ms":16930,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One set of equations links star formation history shape to feedback","keywords":["galaxy star formation histories","stellar feedback","AGN feedback","cosmological simulations","baryon cycling","double power-law SFH","CAMELS","normalizing flows"],"falsifier":"Recompute average SFHs directly from the CAMELS LH simulation particle data, fit double power laws, and compare the resulting alpha, beta, tau, phi, eta to Eqn 18 on a grid of parameter values; systematic deviations beyond sampling noise would disprove universality. Specifically examine ASTRID galaxies with late-time star-formation tails: if those SFHs are not well described by a double power law, the falling-slope equation is biased.","tokens_in":1963,"feed_emoji":"","tokens_out":2979,"duration_ms":90099,"temperature":0.7,"pith_summary":"This paper tries to show that average galaxy star formation histories can be summarized by a double power-law curve whose five shape parameters depend on a small set of physical quantities: halo mass, baryon fraction, black hole mass, and feedback strength. Using the three CAMELS simulation models, a single set of equations with model-specific coefficients reproduces the average SFHs across all three. In the equations, cosmology sets the early rise, halo mass sets the overall scale, and feedback plus baryon cycling set the late-time decline and width. This matters because it offers a path from observations of galaxy samples back to feedback physics.","feed_headline":"One equation set maps galaxy star formation histories","feed_subtitle":"Halo mass, baryon fraction, black hole mass, and feedback strength predict how galaxies formed stars in three major simulations.","key_machinery":"The load-bearing machinery is a double power-law SFH parameterization plus a normalizing-flow emulator. The double power law, SFH(t) = phi * ([(t-eta)/tau]^alpha + [(t-eta)/tau]^(-beta))^(-1), compresses each average SFH into five interpretable numbers. The normalizing flow, trained on the CAMELS LH simulations, generates average SFHs anywhere in parameter space, enabling regression of those five numbers against halo mass, baryon fraction, black hole mass, and feedback scalings. Symbolic regression term frequencies and generalized additive model losses guide which physical variables enter each equation, with a bias toward galaxy state variables over direct feedback parameters.","core_discovery":"The paper's central discovery is a system of empirical equations, Eqn 18, describing the average SFH shape parameters in terms of Omega_m, sigma_8, halo mass, baryon fraction, relative black hole mass, and the CAMELS feedback parameters. The same functional form applies to IllustrisTNG, SIMBA, and ASTRID, with only coefficients changing (Table 2). The rising slope beta depends mainly on Omega_m; the falling slope alpha on halo mass with AGN and baryon terms; the peak/width tau on cosmology, halo mass, baryon fraction, and black hole mass; the normalization phi on halo mass times feedback corrections; and the start time eta on Omega_m. The paper further finds that stellar feedback is the domi","pith_inferences":["Eqn 18 is fitted on emulated SFHs; applying it to other simulation suites like SWIFT-EAGLE or CAMELS-SAM would test whether the shared functional form reflects physical regularity or a property of these three models.","AGN feedback is almost unconstrained from SFH shape alone, so combining these equations with baryon fraction or black hole mass observations should sharpen late-time feedback constraints.","The double power-law's poor fit to ASTRID's sustained late-time tails suggests a modified form with an added plateau could alter the predicted falling slopes and the conclusion that ASTRID galaxies quench fastest."],"forward_implications":["Observed average SFHs could constrain cosmology and stellar feedback strength using Eqn 18 without rerunning simulations.","The same equation form should apply to other simulation codes, turning cross-model calibration into a coefficient-fitting exercise.","Because stellar feedback changes black hole growth, SFH-based constraints on stellar and AGN feedback will remain partially degenerate.","Reparameterizing winds by mass loading and energy per unit SFR makes the three models' SFH responses qualitatively consistent, clarifying interpretations."],"supporting_citations":[{"why":"Defines the CAMELS simulation suite and its stellar/AGN feedback scalings for TNG and SIMBA; supplies the CV, 1P, and LH runs used throughout.","marker":"Villaescusa-Navarro et al. 2021"},{"why":"Presents the CAMELS data release including the Latin Hypercube training sets for all three models.","marker":"Villaescusa-Navarro et al. 2023"},{"why":"Describes the CAMELS/ASTRID implementation, including black hole seeding and thermal/kinetic AGN feedback, whose parameters are analyzed here.","marker":"Ni et al. 2023"},{"why":"Establishes the SFH computation method and the mass-resolution thresholds used to define the robust galaxy sample.","marker":"Iyer et al. 2020"},{"why":"Supplies the double power-law parametrization of SFH shapes that the paper fits.","marker":"Carnall et al. 2019"},{"why":"Supplies the simulation-based inference software used to train the normalizing-flow emulator.","marker":"Ho et al. 2024"},{"why":"Companion paper documenting the normalizing-flow emulator and its forward modeling; the emulator supplies the SFHs used in the fits.","marker":"Lovell et al. 2024"},{"why":"Identifies the stellar- and AGN-dominated regimes in baryon fraction versus halo mass that the paper uses to interpret its state variables.","marker":"Wright et al. 2024"},{"why":"Provides the minimalist regulator model for energy- versus mass-loaded winds, used to interpret the wind reparametrization and baryon fraction trends.","marker":"Voit et al. 2024c"}],"fun_headline_variants":["One equation set predicts galaxy star formation in three major simulations","Unified formula links feedback and cosmology to galaxy growth across simulations","Star formation histories decoded by a single set of equations","How feedback shapes galaxy formation: one universal equation","A single equation set maps feedback and cosmology to star formation"],"cache_read_input_tokens":50560,"weakest_assumption_plain":"The entire analysis uses SFHs sampled from a machine-learning emulator rather than directly from the simulations, and the emulator is validated only qualitatively against a small single-parameter set; if it is biased, or if the double power-law fails for a substantial fraction of average SFHs, the derived equations and conclusions inherit that distortion.","fun_headline_variants_meta":{"raw":{"variants":["One equation set predicts galaxy star formation in three major simulations","Unified formula links feedback and cosmology to galaxy growth across simulations","Star formation histories decoded by a single set of equations","How feedback shapes galaxy formation: one universal equation","A single equation set maps feedback and cosmology to star formation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000281,"raw_usage":{"total_tokens":1538,"prompt_tokens":815,"completion_tokens":723,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":644}},"tokens_in":559,"tokens_out":723,"duration_ms":8095,"temperature":1.0,"reasoning_tokens":644,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:31:55.856614+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute average SFHs directly from the CAMELS LH simulation particle data, fit double power laws, and compare the resulting alpha, beta, tau, phi, eta to Eqn 18 on a grid of parameter values; systematic deviations beyond sampling noise would disprove universality. Specifically examine ASTRID galaxies with late-time star-formation tails: if those SFHs are not well described by a double power law, the falling-slope equation is biased.","supporting_citations":[],"review_version":1}