{"id":"c39f0fd4-486c-4bab-b239-326669cfc0c7","arxiv_id":"2412.00217","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"No single CAMELS IllustrisTNG simulation reproduces the scaling relations of both early- and late-type galaxies across datasets, and inferred cosmological and feedback parameters vary strongly with the chosen data.","lead":"Using galaxies from the CAMELS IllustrisTNG simulations, this paper searches for the simulation that best reproduces the sizes, dark matter fractions, and dark matter masses of early-type galaxies observed by SPIDER, ATLAS3D, and MaNGA DynPop. It finds that no single simulation matches the full range of observed galaxy types, and that the preferred simulation parameters shift depending on which dataset is used.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Forward-modeling check needed: simulated 3D half-mass radii and DM fractions may not be directly comparable to projected 1.35Re and SIS/JAM-derived observed values, so the claimed ETG/LTG dichotomy failure could be a comparison-space artifact.","rationale":"The reader's weakest-assumption analysis pinpoints the same load-bearing issue: the paper compares simulated 3D subfind quantities with observed quantities derived through projection, half-light-to-half-mass conversions, and dynamical modeling. This is not a minor technicality because the central conclusion is that the observed ETG/LTG dichotomy cannot be reproduced by simulations. If the observational pipeline systematically shifts ETGs and LTGs differently in the R*,1/2, fDM, and MDM relations, then the 'failure' could be an artifact of the comparison space rather than a physical property of IllustrisTNG. The paper itself concedes in Appendix A that forward-modeling would be needed to resolve this and does not perform it, so the concern is tied to a stated limitation rather than an unsupported speculation. I do not see a stronger load-bearing threat: the parameter constraints are explicitly framed as discrete-grid rankings, the chi-square method is transparent, and the resolution checks in Section 4.6 strengthen the qualitative dichotomy claim against numerical-resolution artifacts. Therefore the reader's CONDITIONAL verdict is appropriate, with the condition that the comparison space be validated before the quantitative parameter values are used for cosmological interpretation.","tokens_in":30859,"tokens_out":4625,"duration_ms":48910,"concrete_test":"Project the simulated galaxies from at least one CAMELS IllustrisTNG run (e.g. fiducial 1P_1_0 and joint best-fit LH_531) and, ideally, TNG100-1 along random lines of sight. Fit Sersic/MGE models to the mock images to measure Re, assign stellar masses and M/L using the same Chabrier IMF assumptions as the DynPop catalog, and generate mock stellar kinematics to run the JAMsph+gNFW modeling used by DynPop (or a validated surrogate). Derive R*,1/2, M*,1/2, and fDM with the exact observational recipes, apply the same T-Type/PLTG/VC/VF classification thresholds or the sSFR threshold, and recompute the median ETG and LTG trends. If the simulated ETG-LTG separation after forward-modeling is comparable to the observed MaNGA DynPop separation, the paper's central claim stands; if it shrinks, the claimed failure is an artifact of comparing 3D half-mass radii to 1.35Re and SIS/JAM mass models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central negative claim, that no CAMELS IllustrisTNG run reproduces the observed ETG/LTG dichotomy, rests on a one-to-one mapping between simulated 3D quantities and observationally projected, model-dependent quantities (Sections 2.1.1, 2.2.1-2.2.3, Appendix A). Simulated R*,1/2 is the 3D stellar half-mass radius from subfind; observed R*,1/2 is a deprojected half-light radius estimated as 1.35 Re, and M*,1/2 is set to M*/2. DM fractions and DM masses for SPIDER and ATLAS3D are obtained from SIS Jeans models, while MaNGA DynPop uses JAMsph+gNFW. If the 1.35 Re conversion, the M*/2 approximation, or the dynamical mass model is biased as a function of stellar mass or galaxy type, the apparent 'dichotomy' between observed ETGs and LTGs could be partly manufactured by the projection and modeling pipeline, and the no-single-simulation conclusion would not be a statement about the simulations themselves. The paper explicitly lists these biases in Appendix A and states that full forward-modeling is outside the scope, so this is a genuine unmet condition rather than a speculative one. The comparison with original TNG simulations (Section 4.6) shows the same qualitative non-dichotomy across resolutions, which argues against a pure resolution artifact, but it does not control for the comparison-space bias.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper extends the CASCO framework of Busillo et al. (2023) from late-type galaxies to early-type galaxies, comparing CAMELS IllustrisTNG simulations with three observational datasets (SPIDER, ATLAS3D, and MaNGA DynPop). The authors examine the size--stellar mass, dark-matter-fraction--stellar mass, and dark-matter-mass--stellar mass relations of simulated ETGs, rank the 1061 CAMELS runs with a cumulative reduced chi-squared statistic, and derive constraints on Omega_m, sigma_8, and the feedback parameters ASN1, ASN2, AAGN1, and AAGN2 via bootstrap resampling. They find that the best-fit simulations and the resulting parameter constraints differ between datasets, and that no single simulation reproduces the full MaNGA DynPop sample containing both ETGs and LTGs, because the observed ETG and LTG trends are dichotomous while simulated ETGs and LTGs follow similar trends. The paper also compares with the original IllustrisTNG simulations at different resolutions to test for resolution effects, and includes appendices on observational biases, raw CDF constraints, chi-squared degeneracy maps, and a selection-bias check.","tokens_in":31082,"tokens_out":5190,"duration_ms":53606,"significance":"If the central negative result is robust, it is an important cautionary finding for simulation-based inference: it would show that the current CAMELS IllustrisTNG parameter space cannot jointly reproduce the structural and dark-matter scaling relations of both early- and late-type galaxies, and that inferred cosmological and feedback parameters depend on galaxy type and observational dataset. The paper is systematic in its coverage: it uses three ETG datasets, checks the LTG+ETG joint fit, tests gas-mass corrections in Appendix A, and compares with TNG300, TNG100, TNG50, and lower-resolution TNG100 runs. These resolution and gas-correction checks give credibility to the claim that the failure to reproduce the ETG/LTG dichotomy is not simply a resolution artifact. However, the headline claim is vulnerable to a comparison-space issue: simulated 3D subfind quantities are compared directly with observed quantities derived from projected light and dynamical modeling, and the paper explicitly does not forward-model the simulations into the observational space.","major_comments":[{"comment":"The central claim that no CAMELS IllustrisTNG simulation reproduces the observed ETG/LTG dichotomy rests on a direct one-to-one comparison between simulated 3D quantities (stellar half-mass radius, DM fraction within that radius, DM mass within that radius) and observed quantities derived with R*=1.35 Re, M*,1/2=M*/2, and SIS- or JAM-based dynamical models. The paper lists these projection and modeling biases in Appendix A and states that full forward-modeling is outside the scope, but it does not quantify whether type-dependent biases could create or erase the apparent ETG/LTG dichotomy. Because the dichotomy is the load-bearing element of the paper's main conclusion, the authors should provide a sensitivity test, for example by forward-modeling mock observations of the TNG snapshots or by estimating how large a type-dependent M/L or mass-anisotropy bias would need to be to explain the observed separation. Without such a test, the headline negative result remains vulnerable to a comparison-space artifact.","section":"Section 2.2 and Appendix A"},{"comment":"The reduced chi-squared in Eq. (2) is stochastic: the numerator subtracts a randomly drawn point from a Gaussian centered on the observed median, so repeated evaluations for the same simulation and the same data produce different chi-squared values. The bootstrap procedure resamples the datasets but does not propagate this additional Monte-Carlo noise into the reported parameter uncertainties. In addition, the Gaussian smoothing of the empirical CDFs with standard deviation sigma=1 is ad hoc, and the quoted 16th/50th/84th percentiles change when the smoothing is omitted (Tables 3 and B.1). The paper should quantify the sensitivity of the constraints to the smoothing scale and to the number of bootstrap resamples, and should either justify Eq. (2) or replace it with a deterministic ranking statistic.","section":"Section 4.1, Eq. (2)"},{"comment":"The full MaNGA DynPop ETG+LTG analysis yields a best-fit simulation, LH_531, with cumulative reduced chi-squared 8.77, and the bootstrap constraints in Table 5 have very large and asymmetric uncertainties (for example ASN1 = 1.83^{+0.74}_{-1.19}). The conclusion that no single simulation reproduces the full sample is therefore based on the absence of any good fit, but the paper does not provide a statistical test of whether the observed ETG/LTG dichotomy is significantly different from the simulated one. A quantitative comparison of the slopes or offsets between ETG and LTG median trends as a function of stellar mass would strengthen the central claim and would make the result less dependent on visual inspection of Figs. 8 and 9.","section":"Section 4.5 and Table 5"},{"comment":"The original IllustrisTNG model was calibrated partly against the z=0 stellar mass--stellar size relation and the gas mass content of groups, as stated in Section 2.1.2. Because the CAMELS fiducial run uses the same subgrid model and parameters, the size-mass and DM-fraction relations are not fully independent probes of cosmology or feedback. This does not invalidate the negative result, but it means the derived values for Omega_m and sigma_8 should be interpreted as conditional on the TNG subgrid implementation rather than as independent cosmological measurements. The paper should make this conditioning explicit in the abstract or in the discussion of the SPIDER/ATLAS3D/MaNGA constraints.","section":"Section 2.1.2 and Section 4.6"}],"minor_comments":[{"comment":"The phrase 'mathematicaresource function' appears to be a formatting artifact; it should read 'Mathematica resource function'.","section":"Section 4.1"},{"comment":"There are typographical errors such as 'avaliable' in the discussion of the new CAMELS simulations; the manuscript would benefit from a careful proofreading pass.","section":"Section 3.1"},{"comment":"The selection criteria for MaNGA LTGs are listed as T-Type>=0, PLTG>=0.5, VC=3, and VF=0; it should be stated explicitly whether all four conditions are required simultaneously and how many galaxies are removed by each individual cut.","section":"Section 2.2.3"},{"comment":"In the caption of Fig. 1, the labels 'ASN,1=0.25, ASN,1=1.00, ASN,1=2.30' appear to use 'ASN,1' three times where the second and third entries should refer to the other varied parameters; please correct the caption to match the column labels.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its limitations and includes useful robustness checks, but the main negative claim requires a stronger control of the comparison space. The forward-modeling concern raised by the skeptical reader is real and is acknowledged in Appendix A; it is not a purely speculative issue. I see no grounds for rejection, but the manuscript needs a quantitative sensitivity analysis or mock-observation test before the central claim can be considered established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a solid second paper in the CASCO series, and it does what Paper I didn't: it takes the same ranking method to early-type galaxies, using three independent datasets, and it documents a clear failure of the CAMELS IllustrisTNG simulations to reproduce the observed dichotomy between ETG and LTG scaling relations in MaNGA DynPop. The resolution check with the original TNG runs (including the same-volume different-resolution TNG100 series) shows the same qualitative behavior, so the negative result is not just a low-resolution artifact. The gas-correction variant in Appendix A also does not rescue the dichotomy. That is real evidence, and I take the central negative claim seriously.\n\nThe paper's main soft spot is one the authors acknowledge in Appendix A: the simulated quantities are 3D (half-mass radius, DM fraction within it), while the observed quantities are projected and model-derived (1.35 Re, M*/2, SIS/JAM masses). If those conversions are biased as a function of mass or galaxy type, the apparent dichotomy could be partly manufactured by the comparison space. The authors list the biases but do not forward-model the simulations into the observational space. This is a genuine unmet condition, not a speculation. The same caveat applies to the absolute parameter values: they come from a discrete grid-ranking with bootstrap resampling, a Gaussian CDF smoothing width chosen by hand, and a chi-squared that includes a random draw. The error bars in Tables 3 and 5 therefore do not include model inadequacy, and the numbers should be read as best-fit simulation selections rather than posterior inferences. The paper is fairly careful about this in the method section, but the abstract and conclusions lean a little harder on the constraints than the method supports.\n\nThat said, this is a useful and honest contribution. The dataset-specificity of the results is itself a finding, and the paper is written clearly enough to reproduce from the public catalogs. It deserves a serious referee. I would recommend conditional acceptance, with the main revision being a reframing of the parameter constraints and, ideally, an explicit forward-modeling or bias-correction exercise for the comparison space. If I were working on simulation-based inference with CAMELS, I would cite this paper.","headline":"A careful, transparent extension of the CASCO method to ETGs, but the headline no-single-simulation claim depends on an un-forward-modeled comparison between 3D simulated quantities and projected observables.","tokens_in":31723,"tokens_out":2611,"would_cite":true,"duration_ms":24189,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that no CAMELS IllustrisTNG simulation can jointly reproduce the observed scaling relations of early- and late-type galaxies, so the inferred cosmological and feedback parameters depend on which dataset is fitted.","keywords":["galaxy scaling relations","early-type galaxies","late-type galaxies","dark matter fraction","cosmological parameters","supernova feedback","CAMELS simulations","IllustrisTNG"],"falsifier":"Apply the same projection and mass-modeling pipeline used for SPIDER, ATLAS3D, and MaNGA DynPop to the simulated galaxies and rerun the bootstrap ranking; if one simulation then fits all three datasets and reproduces the early/late dichotomy, the paper's central negative result is an artifact of comparing unlike quantities, whereas if no simulation still fits, the claim that current simulations miss real galaxy diversity stands.","tokens_in":30556,"feed_emoji":"🌌","tokens_out":7893,"duration_ms":68317,"temperature":0.7,"pith_summary":"The paper extends the CASCO comparison between cosmological simulations and galaxy observations from late-type galaxies to early-type galaxies. It asks whether one CAMELS IllustrisTNG simulation can reproduce the observed relations between stellar mass and stellar half-mass radius, dark-matter fraction within the half-mass radius, and dark-matter mass within the half-mass radius, for early-type galaxies in SPIDER, ATLAS3D, and MaNGA DynPop. The best-fit simulation changes from sample to sample: SPIDER favors a high supernova feedback strength $A_{\\rm SN1}$, ATLAS3D and MaNGA favor low $A_{\\rm SN1}$, and MaNGA requires extreme values of $\\Omega_m$ and $\\sigma_8$. When the full MaNGA DynPop sample is used, so that early- and late-type galaxies come from the same survey and analysis, no simulation reproduces both trends. The paper reads this as evidence that current simulations, and the six varied parameters explored here, do not capture the full diversity of real galaxies.","feed_headline":"No single simulation fits early and late galaxies","feed_subtitle":"CAMELS IllustrisTNG runs give different best-fit cosmology and feedback for each observational dataset.","key_machinery":"The ranking machinery is the cumulative reduced chi-squared. For each of the three scaling relations, each simulated galaxy is compared with the observed median trend interpolated at its stellar mass, with the observed 16th-to-84th percentile scatter used as the uncertainty, and a Gaussian draw from that scatter replaces the fixed observed value in the numerator. Simulations are first filtered to require at least one simulated galaxy in each stellar-mass bin, then ranked by the sum of reduced chi-squared over the three relations. The parameter constraints come from 100 bootstrap resamplings of the observed and simulated samples, with the best-fit simulation selected each time; smoothing the empirical cumulative distributions of the selected parameters gives the reported 16th, 50th, and 84th percentiles.","core_discovery":"The central claim is that the CAMELS IllustrisTNG simulations cannot jointly reproduce the scaling relations of early- and late-type galaxies, and that the inferred cosmological and feedback parameters depend on the chosen observational trends. For early-type galaxies alone, the best-fit simulations are LH_523 for SPIDER, LH_797 for ATLAS3D, and LH_586 for MaNGA DynPop, with $\\Omega_m$ between $0.16$ and $0.25$, $\\sigma_8$ between $0.77$ and $0.97$, and the supernova feedback parameters $A_{\\rm SN1}$ and $A_{\\rm SN2}$ reversing between SPIDER (high $A_{\\rm SN1}$, low $A_{\\rm SN2}$) and ATLAS3D or MaNGA (low $A_{\\rm SN1}$, high $A_{\\rm SN2}$). Constraining one simulation for the full MaNGA DynPop sample, early- and late-type together, gives a best fit with cumulative reduced $\\tilde{\\chi}^2 = 8.77$, still a poor representation; the observed early- and late-type relations have different slopes and normalizations, while the simulated populations of both types follow roughly the same trend. The AGN feedback parameters $A_{\\rm AGN1}$ and $A_{\\rm AGN2}$ are effectively unconstrained, because varying them has almost no effect on the simulated scaling relations.","pith_inferences":["A natural next step, not done by the paper, is to forward-model the simulated galaxies through the same projection and mass-modeling pipelines as the observed samples; if the galaxy-type dichotomy survives that test, the conclusion points to baryon physics in the simulations rather than to the comparison space.","The paper's appendix test, which adds gas mass to the simulated dark-matter mass for the MaNGA comparison, shifts the best-fit supernova parameters toward the SPARC-based values; this suggests the exact definition of “dark matter within the half-mass radius” should be made identical on both sides before drawing physical conclusions.","Because the six varied parameters omit other AGN feedback parameters available in newer CAMELS runs, the null result is a statement about this parameter set, not about all possible feedback implementations.","The strong $\\Omega_m$-$\\sigma_8$ degeneracy shown in the chi-squared heat maps implies that these scaling relations alone cannot rank cosmological models on the $S_8$ tension; a combined probe with other observables would be needed."],"forward_implications":["Single-sample constraints are not universal: a simulation that fits SPIDER early-type galaxies does not fit ATLAS3D or MaNGA, and the SPARC late-type best fit fails for early types.","Supernova feedback constraints flip sign with sample: SPIDER wants stronger $A_{\\rm SN1}$ and weaker $A_{\\rm SN2}$, while ATLAS3D and MaNGA want the reverse, so feedback parameters should be reported together with the galaxy type and sample definition.","The AGN feedback parameters $A_{\\rm AGN1}$ and $A_{\\rm AGN2}$ cannot be constrained from these scaling relations, because varying them leaves the trends nearly unchanged.","Resolution is not the fix: the higher-resolution original IllustrisTNG runs also fail to reproduce the observed early/late dichotomy, and lower resolution sometimes matches the data better.","The observed early/late dichotomy in MaNGA DynPop, with different slopes and normalizations for the two galaxy types, is a sharper test for simulations than any single scaling relation."],"supporting_citations":[{"why":"Supplies the ranking method and the SPARC late-type comparison that this paper extends to early-type galaxies.","marker":"Busillo et al. 2023"},{"why":"Introduces the CAMELS IllustrisTNG simulation suite with varied cosmological and feedback parameters that is the object being tested.","marker":"Villaescusa-Navarro et al. 2021"},{"why":"Defines the IllustrisTNG model and its calibrated fiducial parameters used by the CAMELS fiducial simulation.","marker":"Pillepich et al. 2018"},{"why":"Provides the SPIDER early-type galaxy sample, one of the three observed datasets.","marker":"La Barbera et al. 2010"},{"why":"Provides the ATLAS3D early-type galaxy sample used for the second set of observed trends.","marker":"Cappellari et al. 2011"},{"why":"Provides the MaNGA DynPop catalog with JAM dynamical models used for the third observed dataset.","marker":"Zhu et al. 2023"},{"why":"Provides the SPARC late-type galaxy sample that supplies the LTG constraints in the combined fits.","marker":"Lelli et al. 2016"},{"why":"Gives the conversion $R_{*,1/2} \\approx 1.35 R_e$ used to bring observed sizes into the simulated three-dimensional half-mass-radius space.","marker":"Wolf et al. 2010"},{"why":"Defines the singular-isothermal-sphere Jeans dynamical masses used for the SPIDER observed dark-matter fractions.","marker":"Tortora et al. 2012"},{"why":"Provides the theoretical total-mass-stellar-mass relation used to check the simulated total masses.","marker":"Moster et al. 2013"}],"fun_headline_variants":["No single simulation fits early and late galaxies","Best-fit cosmology depends on chosen galaxy sample","CAMELS sims can't reconcile early and late type scaling","Galaxy type decides the inferred feedback strengths","Early vs late galaxies demand different simulation fits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument relies on simulated galaxies and observed galaxies being measured in the same way: simulations give three-dimensional stellar half-mass radii and dark-matter masses, while observations infer these from projected light using a standard radius conversion and simplified spherical mass models, and the paper does not forward-model the simulations through those observational steps.","fun_headline_variants_meta":{"raw":{"variants":["No single simulation fits early and late galaxies","Best-fit cosmology depends on chosen galaxy sample","CAMELS sims can't reconcile early and late type scaling","Galaxy type decides the inferred feedback strengths","Early vs late galaxies demand different simulation fits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000621,"raw_usage":{"total_tokens":2968,"prompt_tokens":1121,"completion_tokens":1847,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":737,"completion_tokens_details":{"reasoning_tokens":1777}},"tokens_in":737,"tokens_out":1847,"duration_ms":13851,"temperature":1.0,"reasoning_tokens":1777,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:36:52.065702+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the same projection and mass-modeling pipeline used for SPIDER, ATLAS3D, and MaNGA DynPop to the simulated galaxies and rerun the bootstrap ranking; if one simulation then fits all three datasets and reproduces the early/late dichotomy, the paper's central negative result is an artifact of comparing unlike quantities, whereas if no simulation still fits, the claim that current simulations miss real galaxy diversity stands.","supporting_citations":[{"cited_title":"R., et al","cited_arxiv_id":null,"evidence_quote":"Supplies the ranking method and the SPARC late-type comparison that this paper extends to early-type galaxies."},{"cited_title":"2021, ApJ, 915, 71","cited_arxiv_id":null,"evidence_quote":"Introduces the CAMELS IllustrisTNG simulation suite with varied cosmological and feedback parameters that is the object being tested."},{"cited_title":"2023, MNRAS, 522, 6326 Article number, page 16 of 20 V","cited_arxiv_id":null,"evidence_quote":"Provides the MaNGA DynPop catalog with JAM dynamical models used for the third observed dataset."},{"cited_title":"R., de Carvalho, R","cited_arxiv_id":null,"evidence_quote":"Defines the singular-isothermal-sphere Jeans dynamical masses used for the SPIDER observed dark-matter fractions."}],"review_version":1}