{"id":"6d6a769f-9504-4d0e-aaa5-6011dd25394c","arxiv_id":"1908.09844","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An extremely randomized tree model trained on IllustrisTNG predicts baryonic properties from dark matter halos and generates a Gpc-scale catalog compatible with semi-analytic models.","lead":"A machine trained on a detailed galaxy simulation learns to estimate galaxy properties such as stellar mass and star formation rate from dark matter alone. The method can quickly build galaxy catalogs for a huge dark matter simulation, which could speed up survey planning and large-scale structure studies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Transfer from TNG hydro to MDPL2 DM-only is asserted, not validated: Appendix B checks only the halo mass function, while the machine uses per-halo features that baryonic back-reaction and different halo finders can shift.","rationale":"The paper makes two claims: an in-sample accuracy improvement (Section 3, Table 2) and a usable Gpc-scale catalogue (Section 4). The in-sample claim is supported by test-set metrics, though model selection on the test set is a secondary concern. The application claim is the more novel and consequential part, and it depends entirely on the unstated equivalence between TNG hydro halo features and MDPL2 DM-only halo features. The authors provide exactly one check of this equivalence, the DM mass function, which is insensitive to the higher-order input features (Vmax, spin, environment, merger history). Given the known back-reaction effects cited in their own Appendix B, the assumption is not secure. The proposed TNG-Dark test is the natural control: it keeps cosmology, initial conditions, and (if SubFind is used) halo definition fixed, isolating the hydro-versus-DM-only transfer. If the machine is unbiased there, the remaining differences between TNG-Dark and MDPL2 (cosmology, Rockstar) can be addressed separately. The reader's weakest_assumption names exactly this transfer and mass-definition issue, so agreement is 'agree'. The conditional verdict already reflects the need for validation; no adjustment is required, hence UNCHANGED.","tokens_in":22448,"tokens_out":8092,"duration_ms":88614,"concrete_test":"Use the publicly available IllustrisTNG-Dark run (the DM-only counterpart of TNG100, same initial conditions and cosmology): apply the trained MSSM to its halo catalogue, match halos 1-to-1 with TNG100 hydro by ID, and compute binned residuals of predicted minus true stellar mass and SFR versus true Mstar or Mhalo. If the binned median bias or scatter exceeds the in-sample MBE/MBSD reported in Table 2 (e.g., MBE_stellar = 0.0013), the transfer assumption fails and the MDPL2 catalogue should be treated as unvalidated; if residuals are consistent with the in-sample errors across the full mass range, the transferability concern is resolved. As a supplementary check, run Rockstar on TNG-Dark and repeat the comparison to separate baryonic back-reaction from halo-finder and mass-definition differences.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Load-bearing concern: Section 4 assumes that the per-halo DM feature vectors of MDPL2 (Rockstar) live in the same space as those of TNG100 hydro (SubFind), so that the ERT trained on TNG can be applied directly. The only support is Appendix B/Figure B1, which compares the DM halo mass function of TNG100 to TNG-Dark and finds a <1% shift. That is a one-dimensional marginal. The machine's inputs are not just Mvir: they include Vdisp, Vmax, spin, environmental measures over a (2 Mpc)^3 volume, and three merger-tree features (Table 1). Baryonic back-reaction is known to alter halo concentration, Vmax, spin, and subhalo abundance at fixed mass (cited in Appendix B: Duffy et al. 2010; Cui et al. 2012; Chua et al. 2019), and the MDPL2 catalogue is produced by a different halo finder with a different mass definition. A <1% match in the mass function does not constrain these joint feature distributions. No per-halo transfer check is performed. The MDPL2 catalogue has no baryonic ground truth, so 'largely compatible with SAMs' is a comparison to other models, not a validation of the transplanted TNG mapping; it neither confirms nor refutes a systematic bias from domain shift. If this assumption fails, every baryonic property in the (1 h^-1 Gpc)^3 catalogue is biased, which is the paper's main use case. The manuscript itself acknowledges the assumption in Section 6.1 and Appendix B, but leaves the decisive test undone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a machine-learning pipeline, MSSM, that uses extremely randomized trees trained on the IllustrisTNG hydrodynamic simulation to predict galaxy baryonic properties (gas mass, stellar mass, black hole mass, star formation rate, metallicity, and stellar magnitudes) from dark-matter halo features. The authors propose three improvements over a baseline model: logarithmic scaling of the error function, addition of historical and environmental halo features, and a two-stage learning scheme that uses predicted stellar magnitudes as an intermediary to predict SFR or stellar mass. They report substantial accuracy gains on a held-out TNG test set, then apply the trained machine to the much larger MultiDark-Planck (MDPL2) DM-only simulation and compare the resulting galaxy catalogue with those from popular semi-analytic models (Sag, Sage, Galacticus), finding 'largely compatible' distributions. The paper also analyzes feature importances and training-set-size requirements, and releases the MSSM galaxy catalogue.","tokens_in":22872,"tokens_out":4713,"duration_ms":48051,"significance":"If the domain-transfer assumption holds, the MSSM provides a fast, flexible way to 'paint' baryonic properties from a high-resolution hydrodynamic simulation onto a very large DM-only volume, which is scientifically valuable for survey-scale predictions. The technical improvements, especially logarithmic error scaling and two-stage learning, are potentially useful to the broader galaxy-formation machine-learning community. The paper is clearly written, gives a detailed description of the pipeline, and makes the catalogue publicly available. However, the headline accuracy claims are measured only inside the TNG training distribution, and the application to MDPL2 rests on a transferability assumption that is validated only via a one-dimensional mass-function comparison. As a result, the main application-phase claim is not yet established with the same rigor as the in-sample accuracy claim.","major_comments":[{"comment":"The application of the TNG-trained machine to MDPL2 assumes that the per-halo DM feature vectors of the two simulations are drawn from the same distribution. The supporting evidence in Appendix B and Figure B1 is only a comparison of the DM halo mass function, which the authors find to be shifted by less than 1% between IllustrisTNG and IllustrisTNG-Dark. This is a one-dimensional marginal check, while the machine's inputs include halo velocity dispersion, maximum circular velocity, spin, local environmental measures within a (2 Mpc)^3 volume, and three merger-tree quantities (Table 1). Baryonic back-reaction is known to affect halo concentration, Vmax, spin, and subhalo abundance at fixed mass (as the authors themselves cite in Appendix B), and the MDPL2 catalogue is produced with a different halo finder (Rockstar) and a different mass definition than the TNG catalogue (SubFind). A <1% match in the mass function does not constrain these joint feature distributions. The manuscript acknowledges this assumption in Section 6.1 and Appendix B but does not perform the decisive per-halo or joint-distribution test, for example by matching halos between TNG and TNG-Dark and comparing the distribution of the feature vector as a whole. Without such a test, the predicted baryonic properties in the (1 h^-1 Gpc)^3 catalogue may be systematically biased, and the 'largely compatible with SAMs' statement in Section 4.1 is a model-to-model comparison that neither confirms nor refutes such bias.","section":"Section 4.1, Section 6.1, Appendix B"},{"comment":"The best combination of the three proposed improvements is selected separately for each output property by evaluating the MBE and MBSD scores on the same test set that is later used to report the final accuracy values. This is a form of test-set reuse or selection on the test set, which can inflate the reported gains relative to what would be achieved on a truly unseen set. The authors should either use a nested cross-validation procedure (where the test set is used only once, after all model selection is complete) or a separate validation set for selecting among the combinations, and then report the accuracy on the untouched test set. They should also quantify the extent of the inflation, for example by reporting the accuracy of a randomly chosen combination or of the 'all improvements together' model, so that readers can assess the sensitivity to the selection procedure.","section":"Section 3.2.4, Table 2"},{"comment":"The accuracy improvements (e.g., MSE decreasing from 2.0e-2 to 1.9e-4, PCC increasing from 0.971 to 0.987) are measured on a test set drawn from the same TNG100 simulation used for training. This is a legitimate held-out test within the training distribution, but it is interpolation, not out-of-sample prediction. The genuine out-of-sample application, on MDPL2, has no baryonic ground truth, so the comparison with SAMs in Section 4.1 is a comparison between two models and cannot validate the transferred mapping. The abstract's claim of 'significantly increased accuracy compared to prior attempts' should therefore be qualified as accuracy in reproducing the TNG galaxy-halo correlation within the TNG volume, not as accuracy on an independent simulation. The authors should make this distinction explicit in the abstract and conclusions.","section":"Section 3.1, Section 4.1"},{"comment":"The training-set sufficiency argument uses learning curves for the baseline model with only three input features (Figure 8 and Section 5.2). The improved model has about fourteen input features (Table 1), including environmental and historical quantities. The authors themselves note in Section 5.2 that if the machine is built with more important input features, a bigger training set may be needed to converge. The current evidence that ~4e4 halos after pruning are sufficient is therefore not directly applicable to the improved model. The authors should provide learning curves for the improved model, or at least for a model with the full feature set, to support the claim that the training set is large enough.","section":"Section 2.5.1, Section 5.2, Figure 8"},{"comment":"The paper reports that the two-dimensional distribution of predicted stellar mass and sSFR is narrower in the MSSM catalogue than in the original IllustrisTNG data (Section 4.2, Figure 6). This narrowing is attributed to underfitting or the limited number of important input features. This is a scientifically relevant limitation because it means the MSSM catalogue does not preserve the full galaxy diversity present in the simulation, and it affects the interpretation of the 'largely compatible' statement in Section 4.1: a distribution that is narrower than the truth can appear compatible in a one-dimensional comparison while being biased in joint and high-order statistics. The authors should quantify this narrowing (e.g., the ratio of standard deviations or a two-sample test) and discuss the practical consequences for downstream scientific applications, such as clustering or abundance matching.","section":"Section 4.2, Figure 6"}],"minor_comments":[{"comment":"There are several typographical errors: 'Becuase' in Section 2.5.1, 'IllutrisTNG' in Section 2.4.1, 'brining' in Section 3.2.1, and 'Feburary' in the header. These should be corrected.","section":"Throughout"},{"comment":"Equation (2) for MSE appears to be missing the square on the difference term; the text defines it as a mean square error, so the formula should read (1/N) sum (y_pred - y_TNG)^2. If this is a typesetting artifact, please ensure the final version is unambiguous.","section":"Section 3.1, Eq. (2)"},{"comment":"The MBSD values for SFR (36.10 and 20.15) are orders of magnitude larger than the MBE values (1.71 and 1.00) in the same table. This seems unusual and should be explained, or the numbers should be checked for a unit or typographical error.","section":"Table 2"},{"comment":"The two-stage learning scheme uses eight photometric bands as an intermediary. The choice of which bands to use for a given output is described in Appendix A, but in Section 3.2.3 it is stated that for SFR, the g band is used, while for stellar mass, the paper refers to Appendix A without specifying the band set used for the results in Table 2. Please clarify exactly which bands are used for each output property in the 'Best combination' row.","section":"Section 3.2.3"},{"comment":"The residual panels in Figure 4 show the difference between the predicted PDF and the IllustrisTNG PDF, even for the MDPL2 predictions and the SAM catalogue. This is potentially confusing because the residuals do not indicate agreement with MDPL2 ground truth (which does not exist), but rather a comparison to the TNG training data. Please clarify in the caption what the residuals represent.","section":"Figure 4"},{"comment":"The claim that MSSM 'does not require any recipes with fine-tuned parameters or human bias' is an overstatement, since the choice of input features, the pruning thresholds, and the selection of stellar magnitudes as an intermediary are all human decisions. Please rephrase to acknowledge these choices.","section":"Section 6.1"}],"recommendation":"major_revision","confidential_remarks":"This is a generally well-executed machine-learning application with a clear narrative, and the in-sample results are convincing. However, the central application-phase claim—that the TNG-trained mapping transfers to MDPL2—rests on an insufficient validation (only the halo mass function is checked), and the model-selection procedure on the test set may inflate the reported gains. These issues are fixable with additional analysis and should be addressed before publication. The paper is a reasonable candidate for MNRAS after major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work on fast galaxy catalogue generation. The paper takes the Kamdar/Agarwal idea of training a random forest on hydrodynamic simulations to map dark matter halos to baryonic properties, and makes several concrete improvements that actually move the needle: a log-scaled error function that fixes the low-mass end, merger-history and environment features, and a two-stage scheme where predicted stellar magnitudes are fed into the next model. The ablation study is careful and the learning-curve analysis is a nice sanity check. They also ship a catalogue for the MDPL2 volume, which is a useful community resource.\n\nThe soft spots are real but not fatal. The headline accuracies come from a test set that was also used to select the best combination of improvements, so the gains are likely a bit optimistic. That is a standard sin in this literature, and the authors are transparent about the search. More important is the application phase: the machine trained on IllustrisTNG is applied to MDPL2, a different simulation with a different halo finder and mass definition, on the assumption that per-halo DM features live in the same space. The support is a <1% match in the halo mass function. That is a weak check, because the model uses Vmax, spin, environment, and merger-tree features that are known to respond to baryonic back-reaction and to halo finder choices. The paper acknowledges this in Section 6.1 and Appendix B, but the decisive per-halo transfer test is left undone. Comparing the resulting catalogue to SAMs is a sanity check, not a validation.\n\nStill, the central argument holds: within the training distribution, MSSM predicts baryonic properties from DM features more accurately than the previous baseline, and the runtime is trivial. The transfer concern is a limitation, not a conceptual failure. The paper is written by people who know what their method can and cannot do.\n\nMy take: this deserves a serious referee. The main requests would be a separate validation set for model selection, explicit hyperparameters, and either a per-halo cross-simulation transfer check using something like TNG-Dark or a stronger caveat on the MDPL2 catalogue. I would not cite it in my own work in the next year, but I would point a student toward it as a clean example of how to do an ablation study in astro-ML.","headline":"Solid, honest engineering paper that improves ML-based halo painting with real technical contributions; the Gpc application is the weak link because the cross-simulation transfer is asserted rather than validated.","tokens_in":23355,"tokens_out":1292,"would_cite":false,"duration_ms":16445,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A machine trained on a hydrodynamic simulation can estimate baryonic galaxy properties from dark-matter halo properties alone, then paint them onto a gigaparsec dark-matter-only volume.","keywords":["galaxy formation","dark matter halos","machine learning","hydrodynamic simulations","semi-analytic models","galaxy catalogues","extremely randomized trees","IllustrisTNG"],"falsifier":"Run the trained MSSM on IllustrisTNG-Dark, the dark-matter-only counterpart of the training simulation, and compare predicted stellar masses with the actual stellar masses of matched halos in the hydrodynamic run: a systematic bias in the $M_\\star$-$M_{\\mathrm{halo}}$ relation, or a scatter larger than the intrinsic scatter of that relation, would falsify the transferability assumption. A second check is to compare the two-point galaxy clustering of the MDPL2 catalogue with observed clustering at fixed stellar mass; a significant mismatch would show that the machine's galaxy-halo connection is not faithful enough for large-volume statistics.","tokens_in":22272,"feed_emoji":"🌌","tokens_out":14722,"duration_ms":131452,"temperature":0.7,"pith_summary":"This paper claims that the galaxy-halo connection inside a hydrodynamic simulation can be learned as a machine mapping from dark-matter halo properties to baryonic galaxy properties, then applied to a much larger dark-matter-only simulation. If true, this offers a cheap way to build galaxy catalogues at gigaparsec scales, where full hydrodynamic runs are infeasible. The machine-assisted semi-simulation model (MSSM) uses an extremely randomized tree algorithm with three improvements: a logarithmically scaled error function, historical and environmental halo features, and two-stage learning in which predicted stellar magnitudes feed a second model. Trained on IllustrisTNG's $(75\\,h^{-1}\\,\\mathrm{Mpc})^3$ volume, it predicts stellar mass, gas mass, black hole mass, star formation rate, metallicity, and stellar magnitudes with smaller mean binned errors than the baseline; applied to the $(1\\,h^{-1}\\,\\mathrm{Gpc})^3$ MultiDark-Planck run, its galaxy distribution functions are largely compatible with semi-analytic catalogues.","feed_headline":"Predicting galaxies from dark matter alone across a gigaparsec box","feed_subtitle":"Trained on IllustrisTNG, it predicts galaxy properties in a dark-matter-only run, matching semi-analytic models.","key_machinery":"The load-bearing mechanism is the extremely randomized tree (ERT), an ensemble of decision trees that chooses node splits randomly. Three training choices carry the accuracy gain: (1) a logarithmically scaled error function, so that fractional errors at the low-mass end are not overwhelmed by absolute errors at the high-mass end; (2) additional input features extracted from each halo's merger tree and local environment, such as merger counts, last-major-merger mass ratio, local density, and the semi-potential $\\Phi_s=\\sum_i M_i/R_i$ within a $(2\\,\\mathrm{Mpc})^3$ volume; and (3) two-stage learning, in which a first ERT predicts stellar magnitudes from dark-matter features and those predicted magnitudes become inputs to a second ERT that predicts star formation rate. The first machine also outputs feature importances, which show maximum circular velocity dominating at high redshift and halo mass and velocity dispersion taking over by $z=0$.","core_discovery":"The paper's central claim is that the baryonic content of a galaxy is predictable, to useful accuracy, from the bulk properties of its dark-matter halo alone, once a machine has been shown enough examples from a hydrodynamic simulation. In the reported test, stellar mass prediction improves from a mean binned error of $0.0018$ (baseline) to $0.0013$, and star formation rate from $1.71$ to $1.00$; the predicted probability distributions for all six baryonic properties move closer to the IllustrisTNG data. When the trained machine is applied to the MultiDark-Planck dark-matter-only simulation, the resulting catalogue's distribution functions are largely compatible with the SAG semi-analytic model, and for black hole mass, star formation rate, and stellar magnitudes it tracks IllustrisTNG more closely than SAG does, which is exactly the behaviour the model was designed for.","pith_inferences":["Pith inference: the two-stage learning trick should generalise; any baryonic property that is accurately predictable from dark matter and strongly correlated with a harder target could be chained as an intermediary, suggesting a searchable design space for future pipelines of this kind.","Pith inference: because the transferability assumption is only tested on the halo mass function, a natural next experiment is to run the trained machine on IllustrisTNG-Dark and compare per-halo predictions with the hydrodynamic run; this would isolate transfer error from model error.","Pith inference: if the machine's galaxy-halo connection is faithful, the MDPL2 catalogue could be forward-modelled into galaxy clustering, weak-lensing, or CMB-lensing predictions; discrepancies with observations would point to where IllustrisTNG's baryon physics needs revision.","Pith inference: the feature-importance analysis suggests that at low redshift halo mass and velocity dispersion carry most of the information; adding merger-orbit or tidal features could broaden the output diversity that the paper identifies as too narrow."],"forward_implications":["Large-scale galaxy catalogues from dark-matter-only runs become a matter of minutes: the machine paints roughly $10^6$ halos in a $(1\\,h^{-1}\\,\\mathrm{Gpc})^3$ volume in tens of minutes, versus weeks for a hydrodynamic simulation.","The baryon physics of a specific hydrodynamic simulation can be transplanted onto any sufficiently resolved dark-matter-only simulation without analytic recipes or tuned parameters.","Differences between MSSM and SAM catalogues localise where SAM prescriptions deviate from hydrodynamic simulations, pointing to specific subgrid physics to improve.","Because the machine reproduces the IllustrisTNG number densities, the resulting catalogue is suited for volume-limited studies such as baryonic acoustic oscillations and large-scale structure statistics."],"supporting_citations":[{"why":"Supplies the earlier random-forest baseline whose accuracy MSSM is claimed to significantly improve on.","marker":"Kamdar et al. 2016b"},{"why":"Defines the extremely randomized tree algorithm on which the machine is built.","marker":"Geurts et al. 2006"},{"why":"Introduces the IllustrisTNG (TNG100) simulation whose halo and galaxy data form the training set.","marker":"Nelson et al. 2018a"},{"why":"Describes the IllustrisTNG galaxy formation model whose baryonic properties are the prediction targets.","marker":"Pillepich et al. 2018"},{"why":"Provides the MultiDark-Planck (MDPL2) dark-matter-only simulation that the trained machine is applied to.","marker":"Klypin et al. 2016"},{"why":"Supplies the semi-analytic model catalogues used to assess whether MSSM is compatible with SAMs.","marker":"Knebe et al. 2018"},{"why":"The SAG semi-analytic catalogue used as the representative SAM in the PDF comparisons.","marker":"Cora et al. 2018"},{"why":"Provides the halo finder and merger tree construction that produced the MDPL2 halo catalogue and merger histories used as inputs.","marker":"Behroozi et al. 2013"}],"fun_headline_variants":["AI predicts galaxy traits from dark matter halos alone","From dark matter to galaxies: ML learns baryonic physics","Machine learns to fill in galaxy properties from DM-only sims","Trained on hydro sim, AI predicts baryons from dark matter","Fast galaxy predictions from dark matter alone via ML"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that halos in a dark-matter-only simulation match halos in the hydrodynamic simulation closely enough for the learned galaxy-halo mapping to carry over, and that the two simulations' halo mass definitions are compatible after the mass cut; the paper's support for this is a less than 1% agreement in the overall halo mass function, not per-halo comparisons.","fun_headline_variants_meta":{"raw":{"variants":["AI predicts galaxy traits from dark matter halos alone","From dark matter to galaxies: ML learns baryonic physics","Machine learns to fill in galaxy properties from DM-only sims","Trained on hydro sim, AI predicts baryons from dark matter","Fast galaxy predictions from dark matter alone via ML"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000528,"raw_usage":{"total_tokens":2596,"prompt_tokens":1044,"completion_tokens":1552,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":1471}},"tokens_in":660,"tokens_out":1552,"duration_ms":10028,"temperature":1.0,"reasoning_tokens":1471,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:59:34.009075+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained MSSM on IllustrisTNG-Dark, the dark-matter-only counterpart of the training simulation, and compare predicted stellar masses with the actual stellar masses of matched halos in the hydrodynamic run: a systematic bias in the $M_\\star$-$M_{\\mathrm{halo}}$ relation, or a scatter larger than the intrinsic scatter of that relation, would falsify the transferability assumption. A second check is to compare the two-point galaxy clustering of the MDPL2 catalogue with observed clustering at fixed stellar mass; a significant mismatch would show that the machine's galaxy-halo connection is not faithful enough for large-volume statistics.","supporting_citations":[],"review_version":1}