{"id":"92395147-199a-4ae9-8d90-02b9fdce86ac","arxiv_id":"2509.09626","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"In TNG and EAGLE simulations, central and high-mass satellite galaxies are best predicted to quench by black hole mass while low-mass satellites by halo mass, yielding opposite predictions for detecting environmental quenching with MOONRISE.","lead":"This paper uses two major galaxy simulations to predict how galaxies stop forming stars at cosmic noon, and what the upcoming MOONRISE survey should see. It finds that massive galaxies quench via their black holes, while small satellite galaxies quench via their environment, and the two simulations disagree on whether MOONRISE will detect the latter.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"EAGLE 'no environmental quenching in MOONRISE' claim is contradicted by the paper's own Table 2: 406 low-mass satellites (45 quenched) at z=1, yet text says 'essentially devoid'.","rationale":"The reader's weakest assumption concerned sensitivity of the mass threshold and buffer. My concern is sharper and partly internal: the paper's own Table 2 contradicts the textual claim that EAGLE MOONRISE-Like samples at z≥1 are essentially devoid of low-mass satellites. At z=1, Table 2 reports 406 low-mass satellites and 45 quenched low-mass satellites. This directly undermines the headline EAGLE prediction of undetectable environmental quenching in MOONRISE, independent of how the threshold is varied. The omission of RF results for this sample (Table C6) further obscures the inconsistency. This is load-bearing because the 'starkest discrepancy' is the paper's primary novel prediction for MOONRISE. The concern does not invalidate the RF findings for centrals/high-mass satellites or the qualitative environmental quenching of low-mass satellites; it specifically affects the EAGLE-vs-TNG forecast. Thus the verdict stays CONDITIONAL, matching the reader, but for a more concrete reason: the paper must reconcile Table 2 with its text and either report the z=1 EAGLE low-mass RF results or clearly justify their exclusion.","tokens_in":39464,"tokens_out":8403,"duration_ms":82916,"concrete_test":"Using the public EAGLE data, reconstruct the MOONRISE-Like z=1 sample with the paper's selection (M*>10^9.5, 2D+z observer space, 70% completeness) and the paper's low-mass classification (log(M*/M_sun) < 9.65, assuming symmetric 0.2 dex buffer). Count the number of quenched low-mass satellites. If the count is ~45, as in Table 2, scale this by the ratio of MOONRISE survey volume to the 100 Mpc EAGLE box and compare to the survey's expected detection threshold. Separately, rerun the Random Forest classification on this sample using the paper's reported hyperparameters; if it yields meaningful feature importances with environmental parameters dominating, the paper's omission is unjustified and the headline prediction fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central falsifiable prediction—that MOONRISE will see environmentally quenched low-mass satellites if TNG is right but essentially none if EAGLE is right—rests on the claim that the EAGLE MOONRISE-Like sample is 'essentially devoid of low-mass satellites' at z≥1 (Section 4.1). This is inconsistent with Table 2, which lists 406 low-mass satellites and 45 quenched low-mass satellites for EAGLE at z=1 in the MOONRISE-Like sample. The low-mass threshold for EAGLE z=1 is 9.85±0.2 (Table 3); applying the 0.2 dex buffer symmetrically gives a low-mass cut of 9.65, so MOONRISE-like galaxies with 9.5<log(M*/M_sun)<9.65 are included. Table 2 shows this population is not empty. The paper omits RF results for this sample (Table C6) and concludes environmental quenching is 'undetectable' (Section 4.3). If 45 quenched low-mass satellites exist in a 100 Mpc box, scaling to the MOONRISE survey volume suggests a detectable population. The threshold/buffer choice affects the number, but the internal contradiction is independent: the paper's own counts refute 'essentially devoid.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses the IllustrisTNG and EAGLE simulations at z=0, 1, and 2 to identify which physical properties best predict quiescence for centrals, high-mass satellites, and low-mass satellites, with an eye toward VLT-MOONRISE. After constructing a 'MOONRISE-Like' sample (M*>10^9.5 Msun, 2D+z observer space, 70% completeness), the authors apply Random Forest classification to rank M_BH, M_*, M_Halo, delta, and D_cen. They find that M_BH dominates for centrals and high-mass satellites, while M_Halo dominates for low-mass satellites. They also derive redshift-dependent mass thresholds by fitting Schechter functions to quenched satellite stellar mass functions, and use these to claim a key divergence: TNG predicts environmentally quenched low-mass satellites visible to MOONRISE, whereas EAGLE predicts essentially none. The paper closes with a series of testable predictions for MOONRISE, including a possible distinction between EAGLE and TNG in the abundance of low-mass quenched satellites at cosmic noon.","tokens_in":39818,"tokens_out":8546,"duration_ms":97669,"significance":"If the central predictions were robust, this would be a timely and useful paper for planning MOONRISE science. Its strengths are the construction of survey-like samples, transparent reporting of Random Forest hyperparameters (Tables C1-C8), the use of two state-of-the-art simulations, and multiple complementary analyses (quenched fractions, Delta f_Q, radial trends). However, one of the headline predictions is internally contradicted by the paper's own Table 2, and the class-split methodology is at risk of circularity because the satellite mass thresholds are explicitly designed to separate the two quenching mechanisms and are then used to demonstrate that the mechanisms differ. With the internal contradiction corrected and the threshold sensitivity tested, the paper would provide a valuable comparative test of TNG versus EAGLE for MOONRISE.","major_comments":[{"comment":"The text in Sec. 4.1 states that the EAGLE MOONRISE-Like samples at z>=1 are 'essentially devoid of low-mass satellites', and Fig. 6 and Sec. 4.3 repeat that no low-mass satellites are present at z=1,2. This is contradicted by Table 2: EAGLE z=1 MOONRISE-Like has N_LMS=406 and N_LMS,Q=45. With the z=1 bisecting mass of 9.85+/-0.2 (Table 3), the low-mass threshold after the 0.2 dex buffer is 9.65, so the sample contains galaxies with 9.5<log(M*/Msun)<9.65. Table C6 leaves the RF entries for EAGLE ML LM at z=1 and z=2 as ellipses. The claim that environmental quenching is 'undetectable' or that EAGLE predicts 'essentially none' is therefore not supported by the provided counts. Please re-run the z=1 low-mass satellite analysis, quantify expected counts in the MOONRISE volume, and revise the conclusion accordingly.","section":"Sec. 4.1, Table 2, Fig. 6, Table C6"},{"comment":"Satellites are split using thresholds explicitly 'designed to separate environmental and intrinsic quenching mechanisms' (Abstract). The threshold is the minimum gradient of the quenched satellite SMF plus a 0.2 dex buffer (Sec. 4.1). The RF conclusion that low-mass satellites quench environmentally is then obtained from these very classes, so part of the result is baked into the sample definition. This matters because the low-mass upturn may be a simulation artifact (Kukstas et al. 2022, cited in Sec. 4.1). Please test the stability of the conclusions by varying the buffer (e.g., +/-0.1, +/-0.3 dex, and no buffer) and re-applying the RF and Delta f_Q analyses, or by using a continuous approach that does not preselect classes. Without such a test, the central dichotomy is not independent of the class definition.","section":"Sec. 4.1 and Abstract"},{"comment":"The abstract and Sec. 4.4 claim 'strong evidence for the rejuvenation of star formation from z=2 to z=0 in EAGLE'. This is inferred from cross-sectional quenched fractions as a function of M_BH at three separate redshifts, not from tracking individual galaxies. Differential assembly histories could produce the same pattern without galaxies transitioning back to star formation. Please support the claim with merger-tree tracking of individual galaxies, or soften the wording to 'consistent with rejuvenation'.","section":"Sec. 4.4, Figs. 9-10"}],"minor_comments":[{"comment":"Typos: 'uni-model' should be 'unimodal' in Sec. 4.2; 'wains' should be 'wanes' in Sec. 4.3.1.","section":"Sec. 4.2, Sec. 4.3.1"},{"comment":"Eq. (12) appears to have a typo: p_i(n^2) should be p_i(n)^2. Also, the duplicated Pedregosa et al. (2011a,b) references should be merged.","section":"Eq. (12), References"},{"comment":"Random Forest relative importance is referred to as 'causal' in several places. Since RF importance is a predictive/associational measure, I recommend rephrasing to avoid overclaiming causality.","section":"Sec. 3.3, Sec. 5"},{"comment":"Please comment on the sensitivity of the sSFR quenching threshold to the -11.5 dex exclusion boundary and to the 3-sigma choice. This threshold controls all downstream quenched fractions and could affect the satellite counts in Table 2.","section":"Sec. 3.4, Eq. (14)"}],"recommendation":"major_revision","confidential_remarks":"The internal contradiction between the 'essentially devoid' text and Table 2's EAGLE z=1 counts (406 low-mass satellites, 45 quenched) is the most serious issue; it directly undermines the paper's most prominent prediction. The circularity of the mass-threshold split is a more subtle but equally important concern that will require additional sensitivity tests, not just language changes. If these are addressed, the paper would be a solid MNRAS contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper extends the group's Random Forest quenching program to z=1,2 and produces genuinely useful, survey-ready predictions for VLT-MOONRISE. The cross-simulation agreement—BH mass for centrals and high-mass satellites, halo mass for low-mass satellites—is a solid result, and the MOONRISE-Like sample construction is careful and transparent. It deserves a serious referee.\n\nBut the centerpiece discrepancy between TNG and EAGLE is softer than the text claims, and one part is internally inconsistent. The text says the EAGLE MOONRISE-Like sample is \"essentially devoid\" of low-mass satellites at z≥1, yet Table 2 lists 406 low-mass satellites and 45 quenched ones at z=1. That is not \"essentially devoid.\" Unsurprisingly, the RF results for that sample are omitted (Table C6), so the conclusion that environmental quenching would be \"undetectable\" is not actually supported by the paper's own tables. Fixing this matters: 45 quenched low-mass satellites in a 100 Mpc box, scaled to the MOONRISE volume, could be a detectable population.\n\nThe deeper worry is that the headline prediction depends on a hand-chosen buffer (0.2 dex) and on mass thresholds fitted to the quenched satellite mass function, which themselves assume the bimodality separates environmental from intrinsic quenching. That makes the conclusion partly circular, particularly because the authors recognize that low-mass satellite over-quenching in simulations (Kukstas et al. 2022) could create exactly the bimodality they are exploiting. They acknowledge this caveat but do not propagate it into the prediction. If the threshold or buffer shifts by a few tenths of a dex, the EAGLE \"no low-mass environmental quenching\" prediction could reverse.\n\nWhat is new here is the systematic forecast for MOONRISE: the two simulations make a testable, mutually exclusive statement about low-mass satellite quenching at cosmic noon, and the RF methodology is reported with enough detail to reproduce. The causal language—\"environmental quenching\"—still outruns the evidence in places, but the core finding that BH mass drives quenching in centrals and high-mass satellites is consistent with the group's prior work and with other studies.\n\nFor a referee: the paper is important enough to deserve peer review, but I would not accept it as is. The authors need to either correct the \"devoid\" claim, provide RF results or a quantitative detectability estimate for the EAGLE z=1 sample, and show how sensitive the headline prediction is to the buffer and threshold choices. A careful referee could turn this into a stronger paper.","headline":"Useful MOONRISE-era predictions from a familiar RF pipeline, but the boldest EAGLE claim contradicts the paper's own Table 2 and needs a hard look before it is taken at face value.","tokens_in":40349,"tokens_out":3672,"would_cite":true,"duration_ms":44283,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two leading cosmological simulations agree on the physics of galaxy quenching but disagree sharply on what the VLT-MOONRISE survey will see at cosmic noon.","keywords":["galaxy quenching","cosmic noon","VLT-MOONRISE","IllustrisTNG","EAGLE simulation","Random Forest classification","satellite galaxies","AGN feedback"],"falsifier":"If VLT-MOONRISE finds a significant population of quenched low-mass satellites (log M*/M_sun ~9.5-10) in groups at z~1-2, EAGLE's prediction fails; if it finds essentially none, TNG's prediction fails. Concretely, counting quenched low-mass satellites above the survey's mass completeness limit in the first MOONRISE data release would settle which simulated prediction is right.","tokens_in":39367,"feed_emoji":"🔭","tokens_out":3444,"duration_ms":38052,"temperature":0.7,"pith_summary":"This paper argues that in both IllustrisTNG and EAGLE, quenching follows two distinct paths: central galaxies and high-mass satellites stop forming stars because of supermassive black hole growth (AGN feedback), while low-mass satellites quench because of their group environment, best traced by halo mass. This split holds from z=0 to z=2. The authors build 'MOONRISE-Like' samples that mimic the upcoming VLT-MOONRISE survey's mass limit, completeness, and observer geometry, and produce a stark testable disagreement: TNG predicts MOONRISE will see environmentally quenched low-mass satellites at cosmic noon, EAGLE predicts essentially none. A secondary finding is that EAGLE shows galaxies re-igniting star formation at late times while TNG does not, which the paper ties to the presence or absence of a preventative AGN feedback mode. A sympathetic reader should care because MOONRISE can directly rule on which simulation's quenching physics matches nature.","feed_headline":"Two simulations disagree on quenched dwarfs at cosmic noon","feed_subtitle":"VLT-MOONRISE will see them if IllustrisTNG is right, almost none if EAGLE is right.","key_machinery":"The central machinery is a Random Forest classifier applied to simulated galaxies, with input features including black hole mass, halo mass, stellar mass, local over-density, and distance to central. The classifier identifies the most causally predictive parameter for quiescence while controlling for inter-correlations. This is combined with a bespoke mass-threshold method that locates the local minimum of the gradient of a Schechter fit to the quenched satellite stellar mass function, adds a 0.2 dex buffer, and splits satellites into low- and high-mass classes. The 'MOONRISE-Like' sample construction applies the survey's mass completeness limit, a 2D+z observer space, and 70% completeness t","core_discovery":"Using Random Forest classification on simulated galaxies at z=0, 1, and 2, the authors find that supermassive black hole mass is the best predictor of quiescence for central galaxies and high-mass satellites, while group halo mass is the best predictor for low-mass satellites, at every epoch. The split between low- and high-mass satellites is defined by a bespoke mass threshold derived from the bimodality of the quenched satellite stellar mass function, with a 0.2 dex buffer. The starkest discrepancy between the two simulations is in the mass threshold analysis: IllustrisTNG predicts environmentally quenched low-mass satellites that would be visible within VLT-MOONRISE's survey limits, where","pith_inferences":["If MOONRISE finds a significant abundance of quenched low-mass satellites, it would indirectly support TNG-style preventative AGN feedback and challenge EAGLE's thermal feedback; the opposite would support EAGLE. This is an editorial inference, not stated in the paper.","The paper implicitly suggests that density-only environmental studies at cosmic noon may misclassify high-mass satellite quenching as environmental when the true driver is black hole mass; one could test this by running the same Random Forest analysis with and without density as an input on MOONRISE data.","The mass-threshold method could be applied to other cosmological simulations or to the observed quenched mass function to check whether the bimodality is a robust feature or an artifact of simulation over-quenching; if it is artificial, the threshold-based predictions would need revision.","A testable extension is to measure the quenched fraction of low-mass satellites as a function of distance to the central galaxy and halo mass in MOONRISE, which would directly probe the ram-pressure, starvation, and interaction mechanisms the paper identifies as likely drivers."],"forward_implications":["If VLT-MOONRISE detects a population of quenched low-mass satellites at z~1-2, it would validate the TNG prediction and challenge EAGLE's quenching prescription.","If VLT-MOONRISE finds essentially no environmentally quenched low-mass satellites above its mass limit, the EAGLE prediction would be favored over TNG.","Measurements of quenched fraction versus black hole mass can distinguish TNG's preventative 'kinetic mode' AGN feedback from EAGLE's thermal feedback, especially through the presence or absence of rejuvenation.","The Random Forest results imply that halo mass, rather than local density alone, is the superior environmental predictor; accurate halo mass estimates in MOONRISE will be essential.","The study's mass-threshold method yields a specific, testable feature in the quenched stellar mass function: a dip separating the environmental and intrinsic quenching regimes that should appear if the bimodality is real."],"fun_headline_variants":["Simulations split on whether MOONRISE will find quenched dwarfs","Black hole mass predicts quiescence, but dwarfs hinge on halo mass","MOONRISE will test if quenched dwarfs exist: TNG vs EAGLE","TNG predicts quenched dwarfs for MOONRISE, EAGLE doesn't"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper assumes that the split between low- and high-mass satellites, derived from the fitted dip in the quenched satellite stellar mass function plus a 0.2 dex buffer, cleanly separates environmental from intrinsic quenching; if that dip is an artifact of simulations over-quenching low-mass satellites, the EAGLE/TNG prediction difference could reverse.","fun_headline_variants_meta":{"raw":{"variants":["Simulations split on whether MOONRISE will find quenched dwarfs","Black hole mass predicts quiescence, but dwarfs hinge on halo mass","MOONRISE will test if quenched dwarfs exist: TNG vs EAGLE","TNG predicts quenched dwarfs for MOONRISE, EAGLE doesn't"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000921,"raw_usage":{"total_tokens":3838,"prompt_tokens":843,"completion_tokens":2995,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":2905}},"tokens_in":587,"tokens_out":2995,"duration_ms":27145,"temperature":1.0,"reasoning_tokens":2905,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T18:44:29.025665+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If VLT-MOONRISE finds a significant population of quenched low-mass satellites (log M*/M_sun ~9.5-10) in groups at z~1-2, EAGLE's prediction fails; if it finds essentially none, TNG's prediction fails. Concretely, counting quenched low-mass satellites above the survey's mass completeness limit in the first MOONRISE data release would settle which simulated prediction is right.","supporting_citations":[],"review_version":1}