{"id":"6b1942f9-e203-4878-ad42-7bb883af76cc","arxiv_id":"2512.06071","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Satellite quenched fractions rise toward lower stellar mass in all three simulations and observed samples, while radial trends differ by host environment.","lead":"This paper checks whether computer simulations of galaxy formation reproduce how often small satellite galaxies around Milky Way-sized hosts stop forming stars. It finds the stellar-mass trend is robust across three simulations and three observational surveys, while radial quenching patterns depend on host environment.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-dataset quenching definitions (sSFR vs Hα EW vs colors) may not trace the same physical state; the claimed robust mass trend could be a tracer artifact.","rationale":"The paper's headline claim – that the stellar-mass dependence of satellite quenching is a robust outcome of galaxy formation models – is exactly the claim that would be invalidated if the cross-dataset quenching definitions are not equivalent. The reader's weakest assumption identifies this same issue, and I agree it is the most load-bearing. The concern is concrete: the thresholds in Eqs. 5–6 and the ELVES color flag are not calibrated to one another, and the tracers' differing timescales (10 Myr vs 100 Myr+ ) are known to matter for bursty low-mass dwarfs. If the mismatch is mass- or radius-dependent, the 'closely matching' curves in Fig. 5 could be an artifact of using inconsistent rulers. The paper does some good work: it selects hosts and satellites in a unified aperture, uses a common mass range, and cites Samuel et al. (2022) for definition-robustness within FIRE-2. But that previous test does not cover the SAGA/ELVES thresholds or the instantaneous TNG50 SFR, and no forward-modeling of tracers is performed here. The proposed test – recomputing quenched fractions in TNG50 and FIRE-2 with mock Hα and 100 Myr-averaged sSFR – would settle the issue. If the trends survive, the central claim is genuinely robust; if not, the conclusion needs to be restricted to the specific tracer adopted. The secondary concern about the FIRE-2 paired-host inversion (six hosts, no significance test) affects a peripheral claim and is already hedged by the authors, so it does not change the verdict. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":21398,"tokens_out":4910,"duration_ms":53230,"concrete_test":"Recompute the FIRE-2 and TNG50 quenched fractions using three parallel tracer definitions: (a) the current log10(sSFR10Myr) < −11; (b) log10(sSFR100Myr) < −11; and (c) an Hα-equivalent-width proxy computed from star formation histories (e.g., Starburst99 spectral mapping) with the SAGA threshold EWHα < 2 Å. If the mass–quenched fraction curves shift by more than the Poisson uncertainties, or if the rank order across simulation suites changes, the 'robust mass trend' is tracer-dependent. The test isolates tracer effects from code-to-code differences within a single simulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the mass dependence of satellite quenching is robust across simulations and observations (§4.1, Fig. 5) rests on the assumption that the adopted quenching definitions are physically equivalent. In §2.9, FIRE simulations use log10(sSFR10Myr) < −11 (Eq. 5) and TNG50 uses an instantaneous SFR with the same threshold, while SAGA uses EWHα − σEW < 2 Å (Eq. 6) and ELVES relies on a color-based flag. These tracers act on different timescales: Hα probes ~10 Myr, UV/optical colors probe ~100 Myr–1 Gyr. For low-mass, bursty satellites, a 10 Myr sSFR can be zero while the 100 Myr sSFR is substantial (or vice versa), so the binary quenched/sterving classification can differ systematically with stellar mass and radius. The paper cites Samuel et al. (2022) to argue definition-insensitivity, but that test was performed only within FIRE-2, not across the full mass range and not against the specific SAGA/ELVES thresholds. No mock Hα or UV SFR measurements are constructed for TNG50 or FIREbox, so the quantitative agreement in the left panel of Fig. 5 has not been shown to reflect a common physical quenching state rather than a coincidence of tracer choices.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares the quenched fraction of satellites around Milky Way–mass hosts across three cosmological simulations (FIREbox, FIRE-2 zoom-ins, IllustrisTNG50) and three observational samples (SAGA, ELVES, MW+M31). A common host-halo mass interval, satellite stellar-mass range, and radial aperture are adopted, and a unified quenching classification is attempted. The central claim is that the stellar-mass dependence of satellite quenching is a robust, simulation-independent outcome, while radial trends are more sensitive to host environment and assembly history. The paper identifies the paired FIRE-2 MW–M31 analogs as the driver of an inverted radial quenched-fraction profile and interprets this as an environmental imprint.","tokens_in":21768,"tokens_out":5725,"duration_ms":64757,"significance":"If the central claim holds, the mass–quenched-fraction relation for low-mass satellites would be a useful benchmark for galaxy-formation models, and the paper would be a valuable cross-code/cross-survey comparison. The study has clear strengths: it uses publicly available observed catalogs, avoids fitting parameters to the target quenched fractions, and imposes explicit selection criteria across all datasets. The paper is also honest about many limitations, explicitly deferring mock-observation and orbit analyses to future work. However, the quantitative mass-trend claim depends on the assumed equivalence of very different quenching tracers, and the distinctive radial/environment claim rests on a small number of paired hosts without uncertainty quantification or a mass-matched control. These issues are load-bearing and need to be addressed before the paper can be accepted.","major_comments":[{"comment":"The central claim that all three simulations 'closely match' the observed mass dependence rests on the equivalence of four quenching diagnostics: FIRE uses log10(sSFR10Myr)<−11, TNG50 uses instantaneous sSFR with the same threshold, SAGA uses EWHα−σ<2 Å, and ELVES uses a catalog color-based flag. These tracers probe different timescales, and for bursty, low-mass satellites a 10 Myr sSFR can differ markedly from a 100 Myr or Hα-based measure. The cited Samuel et al. (2022) test was performed within FIRE-2 only and does not establish cross-code or cross-observable equivalence. Because no mock Hα or UV SFRs are constructed for TNG50/FIREbox, the quantitative agreement in Fig. 5 left could partly be a tracer artifact. Please add mock-tracer tests or substantially qualify the 'closely matching' claim.","section":"§2.9, Eqs. (5)–(6); Fig. 5 left"},{"comment":"The inverted radial quenched-fraction trend in the FIRE-2 paired hosts is based on only six systems, and Fig. 6 shows no uncertainty intervals or significance tests. More importantly, the paired systems lack satellites above M⋆≈10^8.5 M⊙ and show strong radial mass segregation: star-forming satellites are closer and more massive, while quenched satellites are farther and less massive. Given the steep mass dependence of quenching, the radial reversal could simply reflect mass segregation rather than a distinct environmental quenching mechanism. The conclusion that the paired hosts 'entirely' drive the discrepancy and imprint an environmental signal requires a mass-matched radial comparison or a regression that controls for stellar mass, with host-level confidence intervals.","section":"§3.3, Figs. 6 and 7"},{"comment":"The projected analysis treats each of the three orthogonal sightlines as an independent realization, as indicated by labels such as '17 × 3 orientations' and '13 × 3 orientations,' and the plotted uncertainties are Poisson errors on the stacked fraction. These three projections are highly correlated, so Poisson errors understate the true host-to-host variance. This is particularly important for the radial-trend comparisons where consistency is judged 'within uncertainties.' Please provide host-level bootstrap or jackknife confidence intervals, or at least show the host-to-host scatter band alongside the Poisson band.","section":"§2.8 and Fig. 5 caption; Fig. 6"}],"minor_comments":[{"comment":"FIREbox contributes only 25 satellites across 17 hosts, with a median of 1 satellite per host. The mass- and radius-binned quenched fractions therefore have very low statistical power. Please report per-bin satellite counts and explicitly state the resulting limitation on the FIREbox comparison.","section":"§2.7–2.8; Fig. 5"},{"comment":"The mapping between log10(sSFR)<−11 and the SAGA Hα EW criterion is only described as 'consistent with the approximate division used in Geha et al. (2024).' Since SAGA's primary criterion is Hα EW, please clarify the physical mapping and note whether the adopted sSFR threshold corresponds to a similar Hα-derived cut for the simulated galaxies.","section":"§2.9"},{"comment":"The shaded bands are described as 1σ host-to-host scatter for the simulations, but no equivalent uncertainty is shown for SAGA or ELVES. Adding bootstrap or jackknife bands for the observed samples would make the comparison more symmetric.","section":"Fig. 3"},{"comment":"Minor typographical issues: §2.2 'enatbles' should be 'enables'; §2.8 'atellites' should be 'satellites'; Fig. 6 axis label uses 'log(M⋆,sat/M⊙)' while Fig. 5 uses 'log10(M⋆,sat/M⊙)'. Please use consistent notation.","section":"Throughout"},{"comment":"The abstract describes 'nearly uniform radial selections,' but the methodology applies the same 300-kpc aperture and inner 10-kpc exclusion to all datasets. Consider removing 'nearly' or explaining the specific non-uniformities that remain (e.g., ELVES radial-coverage cuts).","section":"Abstract and §5"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and useful question, and the mass-trend result is likely to be robust, but the current version overinterprets the paired-host radial signal and relies on unvalidated cross-tracer equivalence. The requested tests—mock Hα/UV tracers and a mass-matched or regression-based radial analysis—are feasible and should be doable within the scope of a revision. I would not recommend rejection, but the distinctive claims need stronger support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline is that the mass-trend result survives a close read. The standardized host selection and common mass range across FIREbox, FIRE-2, TNG50, SAGA, ELVES and MW+M31 is genuinely new work, and the agreement across three independent simulation codes and three observational samples is striking. I'd be comfortable quoting this as evidence that satellite quenched fraction versus stellar mass is a robust benchmark for galaxy formation models.\n\nThe radial part is where I get cautious. The claim that the FIRE-2 zoom-ins suppress the inner quenched fraction because of their paired MW–M31 analogs is supported by splitting hosts into isolated and paired groups (Figs. 6–7). But that split is six hosts. Fig. 6 shows no error bars or confidence intervals, and the segregation in Fig. 7 is descriptive. In a small sample like this, one or two atypical systems can produce the inversion. I don't think it's a fabricated effect—the trend is stark and matches Samuel et al. 2022's earlier finding—but it needs bootstrap or jackknife uncertainties before it becomes a robust environmental imprint.\n\nThe stress-test worry about tracer equivalence (sSFR vs Hα EW vs colors) is real but probably not fatal for the mass trend. A 10 Myr sSFR and a 100 Myr UV color can disagree for bursty dwarfs, and the paper's citation of Samuel et al. 2022 only shows definition-insensitivity inside FIRE-2, not against the SAGA/ELVES thresholds. Still, the mass trend is so monotonic and consistent that I'd be surprised if tracer choice is driving it. For the radial comparison, tracer differences could matter more, especially if star-forming and quenched satellites separate by radius. A mock Hα/UV measurement for TNG50 and FIREbox would close this gap. I'd ask for that in revision.\n\nWorth noting: the paper is honest about its limits and says forward modeling and mock surveys are needed. That's the right tone. The absence of released code or catalogs is a minor annoyance.\n\nRecommend: send to a serious referee. The mass-trend robustness is a useful community benchmark, and the paired-host environmental imprint is worth pinning down with proper uncertainties. With that added, this is a solid contribution.","headline":"A credible cross-suite demonstration that the mass trend in satellite quenching is robust, with a small-N radial claim that still needs error bars.","tokens_in":22277,"tokens_out":3455,"would_cite":true,"duration_ms":35954,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that the stellar-mass dependence of satellite quenching is a robust outcome across three independent cosmological simulations and three observational datasets, while the radial dependence is environment-sensitive.","keywords":["satellite galaxies","quenched fraction","dwarf galaxies","galaxy formation simulations","FIREbox","IllustrisTNG50","SAGA survey","ELVES survey"],"falsifier":"Recompute all quenched fractions with a single common tracer, such as a UV-based specific star-formation rate averaged over 100 million years, applied consistently to every simulated and observed galaxy; then check whether the low-mass rise persists and whether the FIRE-2 paired-host radial inversion survives a bootstrap resampling test on the six paired systems. If the mass trend disappears or the inversion is not statistically significant, the paper's central claim loses support.","tokens_in":21307,"feed_emoji":"🔭","tokens_out":6213,"duration_ms":61528,"temperature":0.7,"pith_summary":"The paper asks whether modern cosmological simulations reproduce the observed pattern that low-mass satellite galaxies around Milky Way–mass hosts are more likely to have stopped forming stars. It compares three simulation suites with two observational surveys and the combined Milky Way + M31 satellite population, using uniform host-mass, satellite-mass, and radial cuts. The central finding is that the rise in quenched fraction toward low stellar mass is robust across all datasets, while the radial dependence varies, with the FIRE-2 paired MW–M31 analogs producing an inverted profile. The authors conclude that the mass trend is a stable prediction of current galaxy-formation models, and that radial trends carry information about host environment and assembly history. A sympathetic reader would take this as evidence that the mass-dependent quenching trend can serve as a benchmark, while radial trends require more careful modeling.","feed_headline":"Simulations agree on the mass trend in satellite quenching","feed_subtitle":"If true, the low-mass trend is a stable benchmark; radial profiles still depend on host environment.","key_machinery":"The load-bearing tool is the stacked quenched fraction—the total number of quenched satellites divided by the total number of satellites in a stellar-mass or projected-radius bin—computed under a common selection: hosts with halo mass 10^11.9–10^12.2 solar masses, satellites with stellar masses 10^7–10^10 solar masses within 300 kiloparsecs, and a shared quenching threshold. The argument is carried further by the paired-versus-isolated split of the FIRE-2 hosts, which isolates the environmental origin of the radial anomaly. This split shows that the inverted radial profile is not a property of the simulation code but of the specific paired-host environment.","core_discovery":"On the paper's own terms, the discovery is that the stellar-mass dependence of satellite quenching—lower-mass satellites being more likely to be quenched—is reproduced quantitatively by three independently built cosmological simulations (FIREbox, the FIRE-2 zoom-ins, and TNG50) when compared to SAGA, ELVES, and the Milky Way + M31 system. The radial dependence is not similarly universal: SAGA and ELVES show gently declining quenched fractions with projected radius, TNG50 matches that behavior, FIREbox is consistent with a nearly flat trend within uncertainties, and the FIRE-2 zoom-ins show suppressed inner quenching. Tracing this discrepancy, the paper finds it originates entirely from the s","pith_inferences":["If the mass trend is a genuine benchmark, then the physical drivers of low-mass satellite quenching—weak internal feedback, gas stripping, reionization—are captured similarly by all three models; next-generation simulations should therefore prioritize reproducing radial trends, where host-specific history matters.","A natural testable extension is to apply the same isolated-versus-paired environment classification to TNG50 and FIREbox hosts; if the inverted radial profile appears there too, the environmental imprint is general rather than unique to the FIRE-2 zoom-ins.","The paired-host inversion suggests an observational prediction: Milky Way–M31-like pair environments may show a suppressed central quenched fraction compared to isolated Milky Way analogs, and existing survey samples could be split by the presence of a comparably massive companion to test this.","The paper's own caution about halo finders and interlopers implies that forward-modeled mock surveys—not just raw catalogs—are the natural next test; without them, part of the radial differences could be attributed to observational selection rather than astrophysics."],"forward_implications":["The stellar-mass dependence of satellite quenching can serve as a robust benchmark for galaxy-formation models, since three independent simulation suites match SAGA, ELVES, and the Milky Way + M31 system.","Radial quenched-fraction profiles are environment-sensitive: the observed gentle decline with radius is reproduced by TNG50, FIREbox is consistent with a nearly flat trend, and the FIRE-2 zoom-ins show an inverted profile driven by paired MW–M31 analogs.","The FIRE-2 paired hosts lack satellites above roughly 10^8.5 solar masses and show strong radial segregation between star-forming (inner) and quenched (outer) satellites, which explains their suppressed central quenched fraction.","Host environment and assembly history can leave an imprint on satellite quenching statistics, so analyses that average over all Milky Way–mass centrals may dilute such environmental signals.","Comparisons between simulations and surveys require forward modeling of projection effects, interlopers, and surface-brightness limits before quantitative conclusions about radial trends can be drawn."],"fun_headline_variants":["Mass trend in satellite quenching crosses simulation divide","Satellite quenching: mass trend robust, radial trend not","Satellite quenching mass trend robust, radial varies","Mass quenching trend consistent across simulations and data","Satellite quenching mass dependence: simulations side with data"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The analysis assumes that the different ways of declaring a satellite quenched—star-formation rate averaged over the last 10 million years in FIRE, instantaneous star-formation rate in TNG50, hydrogen-alpha equivalent width in SAGA, and color-based flags in ELVES—trace the same underlying physical state; if they respond to different timescales in a mass- or radius-dependent way, the claimed consensus could be a definitional artifact.","fun_headline_variants_meta":{"raw":{"variants":["Mass trend in satellite quenching crosses simulation divide","Satellite quenching: mass trend robust, radial trend not","Satellite quenching mass trend robust, radial varies","Mass quenching trend consistent across simulations and data","Satellite quenching mass dependence: simulations side with data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001,"raw_usage":{"total_tokens":4118,"prompt_tokens":845,"completion_tokens":3273,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":3201}},"tokens_in":589,"tokens_out":3273,"duration_ms":21539,"temperature":1.0,"reasoning_tokens":3201,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T18:15:41.262069+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute all quenched fractions with a single common tracer, such as a UV-based specific star-formation rate averaged over 100 million years, applied consistently to every simulated and observed galaxy; then check whether the low-mass rise persists and whether the FIRE-2 paired-host radial inversion survives a bootstrap resampling test on the six paired systems. If the mass trend disappears or the inversion is not statistically significant, the paper's central claim loses support.","supporting_citations":[],"review_version":1}