{"id":"5175ff48-cdeb-42a2-b9b6-ea500f3893ac","arxiv_id":"2607.27319","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"The halo gas fraction–halo mass relation in COLIBRE is non-monotonic and depends strongly on the subgrid AGN feedback model, with jet-hybrid feedback producing lower group gas fractions that better match eROSITA and kSZ constraints.","lead":"Using the COLIBRE cosmological simulations, this paper maps how much gas haloes of different masses retain and shows that supernova and AGN feedback produce a non-monotonic relation. It finds that a hybrid AGN model with jets lowers group gas fractions into better agreement with eROSITA and kSZ measurements, while the fiducial model matches older X-ray data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Resolution dependence of fgas makes the headline claim of 'better agreement with observations' selection-dependent; the claim needs to be conditioned on resolution and dataset.","rationale":"The paper's strongest empirical contribution is that COLIBRE's fgas-M200 relation is low and matches some recent constraints. But the relation is resolution-dependent by an amount comparable to the spread among datasets: Fig. 1 shows shifts of ~0.1-0.2 in fgas/fcosmic between m7 and m6, and Fig. 3 shows the two resolutions bracket the pre- vs post-eROSITA constraints. Since there is no converged prediction, 'COLIBRE predicts' depends on a resolution choice. This is particularly acute because the model was calibrated at each resolution separately with resolution-dependent parameters, so each resolution is effectively a different model. The authors acknowledge this in §3.1.2 and in the summary, but the abstract and reader's strongest claim treat COLIBRE as a single entity. This does not invalidate the paper: the hybrid-vs-fiducial comparison at fixed resolution is a valid sensitivity test, and the claim that similar galaxy populations can yield different fgas is supported. The correct response is to condition the headline on resolution and dataset, which is a revision, not a rejection. The attribution to ΔT_AGN is partially supported by the non-calibrated right-hand panel of Fig. C1, so I do not treat it as the primary issue; the recalibration caveat remains but is secondary. I therefore keep the reader's CONDITIONAL verdict.","tokens_in":44591,"tokens_out":5923,"duration_ms":52951,"concrete_test":"Construct a single quantitative comparison of all available COLIBRE resolutions (L400m7, L200m6, L025m5 and hybrid runs) against each observational dataset in Fig. 3 over the overlapping mass range 10^13 < M500/Msun < 10^14.5, using a common metric such as median offset or chi-square. If the best-matching resolution/model changes from one dataset to another (e.g., m6 for Kugel+23 but m7 for eROSITA/kSZ), the paper should state this explicitly and downgrade the claim. Additionally, run L025m5 to z=0 with the hybrid model or a higher-resolution zoom for groups; if the resolution trend continues, the low fgas at m7 is not converged.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In §3.1.2 the authors show fgas at fixed M200 increases with resolution (L400m7 < L200m6 < L025m5) and explicitly state the relation is not converged. Fig. 3 then shows that at m6 the fiducial model agrees with the pre-eROSITA X-ray compilation, while at m7 the hybrid model agrees with eROSITA/kSZ constraints. Because the observational constraints are mutually inconsistent (§3.2.3), 'COLIBRE produces lower gas fractions... and better agreement with observational constraints' is not a single claim: it selects one resolution and one dataset. The paper is transparent about both issues, but the abstract and summary do not attach these caveats. Thus the central empirical result is currently a family of results; a reader cannot determine whether COLIBRE has made a robust prediction or whether the match is an artifact of choosing m7 + hybrid for eROSITA/kSZ data and m6 + fiducial for pre-eROSITA data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents predictions for the z=0 halo gas mass fraction relation f_gas-M_200 in the COLIBRE cosmological simulations, using three resolution levels and two AGN feedback prescriptions (fiducial thermal and hybrid thermal/jet). It reports a non-monotonic relation, a multiphase census of halo gas, and significant resolution dependence. Comparisons with X-ray and kSZ constraints show that the m6 fiducial model matches pre-eROSITA X-ray data while the m7 hybrid model matches eROSITA and kSZ constraints. The paper argues that COLIBRE generally produces lower group/cluster gas fractions than EAGLE and most contemporary simulations, attributes this mainly to the BH-mass-dependent AGN heating temperature ΔT_AGN (and, in the hybrid model, to jets), and explores the redshift history of gas expulsion and re-accretion.","tokens_in":44884,"tokens_out":4699,"duration_ms":42054,"significance":"The strength of the paper is that COLIBRE was calibrated to galaxy stellar mass functions and size-mass relations, not to halo gas fractions, so the f_gas-M_200 relation is a genuine prediction. The use of multiple resolutions and two calibrated AGN prescriptions, plus the comparison with a broad set of observational constraints, makes this a valuable contribution. The demonstration that two models with similar galaxy populations can have substantially different halo gas content is an important result for feedback modelling. The paper is also transparent about the non-convergence with resolution and the inconsistency between observational datasets. If the claims are appropriately qualified, this will be a useful reference for interpreting eROSITA and kSZ constraints.","major_comments":[{"comment":"The headline claim that \"COLIBRE produces lower gas fractions for groups and clusters than EAGLE and other contemporary simulations, and better agreement with observational constraints\" is not a single, resolution-independent claim. §3.1.2 and Fig. 1 show that f_gas at fixed M_200 increases with resolution and is not converged. In Fig. 3, the m6 fiducial simulation agrees with the pre-eROSITA X-ray compilation but is high relative to eROSITA/kSZ constraints, while the m7 simulation falls below the pre-eROSITA data. In Fig. 4, the m7 hybrid model matches eROSITA/kSZ but is low relative to pre-eROSITA data. Since §3.2.3 states that the observational constraints are mutually inconsistent, \"better agreement\" is conditional on choosing one resolution and one dataset. The body is transparent about this, but the abstract and the summary bullets present the result as robust and unique. Please qu","section":"Abstract; §3.1.2, Fig. 3"},{"comment":"The causal attribution of COLIBRE's lower group/cluster gas fractions relative to EAGLE to the BH-mass-dependent ΔT_AGN (Eq. 7) is only partially supported. The model variants in Appendix C were each independently recalibrated, and the paper itself states that \"the removal of individual model features ... are therefore not strictly the only changes made.\" The non-calibrated AGN parameter variations in Fig. C1 do show that ΔT_AGN affects f_gas, and Appendix D shows that in L200m6 the bulk of AGN-driven expulsion occurs at ΔT_AGN higher than EAGLE's fixed 10^8.5 K. However, other BH modelling changes (repositioning, super-Eddington accretion, changed energy injection method) are acknowledged as potential contributors. The conclusion that the improvement \"can be attributed to\" the ΔT_AGN scaling is therefore stronger than the evidence. Please soften to \"is consistent with\" or, ideally, add","section":"§3.4.1, Appendix C/D"},{"comment":"The comparison between the fiducial and hybrid AGN models is not a controlled experiment: the hybrid simulations also use different calibrated values of the BH seed mass and feedback efficiencies, and §3.3 states that the seed-mass difference amplifies the m7 difference. Thus the statement \"the hybrid AGN feedback model produces lower gas fractions\" describes the effect of the whole recalibrated model variant, not of the jet prescription alone. This distinction matters for the abstract's claim that a \"hybrid AGN feedback model\" produces lower gas fractions. Either present these as model-level comparisons, or add a run in which only the jet/wind mechanism is changed while all other calibration parameters are held fixed.","section":"§3.3, Fig. 4"}],"minor_comments":[{"comment":"The stacked temperature-bin bars in Fig. 2 and Fig. A1 are informative, but the stacking order is not stated. A sentence in the caption explaining the order (e.g. coldest at bottom) would help.","section":"Fig. 2 / Fig. A1"},{"comment":"The grey shaded bands from the two baryonification models are easily confused with individual data points. Consider using labelled filled bands with distinct edge styles.","section":"Fig. 3"},{"comment":"The definition of f_gas includes all gas within r_500, while the X-ray/kSZ constraints largely trace hot gas. The temperature-cut comparison is given later, but a one-sentence reminder in the Fig. 3/4 captions would improve readability.","section":"§2.4"},{"comment":"Typo: \"observational contraints\" should be \"observational constraints.\"","section":"§3.2.1"},{"comment":"The model variant names (ThermalKinetic_varΔT_SN_varfE etc.) are hard to parse. A small table or a list with the deactivated features would make the comparison much easier to follow.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and technically solid; the main issues are framing and causal attribution rather than a fundamental error. The authors can address the concerns by qualifying the abstract/summary claims by resolution and dataset, and by softening or better supporting the attribution to ΔT_AGN. An additional controlled run with fixed ΔT_AGN would strengthen the causal claim considerably, but I would not require it if the wording is made appropriately cautious."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you care about halo gas fractions as a feedback diagnostic. The COLIBRE fgas–M200 relation is new, the hybrid-vs-fiducial AGN comparison is new, and the quantitative contrast with EAGLE is new and instructive. The paper is honest about most of its caveats, which is rarer than it should be.\n\nThe strongest part is the controlled parameter variation in Appendix C. The reader's worry about recalibrated variants is real but partially answered there: the non-calibrated ±0.5 dex shifts in ΔT_AGN show a direct, monotonic effect on fgas, and the fixed-temperature variants are consistent with the EAGLE comparison. So the causal story is plausible, not just asserted. The comparison with observations is also careful on temperature cuts and selection effects, and the paper explicitly refuses to privilege one observational dataset over another. That restraint is earned.\n\nWhere the paper is soft is the abstract and summary. The fgas relation is not converged in resolution: at fixed mass, fgas increases from m7 to m6, and the paper says so in §3.1.2. But then the headline claim—\"COLIBRE produces lower gas fractions and better agreement with observational constraints\"—silently picks m6+fiducial for pre-eROSITA and m7+hybrid for eROSITA/kSZ. The authors are transparent that at least one simulation matches any given dataset, but that is not the same as a robust prediction. The abstract needs to condition the claim on resolution and dataset, or the paper will be cited for a conclusion it does not actually establish.\n\nMinor point: the EAGLE difference is attributed mainly to the variable ΔT_AGN, but other BH modelling changes (repositioning, super-Eddington accretion, injection method) also changed. The paper acknowledges this, but the summary leans harder on ΔT_AGN than the evidence strictly allows.\n\nOverall: worth a serious referee. I would accept it with minor-to-moderate revision, mainly asking for the abstract and conclusions to match the paper's own honest caveats. The analysis is reproducible, the model variants are public, and the field needs this kind of clean comparison.\n\nI'd bring it to reading group, and I'd cite it—with the resolution caveat attached.","headline":"Solid, transparent simulation paper with genuinely new COLIBRE gas-fraction predictions, but the abstract oversells the resolution/dataset dependence of the 'better agreement' claim.","tokens_in":45580,"tokens_out":1318,"would_cite":true,"duration_ms":14206,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Halo gas fractions are a sensitive probe of feedback: COLIBRE and its predecessor produce similar galaxies but very different gas contents in groups and clusters, and the difference traces to a black-hole-mass-scaled AGN heating temperature","keywords":["halo gas fractions","AGN feedback","supernova feedback","cosmological simulations","galaxy groups and clusters","circumgalactic medium","kinetic Sunyaev-Zel'dovich effect","eROSITA"],"falsifier":"The cleanest check is a single-change simulation: take the fiducial model at fixed resolution and replace the BH-mass-scaled ΔT_AGN with a fixed value (the predecessor's 10^8.5 K) without re-calibrating anything; if group gas fractions do not rise back toward the predecessor's values, the scaling is not the cause. Observationally, the available group gas fractions are mutually inconsistent: a cross-calibration of eROSITA stacks, XMM-Newton profiles, and kSZ baryonification constraints at M500 ~ 10^13-10^14 M_sun that converged on one value would decide whether the low fractions the hybrid mode","tokens_in":44494,"feed_emoji":"🌌","tokens_out":21028,"duration_ms":147864,"temperature":0.7,"pith_summary":"This paper uses the COLIBRE suite of cosmological simulations to establish how feedback from supernovae and active galactic nuclei sets the gas content of dark-matter haloes, from dwarf galaxies to galaxy clusters. Its central result is that the halo gas fraction versus halo mass relation is a far more discriminating test of feedback physics than the galaxy population alone: COLIBRE and its predecessor EAGLE reproduce similar galaxies, yet COLIBRE keeps substantially less gas in group- and cluster-scale haloes, agreeing better with X-ray and kinetic Sunyaev-Zel'dovich measurements. The paper traces this improvement to a single design choice, scaling the AGN heating temperature linearly with black hole mass, so feedback events become more energetic precisely where they must overcome deep potential wells. A second, more realistic 'hybrid' AGN model that adds jets lowers group gas fractions further, matching the newest eROSITA stacks and kSZ constraints without degrading the agreement with galaxy observations. If this is right, measuring the gas content of haloes, not just the galaxies inside them, can break the degeneracy between feedback models that currently pass the same galaxy tests.","feed_headline":"Halo gas exposes what galaxy counts hide: AGN feedback strength","feed_subtitle":"A jet-driven AGN model leaves group haloes gas-poor, matching new eROSITA and kSZ data.","key_machinery":"The load-bearing object is the AGN heating temperature increment ΔT_AGN, which scales with black hole mass, so the energy per feedback event grows with it. This scaling lets AGN feedback expel gas from the deep potential wells of groups and clusters, where a fixed heating increment fails. The hybrid model applies the same idea through a jet velocity scaling with the square root of black hole mass, half the energy traveling in collimated jets that couple to gas efficiently. A second mechanism, a supernova heating temperature scaling with gas density, makes stellar feedback gentler and raises gas fractions in dwarf haloes. These scalings drive the non-monotonic f_gas-M relation and COLIBRE's l","core_discovery":"COLIBRE's claim: halo gas fractions are non-monotonic in halo mass, peaking near 10^11.5-12 M_sun; supernova feedback strips dwarf haloes, AGN feedback depletes groups and clusters. The fiducial thermal model, whose heating increment ΔT_AGN scales with black hole mass, matches Chandra/XMM-Newton gas fractions but runs high versus eROSITA stacks and kSZ-derived values; the hybrid thermal-plus-jet model runs lower and matches those newer data. The paper attributes the reduction mostly to the BH-mass scaling, which makes feedback events more energetic where potential wells are deepest; most expulsion occurs above the predecessor's fixed 10^8.5 K heating. Both variants were calibrated to observe","pith_inferences":["A controlled experiment the paper does not run: replace ΔT_AGN with a fixed value inside the fiducial model without re-calibrating anything, and re-measure group-scale gas fractions; the paper's re-calibrated variants leave the scaling entangled with other parameter changes, so this single-change test would isolate the cause.","The resolution trend means inferred feedback strength from kSZ and X-ray 'missing baryons' analyses is likely resolution-degenerate in the same direction: higher-resolution runs of one model bracket observations differently, so observational claims of strong feedback should be quoted against a specific resolution.","COLIBRE's cold-gas census — up to half of halo gas below 10^4.5 K at M200 near 10^11.4 M_sun, down to 10 K — is invisible to the X-ray and kSZ constraints this paper compares against; CGM absorption-line surveys of low-mass haloes could test that temperature breakdown directly.","The lower group gas fractions imply stronger baryon-feedback suppression of the matter power spectrum on group scales than the predecessor predicted; plugging COLIBRE's f_gas relation into baryonification fits to kSZ data is a direct way to test whether a COLIBRE-like baryon model also addresses the current S8 tension."],"forward_implications":["Simulations whose AGN heating does not grow with black hole mass will tend to over-predict the gas content of groups and clusters, and hence under-predict the feedback-driven suppression of the matter power spectrum on group scales; Appendix D shows COLIBRE expels most baryons with a heating temperature well above the predecessor's fixed value.","Galaxy-scale observables alone cannot fix the baryon content of haloes: two calibrated COLIBRE variants pass the same galaxy tests yet differ significantly in halo gas, a degeneracy the paper argues only halo-gas observations can break.","The hybrid jet model reaches the low group gas fractions suggested by eROSITA and kSZ data within a model that still matches galaxy populations, showing the stronger feedback those data appear to require is compatible with a successful galaxy formation model (with the paper's caveat that comparably strong models can fail like-for-like X-ray comparisons of cluster thermodynamics).","Gas fractions increase with resolution at fixed halo mass in COLIBRE, so the same physics gives different observed-level gas fractions at different resolutions; any simulation-observation comparison must be read at a specified resolution.","The expulsion-then-re-accretion sequence — clusters re-accrete gas and end up gas-rich while groups stay depleted — explains how strong-feedback models can lower group gas fractions without over-depleting clusters, because cluster progenitors were depleted less and replenish later."],"fun_headline_variants":["Hybrid AGN feedback trims halo gas to fit eROSITA and kSZ","AGN feedback mode shapes halo gas: hybrid fits new data","Halo gas fractions favor hybrid AGN feedback over thermal alone","COLIBRE: hybrid AGN feedback aligns with eROSITA and kSZ gas data","Gas-poor groups and clusters point to hybrid AGN feedback"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim that black-hole-mass-scaled AGN heating is what makes COLIBRE's groups and clusters gas-poor rests on model variants that removed that feature while also re-calibrating other parameters, so the scaling itself, rather than a correlated choice such as black hole seed mass or coupling efficiency, is the assumed cause—and separately, gas fractions rise with resolution, so the observed-level comparison is not unique.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid AGN feedback trims halo gas to fit eROSITA and kSZ","AGN feedback mode shapes halo gas: hybrid fits new data","Halo gas fractions favor hybrid AGN feedback over thermal alone","COLIBRE: hybrid AGN feedback aligns with eROSITA and kSZ gas data","Gas-poor groups and clusters point to hybrid AGN feedback"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000778,"raw_usage":{"total_tokens":3332,"prompt_tokens":858,"completion_tokens":2474,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":2377}},"tokens_in":602,"tokens_out":2474,"duration_ms":17300,"temperature":1.0,"reasoning_tokens":2377,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:33:31.130081+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The cleanest check is a single-change simulation: take the fiducial model at fixed resolution and replace the BH-mass-scaled ΔT_AGN with a fixed value (the predecessor's 10^8.5 K) without re-calibrating anything; if group gas fractions do not rise back toward the predecessor's values, the scaling is not the cause. Observationally, the available group gas fractions are mutually inconsistent: a cross-calibration of eROSITA stacks, XMM-Newton profiles, and kSZ baryonification constraints at M500 ~ 10^13-10^14 M_sun that converged on one value would decide whether the low fractions the hybrid mode","supporting_citations":[],"review_version":1}