{"id":"b9a6c5d3-aca2-4ff2-9c06-bd8f3744e49b","arxiv_id":"2505.01502","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using stacked eROSITA spectra of optically selected galaxy groups, the authors measure the mass-temperature relation down to about 10^13 solar masses and find a single power law with slope 1.65 ± 0.11, consistent with cluster scaling and with self-similarity.","lead":"Astronomers stacked X-ray spectra of thousands of optically selected galaxy groups from the eROSITA all-sky survey and measured the average gas temperature in seven mass bins from 10^13 to near 10^15 solar masses. They find galaxy groups follow the same mass-temperature power law as massive clusters, supporting temperature-based mass estimates across the full range.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Y07 luminosity-based masses converted to M500 via a Magneticum-derived ratio are the least secure link; a mass-dependent bias in the group-mass calibration or the M500/M200 conversion would directly shift the slope and intercept of Eq. (2).","rationale":"The paper attempts to measure the mean M-T relation from optically selected galaxy groups using spectral stacking, with validation on a hydrodynamic simulation. The central claim is Eq. (2), a single power law from ~1e13 to ~1e15 Msun. The methodology is generally sound: mock tests show the stacking recovers input temperatures, the background treatment is carefully checked, and the fit is repeated with both ODR and MCMC. I give credit for the Magneticum mock validation, the background comparison, and the explicit tests of metallicity and column density assumptions. However, the x-axis masses are not directly measured; they inherit the Y07 luminosity-mass calibration and a simulation-based M500/M200 conversion. The paper validates the conversion against Colossus, but that check assumes a density profile, and the mock group-finder test validates the algorithm within Magneticum rather than the absolute Y07 mass scale on SDSS data. If the luminosity-to-mass calibration is biased at group masses, the slope and intercept of Eq. (2) shift, potentially changing the conclusion about the group-regime slope. This is not an internal inconsistency; it is a calibration risk that can be quantified. The proposed re-fit with the alternative stellar-mass proxy is feasible with the existing catalog and would directly test the sensitivity of the result to the mass-proxy choice. I therefore do not change the reader's CONDITIONAL verdict.","tokens_in":23294,"tokens_out":7195,"duration_ms":76895,"concrete_test":"Re-derive Eq. (2) using the Y07 stellar-mass-based halo masses (also provided in the Y07 catalog) instead of the luminosity-based masses, keeping the same spectral stacking, temperature estimates, and M500/M200 conversion. If the resulting slope and intercept differ from the quoted 1-sigma uncertainties (0.11 in slope, 0.05 in intercept), the mass-proxy choice is a dominant systematic and the single-power-law conclusion is not yet robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is the mass scale of the x-axis. In Section 2.1, Y07 masses (defined at M180) are converted to M500 by assuming M180 ~ M200 and multiplying by a mass-dependent M500/M200 ratio of ~0.7 (0.046 dex scatter) taken from the Magneticum simulation (Fig. 1). This conversion is applied to the mass bins used in the fit of Eq. (2). The Y07 luminosity-based mass calibration itself is not independently anchored at group masses: the Colossus comparison in Section 2.1 validates the conversion only under an assumed density-profile/concentration model, and the mock validation of Marini et al. (2025b) tests the group-finder algorithm within Magneticum, not the absolute Y07 mass-to-light calibration on real SDSS data. Consequently, any mass-dependent bias in the Y07 luminosity-to-M180 calibration, or a mass-dependent offset in the simulated M500/M200 ratio, enters directly as a bias in both the slope and intercept of the M-T relation. This could mimic or hide a change of slope in the group regime, which is exactly the paper's central claim. The quoted 1-sigma statistical uncertainties do not include this systematic.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper measures the X-ray temperature of low-mass galaxy groups and clusters by stacking eRASS1 spectra of optically selected Yang et al. (2007) groups, using the Magneticum simulation to validate the stacking pipeline. From seven stacked mass bins spanning roughly 10^13 to 7x10^14 Msun, the authors fit a single power-law mass-temperature relation, Eq. (2): log10(M500/Msun) = (1.65 +/- 0.11) log10(TX/1 keV) + (13.38 +/- 0.05), with intrinsic scatter 0.13 +/- 0.03 dex. They report that this relation is consistent with the self-similar prediction and with the cluster-based relation of Lovisari et al. (2015), and they argue that a single power law holds across the entire sampled range, so that temperature remains a reliable mass proxy down to group scales. The paper also compares the measured relation with Magneticum, FLAMINGO, and EAGLE predictions and discusses the implications of AGN feedback for the intragroup medium.","tokens_in":23454,"tokens_out":5731,"duration_ms":60766,"significance":"If the central result holds, the paper significantly extends the observationally calibrated mass-temperature relation into the poorly populated group regime using an X-ray unbiased, optically selected sample. The methodological contribution is also valuable: stacking spectra of optically selected groups is a promising route to extracting thermodynamic information from systems that are individually undetected in shallow surveys. Strengths of the work include the use of a sample selected independently of X-ray properties, a detailed mock-based validation of the stacking and spectral-fitting pipeline, explicit background checks against blank fields, bootstrap-based uncertainty estimates, and the use of public eRASS1 and Y07 data. The central conclusion, however, rests on the absolute calibration of the Y07 mass scale and on several modeling assumptions inherited from the Magneticum simulation, and the quantitative validation is performed at eRASS:4 depth rather than eRASS1 depth; these issues need to be addressed before the claim of a universal single power law can be accepted without qualification.","major_comments":[{"comment":"The x-axis masses in Eq. (2) are derived from Y07 luminosity-based halo masses defined at M180, converted to M500 by assuming M180 ~ M200 and multiplying by a Magneticum-derived M500/M200 ratio of ~0.7 with 0.046 dex scatter (Fig. 1). The central claim that the M-T relation is a single power law without a slope change in the group regime is therefore only as secure as the Y07 mass-to-light calibration and the simulated conversion ratio. The Colossus comparison in §2.1 validates the conversion only under an assumed dark-matter density profile and concentration model; it does not test the absolute Y07 mass scale at M500 ~ 10^13 Msun. A mass-dependent bias of the order of plausible M/L calibration errors would directly bias both the slope and intercept of Eq. (2) and could mimic or hide a group-regime break. Please propagate this systematic into the fit (for example, as a covariance term) or demonstrate robustness by repeating the fit with an independent mass calibration, such as the Y07 stellar-mass proxy, updated Yang et al. catalogs, or stacked weak-lensing masses.","section":"§2.1, Fig. 1, Eq. (2)"},{"comment":"The gadem temperature model fixes the width of the Gaussian emission-measure distribution to T_sigma = 0.2 keV, stated to be 'based on the distribution of mass-weighted temperatures from the simulations.' Because the same Magneticum simulation is used both to set this prior and to validate the temperature recovery, the validation is partly circular for this parameter. If the true temperature dispersion in low-mass groups differs from the simulated one, for example because of a wider multiphase gas distribution or AGN-driven outflows, the fitted mean temperature will be biased in a temperature-dependent way, and that bias will propagate into the slope of Eq. (2). Please add a sensitivity test that varies T_sigma over a plausible range and reports the induced change in recovered temperatures, and ideally compare the assumed dispersion with measured temperature distributions of high-S/N eRASS1 or XMM-Newton groups.","section":"§5.2 and §6.1"},{"comment":"The quantitative validation of the stacking and spectral fitting, including the consistency of stacked and input spectra and the comparison of recovered versus input temperatures, is carried out on mock observations at eRASS:4 depth, while the observed analysis uses eRASS1 data. Appendix B shows eRASS1 stacked images but does not repeat the temperature-recovery tests at eRASS1 depth. Since eRASS1 has roughly four times fewer photons and a correspondingly larger background and unresolved-AGN contribution, the validation as presented does not directly cover the actual data conditions. Please repeat the mock temperature-recovery test at eRASS1 depth, or at least quantify the expected bias and increased uncertainty from the lower S/N in each mass bin.","section":"§3.3, Figs. 8-9, Appendix B"},{"comment":"The conclusion that the M-T relation 'does not change slope' in the group regime is supported only by fitting a single power law and by the statement that the data agree with Lovisari et al. (2015) within 1 sigma. No comparison is shown against a two-slope or broken power-law model, and the uncertainties in Table A.1 are large; for example, the bin at log10(M500/Msun)=14.09 has kT=2.47(+2.38/-0.81) keV at 3 sigma. A single power law will almost necessarily remain consistent when the error bars are this large, so 'no significant slope change' should be quantified with a model-comparison statistic or by placing an explicit upper limit on the slope difference between the low-mass and high-mass bins.","section":"§7.1, Eq. (2)"}],"minor_comments":[{"comment":"Equation (2) should be typeset as log10(M500/Msun) = (1.65 +/- 0.11) log10(TX/1 keV) + (13.38 +/- 0.05) to avoid the impression that the uncertainty multiplies the whole temperature term.","section":"§7.1"},{"comment":"The paper quotes 1-sigma uncertainties for Eq. (2) but reports 3-sigma uncertainties in Table A.1; please state this difference explicitly in the table caption and in the text.","section":"§7.1 and Table A.1"},{"comment":"The gadem model reference appears as a broken inline link ('gadem3https://...'); the citation and the surrounding formatting need to be fixed.","section":"§5.2"},{"comment":"The Fig. 8 caption refers to point-source masking based on the eRASS1 catalog, while the mock analysis described in §6.1 uses eRASS:4-equivalent mocks; please clarify which catalog was actually used.","section":"§6.1 and Fig. 8"},{"comment":"The appendix heading uses 'eRASS4' while the rest of the paper uses 'eRASS:4'; please unify the notation.","section":"Appendix B"},{"comment":"The abstract states that the relation spans up to about 10^15 Msun, but the highest bin in Table A.1 is log10(M500/Msun) ~ 14.85, corresponding to roughly 7x10^14 Msun; please reconcile the wording with the actual range of the fitted bins.","section":"Abstract and Table A.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of A&A and the stacking methodology is a potentially important step for eROSITA-era group studies. The main issue for the referee is not the statistical quality of the temperature measurements but the absolute mass scale of the x-axis and the mismatch between the eRASS:4 validation and the eRASS1 data; a convincing treatment of these two systematics would make the central claim much stronger."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper delivers what it claims — a spectral-stacked M–T relation from eRASS1 for optically selected groups down to ~1e13 Msun, with a slope consistent with cluster scaling. The main caveat is that the mass axis rests on the Y07 luminosity calibration and a Magneticum-based M500/M200 conversion, so the intercept and low-mass slope are only as good as that external mass scale.\n\nWhat's genuinely new: previous stacking (Bulbul+14) stopped at cluster masses; this extends the approach to the group regime using ~18,800 Y07 groups. The method is careful: point-source masking, two background treatments, bootstrap errors, mock validation with Magneticum/PHOX/SIXTE, and the blank-field background check is solid. The best-fit slope 1.65±0.11 and intrinsic scatter 0.13 dex are consistent with Lovisari+15 and with self-similar within ~1.7σ, so the result is a confirmation rather than a revision, but it's the first time the relation is anchored by an optically selected, X-ray unbiased sample at these masses.\n\nSoft spots, in rough order of seriousness:\n\n1. The mass calibration is the load-bearing assumption. Y07 masses are luminosity-based; converting M180 to M500 assumes M180≈M200 and uses a Magneticum-derived ratio ~0.7 with 0.046 dex scatter. The Colossus comparison checks the conversion under an NFW assumption, but it does not validate the absolute Y07 mass-to-light calibration at group masses. If Y07 masses are biased in a mass-dependent way at M~1e13–1e14, both slope and intercept shift. The authors acknowledge this only indirectly. This is a real limitation, but it is not a flaw unique to this paper — it's the standard state of affairs for optically selected group masses.\n\n2. The mock validation is done at eRASS:4 depth, not eRASS1. Appendix B shows eRASS1 stacks are detectable, but they don't show the full temperature-recovery test at eRASS1 depth. The pipeline should be less precise, not biased, but a quantified check would be better. Minor.\n\n3. Excluding all groups with point sources inside R500 removes ~30% of the sample. The simulation test in Fig. 13 suggests the M–T relation is unaffected by X-ray AGN presence, which is good, but it's still a selection that could interact with mass. Minor to moderate.\n\n4. The conclusion that temperature can serve as a mass proxy across the entire mass range is a bit stronger than what was measured — they measured mean temperatures in stacked bins, not per-object scatter. A more careful statement would say 'average temperature' or 'the mean relation'. This is a wording issue, not a fundamental flaw.\n\nOverall, I think the central result holds up: a single power law from 1e13 to 1e15, with no evidence of a slope break. The systematic uncertainties are dominated by the mass calibration, and the quoted errors probably understate the total error budget. But that's typical for this field.\n\nWho this is for: anyone using X-ray temperature as a mass proxy for cosmological analyses, or working on group scaling relations. It's also a useful methodological template for stacking faint X-ray emission.\n\nRecommendation: send to a serious referee. The paper deserves full review; the mass-calibration issue should be discussed but is not fatal.","headline":"Solid, careful stacking measurement that extends the M–T relation to optically selected groups; main caveat is the external mass calibration, not the X-ray analysis.","tokens_in":24169,"tokens_out":2693,"would_cite":true,"duration_ms":26050,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The mass–temperature relation is one power law from galaxy groups to massive clusters, with slope $1.65$, making temperature a workable mass proxy across two decades in halo mass.","keywords":["galaxy groups","galaxy clusters","mass-temperature relation","X-ray spectral stacking","eROSITA","intracluster medium","mass proxy","self-similar scaling"],"falsifier":"Measure independent masses for the same stacked groups, for example with weak gravitational lensing or thermal Sunyaev–Zel'dovich observations, and refit the mass–temperature relation; a mass-dependent offset between lensing masses and the luminosity-calibrated $M_{500}$ values would break the single-power-law claim. Alternatively, the deeper eRASS:4 survey should reproduce the same slope with smaller scatter rather than a steeper low-mass end.","tokens_in":22963,"feed_emoji":"🔭","tokens_out":8581,"duration_ms":73710,"temperature":0.7,"pith_summary":"This paper sets out to prove that the relation between X-ray gas temperature and halo mass is one continuous power law from small galaxy groups of about $10^{13}\\,M_\\odot$ up to massive clusters near $10^{15}\\,M_\\odot$, with no change of slope in the poorly explored group regime. The evidence is built by stacking thousands of very faint spectra from the first eROSITA all-sky survey for optically selected galaxy groups, an approach that extracts an average temperature where individual group detections are impossible. The paper validates the stacking on mock X-ray observations of a hydrodynamical simulation and finds that the same procedure recovers the input temperatures. If the central claim is right, gas temperature is a reliable halo-mass proxy across two decades in mass, which would let temperature-derived masses be used in cluster mass-function and cosmological analyses.","feed_headline":"One power law links group and cluster masses to gas temperature","feed_subtitle":"Stacking thousands of faint eROSITA spectra shows temperature works as a mass proxy across two decades in halo mass.","key_machinery":"The carrying mechanism is spectral stacking of optically selected groups: groups are sorted into halo-mass bins, their eRASS1 event lists are masked of point sources, spectra are extracted within $R_{500}$, shifted to a common rest frame, and co-added; each stacked spectrum is then fit with a multi-temperature plasma model (gadem) to recover a mean gas temperature per bin. The optical luminosity-based halo masses are converted to $M_{500}$ by assuming $M_{180}\\sim M_{200}$ and a mass-dependent $M_{500}/M_{200}\\simeq0.7$ with $0.046$ dex scatter taken from the hydrodynamical simulation. The same pipeline is applied to mock eROSITA observations built from that simulation and an X-ray telescope simulator; agreement between input and recovered temperatures is what turns the stacked temperature into a claimed unbiased measurement.","core_discovery":"The paper's central result is the best-fit relation $\\log_{10}(M_{500}/M_\\odot) = (1.65\\pm0.11)\\,\\log_{10}(T_X/1\\,\\mathrm{keV}) + (13.38\\pm0.05)$, with intrinsic scatter $0.13\\pm0.03$ dex. The claim is that this single power law, statistically consistent with the previously established cluster relation and within about $1.7\\sigma$ of the self-similar prediction $M\\propto T^{1.5}$, describes galaxy groups and clusters alike. The low-temperature end is anchored by seven stacked mass bins in which average temperatures of $0.7$–$1$ keV are measured from co-added spectra, while the mock tests show that contamination from unresolved AGNs and spurious optical detections does not systematically bias these temperatures. On this basis the authors conclude that AGN feedback and cooling redistribute baryons or alter the group core but do not change the overall temperature of the hot gas, so the temperature can serve as a mass proxy over the entire sampled range.","pith_inferences":["Editorial inference: if the single power law survives independent mass calibration, the apparent group/cluster differences reported by some X-ray-selected samples are probably selection effects, with X-ray selection favoring low-entropy systems, rather than a real break in the relation.","Editorial inference: re-running the identical stacking with weak-lensing masses for the same mass bins would directly test the assumed $M_{500}$ conversion and the luminosity-based mass calibration, a test the current data cannot provide.","Editorial inference: the mock result that current AGN activity does not change the relation suggests the gas response time exceeds the AGN duty cycle; deeper data with radio-mode feedback indicators could test this by comparing active and inactive systems at fixed mass."],"forward_implications":["Temperature-based mass estimates can be extended to optically selected groups roughly two decades below the mass range of current X-ray-selected cluster samples.","Cluster and group mass-function studies can adopt a single calibrated $M$–$T$ relation over this larger range, improving constraints on cosmological parameters such as $S_8$.","With the deeper eRASS:4 data, per-bin statistics will grow enough to measure temperature profiles rather than only average temperatures, giving a sharper view of AGN feedback.","The same stack-and-measure approach can be applied to any X-ray-faint population selected by other means, such as low-luminosity AGNs or high-redshift cluster candidates."],"supporting_citations":[{"why":"Provides the optically selected SDSS group catalog whose luminosity-based halo masses, centers, and redshifts define the sample.","marker":"Yang et al. (2007)"},{"why":"Tests the group finder on a mock galaxy catalog and establishes completeness, contamination, and mass-proxy reliability.","marker":"Marini et al. (2025b)"},{"why":"Builds the mock eROSITA observations from the hydrodynamical lightcone and quantifies X-ray catalog completeness.","marker":"Marini et al. (2024)"},{"why":"Shows that stacking optically selected groups in mock eROSITA data recovers the input X-ray surface brightness profile, motivating the spectral stacking.","marker":"Popesso et al. (2024d)"},{"why":"Supplies the PHOX photon simulation of X-ray components and validates the AGN luminosity function used for mock contamination estimates.","marker":"Biffi et al. (2018)"},{"why":"Provides the instrument simulator used to turn simulated photon lists into mock eROSITA event files.","marker":"Dauser et al. (2019)"},{"why":"Serves as the X-ray-selected group/cluster mass–temperature baseline that the new relation is compared against.","marker":"Lovisari et al. (2015)"},{"why":"Supplies the eRASS1 cluster catalog used for comparison and the survey data description.","marker":"Bulbul et al. (2024)"},{"why":"Provides the eROSITA-to-XMM temperature cross-calibration correction applied before comparing with literature relations.","marker":"Migkas et al. (2024)"}],"fun_headline_variants":["One power law ties halo mass to gas temperature, groups to clusters","eROSITA stacking: temperature reliably proxies mass for all halos","Stacked spectra show one mass-temperature law from groups to clusters","Temperature is a reliable mass proxy across groups and clusters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The relation rests on the assumption that the optical luminosity-calibrated halo masses can be converted to $M_{500}$ with a fixed, mass-dependent ratio of about $0.7$ taken from a hydrodynamical simulation; if that conversion is biased as a function of mass, the fitted slope and intercept shift.","fun_headline_variants_meta":{"raw":{"variants":["One power law ties halo mass to gas temperature, groups to clusters","eROSITA stacking: temperature reliably proxies mass for all halos","Stacked spectra show one mass-temperature law from groups to clusters","Temperature is a reliable mass proxy across groups and clusters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001057,"raw_usage":{"total_tokens":4519,"prompt_tokens":1112,"completion_tokens":3407,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":728,"completion_tokens_details":{"reasoning_tokens":3336}},"tokens_in":728,"tokens_out":3407,"duration_ms":22226,"temperature":1.0,"reasoning_tokens":3336,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:18:08.403291+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure independent masses for the same stacked groups, for example with weak gravitational lensing or thermal Sunyaev–Zel'dovich observations, and refit the mass–temperature relation; a mass-dependent offset between lensing masses and the luminosity-calibrated $M_{500}$ values would break the single-power-law claim. Alternatively, the deeper eRASS:4 survey should reproduce the same slope with smaller scatter rather than a steeper low-mass end.","supporting_citations":[{"cited_title":"2024, , 688, A107","cited_arxiv_id":null,"evidence_quote":"Provides the eROSITA-to-XMM temperature cross-calibration correction applied before comparing with literature relations."}],"review_version":1}