{"id":"2ab9603f-1486-4f71-9e35-65210970f128","arxiv_id":"2501.07635","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Four weak-lensing-selected massive clusters show an excess of rare, gas-poor ICM features compared to clusters selected by their hot gas, indicating a fundamental selection bias in ICM-selected cluster samples.","lead":"Astronomers measured gas, X-ray, and pressure properties of four massive galaxy clusters chosen by their gravity (weak lensing) rather than by their hot gas. Compared with clusters selected by their gas, these four show far more unusual, faint, and gas-poor features than expected, suggesting standard cluster samples miss an important population.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Expected-outlier baseline in Fig. 13 uses the small, selection-biased scatter of ICM-selected samples that Sec. 1 itself argues is underestimated; the headline 10x excess may be an artefact of comparing to a biased null rather than a fundamental bias.","rationale":"Good-faith reading: the paper does a careful multi-wavelength follow-up and the data on the four clusters appear solid. The Eddington correction, the caustic check for id5, the KiDS cross-check, and the explicit treatment of the two merging systems are evidence of care. The scientific question - whether ICM-selected samples miss a population of gas-poor massive clusters - is well posed, and this pilot is a reasonable first look. However, the headline claim is a statistical statement about the frequency of outliers, and the null is built from the very samples whose selection bias is the subject. The reader's weakest assumption concerns shear-peak selection; I think the more load-bearing problem is the baseline against which 'rare' is defined. The two are related: if shear-peak selection also correlates with merger state, the outliers could be doubly non-representative. But the biased null is sufficient to undermine the quantitative '0.2 vs 2' claim even before asking whether the four objects represent all mass-selected clusters. I therefore propose the Monte Carlo test above. If the corrected null still makes 12 outliers rare, the conditional verdict should stand and the claim is supported. If not, the paper should either soften the abstract or add the corrected expected counts. This is why I keep the verdict UNCHANGED relative to the reader's CONDITIONAL: the concern does not change the need for a larger sample and corrected statistics, but it sharpens what has to be demonstrated.","tokens_in":31551,"tokens_out":10021,"duration_ms":111411,"concrete_test":"Run a Monte Carlo null that uses selection-corrected scatter for all seven properties: take XUCS for L_X, the Nagarajan et al. (2019) selection-corrected Y-M relation, and the 0.4 dex scatter reported by Ghirardini et al. (2024)/eROSITA for density and pressure proxies, with the mean relations from those unbiased or selection-corrected analyses. Draw 10^4 realizations of four clusters at the Table 1 masses, count >2σ outliers per property using this null, and compare with Fig. 13. If the observed 12 outliers (or 3/4 for ne(0)) fall within the simulated distribution, the central claim fails; if they remain a <1% tail, the claim survives. As a cross-check, repeat with the full Miyazaki+18 parent sample of 11 clusters to test whether the four selected objects are representative.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (abstract; Sec. 3.1, Fig. 13) is that four shear-selected clusters show ~12 '>2σ' outliers across seven properties when only ~1.4 are expected if the features were independent (or 0.2 per property). The null distribution for this calculation is taken from the ICM-selected comparison samples themselves: Pratt et al. (2022) for L_X, Ghirardini et al. (2020) for density/pressure, Nagarajan et al. (2019) for Y. But Sec. 1 argues, and Secs. 2.6/3.2 repeat, that such ICM-selected samples are biased in exactly the quantity being tested and have severely underestimated intrinsic scatter: values of 0.02-0.17 dex in the literature versus ~0.4-0.5 dex in samples selected without ICM (XUCS, eROSITA). A 2σ threshold defined by an underestimated scatter is not a valid null for the mass-selected population. The paper itself shows one concrete example: id5 is 0.45 dex low in L_X, quoted as 6σ_intr below the SZ-selected relation, but only 0.4σ_intr from the XUCS relation (Sec. 2.6, Fig. 9) - i.e., this counted outlier disappears with an unbiased null. It also notes that the low-Y outliers of id5/id34 sit exactly where the selection-corrected Nagarajan relation postulates clusters. If the same correction were applied to the other six properties, the expected number of 'rare' features would rise substantially, and the observed count may no longer be anomalous. The non-independence of the seven properties (admitted in Sec. 3.1) makes the claimed 10x excess even less probative. Thus the headline inference overstates what the data demonstrate; what remains is a suggestive pilot result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a pilot study of four massive galaxy clusters selected from the HSC weak-lensing survey, i.e. independently of their baryon content. For each cluster the authors derive shear-based masses (with Eddington-bias correction), core-excised X-ray luminosities, electron density and pressure profiles, Compton-Y parameters, and optical richness. They compare seven of these properties with those of ICM-selected comparison samples and report an excess of rare (>2 sigma) features: they expect on average 0.2 such features per property but observe about two per property. The two most extreme objects, id5 and id34, are dynamically complex, and id5 is scrutinized in detail with DESI spectroscopy, caustic masses, KiDS shear, and contamination tests. The paper concludes that ICM-selected samples are biased in our knowledge of cluster thermodynamic properties.","tokens_in":1758,"tokens_out":1751,"duration_ms":70442,"significance":"If the central statistical claim were fully supported, the paper would be important: it would demonstrate that X-ray and SZ selected samples miss a real population of massive, gas-poor clusters and that scaling relations built on those samples are correspondingly biased. The strengths of the paper are substantial: the Eddington-bias treatment is explicit and careful; the X-ray background and PSF handling is detailed; the mass of id5 is cross-checked with caustics and independent KiDS ellipticities; and hydrostatic equilibrium is avoided for the two clusters where it is risky. The data themselves, including the new Swift/Chandra and NIKA2/ACT measurements, are valuable. However, the headline '10x excess' calculation in Sec. 3.1 depends on a null model whose scatter is taken from the very ICM-selected samples that Sec. 1 argues are biased, and several of the counted outliers are shown elsewhere in the paper to disappear once a selection-unbiased comparison is used. The claim therefore needs substantial reworking before it can support the stated conclusion.","major_comments":[{"comment":"The central statistical claim is not a valid null test as presented. The expected number of >2 sigma outliers is computed using the intrinsic scatter of the ICM-selected comparison samples (Pratt et al. 2022; Ghirardini et al. 2020; Nagarajan et al. 2019), but Sec. 1 argues that those samples have severely underestimated scatter (0.02-0.17 dex versus 0.4-0.5 dex in samples selected without the ICM). A 2 sigma threshold defined by an underestimated scatter is not a null distribution for a mass-selected population. The paper itself provides the concrete counterexample: Sec. 2.6 and Fig. 9 state that id5 is 6 sigma below the SZ-selected L_X relation but only 0.4 sigma from the X-ray-unbiased XUCS relation, so the counted id5 luminosity outlier disappears with an unbiased null. Similarly, Sec. 2.7.2 and Fig. 12 show that the low-Y outliers of id5 and id34 fall exactly where the selection-corrected Nagarajan et al. (2019) relation postulates clusters. The expected-outlier calculation must either use scatter estimates from unbiased samples (e.g. XUCS or the recent eROSITA-based scatter of Ghirardini et al. 2024) or be reframed as 'rare relative to a biased ICM-selected sample', which would not support the abstract's conclusion of a fundamental bias.","section":"Sec. 3.1, Fig. 13"},{"comment":"The four clusters are not a random draw from a mass-selected population, so the binomial or Poisson statistics in Sec. 3.1 do not directly apply. The selection is described as 'the two most massive clusters' and 'randomly two, out of three clusters with largest signal-to-noise visible in the spring nights', and after Eddington correction the selected masses shift substantially. Because shear-peak S/N correlates with concentration, dynamical state, and line-of-sight projection, the sample may be biased in exactly the properties being tested. The two most anomalous objects, id5 and id34, are both dynamically complex with infalling groups (Sec. 3.1), and id5 is the highest-S/N massive cluster in the parent sample; the observed outlier excess could then be a property of the selection rather than of the general cluster population. The authors should either demonstrate that the sample is representative by resampling from the parent shear-selected catalog, or explicitly present the paper as a pilot study whose conclusion is limited to the four objects and not yet a population statement.","section":"Sec. 2.1, Sec. 3.1"},{"comment":"The analysis treats the seven non-independent features as if they provide seven independent tests. The text admits that the features are non-independent, and then tries to mitigate this by 'focusing on just one ICM-based feature', but the abstract and the per-feature average still quote the aggregated number '12 outliers among 7 non-independent features' and 'two rare features in each one of the seven properties'. Since L_X, Y, pressure, and electron density are physically correlated, and since id5 and id34 contribute the majority of the outliers, the effective number of independent trials is much smaller than 28 (4 clusters x 7 properties). The p-values quoted in Fig. 13 should be replaced by a joint test that accounts for covariance among properties, or by a conservative count based on the number of independent clusters showing any rare feature. Without this, the reported significance is overstated.","section":"Sec. 3.1, Fig. 13"}],"minor_comments":[{"comment":"The text and figure should state explicitly whether '>2 sigma' means a one-sided 2.3% threshold or a two-sided 5% threshold; the expected count of 0.2 per four objects per property corresponds to a two-sided 95% interval, whereas many readers will interpret '2 sigma' as one-sided 2.3%, which changes the expected counts by a factor of about two.","section":"Sec. 3.1, Fig. 13"},{"comment":"There are several typos and spacing issues: 'Miyakasi et al. (2018)' should be 'Miyazaki et al. (2018)', 'Andreon & Weaver 2015abouthowtodealwiththiscase' in Sec. 1 is missing a space, and 'core-excised X-ray ray luminosity' in Sec. 2.6 has a duplicated word.","section":"Sec. 2.1, Sec. 2.6"},{"comment":"The text says that the quoted Y_sph,500 values in Table 1 are as measured and that a 0.06 dex correction is applied only for comparison, but the abstract and Fig. 12 do not make this distinction clear; a short statement in the table caption would prevent misinterpretation.","section":"Table 1, Sec. 2.7.2"}],"recommendation":"major_revision","confidential_remarks":"The paper contains genuinely useful measurements and a careful treatment of several systematics, but the headline claim currently rests on a null model built from the very samples whose bias the paper aims to demonstrate. I think the central claim is defensible only after the outlier statistics are recomputed with an unbiased scatter, or after the conclusion is explicitly weakened to a statement about rarity relative to ICM-selected samples. The manuscript fits the journal's scope and the observations are worth publishing, but not in the present form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is the data: four weak-lensing-selected massive clusters with X-ray, SZ, and optical follow-up, and the first radial ICM profiles for such a sample. The individual measurements are done carefully. Eddington bias is treated explicitly, id5 is cross-checked with caustics and KiDS shear, and the joint/disjoint fits are sensible. The authors also openly flag where hydrostatic equilibrium is risky. That is honest, reproducible work and it will be useful to the cluster community.\n\nThe soft spot is the central statistical claim. The abstract says the sample shows a 10x excess of rare (>2 sigma) features compared to ICM-selected samples, and calls this evidence for a fundamental bias. But the null for that comparison is built from the scatter of the very ICM-selected samples the paper argues are biased and have underestimated intrinsic scatter. A 2-sigma threshold defined by 0.02-0.17 dex scatter is not a valid null for the mass-selected population. The paper itself shows the problem: id5 is 6 sigma below the SZ-selected L_X relation but only 0.4 sigma from the XUCS relation, and the low-Y outliers of id5 and id34 sit exactly where the selection-corrected Nagarajan relation postulates clusters. So a large fraction of the counted outliers disappears when the null is corrected.\n\nThe seven properties are also non-independent (admitted in Sec. 3.1), and the four clusters are a post hoc selection of the two most massive and two highest-S/N objects, not a predefined statistical sample. The two most anomalous clusters are both dynamically complex with infalling groups. The claim as stated in the abstract overstates what the data show. What remains is a suggestive pilot result: three or four objects with genuinely low gas content relative to the usual ICM-selected relations, worth investigating further.\n\nThis paper deserves peer review. The measurements are careful, the analysis is transparent, and the question is important for eROSITA completeness and scaling relations. But the referee should push the authors to either build a null using unbiased scatter and full covariance, or soften the conclusions to a pilot-level statement. I would not let the headline claim stand as written.","headline":"Careful multi-wavelength follow-up of four shear-selected clusters, but the claimed 10x excess of rare ICM features rests on a biased null and needs to be softened or re-derived.","tokens_in":32557,"tokens_out":1275,"would_cite":true,"duration_ms":14886,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Four clusters selected by mass, not gas, show far more rare gas features than gas-selected samples predict, evidence that our picture of cluster gas is systematically biased.","keywords":["galaxy clusters","weak gravitational lensing","intracluster medium","X-ray luminosity","Sunyaev-Zeldovich effect","selection bias","mass-observable scaling relations","mass-selected samples"],"falsifier":"Take a complete sample of 50 or more clusters selected purely by weak lensing in the same mass and redshift range and measure the same seven ICM properties; if the fraction of >2σ outliers per property is close to the expected 5% rather than the roughly 50% seen here, the anomaly is a property of these four objects rather than of ICM-selected sampling.","tokens_in":31319,"feed_emoji":"🔭","tokens_out":6369,"duration_ms":57606,"temperature":0.7,"pith_summary":"The paper argues that most knowledge of the hot gas (intracluster medium, ICM) in galaxy clusters comes from clusters that were themselves selected by that gas, which biases the resulting picture. It studies four massive clusters chosen purely by weak gravitational lensing, meaning selection depends on total mass rather than baryon content, and measures seven ICM-related properties for them. Relative to gas-selected clusters of the same mass, the four clusters show on average about two rare (more than 2σ) features per property, while only about 0.2 are expected if gas-selected samples faithfully represent the population. The paper concludes that X-ray- and Sunyaev-Zeldovich-selected catalogs are missing a real population of massive, gas-poor clusters, and that scaling relations built from those catalogs are biased.","feed_headline":"Four massive clusters show 10x too many rare gas features","feed_subtitle":"A lensing-selected sample finds gas-poor, X-ray-faint clusters that standard X-ray and SZ surveys miss.","key_machinery":"The load-bearing device is a sample selected by weak gravitational lensing from the HSC shear-selected catalog, whose inclusion probability depends on total mass rather than on baryon content. The comparison machinery is a set of seven ICM observables—richness, core-excised X-ray luminosity, Compton Y, electron density and pressure profiles, and their central values—measured for the four clusters from shear, X-ray, and SZ data and compared with gas-selected reference samples, notably SZ-selected profile libraries and X-ray-selected scaling relations. A Bayesian forward-modeling code (MBProj2 extended to shear) derives masses and profiles while applying a Tinker mass-function prior to correct for Eddington bias; for the two out-of-equilibrium clusters the analysis deliberately avoids assuming hydrostatic equilibrium.","core_discovery":"The central discovery is that a small, shear-selected sample of four massive clusters is dramatically more unusual in its ICM properties than gas-selected samples would predict. Cluster id5, the most striking object, has a weak-lensing mass of log M500/Msun = 14.68 ± 0.10 but a core-excised X-ray luminosity and richness that sit far below the relations defined by X-ray- and SZ-selected clusters, and its pressure and density profiles fall below the ±2σ range of SZ-selected clusters; an independent caustic analysis confirms its mass. The other unusual object, id34, is similarly low in pressure and Compton Y. Across seven explored properties, the sample shows 12 outliers beyond 2σ where about 1.4 would be expected by chance in four objects, or an average of two rare features per property against an expected 0.2. The paper interprets this excess as evidence that the thermodynamic properties of clusters derived from ICM-selected samples are fundamentally biased: massive clusters with low gas content exist and are largely absent from those samples.","pith_inferences":["My inference: if the anomaly is driven by the infalling groups seen in id5 and id34, then a mass-selected sample split by dynamical state should show the outliers concentrated in merging clusters, making the bias partly a dynamical-state bias rather than purely a gas-fraction bias.","My inference: the true intrinsic scatter of X-ray luminosity at fixed mass may be sample-dependent, so calibrating cluster scaling relations for cosmology may require explicitly modeling the population that gas-selected surveys cannot see.","A testable extension: applying the same seven-property comparison to roughly 50 to 100 lensing-selected clusters would directly measure the outlier rate and reveal which of the seven properties suffer the strongest selection bias, guiding which scaling relations need re-derivation.","My inference: if low-richness massive clusters like id5 are common, richness-based mass estimates can be wrong by large factors for a minority of objects, which would affect optically selected cluster cosmological samples."],"forward_implications":["If the paper is right, X-ray and SZ cluster catalogs are incomplete not only at the faint end but across a population of massive gas-poor clusters, so mass-observable scaling relations derived from them underpredict scatter and overpredict gas content at fixed mass.","A massive cluster like id5 can be absent from an X-ray cosmological survey such as eROSITA DR1 even at redshift 0.25, implying that X-ray mass completeness is lower than commonly assumed.","The population of low-Compton-Y clusters that a previous analysis of X-ray-selected data had to postulate in order to correct its scaling relation is directly observed in id5 and id34, supporting selection-effect corrections of that type.","Baryon-independent, lensing-based selection becomes a necessary tool for measuring cluster thermodynamics and for calibrating the mass-observable relations used in cluster cosmology."],"supporting_citations":[{"why":"Supplies the HSC weak-lensing shear-selected cluster catalog from which the four clusters are drawn.","marker":"Miyazaki et al. (2018)"},{"why":"Provides the SZ-selected electron density and pressure profiles, including the mean and ±2σ range, used as the comparison baseline.","marker":"Ghirardini et al. (2020)"},{"why":"Defines the Compton Y-mass scaling of an X-ray-selected sample and postulates the low-Y population that id5 and id34 occupy.","marker":"Nagarajan et al. (2019)"},{"why":"Provides the SZ-selected luminosity-mass relation against which id5's core-excised luminosity is measured.","marker":"Pratt et al. (2022)"},{"why":"Defines the X-ray-unbiased comparison sample (XUCS) that makes id5's low luminosity look typical.","marker":"Andreon et al. (2016)"},{"why":"Supplies ACT cluster detections and redshifts; id5 is absent from its shallower maps, motivating the deeper-search analysis.","marker":"Hilton et al. (2021)"},{"why":"Releases the deeper ACT Compton maps from which the paper measures spherical Compton Y for the four clusters.","marker":"Coulton et al. (2024)"},{"why":"Provides Eddington-bias-corrected mass estimates for these same clusters, used as a cross-check.","marker":"Hamana et al. (2023)"},{"why":"Supplies the mass function used as a prior to correct for Eddington bias in the mass and profile fits.","marker":"Tinker et al. (2008)"},{"why":"MBProj2 forward-modeling code that the paper extends to jointly fit shear and X-ray data.","marker":"Sanders et al. (2018)"}],"fun_headline_variants":["Weak-lensing cluster sample shows 10x rare gas features","Gas-selected clusters miss gas-poor giants found by lensing","Baryon-independent selection exposes cluster gas bias","Mass-selected clusters reveal 10x more unusual gas features","Gas-poor giant clusters are 10x more common than gas surveys show"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The four clusters are taken from a weak-lensing shear-peak catalog that is assumed to be unbiased with respect to baryon content and ICM properties, and the two most massive plus two high-signal-to-noise clusters are treated as representative of the larger mass-selected population.","fun_headline_variants_meta":{"raw":{"variants":["Weak-lensing cluster sample shows 10x rare gas features","Gas-selected clusters miss gas-poor giants found by lensing","Baryon-independent selection exposes cluster gas bias","Mass-selected clusters reveal 10x more unusual gas features","Gas-poor giant clusters are 10x more common than gas surveys show"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001299,"raw_usage":{"total_tokens":5360,"prompt_tokens":1063,"completion_tokens":4297,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":679,"completion_tokens_details":{"reasoning_tokens":4210}},"tokens_in":679,"tokens_out":4297,"duration_ms":28351,"temperature":1.0,"reasoning_tokens":4210,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:38:35.967174+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a complete sample of 50 or more clusters selected purely by weak lensing in the same mass and redshift range and measure the same seven ICM properties; if the fraction of >2σ outliers per property is close to the expected 5% rather than the roughly 50% seen here, the anomaly is a property of these four objects rather than of ICM-selected sampling.","supporting_citations":[{"cited_title":"2019, MNRAS, 488,","cited_arxiv_id":null,"evidence_quote":"Defines the Compton Y-mass scaling of an X-ray-selected sample and postulates the low-Y population that id5 and id34 occupy."},{"cited_title":"W., Arnaud, M., Maughan, B","cited_arxiv_id":null,"evidence_quote":"Provides the SZ-selected luminosity-mass relation against which id5's core-excised luminosity is measured."},{"cited_title":"2021, ApJS, 253,","cited_arxiv_id":null,"evidence_quote":"Supplies ACT cluster detections and redshifts; id5 is absent from its shallower maps, motivating the deeper-search analysis."},{"cited_title":"S., Duivenvoorden, A","cited_arxiv_id":null,"evidence_quote":"Releases the deeper ACT Compton maps from which the paper measures spherical Compton Y for the four clusters."},{"cited_title":"2023, PASJ, 75,","cited_arxiv_id":null,"evidence_quote":"Provides Eddington-bias-corrected mass estimates for these same clusters, used as a cross-check."},{"cited_title":"Sayers, J., Mantz, A","cited_arxiv_id":null,"evidence_quote":"MBProj2 forward-modeling code that the paper extends to jointly fit shear and X-ray data."}],"review_version":1}