{"id":"5f38c39d-06fd-4a35-9e8b-d5b8bf1d85b5","arxiv_id":"2501.14019","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Hydrodynamical simulations of the Three Hundred project produce weak-lensing mass-richness calibrations that are consistent across two galaxy formation models and broadly match SDSS redMaPPer when a stellar mass threshold of 10^10 solar masses is adopted.","lead":"This paper uses two hydrodynamical simulations of galaxy clusters to measure how weak-lensing mass estimates are biased and to calibrate the relation between cluster richness and lensing mass. It finds that the two different galaxy formation models give consistent mass-richness relations, a result that can serve as a calibration guideline for cluster cosmology surveys such as Euclid.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equating Eq. (14) cylinder counts with redMaPPer richness is unvalidated; the Mstar,min=10^10 match to redMaPPer may be post-hoc, so the survey-facing calibration is conditional.","rationale":"The reader's weakest assumption correctly identifies the simulated richness definition as the point where the observational comparison is least secure. My reading agrees: the paper's derivation of the weak-lensing mass bias and the internal consistency between GadgetX and GIZMO-SIMBA are supported by the presented fits and the controlled simulation setting, and the WL mass bias analysis is a solid calibration contribution. However, the claim that the relation 'aligns well with SDSS redMaPPer cluster analyses' is load-bearing for the paper's stated aim of providing guidance for observational experiments. The alignment is obtained by choosing Mstar,min=10^10 h^-1 M_sun, a free parameter, and by comparing Eq. (14) counts with redMaPPer richness after arbitrary rescaling. No redMaPPer or AMICO algorithm is applied to simulated galaxies, so the mapping between the two richness definitions is unvalidated. This does not invalidate the hydro-model comparison or the WL bias results, but it should move the survey-facing conclusion to conditional status, which is exactly what the reader already recommended. The most useful next step is a mock-catalogue end-to-end test of redMaPPer on the simulations; until that is done, the observational alignment should be described as a model-dependent choice rather than a demonstrated reproduction of redMaPPer richness. The other concerns noted by the reader (diagonal covariance, mass-selected progenitor sample, fixed quadratic coefficients) are secondary to this mapping issue for the central claim.","tokens_in":23529,"tokens_out":3740,"duration_ms":38518,"concrete_test":"Construct mock galaxy catalogues from The Three Hundred snapshots with realistic photometric properties (rest-frame magnitudes, colours, and photometric redshift errors), run a public redMaPPer implementation on these mocks over the redshift range z=0.2-0.6, and compare the recovered redMaPPer richness with the Eq. (14) cylinder richness for the same clusters. If the two richness estimates are not related by a stable, nearly mass- and redshift-independent mapping with scatter consistent with the reported sigma_log_lambda, then the claimed Mstar,min=10^10 match to redMaPPer does not validate the mass-richness calibration for observational use.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central survey-facing claim is that the simulated weak-lensing mass-observed richness relation, with Mstar,min=10^10 h^-1 M_sun, aligns with SDSS redMaPPer analyses. This rests on Eq. (14): observed richness is defined as the number of (sub)haloes above a stellar mass threshold in a cylinder of radius R200 and height 10 Mpc, minus 4/33 of haloes in a ring between 2 and 3.5 R200. That is not redMaPPer richness. redMaPPer applies red-sequence membership probabilities, an evolving luminosity threshold (0.2 L*), percolation, photometric redshift weighting, and a central galaxy treatment, none of which appear in Eq. (14). The paper does not run redMaPPer or AMICO on simulated galaxies; it simply selects by stellar mass and then rescales literature richnesses using McClintock et al. (2019) Table 5. Because Mstar,min is a free parameter and the adopted value was chosen partly to produce agreement, the redMaPPer alignment could be a tuning effect rather than evidence that the simulated richness reproduces the observed one. If the mapping between Eq. (14) counts and redMaPPer richness is redshift-dependent or mass-dependent, the calibrated mass-richness relation cannot be applied to real surveys even if the weak-lensing masses are correct. The paper itself notes in Section 3 that a stellar mass threshold 'better aligns' with redMaPPer, but provides no quantitative validation of that mapping, which is the load-bearing step for the comparison with observations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses the Three Hundred project's zoom hydrodynamical simulations (GadgetX and GIZMO-SIMBA) together with a dark-matter-only run, for 324 massive clusters selected by M200 > 8e14 h^-1 Msun at z=0, to quantify weak-lensing (WL) cluster mass biases and to calibrate mass-richness relations up to z=0.94. The authors build simulated convergence/shear maps from random projections, fit smoothly truncated NFW profiles to the excess surface mass density, define an observed richness via cylinder counts of subhaloes above a stellar mass threshold with a geometric background subtraction (Eq. 14), and fit forward and inverse WL mass-richness relations. They report that the two hydrodynamical codes give WL mass-richness relations consistent within 1 sigma, that the intercept is redshift-independent, that the slope is approximately constant below z~0.55 and follows a quadratic in redshift, and that the scatter grows linearly with redshift. They further claim that, with a minimum stellar mass of Mstar,min = 1e10 h^-1 Msun, their relation aligns with SDSS redMaPPer cluster analyses.","tokens_in":23845,"tokens_out":4136,"duration_ms":40100,"significance":"If the central claims hold, the paper provides a valuable controlled comparison of baryonic effects on weak-lensing mass calibration and a useful set of fitting tables for mass-richness calibration. The main strengths are the use of two different hydrodynamical codes with identical initial conditions, a DM-only baseline, nine snapshots, projection-averaged weak-lensing profiles, explicit MCMC regression fits, and compact parameter tables (Tables 1-3) that can be used by future studies. The comparison of scatter at fixed true mass versus fixed weak-lensing mass is also a useful result. However, the survey-facing significance is conditional on the assumed mapping between the simulated cylinder richness and observed richness estimators such as redMaPPer, and on the idealized weak-lensing setup (diagonal covariance, fixed source redshift).","major_comments":[{"comment":"The central claim that the simulated relation aligns with SDSS redMaPPer cluster analyses is not quantitatively validated. The observed richness in Eq. (14) is a cylinder count of subhaloes above a stellar mass threshold with a geometric 4/33 background subtraction, whereas redMaPPer uses red-sequence membership probabilities, an evolving 0.2 L* luminosity threshold, percolation, photometric redshift weighting, and a central galaxy treatment. The paper states in §3.2 that a stellar mass threshold 'better aligns' with redMaPPer, but no test of this mapping is provided. Since Mstar,min is a free parameter and the adopted value 1e10 h^-1 Msun was chosen partly to produce agreement, the claimed alignment may be a tuning artifact rather than evidence that the simulated richness reproduces the observed one. This is load-bearing for the survey-facing calibration. Please either validate the mapping by applying redMaPPer or AMICO to simulated galaxy catalogs, or explicitly reframe the comparison as a test of a simplified stellar-mass-based richness proxy and remove the 'aligns well' claim.","section":"§3, Eq. (14) and §3.2"},{"comment":"The weak-lensing analysis assumes a diagonal covariance matrix and a fixed source redshift zs=3. The diagonal covariance is acknowledged in the text, but it directly affects the reported 1-sigma consistency between GadgetX and GIZMO-SIMBA and the inferred scatter in the richness-mass relation: off-diagonal terms from correlated large-scale structure can substantially increase the uncertainties on the fitted masses. A fixed source plane at zs=3 is also not representative of typical surveys with broad source redshift distributions and photo-z errors, and it can bias the derived masses differently with redshift. The authors should test the sensitivity of their central WL mass bias and mass-richness scatter results to a more realistic source redshift distribution and to at least a block-diagonal covariance including large-scale structure terms, or restrict the claims accordingly.","section":"§2.2–§2.3, Eqs. (8), (12)–(13)"},{"comment":"The sample is selected by M200 > 8e14 h^-1 Msun at z=0, so the richness-mass relations are calibrated on a mass-selected sample rather than on a sample selected by the richness observable. The paper notes that a true mass-selected sample is not strongly affected by Malmquist-Eddington biases, but for application to optically selected cluster surveys the selection function matters and can change both the slope and scatter. The calibration should either be convolved with realistic selection functions or the claims should be limited to the simulated mass-selected sample. This is especially relevant because the comparison with redMaPPer in Fig. 10 assumes that observed optical selection is equivalent to the simulated one.","section":"§2.1, §3.1, §4"}],"minor_comments":[{"comment":"The text referring to Fig. 7 says that the Chen et al. (2024) model fails for Mstar,min = 10^12.5 h^-1 Msun, but the four panels in Fig. 7 show stellar mass cuts up to 10^10.75 h^-1 Msun; this appears to be a typo and should be corrected.","section":"§3"},{"comment":"The last column of Table 2 is missing formatting for the reported errors: entries such as '0.1260.003 + 0.0270.006z' should read 0.126 +/- 0.003 + (0.027 +/- 0.006) z.","section":"Table 2"},{"comment":"The comparison in Fig. 10 uses richnesses rescaled according to Table 5 of McClintock et al. (2019), but the text does not specify whether the simulated Eq. (14) richness is on the same scale as the rescaled redMaPPer richness before the fit; please state explicitly how the rescaling is applied to the simulated values or why it is not needed.","section":"§3.2"},{"comment":"The likelihood in Eqs. (16)–(17) includes the propagated mass uncertainty via B^2 sigma^2_logMwl but does not include an explicit intrinsic scatter parameter; the reported scatter sigma_log lambda_obs is therefore a residual scatter. The paper should clarify whether this residual scatter is meant to include intrinsic scatter and how the absence of an intrinsic scatter term affects the quoted uncertainties.","section":"§2.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's internal comparison of two hydrodynamical codes and the DM-only run is solid and useful. The main weakness is the redMaPPer comparison: it rests on an unvalidated mapping between a stellar-mass cylinder count and redMaPPer richness, and the choice of Mstar,min appears to be partly tuned to produce agreement. I would suggest the authors either run a member-assignment algorithm on the simulations or substantially soften the observational claims. The fixed source redshift and diagonal covariance are additional reasons why the quoted uncertainties and the 1-sigma consistency between hydro runs should be re-examined before the paper is accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid calibration paper from the Three Hundred project. The genuinely new piece is the comparison of weak-lensing mass biases and mass-richness relations across GadgetX, GIZMO-SIMBA, and DM-only runs of the same initial conditions. The earlier Euclid Collaboration paper used only GadgetX, so this is a real extension. The central result—consistency within 1σ between the two hydro codes for the WL mass-richness relation—is defensible and clearly presented. The fitting coefficients in Tables 1–3 and A.1 are the practical deliverable, and they will be directly usable for cluster cosmology.\n\nThe paper is honest about its technical choices. It uses a diagonal covariance and a fixed source redshift in the WL analysis, and it explains why; that is acceptable for a controlled comparison, though the quoted uncertainties are likely understated. The weaker spot is the richness side. Equation (14) is a cylinder count with geometric background subtraction, not redMaPPer or AMICO. The authors say a stellar mass threshold 'better aligns' with redMaPPer, but they do not validate that mapping. The Mstar,min = 10^10 h^-1 M_sun value, where the agreement with SDSS redMaPPer appears, is a free parameter, so the comparison is partly by construction. That does not invalidate the calibration, but it means the survey-facing claim is conditional. The redshift evolution is also affected by the z=0 mass selection of the sample; at z > 0.55 the slope changes partly reflect progenitor bias. The hand-fixed quadratic coefficients are a fitting choice, reasonable but not physical.\n\nThe data are simulation outputs from a public project, the methods are standard, and the citation pattern looks appropriate for this field. I would accept this for peer review. It is a workmanlike contribution from an established collaboration, the coefficients are useful, and the limitations are acknowledged. The main things to push on in revision: quantify the diagonal-covariance effect and discuss the Mstar,min freedom more explicitly. I'd cite it for the coefficients.","headline":"Solid extension of the Three Hundred weak-lensing calibration to a second hydro code; the fitting tables are useful, but the redMaPPer agreement is partly tuned by the stellar-mass cut and should be read as conditional.","tokens_in":24468,"tokens_out":4150,"would_cite":true,"duration_ms":34951,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two simulation codes with different baryon treatments give consistent weak-lensing mass–richness relations for clusters.","keywords":["galaxy clusters","weak gravitational lensing","mass-richness relation","hydrodynamical simulations","baryonic feedback","projection effects","cluster cosmology","mass-observable relations"],"falsifier":"Run the paper's cylinder-count richness estimator on a realistic mock galaxy catalogue that includes photometric noise and selection, then run the redMaPPer cluster finder on the same catalogue: if the redMaPPer richness differs from $\\lambda_{\\rm obs}$ systematically with mass or redshift beyond the quoted scatter, the calibration would not transfer to real surveys. A direct observational check is to measure the slope and scatter of the mass–richness relation for SDSS redMaPPer clusters at $z<0.3$ with a stellar-mass-complete sample and compare with the $M_{\\mathrm{star,min}}=10^{10}\\,h^{-1}\\,M_\\odot$ prediction.","tokens_in":23317,"feed_emoji":"🔭","tokens_out":15152,"duration_ms":123608,"temperature":0.7,"pith_summary":"This paper asks whether the way baryons are treated in simulations changes the masses weak lensing recovers for galaxy clusters, and whether a galaxy-count proxy called richness can still be calibrated to those masses. The answer it defends is no for the first question and yes for the second: the weak-lensing mass bias is the same within uncertainties for the GadgetX and GIZMO-SIMBA hydrodynamical runs and a dark-matter-only run, when the lensing mass is compared with the true mass of the same simulation. The paper then measures the weak-lensing mass–observed richness relation, $\\langle\\log\\lambda_{\\rm obs}|M_{\\rm wl}\\rangle = A(z)+B(z)\\log(M_{\\rm wl}/M_p)$, and finds both hydro runs consistent within $1\\sigma$; in the combined fit the intercept is redshift-independent, the slope is a second-order polynomial in $z$ but nearly flat below $z\\simeq0.55$, and the scatter grows linearly with redshift. With a minimum stellar mass of $10^{10}\\,h^{-1}\\,M_\\odot$, the simulated relation lines up with SDSS redMaPPer calibrations, so the fitted relations are offered as model priors for cluster-count cosmology.","feed_headline":"Two hydro codes agree on cluster mass–richness relation","feed_subtitle":"Calibrated with a 10^10 solar-mass stellar cut, it matches SDSS redMaPPer clusters.","key_machinery":"The load-bearing tool is a forward-model weak-lensing pipeline applied to the same 324 clusters run with GadgetX, GIZMO-SIMBA, and a dark-matter-only version. Particles within $\\pm5$ Mpc along the line of sight are collapsed onto lens planes for three orthogonal projections, shear maps are sampled at 30 background galaxies per square arcminute, and the excess surface mass density profiles are fitted with a smoothly truncated Navarro-Frenk-White (BMO) profile using a Bayesian Monte Carlo Markov Chain to obtain $M_{\\mathrm{wl}}$ and concentration. Richness is defined as the count of haloes and subhaloes above a stellar mass threshold in a cylinder of radius $R_{200}$ and height 10 Mpc, corrected for projected interlopers by subtracting $4/33$ of the halo count in an outer annulus (Eq. 14). The mass–richness relations are extracted by Bayesian linear regression with a Gaussian likelihood that propagates errors on both axes, producing the redshift- and stellar-mass-cut-dependent regression parameters in Tables 1 and 2.","core_discovery":"The paper establishes that baryonic physics does not break the weak-lensing calibration of cluster masses. Comparing the average weak-lensing mass bias of the two hydro runs with the dark-matter-only reference, the biases agree when each hydro run is compared with its own true mass, while relative to the DM-only mass GadgetX masses run a few percent high and GIZMO-SIMBA a few percent low, with the offsets closing above $M_{200,\\mathrm{DM}}\\simeq10^{15}\\,h^{-1}\\,M_\\odot$. It then constructs the observed richness from projected halo and subhalo counts with background subtraction, fits $\\langle\\log\\lambda_{\\rm obs}|M_{\\rm wl}\\rangle$ with a Bayesian linear regression, and shows that the two hydro codes give regression parameters consistent within $1\\sigma$. In the combined model, the intercept depends only on the stellar mass threshold, the slope follows a second-order polynomial in redshift and stays roughly constant up to $z\\simeq0.55$, and the scatter in richness at fixed weak-lensing mass grows linearly with redshift; the scatter is smaller when richness is tied to the true mass. At $M_{\\mathrm{star,min}}=10^{10}\\,h^{-1}\\,M_\\odot$ the observed-richness–weak-lensing-mass relation matches SDSS redMaPPer clusters, which is the paper's basis for offering the fits as priors for survey cluster cosmology.","pith_inferences":["Because the paper fixes one source redshift distribution, ignores off-diagonal shape-noise covariance, and does not model mis-centering, the precise redshift trends of the slope and scatter may shift when those survey effects are included; the cross-code consistency is more robust than the absolute parameter values.","The $4/33$ geometric background subtraction and hard stellar-mass cuts approximate redMaPPer's probabilistic red-sequence membership; applying the same pipeline inside a mock with photometric errors and running redMaPPer on it would test whether the calibration transfers directly.","The small hydro-code-dependent difference in weak-lensing mass relative to the DM-only reference suggests that dark-matter-only halo mass functions could remain usable for very massive clusters, while surveys probing lower masses may need baryon-dependent mass corrections.","The match to redMaPPer at one stellar mass cut implies a testable program: add photometric and membership-selection models to the simulations and predict full richness distributions, including the tails that dominate cluster-count systematics."],"forward_implications":["The weak-lensing mass bias of clusters can be calibrated without depending strongly on the galaxy-formation code: GadgetX, GIZMO-SIMBA, and the dark-matter-only run give the same bias and scatter when compared with their own true masses.","A single combined weak-lensing mass–observed richness relation, with the fitted redshift and stellar-mass-cut dependence, can be used for mapping survey richnesses to masses for cluster-count cosmology.","The slope of the relation stays nearly constant up to $z\\simeq0.55$, so low-redshift survey analyses do not need a strongly evolving mass–richness calibration.","The scatter of richness at fixed weak-lensing mass grows linearly with redshift and is larger than the scatter at fixed true mass, so redshift-dependent scatter must be included in mass-observable likelihoods.","With a $10^{10}\\,h^{-1}\\,M_\\odot$ stellar mass cut, the simulated observed-richness–weak-lensing-mass relation agrees with SDSS redMaPPer calibrations, giving an observational anchor for the simulation-based relation."],"supporting_citations":[{"why":"defines the Three Hundred sample of 324 massive clusters and the resimulated regions used for all runs.","marker":"Cui et al. (2018)"},{"why":"specifies the GIZMO-SIMBA hydrodynamical runs of the same clusters.","marker":"Cui et al. (2022)"},{"why":"supplies the SIMBA galaxy-formation subgrid models used by the GIZMO-SIMBA run.","marker":"Davé et al. (2019)"},{"why":"supplies the GadgetX smoothed-particle-hydrodynamics code that produces the other baryonic run.","marker":"Rasia et al. (2015)"},{"why":"provides the weak-lensing simulation and truncated-NFW profile-fitting procedure that this paper extends to GIZMO-SIMBA and dark-matter-only runs.","marker":"Euclid Collaboration: Giocoli et al. (2024)"},{"why":"supplies the smoothly truncated NFW (BMO) density profile used to fit the simulated weak-lensing shear profiles.","marker":"Baltz et al. (2009)"},{"why":"supplies the local-background richness-correction approach that underlies Eq. (14).","marker":"Andreon & Bergé (2012)"},{"why":"gives the DES redMaPPer mass–richness calibration that the simulated relation is rescaled against and found to agree with.","marker":"McClintock et al. (2019)"}],"fun_headline_variants":["Hydro codes agree on mass–richness relation","Baryons don't break weak-lensing mass calibration","Hydro sims match redMaPPer cluster richness–mass","Cluster mass–richness relation validated by hydro runs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that counting simulated galaxies above a stellar mass limit in a cylinder and subtracting a fixed background fraction gives the same number that real survey algorithms such as redMaPPer record as richness; if real membership probabilities or the stellar-mass calibration do not match that count, the calibrated mass–richness relation would not apply to observed clusters.","fun_headline_variants_meta":{"raw":{"variants":["Hydro codes agree on mass–richness relation","Baryons don't break weak-lensing mass calibration","Hydro sims match redMaPPer cluster richness–mass","Cluster mass–richness relation validated by hydro runs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1882,"prompt_tokens":1185,"completion_tokens":697,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":801,"completion_tokens_details":{"reasoning_tokens":632}},"tokens_in":801,"tokens_out":697,"duration_ms":6221,"temperature":1.0,"reasoning_tokens":632,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:27:25.311054+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's cylinder-count richness estimator on a realistic mock galaxy catalogue that includes photometric noise and selection, then run the redMaPPer cluster finder on the same catalogue: if the redMaPPer richness differs from $\\lambda_{\\rm obs}$ systematically with mass or redshift beyond the quoted scatter, the calibration would not transfer to real surveys. A direct observational check is to measure the slope and scatter of the mass–richness relation for SDSS redMaPPer clusters at $z<0.3$ with a stellar-mass-complete sample and compare with the $M_{\\mathrm{star,min}}=10^{10}\\,h^{-1}\\,M_\\odot$ prediction.","supporting_citations":[{"cited_title":"N., Gruen, D., et al","cited_arxiv_id":null,"evidence_quote":"gives the DES redMaPPer mass–richness calibration that the simulated relation is rescaled against and found to agree with."}],"review_version":1}