{"id":"e8ebb067-abbc-41b1-ab74-94eeac7d10c5","arxiv_id":"2502.08444","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Joint counts plus weak lensing on DC2 redMaPPer clusters recovers the simulated mass-richness relation within 2 sigma, with photo-z biases reduced by a fitted correction and shear-richness covariance shifting results by about 0.5 sigma.","lead":"Using simulated LSST-like data, this paper measures the mass-richness relation of redMaPPer galaxy clusters by combining cluster counts with weak lensing, and tests how modeling choices and observational errors shift the result. A smart generalist would read it to see whether LSST cluster cosmology can calibrate cluster masses well enough for dark energy measurements.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Selection function is the load-bearing link, but its ClEvaR calibration is self-referential: completeness and the fiducial relation come from the same matching, so internal consistency cannot detect matching bias.","rationale":"The reader's weakest assumption identifies the selection function, and I agree it is the most load-bearing. My stress-test sharpens this: the selection function and the reference truth are measured from the same matched catalog, so the paper's internal consistency checks cannot detect a matching bias. The paper itself acknowledges the spurious-detection neglect as a strong assumption but provides no external test. A mock-injection calibration is the standard remedy (the paper lists it in Sect. 2.1 as a valid approach) and would settle whether the completeness/purity are accurate. Without such a test, the central claim is conditional on an untested calibration, which matches the CONDITIONAL verdict the reader reached. I recommend no change to the verdict.","tokens_in":41017,"tokens_out":5065,"duration_ms":52912,"concrete_test":"Inject synthetic redMaPPer-like clusters with known M200c and z into the cosmoDC2 galaxy catalog, rerun redMaPPer, and measure the recovered fraction as a function of mass, richness, and redshift. Compare this injection-based completeness c_inj(m,z) with the ClEvaR-based completeness used in Eqs. (1) and (13). Then repeat the baseline N+DeltaSigma inference using c_inj(m,z) and quantify the shift in the six scaling parameters. If the shift exceeds the reported 1-sigma errors, the claim of unbiased recovery is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that counts+lensing recovers the underlying mass-richness relation rests on the selection function c(m,z) entering Eqs. (1) and (13). This selection function is calibrated exclusively from the ClEvaR matched catalog, and the same catalog provides the fiducial P(lambda|M,z) used as the 'truth' reference (Sect. 3.4). A systematic error in the matching—e.g., redshift-dependent incompleteness or residual projection contamination—would bias the model predictions and the reference relation in tandem. The internal consistency reported in Sect. 6.1 is therefore not a valid external validation; it only shows the pipeline is self-consistent. The paper explicitly flags the neglect of spurious detections as a 'strong assumption' (Sect. 2.2, footnote 9) and defers to a purity estimate of ~99%, but that purity is also derived from the same matching. Because the selection function multiplies the halo mass function in the count model and the lensing model, any fractional error in c(m,z) propagates almost linearly into the inferred normalization ln lambda_0 and mass slope mu_m. None of the robustness tests in Sect. 5 vary the selection function; they only vary the density profile, c(M) relation, photo-z, and covariance. Thus the weakest link is the untested, self-referential calibration of the selection function.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constrains the redMaPPer cluster mass-richness relation in the cosmoDC2/DC2 simulation using cluster number counts and either stacked weak-lensing profiles or mean lensing masses, across richness 20<lambda<200 and redshift 0.2<z<1. The analysis uses a six-parameter log-normal scaling relation, a selection function calibrated from the ClEvaR matched cluster-halo catalog, and a fixed simulation cosmology. The authors report that the joint count+lensing analysis recovers the fiducial relation inferred from the matched catalog, that the constraints are robust to concentration-mass relation and density profile choices, that photometric redshift uncertainties introduce about 1 sigma biases which can be mitigated by a fitted multiplicative correction, and that shear-richness covariance shifts results by up to about 0.5 sigma. The paper also provides a quantitative tension metric and makes use of several public DESC software tools.","tokens_in":41263,"tokens_out":5030,"duration_ms":54555,"significance":"If the central claim holds, this is a useful validation of an LSST-era cluster scaling relation pipeline: it is the first redMaPPer mass-richness constraint in cosmoDC2, it explores a broad set of modeling and observational systematics (concentration-mass relations, NFW/Einasto/Hernquist profiles, BPZ/FlexZBoost photo-zs, shear-richness covariance), and it provides a transparent comparison against a fiducial relation. The strengths include the explicit use of public DES C tools, the side-by-side comparison of stacked profiles versus mean masses, and the candid discussion of limitations in Sect. 6.2. Nonetheless, the main validation is against a fiducial relation derived from the same matching that calibrates the selection function, so the external reach of the 'recovery' claim is weaker than the abstract implies; this, together with the diagonal-only lensing covariance, is the main reason the central claim needs additional support.","major_comments":[{"comment":"The central validation claim in Sect. 6.1 rests on agreement with a fiducial relation that is fitted to the same ClEvaR matched catalog (Sect. 3.4.1) that also supplies the selection function c(m,z) entering the count and lensing predictions (Sect. 3.4.2). This makes the agreement an internal consistency test: a systematic error in the matching, such as redshift-dependent incompleteness or residual projection contamination, would shift the model predictions and the reference relation together and would not be detected by the posteriors in Fig. 6. None of the robustness tests in Sect. 5 vary c(m,z), and the purity argument in footnote 9 relies on the same matched catalog. I request a perturbation test of c(m,z) within the uncertainties quoted in Sect. 3.4.2 (for example, using the stated 80% completeness level at M200c > 10^14 Msun) with the resulting shifts in ln lambda0 and mu_m reported, or, failing that, a revised wording that limits the claim to pipeline self-consistency rather than an external recovery of the mass-richness relation.","section":"Sect. 3.4 and Eqs. (1), (13), (14)"},{"comment":"The stacked-lensing likelihood uses only the diagonal of the bootstrap covariance, justified by the noisiness of off-diagonal terms. This is an acknowledged approximation, but the paper's quantitative consistency statements, including the tension values in Sect. 6.1 and Fig. E.1, depend on the covariance being correct; if the off-diagonal elements are positive, as expected from correlated large-scale structure and halo-profile variations, the quoted uncertainties are underestimated. I ask the authors to provide a quantitative check of the off-diagonal-to-diagonal ratio in the fitted 1-3.5 Mpc range, for example by comparing the bootstrap result with an analytic covariance prediction (Wu et al. 2019) or with a smoothed/shrunk bootstrap covariance, and to state whether the parameter shifts and tension values are affected.","section":"Sect. 4.2.2, Eq. (26)"},{"comment":"The multiplicative correction factor (1+b) is fitted jointly with the scaling parameters and is common to all redshift-richness bins. This can absorb only a monopole normalization error; it cannot correct a photo-z bias that varies with redshift or richness. In the BPZ case the uncorrected shifts in ln lambda0 and mu_m are of order 1 sigma, and the paper does not demonstrate that the factor is not absorbing a redshift-dependent bias. I suggest a diagnostic that fits b per redshift bin (or per redshift-richness bin) and checks whether the b values are consistent, since this directly affects the claim that the photo-z bias is 'mitigated' by the single factor.","section":"Sect. 5.3.1"}],"minor_comments":[{"comment":"The abstract says the constraints are consistent 'at the 1 sigma level' with the fiducial values, while Sect. 5.1 states the posteriors recover the fiducial values 'at the <2 sigma level' and reports a '>2 sigma tension' in the projected mu_z-sigma_z plane; Sect. 6.1 again says 'consistent at the 1 sigma level.' These statements should be harmonized with the actual tension values reported in Fig. E.1.","section":"Abstract vs. Sect. 5.1 and 6.1"},{"comment":"The figure caption calls the case with concentration left free 'concentration-free,' while the text uses 'free-concentration case'; please use one consistent term throughout.","section":"Fig. 7 and Sect. 5.2.1"},{"comment":"The block beginning 'Impact of source photometric redshifts, Sect. 5.3.2' appears to refer to the shear-richness covariance analysis rather than photometric redshifts; the section label should be corrected.","section":"Table E.1"},{"comment":"The lower integration limit m_min is not specified in Eq. (13); for consistency with Eq. (1) and Sect. 4.2.1, the adopted value (10^12 Msun) should be stated where the lensing model is first introduced.","section":"Eq. (13)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well within the scope of A&A and represents useful DESC pipeline validation. The main reason for a major revision is the self-referential calibration of the selection function and fiducial relation from the same matched catalog, which weakens the central 'recovery' claim; the requested perturbation tests should address this without requiring a fundamentally new analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid, carefully executed calibration study for LSST cluster cosmology. It doesn't produce new physics, but it does give the first redMaPPer mass-richness constraints in cosmoDC2 and quantifies shear-richness covariance and photo-z systematics in that mock. If you work on DESC cluster analysis, this is directly useful; if not, it's a well-done instance of established methods.\n\nWhat's genuinely good: the pipeline is built from public DESC tools (CLMM, CCL, PySSC, ClEvaR), the data vectors are clearly described, and the authors run a large set of robustness tests: concentration-mass relations, Einasto/Hernquist profiles, two photo-z codes, and a shear-richness covariance correction. They also state their limitations openly—diagonal-only lensing covariance, fixed halo mass function, approximate mean-mass mapping—rather than burying them. The consistency of the joint counts+lensing posteriors with the fiducial matched relation at the 1-2 sigma level is the main evidence, and it does hold up.\n\nThe soft spots are real but not fatal. The biggest one is the selection function. Completeness comes from the same ClEvaR matching that defines the fiducial mass-richness relation used as 'truth'. So the central consistency test cannot detect a systematic in the matching itself—if the matching is biased, both the model predictions and the reference move together. The paper acknowledges the purity assumption explicitly (footnote 9, Section 2.2), but it never varies the selection function or checks against an independent calibration. That makes the validation somewhat self-referential. I don't think this sinks the paper, because the goal is pipeline validation within a specific simulation, and the matching is the best available definition of truth there. But a referee should ask for an alternate selection function estimate (e.g., injection-based) or at least a sensitivity test to completeness perturbations.\n\nSecondary issues: the lensing covariance is diagonal-only (they justify it, but off-diagonal terms could shift error bars), the halo mass function and selection function are fixed rather than marginalized, and the two-step mean-mass mapping uses a simple average mass rather than the more correct <M^Gamma>^(1/Gamma). These are acknowledged in the conclusions, and they individually move results by less than a sigma, so they're addressable, not disqualifying.\n\nWho is this for? The DESC cluster working group and anyone building mass-richness inference pipelines for LSST. It is a calibration paper, not a cosmological result. The novelty is incremental but real, and the execution is careful.\n\nRecommendation: send it to peer review. A good referee will push on the selection function self-reference and the covariance treatment, but the paper is worth engaging with and the limitations are honestly stated.","headline":"A careful, useful simulation-based calibration of a counts+lensing pipeline for LSST cluster cosmology; the selection function is the load-bearing but self-referential piece.","tokens_in":41895,"tokens_out":2564,"would_cite":false,"duration_ms":27759,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A simulated LSST-like sky shows that combining cluster counts with weak lensing recovers the mass-richness relation of redMaPPer clusters.","keywords":["galaxy clusters","mass-richness relation","weak gravitational lensing","redMaPPer","cluster number counts","DC2 simulation","photometric redshifts","shear-richness covariance"],"falsifier":"Run the same count+lensing fit after adding the stacked shear of unmatched redMaPPer detections to the signal, or after replacing the matching-based completeness with an independently calibrated completeness from injected mock clusters; if the recovered $\\ln\\lambda_0$ or $\\mu_m$ move by more than the reported uncertainties, the selection-function treatment is the source of the bias.","tokens_in":40783,"feed_emoji":"🔭","tokens_out":6473,"duration_ms":61192,"temperature":0.7,"pith_summary":"This paper tries to establish that the mass-richness relation of optically selected galaxy clusters can be inferred reliably from the combination of cluster number counts and weak gravitational lensing, using the mock universe of the DC2 simulation as a testbed for the future LSST survey. The authors fit a six-parameter log-normal scaling relation between halo mass and redMaPPer richness for 3,600 clusters in the range $20<\\lambda<200$ and $0.2<z<1$, using either stacked lensing profiles or mean lensing masses together with counts. They find that the joint analysis matches a fiducial relation derived from direct halo-richness matching at the 1-2 $\\sigma$ level, and that the results are stable under changes in the concentration-mass relation and dark matter density profile. Photometric redshift errors introduce at most about 1 $\\sigma$ bias, which a fitted multiplicative correction absorbs, and shear-richness covariance shifts the mass slope by up to 0.5 $\\sigma$. If correct, this validates a pipeline that can be applied to LSST cluster cosmology.","feed_headline":"Mock-sky test recovers cluster mass-richness relation","feed_subtitle":"Joint counts and lensing in a 440 sq-deg simulated sky match the fiducial scaling relation within about one sigma.","key_machinery":"The load-bearing object is the forward model connecting halo mass and richness through $P(\\lambda|m,z)$, a six-parameter log-normal scaling relation, combined with the redMaPPer selection function $\\Phi(\\lambda,m,z)=c(m,z)/p(\\lambda,z)$, where completeness $c$ is measured by matching detected clusters to simulated halos and purity is assumed near unity for $\\lambda>20$. The lensing side uses the stacked excess surface density $\\Delta\\Sigma(R)$ predicted by integrating $\\Delta\\Sigma(R|m,z)$ over the same mass function, completeness, and richness distribution, with the 1-halo term from an NFW profile (or Einasto/Hernquist) plus a concentration-mass relation; a two-step variant fits a mean mass per richness-redshift bin instead. The machinery works by feeding the same selection-function-weighted halo population into both the count prediction and the lensing prediction, so that the complementary degeneracy directions of counts and lensing can break the scaling-relation parameter degeneracies.","core_discovery":"The central claim is that a joint count+lensing analysis of redMaPPer clusters in the DC2 simulation recovers the underlying mass-richness relation without strong modeling bias. The authors model the observed richness as a log-normal variable around a mean $\\langle\\ln\\lambda|m,z\\rangle = \\ln\\lambda_0 + \\mu_z \\ln((1+z)/(1+z_0)) + \\mu_m \\log_{10}(m/m_0)$, with scatter $\\sigma_{\\ln\\lambda|m,z}$, and predict both cluster counts and stacked excess surface density $\\Delta\\Sigma(R)$ through the halo mass function weighted by the redMaPPer completeness. Combining counts with either stacked profiles or mean masses tightens constraints on the mean scaling parameters by a factor of about seven and on the scatter parameters by a factor of three to four, and the recovered parameters agree with the fiducial relation from halo-richness matching at the 1-2$\\sigma$ level. The paper also claims that the constraints are robust to the choice of concentration-mass relation and to NFW, Einasto, or Hernquist density profiles in the radial fit range $1<R<3.5$ Mpc, and that photo-z systematics and shear-richness covariance produce only small, partially correctable shifts.","pith_inferences":["The $R>1$ Mpc cut imposed by ray-tracing resolution means the inner region, where profile differences and miscentering matter most, is untested; real LSST data with better small-scale resolution could reveal larger modeling sensitivity than this paper measures.","The measured shear-richness covariance is likely suppressed by the simulation's spherical assignment of galaxies in halos, so the 0.5 sigma shift should be treated as a lower bound until triaxial galaxy placement or real data are used.","A natural next step is to free the cosmological parameters in the same counts+lensing fit; the complementarity demonstrated here suggests joint cluster cosmology constraints could be substantially tighter than counts alone.","The photo-z bracket between pessimistic and optimistic estimators offers a practical calibration strategy for LSST: adopt a fitted multiplicative lensing correction and validate it with spectroscopic subsamples."],"forward_implications":["Using either stacked profiles or mean masses in combination with counts gives comparable constraints; errors on $\\ln\\lambda_0$, $\\mu_z$, and $\\mu_m$ shrink by roughly a factor of seven relative to single-probe fits.","Adopting different concentration-mass relations or dark matter density profiles (NFW, Einasto, Hernquist) shifts the inferred scaling relation by less than about $1\\sigma$ when fitting in $1<R<3.5$ Mpc.","Photometric redshift errors from a pessimistic template-based estimator bias the normalization by about $1\\sigma$; fitting a global multiplicative factor $(1+b)$ jointly with the scaling parameters removes most of this bias.","Shear-richness covariance shifts the mass slope $\\mu_m$ by up to $0.5\\sigma$, an effect comparable to the photo-z shift and larger than the shift from an optimistic machine-learning photo-z estimator.","Across all tested analysis configurations, posterior shifts relative to the fiducial relation stay below $2\\sigma$, supporting the reliability of the joint count+lensing pipeline for LSST-era cluster cosmology."],"supporting_citations":[{"why":"Defines redMaPPer and the richness estimator whose mass-richness relation is the target of the analysis.","marker":"Rykoff et al. 2014"},{"why":"Supplies the cosmoDC2 galaxy catalog and ray-traced shears used to build all lensing data vectors.","marker":"Korytov et al. 2019"},{"why":"Documents the DC2 simulation suite, including the cluster-catalog and selection-function inputs.","marker":"Abolfathi et al. 2021"},{"why":"Provides the NFW profile used to model the one-halo lensing signal.","marker":"Navarro et al. 1997"},{"why":"Provides the baseline concentration-mass relation in the lensing profile model.","marker":"Duffy et al. 2008"},{"why":"Provides the halo mass function that weights both count and lensing predictions.","marker":"Despali et al. 2015"},{"why":"Supplies the log-normal mass-richness parametrization and the flat priors adopted for the fit.","marker":"Murata et al. 2019"},{"why":"Provides the shear-richness covariance formalism used to correct stacked profiles for selection bias.","marker":"Zhang et al. 2024"},{"why":"Introduces the multiplicative (1+b) photo-z calibration correction fitted jointly with the scaling parameters.","marker":"Simet et al. 2017"},{"why":"Establishes the two-step mean-lensing-mass approach and its systematics treatment.","marker":"McClintock et al. 2019"}],"fun_headline_variants":["Mock-sky test recovers cluster mass-richness","Lensing recovers cluster scaling in simulation","Simulated sky test validates cluster mass-richness","Joint counts and lensing probe cluster masses","DC2 mock sky pins down cluster scaling relation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis assumes that the redMaPPer completeness measured by matching detected clusters to dark matter halos is accurate and that the roughly 1% spurious detections contribute negligible lensing signal; if unmatched clusters have non-negligible shear or the matching misses real clusters, both the count and stacked-lensing predictions shift and bias the inferred mass-richness relation.","fun_headline_variants_meta":{"raw":{"variants":["Mock-sky test recovers cluster mass-richness","Lensing recovers cluster scaling in simulation","Simulated sky test validates cluster mass-richness","Joint counts and lensing probe cluster masses","DC2 mock sky pins down cluster scaling relation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1492,"prompt_tokens":1121,"completion_tokens":371,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":737,"completion_tokens_details":{"reasoning_tokens":301}},"tokens_in":737,"tokens_out":371,"duration_ms":4304,"temperature":1.0,"reasoning_tokens":301,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T05:04:31.014747+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same count+lensing fit after adding the stacked shear of unmatched redMaPPer detections to the signal, or after replacing the matching-based completeness with an independently calibrated completeness from injected mock clusters; if the recovered $\\ln\\lambda_0$ or $\\mu_m$ move by more than the reported uncertainties, the selection-function treatment is the source of the bias.","supporting_citations":[{"cited_title":"S., Rozo, E., Busha, M","cited_arxiv_id":null,"evidence_quote":"Defines redMaPPer and the richness estimator whose mass-richness relation is the target of the analysis."},{"cited_title":"2019, , 71, 107","cited_arxiv_id":null,"evidence_quote":"Supplies the log-normal mass-richness parametrization and the flat priors adopted for the fit."},{"cited_title":"Impact of Property Covariance on Cluster Weak lensing Scaling Relations","cited_arxiv_id":"2310.18266","evidence_quote":"Provides the shear-richness covariance formalism used to correct stacked profiles for selection bias."},{"cited_title":"2017, , 466, 3103","cited_arxiv_id":null,"evidence_quote":"Introduces the multiplicative (1+b) photo-z calibration correction fitted jointly with the scaling parameters."}],"review_version":1}