Pith. sign in

REVIEW 3 major objections 4 minor 11 references

This paper presents a validation pipeline that cross-matches optical cluster catalogs to X-ray and SZ data, and shows that both redMaPPer and WaZP recover 100% of the most massive SPT clusters while centering correctly in about 93% of unamb

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A validation pipeline reports that redMaPPer and WaZP recover over 88% of SPT-selected clusters and correctly center roughly 93% of X-ray-matched clusters.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection Useful validation pipeline with new WaZP measurements, but the 100% completeness claim and the mis-centering conclusion need tighter quantification and scope. the 3 major comments →

arxiv 2509.07268 v1 pith:HY77JBKE submitted 2025-09-08 astro-ph.CO

Cluster Catalog Validation with Multiwavelength Data

classification astro-ph.CO
keywords galaxy clustersvalidation pipelineredMaPPerWaZPSunyaev-Zel'dovich effectX-ray clustersLSSTDES
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper builds a reusable pipeline for checking whether optically detected galaxy clusters are real, complete, and correctly centered, by comparing them against X-ray and microwave (Sunyaev-Zel'dovich) catalogs. Applied to two DES cluster catalogs, the pipeline shows both redMaPPer and WaZP recover essentially every massive cluster that the South Pole Telescope sees, with 100% recovery above a detection significance of about 10. Centering is good in roughly 93% of clusters with unambiguous X-ray centers, and richness correlates tightly with X-ray temperature and SZ significance. The point is to catch data and algorithm problems before LSST starts producing thousands of cluster catalogs.

Core claim

The central claim is that a validation pipeline cross-matching optical cluster catalogs to SZ and X-ray samples can simultaneously test completeness, centering, and observable scatter, and that on DES data both redMaPPer and WaZP pass these tests. In particular, out of 151 SPT clusters in the redshift range and footprint, redMaPPer matches 134 and WaZP 138, and every SPT cluster above ξ~10 is recovered. Centering fits give well-centered fractions of 0.94±0.07 and 0.92±0.07, and the intrinsic scatter of the richness–X-ray temperature relation agrees between the two finders, while WaZP shows larger scatter in the richness–SZ relation. The authors argue the pipeline is broadly applicable and wo

What carries the argument

The pipeline's core is a cross-matching step using the ClEvaR library, which pairs optical clusters to SPT and Chandra-X-ray clusters within 2 Mpc and redshift 0.05, followed by three tests: (1) recovery fraction versus SZ significance; (2) a two-component gamma distribution fit to optical–X-ray position offsets, with an MCMC yielding the well-centered fraction ρ and the centered/miscentered scales σ and τ; (3) Bayesian scaling-relation fits (Kelly 2007) between richness and X-ray temperature/SZ significance, with intrinsic scatter as the diagnostic.

Load-bearing premise

The reference catalogs are unbiased: that the SPT sample with ξ>5 contains every massive cluster in the footprint, and that the visually cleaned Chandra X-ray subset gives an unbiased measure of where the true cluster centers are.

What would settle it

Take a large sample of SPT clusters just above ξ~5 and check each with deep X-ray follow-up; if a nontrivial fraction of these high-significance clusters are absent from redMaPPer or WaZP, the claimed completeness fails. Conversely, if a synthetic catalog with known injected center offsets is run through the pipeline and the recovered well-centered fraction ρ is systematically off from the input, the centering measurement is biased.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Both redMaPPer and WaZP are reliable for cosmology over the DES footprint in the redshift range 0.2–0.65, at least for massive clusters.
  • The 100% recovery above SZ significance ξ~10 means SZ-selected massive clusters can serve as a completeness benchmark for LSST cluster catalogs.
  • Centering fractions of roughly 0.92–0.94 imply that central galaxy identification is correct in nearly all unambiguous cases, so miscentering will not dominate the cluster mass calibration error budget.
  • The pipeline can be rerun on LSST catalogs to catch selection bugs early, since earlier DES catalog versions had features these tests would reveal.
  • WaZP's higher scatter in the richness–SZ relation suggests that redshift and richness estimation differences matter for scatter, not just for completeness.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the same pipeline were applied to simulated cluster catalogs with injected miscentering and known completeness, it would calibrate the accuracy of the validation metrics themselves—something the paper does not do.
  • The reliance on a curated, visually cleaned X-ray sample means the quoted well-centered fraction applies only to clusters with unambiguous X-ray peaks; the true miscentering fraction for the full cluster population could be higher because faint or disturbed systems are excluded.
  • The SPT ξ>5 cut does not test completeness for lower-mass clusters; LSST science will depend on clusters at much lower richness, where the algorithms might be less complete, and deeper X-ray or SZ follow-up would be needed to check that regime.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper describes a validation pipeline for optically selected galaxy cluster catalogs that cross-matches against SZ and X-ray catalogs and runs three tests: completeness relative to SPT, centering offsets relative to Chandra-observed RASS-MCMF clusters, and intrinsic scatter in richness–X-ray temperature and richness–SZ significance scaling relations. The pipeline is applied to DES Y3 redMaPPer and DES Y1 WaZP catalogs. The reported results are that both algorithms recover SPT clusters with 100% efficiency above ξ∼10, have well-centered fractions ρ≈0.92–0.94, and have comparable or modestly different intrinsic scatters.

Significance. If the claims hold, the pipeline is a useful, reusable tool for the LSST-DESC cluster working group and for the broader cluster community. The manuscript ships reproducible code on GitHub and uses public DES, SPT, and Chandra/X-ray catalogs, which is a clear strength. The quantitative results are modest in scope but directly relevant to catalog validation for DES and LSST. The main value is the demonstration of a standardized multiwavelength validation workflow. However, the headline completeness and mis-centering claims are not fully supported by the statistics as presented, so the conclusions need strengthening before the paper can serve as a reliable reference.

major comments (3)
  1. [Section 3] The claim that 'both redMaPPer and WaZP show excellent completeness ... with 100% recovery of SPT clusters above ξ∼10' is not quantitatively supported. The number of SPT clusters with ξ>10 in the chosen redshift range and footprint is not stated, no confidence interval is given, and the statement is made after applying λ>20 and N_gals>25 cuts (Section 2.1). An SPT cluster with redMaPPer richness below 20 would be counted as unmatched even if the cluster finder detected its member galaxies, so the recovery fraction conflates catalog selection with algorithmic detection. Please report N_ξ>10, per-cluster richness/N_gals for matched and unmatched SPT clusters, and a binomial or bootstrap confidence interval. If the ξ>10 subsample is only ∼10–20 clusters, 100% recovery is compatible with a substantially lower true completeness; the text should state this explicitly.
  2. [Table 1 / Section 3] The completeness percentages in Table 1 (88.7% for redMaPPer, 91.4% for WaZP, for ξ>5) are quoted without uncertainties. Given N=151 SPT clusters, the 95% binomial confidence intervals are roughly ±5–6 percentage points, so the difference between redMaPPer and WaZP is not significant as presented. Adding uncertainties is necessary for the stated comparison and for the 'excellent completeness' language. The table should also clarify that completeness is relative to the SPT ξ>5 sample within the DES Y1 footprint, not absolute completeness (as Section 2.1 itself notes).
  3. [Section 4] The conclusion 'In both cases 8% or less of the clusters were found to be miscentered' overstates the result. The fitted well-centered fractions are ρ=0.94±0.07 and ρ=0.92±0.07 (Table 1), so the 1−ρ values are 0.06±0.07 and 0.08±0.07; the data are consistent with a wide range of mis-centering fractions including values above 8%. Moreover, Section 2.1 explicitly states that the X-ray sample is deliberately limited to well-centered, visually clean clusters, so the fitted parameters are not representative of the optical catalogs overall. The conclusion should state the conditional nature of the estimate and include the quoted uncertainties or an upper limit.
minor comments (4)
  1. [Section 4] Typo: 'algorithims' should be 'algorithms'. Also in Section 3, 'higher then' should be 'higher than'.
  2. [Table 1] The column headers reuse σ for both the centering scale and the intrinsic scatter of the scaling relations, which is confusing. Suggest explicit labels such as σ_center, σ_TX, σ_ξ and a note that the centering parameters are in Mpc.
  3. [Section 2.1] The X-ray sample selection includes z>0.1 but the analysis is restricted to 0.2<z<0.65. Please clarify whether any X-ray clusters outside this redshift range enter the matching, or whether the z>0.1 cut is simply inherited from the parent catalog.
  4. [Section 2.2] The scaling relation notation in the text is typeset inconsistently (e.g., 'E(z) − 2 3 kBTX'). Please use a consistent form such as $E(z)^{-2/3} k_B T_X$ and define r2500.

Circularity Check

0 steps flagged

No significant circularity: validation tests compare against independent SZ and X-ray benchmarks; no fitted parameter is renamed as prediction.

full rationale

The paper's derivation chain is fully comparative: it measures recovery of SPT SZ-selected clusters, fits centering offsets to X-ray centers, and fits scaling relations between richness and X-ray temperature/SZ significance. None of these quantities is defined in terms of the input observable in a way that forces the reported result. The completeness fraction is simply the matched fraction of external SPT clusters; the paper explicitly acknowledges the caveat that 'the completeness of the optical catalog is not determined overall' (Sec. 2.1), showing the metric is bound to the SPT benchmark and the chosen cuts, not a self-consistent definition. The mis-centering model (two-component gamma, Zhang et al. 2019) and the regression method (Kelly 2007) are standard statistical models adopted from the literature; the fitted parameters ρ, σ, τ, and scatters are estimated from the data, not imported from the citations, and the paper compares them to independent Kelly et al. (2024) results. Self-citations (e.g., Zhang et al. 2019, Hollowood et al. 2019) provide processing tools and an ansatz, but are not load-bearing in the sense of containing the target claims. No prediction reduces to a fit by construction, and no 'uniqueness' theorem is invoked. The '100% recovery above ξ∼10' is a small-sample empirical claim whose statistical robustness is a correctness concern, not a circularity concern.

Axiom & Free-Parameter Ledger

7 free parameters · 5 axioms · 0 invented entities

The central claims are the fitted centering and scatter parameters, which are outputs of the pipeline rather than ad hoc inputs. The load-bearing assumptions are the reliability of the SPT and Chandra benchmarks and the functional forms used for the fits. No new physical entities are introduced.

free parameters (7)
  • Well-centered fraction (rho) = 0.94 +/- 0.07 (redMaPPer), 0.92 +/- 0.07 (WaZP)
    Output of the MCMC fit to optical-X-ray offset distribution (Table 1). This is the central centering result.
  • Well-centered scale (sigma) = 0.067 +/- 0.013 (redMaPPer), 0.050 +/- 0.013 (WaZP)
    Scale parameter of the centered gamma component in the two-component model (Section 2.2).
  • Mis-centered scale (tau) = 0.32 +/- 0.21 (redMaPPer), 0.34 +/- 0.20 (WaZP)
    Scale parameter of the mis-centered gamma component; weakly constrained by small X-ray samples.
  • Intrinsic scatter of (E(z)^-2/3 kBT_X - richness) = 0.27 +/- 0.03 (redMaPPer), 0.29 +/- 0.05 (WaZP)
    Posterior scatter from Kelly (2007) regression, reported in Table 1.
  • Intrinsic scatter of (richness - SPT xi) = 0.33 +/- 0.02 (redMaPPer), 0.46 +/- 0.03 (WaZP)
    Posterior scatter from Kelly (2007) regression, reported in Table 1.
  • Cross-matching radius = 2 Mpc
    Hand-chosen in Section 2.1; affects which clusters are matched and therefore all derived fractions.
  • Cross-matching redshift offset = 0.05
    Hand-chosen in Section 2.1; affects matching completeness and contamination.
axioms (5)
  • domain assumption Optical-to-X-ray offsets follow a two-component gamma distribution (Zhang et al. 2019).
    Used to interpret centering fits in Section 2.2; if the functional form is wrong, the derived fractions are biased.
  • domain assumption SPT SZ significance xi is a reliable mass proxy, and the xi>5 cut gives a complete high-mass cluster sample in the footprint.
    Used for completeness and scatter tests in Section 2.1; completeness is only relative to this sample.
  • domain assumption Chandra X-ray peak and MATCha temperatures are unbiased cluster centers and mass proxies for the selected sample.
    Used as truth for centering and scatter in Section 2.1.
  • standard math Kelly (2007) Bayesian regression correctly separates intrinsic scatter from measurement noise.
    Used for the scaling relation fits in Section 2.2.
  • domain assumption Cross-matching with 2 Mpc and Delta z=0.05 correctly identifies physical counterparts.
    Used in the matching step of Section 2.1; wrong matching biases all results.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Cluster Catalog Validation with Multiwavelength Data." pith.science (2026). https://pith.science/paper/HY77JBKE

@misc{pith2026250907268,
  author       = {Pith},
  title        = {Pith review of: Cluster Catalog Validation with Multiwavelength Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HY77JBKE}},
  note         = {Machine review of arXiv:2509.07268}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The Legacy Survey of Space and Time (LSST) will provide a ground-breaking data set for cosmology, but to achieve the precision needed, the data, data reduction, and algorithms measuring the cosmological data vectors must be thoroughly validated and calibrated. In this note, we focus on clusters of galaxies and present a set of validation tests for optical cluster finding algorithms through comparison to X-ray and Sunyaev-Zel'dovich effect cluster catalogs. As an example, we apply our pipeline to compare the performance of the redMaPPer (red-sequence Matched filter Probabilistic Percolation) and WaZP (Wavelet Z Photometric) cluster finding algorithms on Dark Energy Survey (DES) data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

11 extracted references · 1 canonical work pages

  1. [1]

    2025, Cluster Evaluation Resources (ClEvaR), https://github.com/LSSTDESC/ClEvaR, GitHub

    Aguena, M. 2025, Cluster Evaluation Resources (ClEvaR), https://github.com/LSSTDESC/ClEvaR, GitHub

  2. [2]

    N., et al

    Aguena, M., Benoist, C., da Costa, L. N., et al. 2021, Monthly Notices of the Royal Astronomical Society, 502, 4435–4456, 10.1093/mnras/stab264

  3. [3]

    2025, arXiv e-prints, arXiv:2501.05739, 10.48550/arXiv.2501.05739

    Bechtol , K., Sevilla-Noarbe , I., Drlica-Wagner , A., et al. 2025, arXiv e-prints, arXiv:2501.05739, 10.48550/arXiv.2501.05739

  4. [4]

    E., Stalder , B., de Haan , T., et al

    Bleem , L. E., Stalder , B., de Haan , T., et al. 2015, , 216, 27, 10.1088/0067-0049/216/2/27

  5. [5]

    DES Collaboration , Abbott , T. M. C., Aguena , M., et al. 2025, arXiv e-prints, arXiv:2503.13632, 10.48550/arXiv.2503.13632

  6. [6]

    L., Jeltema , T., Chen , X., et al

    Hollowood , D. L., Jeltema , T., Chen , X., et al. 2019, , 244, 22, 10.3847/1538-4365/ab3d27

  7. [7]

    Kelly, B. C. 2007, The Astrophysical Journal, 665, 1489, 10.1086/519947

  8. [8]

    M., Jobel, J., Eiger, O., et al

    Kelly, P. M., Jobel, J., Eiger, O., et al. 2024, Monthly Notices of the Royal Astronomical Society, 533, 572, 10.1093/mnras/stae1786

  9. [9]

    J., Bocquet , S., & Singh , A

    Klein , M., Hern \'a ndez-Lang , D., Mohr , J. J., Bocquet , S., & Singh , A. 2023, , 526, 3757, 10.1093/mnras/stad2729

  10. [10]

    S., Rozo, E., Busha, M

    Rykoff, E. S., Rozo, E., Busha, M. T., et al. 2014, The Astrophysical Journal, 785, 104, 10.1088/0004-637x/785/2/104

  11. [11]

    L., et al

    Zhang, Y., Jeltema, T., Hollowood, D. L., et al. 2019, Monthly Notices of the Royal Astronomical Society, 487, 2578, 10.1093/mnras/stz1361

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.