Pith. sign in

REVIEW 3 major objections 3 minor 19 references

Improving the accuracy of observable distributions for galaxies classified in the Projected Phase Space Diagram

T0 review · 3 major / 3 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper proposes and tests a statistical decontamination method that inverts a confusion matrix to recover intrinsic galaxy colour distributions, improving the classes most affected by misclassification in projected phase space.

desk verdict A useful, honest application of confusion-matrix inversion to PPSD galaxy classification — the simulated test is credible, the unvalidated transfer to observed clusters is the real soft spot. read the letter →

arxiv 2502.04446 v2 pith:LKIWP6IU submitted 2025-02-06 astro-ph.GA

classification astro-ph.GA
keywords galaxyclustersprojectedphasespaceconfusionmatrixmisclassificationbacksplashgalaxiescoloursformationmodelsmachinelearningclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Galaxies are often assigned to dynamical classes—cluster members, backsplash galaxies, recent infallers, infalling galaxies, interlopers—from their position in the projected phase-space diagram (cluster-centric distance versus line-of-sight velocity). These assignments are imperfect, and contamination among classes biases any measurement of class properties such as colour. This paper proposes a statistical correction: estimate the confusion matrix (the table of how often each true class is assigned to each predicted class) from a simulated cluster sample, then invert that matrix and apply it to the observed colour distributions. On simulated and observed data the corrected distributions are closer to the intrinsic ones for the three most contaminated classes—cluster members, backsplash galaxies, and recent infallers—while the already well-classified classes are essentially unchanged. The same recipe works for any projected-phase-space classification scheme and any galaxy property, as long as a confusion matrix can be estimated.

What carries the argument

The machinery is the $n\times n$ confusion matrix $C$, with $C_{ij}$ the fraction of galaxies of true class $j$ that are labelled as predicted class $i$. If $f(M)$ is the matrix whose rows give the observed colour distributions of predicted classes at a fixed stellar mass $M$, and $F(M)$ is the corresponding matrix of intrinsic distributions, the classification step implies $f(M)=C\,F(M)$, so $F(M)=C^{-1}f(M)$ whenever $C$ is invertible. The matrix is estimated from a test set of simulated galaxies whose true dynamical classes are known from their orbital histories and whose predicted classes come from a machine-learning classifier plus a threshold criterion. The paper also checks a stellar-mass-dependent version $C(M)$ and finds it unnecessary, since it changes the recovered distributions only slightly.

What would settle it

Derive an independent membership classification for the same observed galaxies, for example by using spectroscopic redshifts inside an escape-velocity boundary rather than the projected-phase-space classifier, and compare the colour distribution of those members with the recovered cluster-member distribution; a disagreement larger than the residual differences reported in the paper would show that the simulation-derived confusion matrix is not transferable to the observed sample.

Watch

Extended reading notes

Core claim

The central claim is that the observed colour distribution of a predicted class is a linear mixture of the intrinsic distributions of all true classes, weighted by the confusion matrix, so multiplying the observed matrix of distributions by the inverse confusion matrix recovers the intrinsic distributions. Tested on a large simulated cluster population, the inversion removes most of the blue contamination that plagues predicted cluster members and backsplash galaxies and the red contamination that plagues recent infallers. The result is a colour distribution significantly closer to the true one for these classes, measured by summed squared residuals, but no improvement for infalling galaxies and interlopers, whose predicted classes are already clean. Applied to an observed sample of X-ray clusters, the same correction indicates that blue, low-mass galaxies in clusters are almost exclusively recent infallers not yet quenched, and that backsplash galaxies are on average redder than raw classifications suggest.

Load-bearing premise

The load-bearing assumption is that the misclassification rates measured in the simulated clusters match the misclassification rates in the observed galaxies; if real clusters mix the classes with different probabilities, inverting the simulation's confusion matrix will move the observed distributions in the wrong direction.

Editorial extensions

If this is right

  • Colour distributions of cluster members, backsplash galaxies, and recent infallers in projected-phase-space studies should be treated as contaminated unless corrected through a confusion-matrix inversion.
  • The blue, low-mass galaxy population inside clusters is essentially made of recent infallers that have not yet been quenched, so its presence traces infall rather than an in-situ star-forming population.
  • Backsplash galaxies are intrinsically redder than they appear: a single passage through the cluster noticeably quenches their star formation.
  • Because the method works for any property and any projected-phase-space classification scheme, it provides a general post-processing step for measurements of specific star formation rate, stellar mass functions, or sizes in and around clusters.
  • For cleanly classified classes such as infalling galaxies and interlopers, the inversion is unnecessary; using it does not degrade the distributions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the same inversion could be used to correct not only property distributions but also class fractions themselves, enabling unbiased estimates of the relative abundances of cluster members versus infallers as a function of mass.
  • Editorial inference: because the confusion matrix depends on cluster mass range and training choices, the method would be strengthened by constructing a separate matrix for the cluster mass and redshift of each observed sample; the single simulated matrix used here may not transfer to lower-mass groups or higher redshifts.
  • Editorial inference: a natural test is to compare the recovered member colours with an independent membership indicator such as an escape-velocity boundary; agreement would validate the simulation-based matrix, while disagreement would pinpoint where the assumption needs recalibration.
  • Editorial inference: the recovered intrinsic distributions could be fed into quenching models as empirical priors, turning the qualitative statement that backsplash galaxies are redder into a quantitative quenching-timescale constraint.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper proposes a method to correct the measured distributions of galaxy properties (specifically colour) for misclassification in classifications based on the projected phase space diagram (PPSD). The method uses a confusion matrix C estimated from simulations: the observed distributions of predicted classes f are related to intrinsic distributions F through f = C F, so the intrinsic distributions are recovered by inverting C (Eq. 5). The authors test the method on simulated clusters from MultiDark-SAG with known intrinsic classes, reporting improvements in the recovered colour distributions for cluster members, backsplash galaxies, and recent infallers. They then apply the method to a sample of observed X-ray clusters and SDSS galaxies.

Significance. If the method is shown to be robust, it provides a simple and useful correction for a known problem in PPSD-based galaxy classification, and it can be applied to any classification scheme and any galaxy property, given an estimated confusion matrix. The authors make their code available and provide a concrete demonstration on simulations where the true intrinsic classes are known. However, the strength of the central claim depends on whether the validation demonstrates performance on independent data, which is currently not the case. The absence of an out-of-sample test and of a treatment of negative bin counts are the main gaps. The observed application also relies on the transferability of a simulation-derived confusion matrix to observed clusters, a step the authors acknowledge is uncertain.

major comments (3)
  1. [Sect. 3.2] The improvement claimed in the abstract is supported only by an in-sample test: the confusion matrix C and the predicted-class distributions f are both computed from the same test set (Sect. 3.1 and 3.2), and the residuals S_rec and S_pred in Eqs. (6)-(7) are evaluated on that same test set. When C is estimated from the same galaxies used to construct f, the inversion in Eq. (5) can absorb sampling fluctuations in f, making the reconstructed F spuriously close to the empirical intrinsic distributions. To establish that the method improves distributions for independent data, the authors should split the test set (or use a separate simulated cluster sample) to estimate C on one subset and evaluate the inversion on a disjoint subset. Without such an out-of-sample test, the central claim is not fully supported.
  2. [Sect. 2, Eq. (5)] The inversion F = C^{-1} f can produce negative bin counts when empirical C and f contain sampling noise. The paper does not discuss whether negative bins occur, how they are treated, or whether the recovered distributions are renormalized after inversion. Negative counts are unphysical and could bias the residuals used to claim improvement. The authors should state whether any negative values appear, and if so, describe the method used to handle them, or alternatively use a constrained inversion that enforces non-negativity.
  3. [Sect. 4] The confusion matrix estimated from MultiDark-SAG clusters in Sect. 3.2 is applied to observed X-ray clusters without calibration against independent spectroscopic classifications. The authors explicitly note in Sect. 3.3 that C may depend on cluster mass range, classification thresholds, and training choices, and that these factors 'are likely to have a considerable effect.' Nevertheless, the observed colour distributions presented in Fig. A.1 are interpreted as physical results (e.g., 'backsplash galaxies are on average redder than expected'). The reliability of these observed conclusions depends on the transferability of C, which is not demonstrated. The paper should either include an explicit caveat that the observed results are illustrative and model-dependent, or provide some validation of C on observed galaxies with known classifications.
minor comments (3)
  1. [Fig. 1] The caption states that columns correspond to intrinsic classes and rows to predicted classes, but the displayed matrix has row sums of unity, which corresponds to the composition of predicted classes (i.e., P(intrinsic | predicted)). This differs from the definition of C in Eq. (1), where columns are intrinsic classes and column sums should be unity. Please clarify the orientation of the matrix in the figure and ensure the matrix used in Eq. (5) is defined consistently.
  2. [Sect. 3.1] The sentence 'For the sake of consistency, we re-trained roger accordingly' is vague; please specify the training and test split sizes explicitly.
  3. [Eqs. (6)-(7)] The residuals S_pred and S_rec are sums of squared differences without uncertainties. Given the finite sample sizes, providing bootstrap or jackknife confidence intervals on these residuals would help distinguish genuine improvement from noise, especially in bins where the differences are small.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the confusion-matrix inversion is an empirical deconvolution, though its simulated validation is in-sample and the C transfer to observations is assumed.

full rationale

The derivation chain is a linear deconvolution: Eq. (2) writes the predicted-class colour distributions f as C×F, and Eq. (5) recovers the intrinsic distributions as F = C^{-1}×f. The confusion matrix C is measured in Sect. 3.2 from MultiDark-SAG test galaxies using their true orbital classes, not from the colour distributions being corrected, so the improvement reported in Figs. 2 and 3 (S_rec < S_pred, Eqs. 6 and 7) is an empirical comparison against the true simulated intrinsic distributions rather than an identity. The observed-sample application in Sect. 4 uses the same simulation-derived C and does not fit anything to the SDSS data. The main caveats are statistical rather than circular: C and f are both derived from the same test set, making the simulated validation in-sample and potentially optimistic, and the transfer of C from MultiDark-SAG to SDSS X-ray clusters is an extrapolation. The paper itself acknowledges in Sect. 3.3 that C depends on the classifier training and cluster mass range, but this is a robustness limitation, not a circular reduction. Self-citations to ROGER (de los Rios et al. 2021) and the Coenda et al. (2022) classification scheme are prior published tools; Sect. 5 explicitly states the method applies to any PPSD classification, so these citations are not load-bearing in a way that forces the conclusions. No step in the derivation reduces to its own input by construction.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The method introduces no new physical entities; it is a statistical post-processing step. Its outputs depend on user-chosen thresholds and analysis bins, and on two unvalidated assumptions: the transferability of the simulation-derived confusion matrix to observations, and its independence from the reconstructed property (colour).

free parameters (3)
  • Classification thresholds = 0.4 (CL), 0.48 (BS), 0.37 (RIN), 0.54 (IN), 0.15 (ITL)
    User-chosen thresholds on ROGER probabilities defining predicted classes (Sect. 3.1, from Coenda et al. 2022). The confusion matrix, and therefore every recovered distribution, depends on these values.
  • Stellar mass bins = Simulation: [9.1,9.4], [9.4,9.7], [9.7,10.1], [10.1,10.5], [10.5,11.5] in log(M/h^-1 M_sun); Observation: four bins in…
    Analysis bins within which the method is applied; they set the resolution of the results but are not part of the inversion itself.
  • ROGER probability source = K-nearest neighbours
    Among the three ROGER techniques (KNN, SVM, random forest), the authors use KNN probabilities for classification (Sect. 3.1). This choice affects C and the recovered distributions.
assumptions (3)
  • domain assumption The MultiDark Planck 2 + SAG simulation reproduces the orbital kinematics and colour-dependent properties of real cluster galaxies well enough that the confusion matrix computed from simulated test galaxies applies to observed SDSS galaxies.
    This transferability is the bridge used in Sect. 4. The paper acknowledges in Sect. 3.3 that 'the choice of the mass range of the clusters used to train roger is likely to have a considerable effect', but does not validate C against observations.
  • ad hoc to paper The confusion matrix C is independent of the galaxy property x (colour) being reconstructed; misclassification probabilities are the same for red and blue galaxies.
    Eq. (2) uses a single C for all colour bins. The paper tests mass dependence (Sect. 3.3) but not colour dependence, even though PPSD position correlates with colour.
  • standard math The solution F = C^{-1} f is a physically valid probability distribution; negative bin counts are either absent or handled.
    Inverting a noisy or inconsistent C can produce negative counts. The paper does not describe clipping or regularisation after Eq. (5).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving the accuracy of observable distributions for galaxies classified in the Projected Phase Space Diagram." pith.science (2026). https://pith.science/paper/LKIWP6IU

@misc{pith2026250204446,
  author       = {Pith},
  title        = {Pith review of: Improving the accuracy of observable distributions for galaxies classified in the Projected Phase Space Diagram},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LKIWP6IU}},
  note         = {Machine review of arXiv:2502.04446}
}
read the original abstract

Studies of galaxy populations classified according to their kinematic behaviours and dynamical state using the Projected Phase Space Diagram (PPSD) are affected by misclassification and contamination, leading to systematic errors in determining the characteristics of the different galaxy classes. We propose a method to statistically correct the determination of galaxy properties' distributions accounting for the contamination caused by misclassified galaxies from other classes. Using a sample of massive clusters and galaxies in their surroundings taken from the MultiDark Planck 2 simulation combined with the semi-analytic model of galaxy formation SAG, we compute the confusion matrix associated to a classification scheme in the PPSD. Based on positions in the PPSD, galaxies are classified as cluster members, backsplash galaxies, recent infallers, infalling galaxies, and interlopers. This classification is determined using probabilities calculated by the code ROGER, along with a threshold criterion. By inverting the confusion matrix, we are able to get better determinations of distributions of galaxy properties such as colour. Compared to a direct estimation based solely on the predicted galaxy classes, our method provides better estimates of the mass-dependent colour distribution for the galaxy classes most affected by misclassification: cluster members, backsplash galaxies, and recent infallers. We apply the method to a sample of observed X-ray clusters and galaxies. Our method can be applied to any classification of galaxies in the PPSD, and to any other galaxy property besides colour, provided an estimation of the confusion matrix. Blue, low-mass galaxies in clusters are almost exclusively recent infaller galaxies that have not yet been quenched by the environmental action of the cluster. Backsplash galaxies are on average redder than expected.

Figures

Figures reproduced from arXiv: 2502.04446 by the authors.

Figure 1
Figure 1. Confusion matrix for our adopted classification scheme (Coenda et al. 2022). Columns correspond to intrinsic classes, rows to predicted classes. These thresholds are 0.4, 0.48, 0.37, 0.54, and 0.15, for CLs, BSs, RINs, INs, and ITLs, respectively. This scheme is an ap￾propriate balance between sensitivity and precision. 3.2. Applying the method As a first step, we constructed the confusion matrix using the test set … view at source ↗
Figure 2
Figure 2. Colour distributions of galaxies. Each class is shown in a different row, as noted to the right of the panels. Each column considers a particular range of galaxy stellar mass, as noted at the top. Grey shaded histograms are the real colour distributions, i.e. they correspond to the intrinsic classes. Green lines are the colour distributions of the predicted classes. Violet lines are the colour distributions recovere… view at source ↗
Figure 3
Figure 3. Sum of square residuals as a function of galaxy stellar mass. Each panel considers a different galaxy class. Residuals between the real and the predicted distributions are shown with green lines, i.e. S pred in Eq. 6. Residuals between the real and the recovered distributions are shown with violet lines, i.e. S rec in Eq. 7. The residuals between the real and the recovered distributions when we use a stellar mass-de… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

19 extracted references · 11 canonical work pages

  1. [1]

    N., Adelman-McCarthy, J

    Abazajian, K. N., Adelman-McCarthy, J. K., Agüeros, M. A., et al. 2009, ApJS, 182, 543

  2. [2]

    Aguerri, J. A. L., Cuomo, V ., Rojas-Roncero, A., & Morelli, L. 2023, A&A, 679, A5 Aldás, F., Gómez, F. A., Vega-Martínez, C., Zenteno, A., & Carrasco, E. R. 2024, arXiv e-prints, arXiv:2408.05305 Aldás, F., Zenteno, A., Gómez, F. A., et al. 2023, MNRAS, 525, 1769

  3. [3]

    P., et al

    Bohringer, H., V oges, W., Huchra, J. P., et al. 2000, VizieR Online Data Catalog, 212, 90435

  4. [4]

    Choi, H. & Yi, S. K. 2017, ApJ, 837, 68

  5. [5]

    2022, MNRAS, 510, 1934

    Coenda, V ., de los Rios, M., Muriel, H., et al. 2022, MNRAS, 510, 1934

  6. [6]

    & Muriel, H

    Coenda, V . & Muriel, H. 2009, A&A, 504, 347

  7. [7]

    A., Vega-Martínez, C

    Cora, S. A., Vega-Martínez, C. A., Hough, T., et al. 2018, MNRAS, 479, 2 de los Rios, M., Martínez, H. J., Coenda, V ., et al. 2021, MNRAS, 500, 1784 Hernández-Fernández, J. D., Haines, C. P., Diaferio, A., et al. 2014, MNRAS, 438, 2186

  8. [8]

    A., Haggar, R., et al

    Hough, T., Cora, S. A., Haggar, R., et al. 2023, MNRAS, 518, 2398 Jaffé, Y . L., Smith, R., Candlish, G. N., et al. 2015, MNRAS, 448, 1715

Show all 19 references
  1. [9]

    2016, MNRAS, 457, 4340

    Klypin, A., Yepes, G., Gottlöber, S., Prada, F., & Heß, S. 2016, MNRAS, 457, 4340

  2. [10]

    A., & Raychaudhury, S

    Mahajan, S., Mamon, G. A., & Raychaudhury, S. 2011, MNRAS, 416, 2882 Martínez, H. J., Coenda, V ., Muriel, H., de los Rios, M., & Ruiz, A. N. 2023, MNRAS, 519, 4360 Muñoz Rodríguez, I., Georgakakis, A., Shankar, F., et al. 2024, MNRAS, 532, 336

  3. [11]

    & Coenda, V

    Muriel, H. & Coenda, V . 2014, A&A, 564, A85

  4. [12]

    Muzzin, A., van der Burg, R. F. J., McGee, S. L., et al. 2014, ApJ, 796, 65

  5. [13]

    Oman, K. A. & Hudson, M. J. 2016, MNRAS, 463, 3083

  6. [14]

    2019, MNRAS, 484, 1702

    Pasquali, A., Smith, R., Gallazzi, A., et al. 2019, MNRAS, 484, 1702

  7. [15]

    Popesso, P., Böhringer, H., Brinkmann, J., V oges, W., & York, D. G. 2004, A&A, 423, 449

  8. [16]

    2017, ApJ, 843, 128

    Rhee, J., Smith, R., Choi, H., et al. 2017, ApJ, 843, 128

  9. [17]

    N., Martínez, H

    Ruiz, A. N., Martínez, H. J., Coenda, V ., et al. 2023, MNRAS, 525, 3048

  10. [18]

    M., de Carvalho, R

    Sampaio, V . M., de Carvalho, R. R., Aragón-Salamanca, A., et al. 2024, MN- RAS, 532, 982

  11. [19]

    2019, ApJ, 876, 145 Article number, page 7 of 8 A&A proofs: manuscript no

    Smith, R., Pacifici, C., Pasquali, A., & Calderón-Castillo, P. 2019, ApJ, 876, 145 Article number, page 7 of 8 A&A proofs: manuscript no. aa51999-24 Appendix A: Additional figure 0 1 2 3Normalised counts log(M/h 1M ) [9.5-10.0] pred rec log(M/h 1M ) [10.0-10.5] log(M/h 1M ) [1...

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.