{"id":"fb0127aa-cf45-41c1-853f-bdd123dd2ddc","arxiv_id":"1908.07593","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Friends-of-Friends galaxy association on the MDPL2 simulation is heavily contaminated for massive halos in dense environments, and mass-dependent projected-distance and velocity cuts reduce the false positives.","lead":"This paper runs a common galaxy-grouping algorithm on a simulated dark matter universe and measures how often it adds fake members or loses real ones. It matters because galaxy surveys rely on such groupings to measure how clusters affect galaxies, and the paper shows that in dense regions the algorithm adds many false members, biasing velocity dispersions and luminosity functions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-sample tuning and lack of independent evaluation weaken the headline contamination fractions; the 62% and 7x/4x numbers have no error bars.","rationale":"The reader's weakest_assumption (Rockstar is truth) is real but not the most damaging flaw, because the paper explicitly states this assumption and even labels it as an assumption in Section 4.2; every validation study needs a ground truth, and Rockstar's host/satellite split is the natural available definition. The more serious issue is that the hyperparameters (Dp,max, VR,max) were selected on the very same halo catalogue that supplies the ground truth, so the success fractions in Figures 3-5 and the contamination fractions in Figures 6-7 are trained-and-tested-on-the-same-data numbers. The paper presents no train/test split, no independent mock, and no analytic estimate of how much of the 62% contamination would persist under a different but reasonable parameter choice. This matters because the headline claim is quantitative ('roughly 62%', '4 times too many', '7 times too many'), and the paper's own Section 4.5 and Section 4.7 show that results are sensitive to the parameter choices and to the distance/velocity approximations. The Virgo application is a nice sanity check but it does not constrain the contamination fractions; it only shows that the largest group contains M87/M49 and the W/M clouds, which is qualitative. No error bars on the key fractions also means a skeptical reader cannot tell whether 62% is 55% or 70%. I agree with the reader that the paper is CONDITIONAL: the direction of the finding is plausible and consistent, but the exact magnitudes need out-of-sample verification. My concern overlaps with the reader's but focuses on in-sample tuning and missing uncertainty quantification rather than on the Rockstar ground-truth definition.","tokens_in":21604,"tokens_out":2320,"duration_ms":39342,"concrete_test":"Re-run the Section 4.2-4.4 pipeline on an independent N-body mock (e.g., the MultiDark Planck 1 simulation or Bolshoi at z=0) using the exact case (A)/(B)/(C) linking lengths, and recompute the high-density RS/O fractions and the >10^15 Msun/h central/satellite overdensities. If the 62% RS/O<1 fraction and the 4x/7x overdensities do not reproduce within reasonable sample variance, the headline contamination numbers are an artifact of in-sample tuning to MDPL2. A cheaper secondary check: regenerate the K-band HOD with a different luminosity function or a different scatter prescription (e.g., a 0.2 dex log-normal scatter in L at fixed M) and see whether the quoted 62%/4x/7x values shift materially.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—FoF association in dense, cluster-like environments yields severe contamination (RS/O<1 for roughly 62% of massive main halos; about 4x too many centrals and 7x too many satellites for 10^15 Msun/h systems)—rests on comparing FoF output against Rockstar's host/satellite classification. That is an honest, explicit definitional choice, not an internal inconsistency, and the Virgo sanity check is real supporting evidence. However, the quantitative headline is weakened by the paper's methodology: (i) the linking-length pairs in cases (A), (B), and (C) are chosen to maximize success on the very same MDPL2 halo catalogue used for the evaluation; Section 4.2 performs the parameter search on Figure 3 and then Section 4.3 measures success on the same simulated halos. This in-sample tuning can inflate the apparent success of the method, and because contamination and success are complementary measures of the same assignments, the quoted contamination fractions inherit the same optimism. (ii) No error bars or sample-variance estimates are given for the key percentages (62%, 7%, 25%, 4x/7x) or for the RS/O distributions; the fractions are quoted as exact values from one Gpc/h simulation at z=0, and the density bins are not independent draws. (iii) The paper's own limitation statements—e.g., that the assigned luminosities are only indicative, that virial masses are not observable, and that direct comparison with surveys is left to future work—mean the quantitative translation to 'observational catalogs are contaminated' depends on untested assumptions about how the K-band HOD magnitudes and eVR errors map to a real survey. The most load-bearing vulnerability is therefore not the Rockstar definition itself but the in-sample optimization plus the absence of any out-of-sample check: the central numbers could shift substantially when the parameters are applied to an independent simulation, a different halo finder, or a mock built with different HOD assumptions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper tests the standard Friends-of-Friends (FoF) linking-length percolation method for associating galaxies into groups, using the MDPL2 dark matter simulation at z=0. The authors build a mock galaxy catalogue by assigning K-band luminosities to Rockstar halos via a simple HOD, add radial-velocity uncertainties calibrated to 2MRS, and compute projected distances in three Cartesian planes. They select two tuned FoF parameter sets (cases A and B) plus the Crook et al. (2007) set (case C), then quantify success as a function of environmental density. The central finding is that FoF association in dense, cluster-like environments is heavily contaminated: for massive main halos in the densest percentile, roughly 62% have more associated members than true satellites (RS/O<1), and for systems with virial mass above 10^15 Msun/h the method produces about four times too many centrals and seven times too many satellites relative to the Rockstar bound-halo populations. The authors also propose mass-dependent cuts on maximum projected distance and radial velocity (Eq. 3 in Section 4.6) that improve the success fraction, and they apply the method to the Virgo cluster as a sanity check.","tokens_in":22070,"tokens_out":5644,"duration_ms":568309,"significance":"If the central result holds, it is important: raw FoF group catalogues in dense environments are not merely incomplete but heavily contaminated, which biases velocity dispersions, harmonic radii, and conditional luminosity functions. The paper uses a large public simulation, makes its benchmark assumption explicit, and gives a transparent, reproducible analysis pipeline. The Virgo application is a useful qualitative sanity check. However, the quantitative headline is weakened by the fact that the linking-length parameters are tuned on the same catalogue used for evaluation, by the absence of uncertainty estimates on the key percentages, and by the reliance on Rockstar's host/satellite classification as the definition of true membership. These issues do not undermine the qualitative direction of the result, but they do prevent the specific contamination rates from being taken as robust, calibration-free measurements.","major_comments":[{"comment":"Cases (A) and (B) are selected by scanning the mesh in Figure 3 and manually choosing parameters that maximize the success fraction on the very same MDPL2 halo catalogue that is then used to measure the success fractions in Figures 4–5 and the contamination statistics in Figures 6–7 and Section 4.4. Because the headline numbers (62% of massive main halos with RS/O<1 in the densest environments, and the 4x/7x central/satellite overcounting for 10^15 Msun/h systems) are measured on the same data used for tuning, they are in-sample estimates rather than independent measures of FoF performance. Please provide an out-of-sample evaluation, for example by splitting the simulation volume into independent subvolumes, using a different simulation, or withholding a subset of halos, and report how the contamination fractions vary with plausible parameter choices.","section":"§4.2, §4.3"},{"comment":"The key percentages and distribution properties in Section 4.3 and Section 4.4 are quoted as exact values from one Gpc/h simulation at z=0, with no error bars or sample-variance estimates. The environmental density bins are not independent draws, and the three projection planes are correlated. Please add bootstrap or jackknife uncertainties over subvolumes (or over the three projection planes) for the quoted fractions such as 62%, 7%, 25%, 4x, and 7x, and for the reported means/dispersions of the RS/O distributions. Without these, the reader cannot assess whether the differences between environmental bins are significant.","section":"§4.3, §4.4"},{"comment":"The constraints in Equation (3) are fitted to the 95th percentiles of DMh,max and ΔVR,max measured from the same MDPL2 simulation, and then the percolation algorithm is rerun on the same halos to demonstrate the improvement in Figure 15. This is an in-sample evaluation: the fitted curves absorb noise and peculiarities of the particular simulation, so the reported improvement is likely overestimated. Please validate the constraints on an independent simulation or an independent subvolume, propagate the fit-parameter uncertainties into the classification, and report how the improvement depends on the choice of percentile (e.g., 90th versus 99th).","section":"§4.6, Eq. (3)"},{"comment":"The evaluation treats Rockstar's host/satellite association as the correct group membership, as stated in Section 4.2 ('Assuming the association of halos in the simulation as the correct one'). This is an explicit and honest definitional choice, but it means the quoted contamination rates are relative to Rockstar's halo-finder classification, not directly to physical galaxy-group membership. Infalling halos near cluster outskirts are particularly ambiguous, and a different halo finder or a different definition of 'true satellite' could shift the quoted percentages. Please discuss the robustness of the contamination fractions to this definitional ambiguity, or provide a side test using a different association definition (for example, splashback-based or phase-space based membership) for a subset of halos.","section":"§4.2, §4.3"}],"minor_comments":[{"comment":"The phrase 'In order to constraint the limitations' should read 'In order to constrain the limitations'.","section":"Abstract"},{"comment":"The y-axis labels 'n [Mpc−3]' appear to describe a count of neighbor halos within 1.75 Mpc/h, which is not a number density. Please relabel the axis as a count (N) or convert to a true number density.","section":"Figure 6"},{"comment":"In the paragraph discussing Figure 9, the text states 'Dp,max = 525 Mpc, i.e., ≈ 355 Mpc h−1'; the units should be kpc, not Mpc, and the conversion as written is inconsistent.","section":"§4.4"},{"comment":"The caption refers to 'Section sec.param' instead of the actual section number where the linking-length parameters are defined.","section":"Figure 17 caption"},{"comment":"The statement that the 95th percentile 'barely matches the 2σ deviation' is imprecise; for a Gaussian distribution the 95th percentile is approximately 1.645σ, not 2σ.","section":"§4.6"},{"comment":"The reference 'York D. G. S. 2000' should be formatted as York et al. (2000), and the entry for Ivezić has a typographical issue in the author name.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of MNRAS and addresses a question of practical importance for group-finder validation. The central qualitative conclusion—FoF group catalogues in dense environments are heavily contaminated—is well supported by the internal comparisons and is reinforced by the Virgo sanity check. The main obstacle to acceptance is the in-sample nature of both the parameter selection and the Section 4.6 constraints; an out-of-sample test or a clear reframing of the numbers as conditional on the tuned parameters is needed. I would not reject the paper on the current evidence, but the quantitative claims should be softened or validated before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline first: this is an honest, well-scoped calibration study of the Friends-of-Friends group finder on MDPL2, and the qualitative result—FoF heavily contaminates the membership of massive halos in dense environments even after the linking lengths are tuned—is almost certainly right. What should not be repeated verbatim are the precise percentages (62% of massive main halos over-associated; roughly 4x/7x excess for 10^15 Msun/h clusters). They carry no error bars and were measured on the same simulation used to pick Dp,max and VR,max.\n\nWhat is genuinely new and useful: quantitative contamination/completeness maps for FoF split by environmental density and mass from a 1 Gpc/h simulation; mass-dependent constraints on projected distance and velocity difference (Eq. 3) that visibly reduce misclassification; and a real-data application to Virgo that recovers M87/M49, the W and M clouds, and NGC 4636. The authors are commendably explicit that Rockstar's host/satellite assignment is their definition of truth, and they run controlled checks: velocity uncertainties change only ~1% of classifications, and the distance-estimation approximation matters only in the nearest velocity bin (~13%). The fact that the 62% RS/O<1 fraction appears independently for both the 10^11 and 10^12 Msun/h cuts (cases A and B) is good internal consistency.\n\nThe soft spots are real but tractable. Linking lengths are hand-picked to maximize success on the same MDPL2 catalogue used for the evaluation, and the Section 4.6 constraints are fitted and re-tested on that same catalogue. Standard practice in this literature, but it means the headline numbers inherit the tuning; an independent simulation or halo finder could shift them. No error bars appear anywhere: the 62%, 4x, 7x, and 25% figures are exact values from one box at one snapshot, with density bins that are not independent. And the mock survey is idealized—complete to a mass limit, no flux limit or geometry—so the translation to SDSS/LSST rests on assumptions left untested. The Rockstar-as-truth caveat is a footnote rather than a flaw: it is explicit, conventional, and the paper's internal logic does not depend on it.\n\nWho gets value: anyone calibrating or validating group finders for upcoming surveys, and anyone needing a compact demonstration that FoF-based velocity dispersions and cluster luminosity functions can be seriously biased. I would not build on the exact numbers without independent confirmation, but the paper deserves a serious referee. The requests should be an out-of-sample sanity check and error bars; with those, the quantitative claims would be supportable.","headline":"Honest FoF calibration study whose qualitative contamination result holds, but whose headline percentages lack error bars and come from in-sample tuning.","tokens_in":22561,"tokens_out":7018,"would_cite":false,"duration_ms":65253,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Friends-of-Friends galaxy grouping, the standard percolation method for identifying groups and clusters in surveys, systematically over-associates members around massive halos in dense environments, inflating cluster-scale systems to…","keywords":["Friends-of-Friends method","percolation","galaxy groups","galaxy clusters","contamination","MDPL2 simulation","halo association","velocity dispersion"],"falsifier":"Redo the same FoF runs on MDPL2 but replace the simulation's host/satellite tree with a different subhalo-tracking definition of true membership, then recompute the fraction of massive main halos with $\\mathrm{RS/O}<1$ in the densest environmental bin; if the fraction drops from roughly 62% to below about 20%, the claimed contamination is an artefact of the chosen truth definition rather than of FoF itself.","tokens_in":21412,"feed_emoji":"🔭","tokens_out":10261,"duration_ms":93923,"temperature":0.7,"pith_summary":"This paper asks whether Friends-of-Friends (FoF) percolation, the standard way to assign galaxies to groups and clusters in redshift surveys, returns the right memberships when observational noise is present. Using the MDPL2 cosmological dark matter simulation at z=0, the authors add realistic radial-velocity errors and projected-distance uncertainties to the halo catalogue, run FoF with several linking-length parameter sets, and compare each output association against the simulation's own bound-halo structure. They find that FoF is reliable in low-density environments but systematically over-associates in dense ones: for cluster-like regions roughly 62% of massive main halos acquire more members than they truly have, and simulated clusters above $10^{15}\\,h^{-1}M_\\odot$ are inflated to about four times as many centrals and seven times as many satellites. The paper also shows that halo-mass-dependent cuts on projected distance and radial velocity, plus substructure tests, can reduce the fake-positive rate, and it applies the tuned method to the Virgo cluster region as a sanity check.","feed_headline":"4x centrals, 7x satellites: group finder overfills clusters","feed_subtitle":"In a simulation test, 62% of massive halos in dense regions gain spurious members.","key_machinery":"The machinery is the Friends-of-Friends linking-length percolation algorithm, which links galaxies whose projected separation and radial-velocity difference both fall below thresholds $D_{\\mathrm{p,max}}$ and $V_{\\mathrm{R,max}}$. The central diagnostic is the ratio $\\mathrm{RS/O}$: the number of satellites that the simulation's halo finder assigns to a main halo divided by the number of halos that FoF associates to it, so values below one count fake positives. The other load-bearing element is the truth definition: the halo finder's host/satellite tree in the MDPL2 simulation, taken as the correct galaxy-group membership. Around this, the paper builds empirical 95th-percentile fits for maximum projected distance and radial velocity difference as functions of main-halo mass, and uses them as post-processing cuts together with a nearest-neighbour substructure test to flag or remove contaminated systems.","core_discovery":"The central discovery is that the percolation/FoF association method, even with linking lengths optimised on the simulation, is not uniformly accurate: its success depends strongly on environment. In the densest environmental percentiles, typical of rich groups and clusters, about 62% of main halos with at least two satellites have $\\mathrm{RS/O}<1$, meaning the method attaches more satellites than are gravitationally bound, with a mean ratio of 0.5 and scatter 0.22. When the FoF output is used to define $>10^{15}\\,h^{-1}M_\\odot$ clusters, the method returns roughly four times the true number of central galaxies and seven times the true number of satellites; the excess members are mostly low-mass field halos that contribute little mass but inflate velocity dispersions and distort conditional luminosity functions. The authors argue that these are not artefacts of the added velocity uncertainties, which change only about 1% of classifications, but consequences of the method's geometry in high-density regions, and that applying cuts derived from the simulation, such as the 95th percentile of maximum projected satellite distance and radial velocity difference as functions of main-halo mass, markedly improves the recovered group properties.","pith_inferences":["A direct test of the paper's diagnosis would be to run the same FoF parameters on the same MDPL2 halo catalogue but replace the simulation's host/satellite labels with a different subhalo-tracking definition of membership; if the dense-environment contamination fraction drops below about 20%, the reported 62% is particular to the chosen truth definition rather than to FoF itself.","The fake positives are mostly low-mass field halos projected near massive cluster halos, which suggests contamination concentrates in the infall region around clusters; a membership criterion based on infall dynamics or splashback radius could remove most spurious members without the paper's mass-dependent cuts.","Because contamination is environment-dependent and strongest exactly where cluster-versus-field comparisons are made, the paper's numbers imply that some observed environmental trends in galaxy properties could be inflated by FoF selection artefacts; testing this would require redoing those comparisons with the paper's cleaned memberships.","The same methodology could be applied to other association algorithms, such as Bayesian or clustering-based group finders, to see whether the contamination is a generic property of projection-based grouping or specific to percolation with fixed linking lengths."],"forward_implications":["FoF-based group and cluster catalogues in dense environments carry a large fraction of fake members, so observables computed from those members, such as velocity dispersions, harmonic radii, and conditional luminosity functions, are biased even when the linking parameters are chosen for maximum overall success.","Counting systems with FoF-defined virial masses above $10^{15}\\,h^{-1}M_\\odot$ overproduces clusters by a factor of roughly four in centrals and seven in satellites, so richness- and luminosity-based cluster samples built this way are substantially contaminated at the high-mass end.","Applying the paper's mass-dependent cuts on maximum projected distance and radial velocity difference, together with substructure tests, can reduce but not eliminate the contamination; the substructure test only flags a third to half of the contaminated systems at moderate confidence.","Low-density environments are much less affected, so FoF results for field galaxies and small groups are more trustworthy than those for clusters.","Radial-velocity uncertainties are not the source of the contamination, since removing them changes only about 1% of classifications; the geometry of the linking thresholds in dense regions drives the fake-positive rate."],"supporting_citations":[{"why":"Provides the halo finder whose host/satellite assignments supply the true group membership against which the FoF output is compared.","marker":"Behroozi et al. 2013"},{"why":"Provides the MDPL2 simulation and its halo catalogue, the data set on which the whole test is built.","marker":"Klypin et al. 2016"},{"why":"Supplies the FoF algorithm implementation and the linking-length parameter pairs the paper adopts, including the low-density-contrast case.","marker":"Crook et al. 2007"},{"why":"Defines the linking-length percolation method for grouping galaxies that the paper is testing.","marker":"Huchra & Geller 1982"},{"why":"Supplies the 2MASS Redshift Survey velocity-uncertainty measurements used to assign realistic radial-velocity errors to simulated halos.","marker":"Huchra et al. 2012"},{"why":"Provides the K-band luminosity function used in the HOD to assign magnitudes, which drives the simulated velocity uncertainties.","marker":"Kochanek et al. 2001"},{"why":"Provides the nearest-neighbour substructure test the paper uses to identify contaminated systems in a fraction of cases.","marker":"Colless & Dunn 1996"},{"why":"An earlier percolation-method accuracy test on a mock SDSS-like catalogue, establishing the baseline of completeness and contamination analysis that this paper extends with environment-resolved statistics.","marker":"Duarte & Mamon 2014"}],"fun_headline_variants":["FoF group finder inflates cluster members 4x-7x","Dense clusters: FoF adds spurious satellites","FoF overfills dense groups: 4x centrals, 7x satellites","Spurious members dominate FoF groups in clusters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole accuracy calculation assumes the simulation's own halo-finder grouping, the host/satellite labels, is the correct galaxy-group membership; if that grouping is ambiguous, especially for infalling galaxies near cluster edges, the quoted contamination fractions measure disagreement with the halo finder rather than error against physical groups.","fun_headline_variants_meta":{"raw":{"variants":["FoF group finder inflates cluster members 4x-7x","Dense clusters: FoF adds spurious satellites","FoF overfills dense groups: 4x centrals, 7x satellites","Spurious members dominate FoF groups in clusters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000843,"raw_usage":{"total_tokens":3651,"prompt_tokens":901,"completion_tokens":2750,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":2676}},"tokens_in":517,"tokens_out":2750,"duration_ms":18878,"temperature":1.0,"reasoning_tokens":2676,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:01:59.871003+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Redo the same FoF runs on MDPL2 but replace the simulation's host/satellite tree with a different subhalo-tracking definition of true membership, then recompute the fraction of massive main halos with $\\mathrm{RS/O}<1$ in the densest environmental bin; if the fraction drops from roughly 62% to below about 20%, the claimed contamination is an artefact of the chosen truth definition rather than of FoF itself.","supporting_citations":[{"cited_title":"A., 2014, @doi [ ] 10.1093/mnras/stu378 , http://adsabs.harvard.edu/abs/2014MNRAS.440.1763D 440, 1763","cited_arxiv_id":null,"evidence_quote":"An earlier percolation-method accuracy test on a mock SDSS-like catalogue, establishing the baseline of completeness and contamination analysis that this paper extends with environment-resolved statistics."}],"review_version":1}