{"id":"cdda5a08-e7b7-44c4-859b-ebeaeb7578c4","arxiv_id":"2411.16455","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"In a GAMA-like simulated survey, all three optical group finders recover over 80% of halos above 10^13 solar masses, with stellar luminosity the most reliable mass proxy.","lead":"This paper runs three widely used galaxy group finders on the same simulated GAMA-like spectroscopic survey to measure how completely and cleanly each recovers the underlying dark matter halos. The results tell astronomers which group catalogues they can trust for studies of galaxy evolution and for stacking X-ray data from eROSITA.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Completeness and contamination claims rest on an unvalidated matching rule (Section 4.1) whose permissiveness is exposed by the paper's own 50% membership-overlap metric, which drops completeness to about 70%; the >80% headline needs a sensitivity test.","rationale":"After reading the paper in good faith, I agree with the reader that the unvalidated matching rule in Section 4.1 is the single most load-bearing assumption. The headline results are not raw detection counts; they are the output of a classification procedure whose thresholds (R200, 3e-3, factor of six) are chosen for plausibility but never varied. The paper's own alternative definition of a 'good' detection, requiring 50% of members to overlap, lowers completeness to about 70%, showing that the 80% figure is sensitive to how a match is defined. Since the same matching rule underlies the UniverseMachine test in Appendix A, that external check does not break the dependence on the permissive matching. The paper is transparent about both definitions, and the qualitative conclusions (optical selection works, luminosity proxies are more stable than velocity dispersion) are likely robust, so I do not recommend rejection. The verdict should remain CONDITIONAL: the authors should add a matching-rule sensitivity analysis, which is inexpensive, or explicitly present the membership-overlap completeness as the primary metric.","tokens_in":24442,"tokens_out":8594,"duration_ms":78676,"concrete_test":"Re-run the matching procedure of Section 4.1 on the same LC30 catalogues with three variants: (i) cylinder radius 0.5 times R200, (ii) redshift window of 1e-3 (about 300 km/s), and (iii) dominance ratios of 3 and 10, recomputing completeness and contamination as in Fig. 3 for all three group finders. If completeness in the bin log10(M200/M_sun) in [13.0, 13.5] stays above 80% and contamination stays below 10% for all variants, the headline claim is robust to the matching definition; if any finder violates these thresholds in a plausible variant, the abstract should be requalified to state the matching-rule dependence explicitly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims of the abstract (greater than 80% completeness, less than 10% contamination for M200 > 10^13 solar masses) are defined through the matching rule of Section 4.1: a detection counts as correct if its centre lies within R200 of a true halo and within a redshift window of 3e-3 (about 900 km/s, comparable to a massive cluster's velocity dispersion), and when multiple halos fall in the cylinder the most massive is chosen unless a factor-of-six dominance holds. This rule is plausible but entirely hand-chosen, with no sensitivity analysis, and it is reused for the UniverseMachine check in Appendix A, so every result in the paper inherits it. The paper itself shows that a stricter, physically motivated definition requiring at least 50% member-galaxy overlap (DET+MEMB, dashed lines in Fig. 3) reduces completeness at M200 > 10^13 to about 70%, below the 80% value used in the headline. Because the matching tolerance is generous, the reported completeness and contamination figures should be treated as upper and lower bounds until the rule's parameter choices are shown to be inconsequential.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper benchmarks three optical group finders (Robotham et al. 2011, Yang et al. 2005, and Tempel et al. 2017) on a GAMA-like mock galaxy survey built from the Magneticum hydrodynamical simulation lightcone LC30. The authors simulate a spectroscopic galaxy sample down to z < 0.2 with stellar mass completeness M_star >= 10^9.8 M_sun, run the three finders, and compare their outputs to the SubFind halo catalogue. They report that all three finders achieve >80% completeness for halos with M200 > 10^13 M_sun, contamination below 10% in that regime, membership accuracy of at least 70% above the group scale, and that total stellar luminosity or total stellar mass are more reliable halo mass proxies than velocity dispersion. They also compare the optical catalogues with an X-ray eROSITA mock from the same simulation and check the impact of Magneticum's top-heavy GSMF using a UniverseMachine lightcone in Appendix A.","tokens_in":24661,"tokens_out":3925,"duration_ms":35062,"significance":"If the headline results are robust, this paper provides a useful, directly comparative calibration of three widely used group finders on the same simulation, which is valuable for interpreting current and upcoming spectroscopic surveys (SDSS, GAMA, DESI, WAVES) and for stacking eROSITA X-ray data on optically selected groups. The use of a realistic lightcone, a ground-truth SubFind halo catalogue, and an external UniverseMachine check are strengths. The authors are also transparent about several caveats, including the GSMF tension in Magneticum and the effect of membership errors on velocity-dispersion masses. However, the central quantitative claims rest on a matching rule that is not sensitivity-tested, and the mass-proxy comparison mixes an in-sample calibration with an external relation; these issues need to be addressed before the conclusions can be considered fully supported.","major_comments":[{"comment":"The definitions of completeness and contamination rely entirely on the hand-chosen matching rule: a detection is correct if its centre lies within R200 of a true halo and within dz < 3e-3, with a factor-of-six dominance criterion for multiple matches. No sensitivity analysis is provided for these parameters. The paper itself shows that the stricter DET+MEMB definition (at least 50% member-galaxy overlap, dashed lines in Fig. 3) reduces completeness at M200 > 10^13 M_sun to roughly 70%, below the headline >80% value. This indicates that the headline numbers are definition-dependent and should not be presented as absolute without demonstrating that the matching thresholds are inconsequential. I recommend reporting both definitions prominently in the abstract and adding a robustness test varying R200, the redshift window, and the dominance factor.","section":"Section 4.1 and Fig. 3"},{"comment":"The comparison of halo mass proxies is unfair because the total stellar mass proxy is calibrated directly on the Magneticum scaling relation, while the r-band luminosity proxy is taken from an external relation (Popesso et al. 2007). Consequently, the smaller scatter reported for luminosity in Table 2 may simply reflect that the luminosity relation is not fitted to the same simulation, whereas the stellar-mass relation absorbs Magneticum's specific scatter. To support the claim that stellar luminosity is the best proxy, both proxies should be calibrated on independent data or on the same simulation with an out-of-sample split. The in-sample nature of the stellar-mass calibration is a load-bearing issue for the central recommendation in the abstract.","section":"Section 4.3 and Table 2"},{"comment":"The UniverseMachine check is presented as evidence that the results are robust to the Magneticum GSMF tension, but it only runs one of the three finders (R11) and only tests completeness, contamination, and the stellar-mass proxy scatter. It does not test the membership accuracy, the luminosity proxy comparison, or the fragmentation behaviour for the other two finders. Additionally, the reported velocity-dispersion scatter in the UniverseMachine test (0.71 dex) is twice that in Magneticum (0.35 dex), a major discrepancy that deserves explicit discussion rather than a one-line remark. The claim that the findings are 'robust' is therefore only partially supported.","section":"Appendix A"}],"minor_comments":[{"comment":"The abstract states the survey is complete for M_star >= 10^9.8 M_sun, but Section 2.2 says 'no magnitude cut is applied.' Please clarify that the completeness refers to the stellar mass resolution limit of the simulation, not a survey selection.","section":"Abstract and Section 2.2"},{"comment":"The summary states the FOF catalogue is 'accurate (~80%)' in assigning galaxy members above 10^13 M_sun, while the abstract and Fig. 3 dashed lines indicate 70%. These numbers should be made consistent.","section":"Section 6, first bullet"},{"comment":"In the last two rows, the T17 column increases from 14,185 (no FRAG_2) to 14,194 (no FRAG_2 and no SPUR_2). Since the latter is a subset of the former, this is impossible and is likely a typo; please check all entries in these rows.","section":"Table 1"},{"comment":"The summary lists SDSS bandpass as 'u, v, g, r, and i'; SDSS uses u, g, r, i, z, and the 'v' should be removed.","section":"Section 6"},{"comment":"The text says 'In the high mass end, contamination is mostly due to SPUR_2 sources,' but Fig. 4 shows that fragmentation also increases strongly with mass. Please clarify how SPUR_2 and FRAG_2 are distinguished in the contribution to contamination at the high-mass end.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the A&A readership, but the editorial decision should hinge on whether the authors can make the matching-rule robustness explicit and re-calibrate the mass-proxy comparison fairly. The Table 1 inconsistency, if it signals a deeper issue in the matching counts, should be resolved before publication. I recommend major revision rather than rejection because the core methodology is sound and the concerns are addressable within the scope of the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, know this: the paper finally puts three widely used group finders – Robotham, Yang, Tempel – on the same GAMA-like mock and gives quantified completeness, contamination, and mass-proxy scatter. That comparison did not exist. The recommendation that total stellar luminosity beats velocity dispersion as a halo mass proxy is consistent across all three finders, and it is actionable.\n\nThe paper is also honest in the right places. The matching is defined explicitly, detections are split into primary, secondary, and tertiary, and the authors show the stricter membership-overlap completeness (DET+MEMB) in the same figure as the headline number. They flag the top-heavy GSMF in Magneticum and test one finder on a UniverseMachine lightcone in Appendix A. That is real engagement with the main caveat.\n\nThe soft spots are real but fixable. The headline \">80% completeness\" depends on a matching rule that accepts a detection within R200 and 3e-3 in redshift (~900 km/s) of a true halo. That cylinder is generous, and the paper's own DET+MEMB metric, requiring at least 50% member overlap, drops completeness to about 70% at M200 > 10^13 M_sun. So the 80% should be read as an upper bound, not a central value. The matching thresholds need a sensitivity test. The stellar-mass proxy is calibrated on Magneticum and then scored on the same halos, so its quoted scatter is partly in-sample; the UniverseMachine check covers only R11. No code or data are released, so the exact numbers cannot be reproduced.\n\nThis paper is for anyone building or using group catalogues for DESI, WAVES, or eROSITA stacking. It deserves a serious referee. The asks are straightforward: vary the matching rule, report completeness and contamination as functions of its thresholds, and run at least one more finder on the independent lightcone. The central conclusion – that optical spectroscopic selection recovers most halos above 10^13 M_sun – holds; just do not quote the 80% without the caveat.","headline":"A genuinely useful head-to-head benchmark of three group finders, though the headline completeness rests on a permissive matching rule that needs a sensitivity test.","tokens_in":25281,"tokens_out":4409,"would_cite":true,"duration_ms":37877,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that three group finders recover at least 80% of simulated galaxy groups and clusters above $10^{13}$ solar masses, with contamination below 10%, and that total stellar luminosity or mass beats velocity dispersion as a…","keywords":["galaxy groups","galaxy clusters","group finder algorithms","mock galaxy surveys","halo mass proxies","spectroscopic surveys","completeness and contamination","conditional luminosity function"],"falsifier":"Rerun the same three finders on an independent mock and match using a stricter cylinder, for example half the virial radius and a redshift window of $10^{-3}$; if completeness at $M_{200}>10^{13}\\,M_\\odot$ drops below 80% or contamination rises above 10%, the headline numbers are artifacts of the matching convention rather than intrinsic properties of the finders.","tokens_in":24204,"feed_emoji":"🔭","tokens_out":9343,"duration_ms":80885,"temperature":0.7,"pith_summary":"This paper tries to establish that optical spectroscopic surveys can reliably find galaxy groups and clusters in the local Universe, despite the systematics that worry observers. The authors run three established group-detection algorithms on the same mock galaxy catalogue, built to mimic a wide deep spectroscopic survey down to redshift 0.2, and measure how well each recovers the known halos. They find that all three detect at least 80% of halos with $M_{200}\\geq10^{13}\\,M_\\odot$, keep contamination below 10% at those masses, and assign member galaxies with at least roughly 70% accuracy above the group mass scale. They also find that total stellar luminosity or total stellar mass recovers halo mass more accurately than velocity dispersion, which matters because velocity-dispersion masses inherit membership errors. If correct, the result validates optical selection for dense regions and gives a practical recipe for using optical group catalogues in combination with X-ray data.","feed_headline":"Simulated survey: optical group finders clear 80% completeness","feed_subtitle":"Velocity dispersion misleads halo mass estimates; total stellar luminosity or mass is the more reliable proxy.","key_machinery":"The mechanism that carries the argument is a controlled mock experiment. A hydrodynamical cosmological simulation provides the true halo population; a lightcone (a redshift-bounded cone of mock sky) is converted into a galaxy catalogue with observed-frame magnitudes, Gaussian redshift errors of $\\sigma=45\\,\\mathrm{km\\,s^{-1}}$, stellar-mass errors of 0.2 dex, and 5% catastrophic redshift failures; and the three group finders are run on this catalogue. Their output is matched to the input halos by a cylinder defined by a projected offset of $R_{200}$ and a redshift window of $3\\times10^{-3}$, with a factor-of-six mass-dominance rule used to decide which halo a detection belongs to when several fall inside. Detections are classified as primary, secondary fragments, or spurious, and these classes feed directly into the completeness and contamination numbers. The halo mass proxies are calibrated scaling relations (BCG stellar mass, total stellar mass, and $r$-band luminosity) whose logarithmic scatters are compared to those of the finders' own velocity-dispersion masses.","core_discovery":"The central claim is a benchmark result, stated per group finder but holding for all three: on a simulated spectroscopic survey matching the depth and completeness of a wide local-Universe survey (redshift below 0.2, stellar masses complete above $10^{9.8}\\,M_\\odot$), the finders recover more than 80% of halos with $M_{200}$ above $10^{13}\\,M_\\odot$ and keep the fraction of spurious detections below 10% in that regime. Membership is at least 70% accurate above the group mass scale, rising further when measured by the brightest galaxies; the recovered conditional luminosity function of members matches the input halo catalogue. Because membership errors bias velocity-dispersion masses, the paper argues that the best halo mass proxy is the total stellar luminosity of the group (scatter 0.24--0.40 dex), closely followed by total stellar mass, and that using either proxy allows even low-richness groups to be kept in the catalogue. The comparison to a simulated X-ray observation of the same mock sky shows the optical catalogues are more complete at group scales, which the authors present as support for stacking X-ray data on optically selected groups.","pith_inferences":["Editorial extension: the reported numbers are tied to the depth of a wide deep local-Universe survey; a direct test would rerun the same three finders on a mock of a shallower survey and measure whether completeness at $10^{13}\\,M_\\odot$ still holds.","Editorial extension: the matching rule could be made probabilistic, assigning detections to halos by membership overlap rather than a six-times-dominance criterion, and the comparison would show how much of the fragmentation statistics is an artifact of the matching convention.","Editorial extension: the luminosity-proxy result suggests a practical recipe for X-ray-era work: stack X-ray photons on optically selected groups chosen by luminosity-based masses, reaching below the X-ray detection threshold.","Editorial extension: the alternative-mock robustness check in the appendix only re-runs one finder; extending it to the other two finders would test whether the completeness result depends on the specific simulation's stellar mass function."],"forward_implications":["For $M_{200}>10^{13}\\,M_\\odot$, all three finders recover at least 80% of halos, so optically selected group catalogues can be used to probe dense environments in the local Universe.","Contamination below $10^{13}\\,M_\\odot$ is dominated by interlopers and fragmentation, so low-mass groups projected near massive halos should be treated with caution.","Velocity-dispersion masses inherit membership errors, while total stellar luminosity or total stellar mass keeps the scatter small even when not all members are recovered; the paper recommends luminosity-based masses.","The recovered conditional luminosity function matches the input, so the finders are reliable for studying galaxy populations and evolution as a function of environment.","Optical selection yields a more complete halo mass distribution at group scales than the X-ray selection from the same simulation, supporting X-ray stacking on optically selected samples."],"supporting_citations":[{"why":"Defines the friends-of-friends group finder (R11) and the GAMA parameter calibration that the mock survey mimics.","marker":"Robotham et al. (2011)"},{"why":"Defines the iterative halo-based group finder (Y05) whose abundance-matched halo masses are benchmarked.","marker":"Yang et al. (2005a)"},{"why":"Defines the SDSS group finder (T17) with FOF linking and membership refinement that is benchmarked here.","marker":"Tempel et al. (2017)"},{"why":"Supplies the lightcone geometry and the synthetic X-ray catalogue used for the optical-versus-X-ray comparison.","marker":"Marini et al. (2024)"},{"why":"Provides the alternative empirical mock catalogue used in the appendix to check that the results are not set by the simulation's stellar mass function.","marker":"Behroozi et al. (2019)"},{"why":"Supplies the r-band luminosity--halo mass scaling relation used as the best-performing mass proxy.","marker":"Popesso et al. (2007)"},{"why":"Independent evidence that luminosity-based mass estimators outperform dynamical estimators for group halos.","marker":"Vázquez-Mata et al. (2020)"},{"why":"Observed stellar mass function used to quantify the massive-end tension in the simulation.","marker":"Bernardi et al. (2013)"}],"fun_headline_variants":["Optical group finders hit 80% completeness in mock survey","Velocity dispersion misleads halo masses; luminosity better","Mock survey: group finders recover >80% halos, <10% false","Stellar luminosity outshines velocity dispersion for halo mass","Test on mock GAMA-like survey validates optical group finders"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The completeness and contamination results depend on the paper's own matching rule, a correct detection being one within one virial radius and a redshift window of $3\\times10^{-3}$ of a true halo, with the most massive halo selected when several fall in the cylinder, and this rule is not checked for sensitivity.","fun_headline_variants_meta":{"raw":{"variants":["Optical group finders hit 80% completeness in mock survey","Velocity dispersion misleads halo masses; luminosity better","Mock survey: group finders recover >80% halos, <10% false","Stellar luminosity outshines velocity dispersion for halo mass","Test on mock GAMA-like survey validates optical group finders"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000683,"raw_usage":{"total_tokens":3189,"prompt_tokens":1123,"completion_tokens":2066,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":739,"completion_tokens_details":{"reasoning_tokens":1977}},"tokens_in":739,"tokens_out":2066,"duration_ms":14385,"temperature":1.0,"reasoning_tokens":1977,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:06:37.270613+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the same three finders on an independent mock and match using a stricter cylinder, for example half the virial radius and a redshift window of $10^{-3}$; if completeness at $M_{200}>10^{13}\\,M_\\odot$ drops below 80% or contamination rises above 10%, the headline numbers are artifacts of the matching convention rather than intrinsic properties of the finders.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the friends-of-friends group finder (R11) and the GAMA parameter calibration that the mock survey mimics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the SDSS group finder (T17) with FOF linking and membership refinement that is benchmarked here."},{"cited_title":"2024, Detecting Galaxy Groups and AGNs populating the local Universe in the eROSITA era, publication Title: arXiv e-prints ADS Bibcode: 2024arXiv240412719M","cited_arxiv_id":null,"evidence_quote":"Supplies the lightcone geometry and the synthetic X-ray catalogue used for the optical-versus-X-ray comparison."},{"cited_title":"2007, Astronomy and Astrophysics, 464, 451, aDS Bibcode: 2007A&A...464..451P","cited_arxiv_id":null,"evidence_quote":"Supplies the r-band luminosity--halo mass scaling relation used as the best-performing mass proxy."},{"cited_title":"A., Loveday, J., Riggs, S","cited_arxiv_id":null,"evidence_quote":"Independent evidence that luminosity-based mass estimators outperform dynamical estimators for group halos."}],"review_version":1}