{"id":"9725d4a6-b233-4dfc-a176-eb51cd5a80f2","arxiv_id":"2501.10622","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A 3D, MPI-parallelized wavelet halo finder whose SIMBA m50n512 catalogs broadly match friends-of-friends after unbinding, at linear time complexity but with no current speed advantage.","lead":"The authors present CWTHF, a new program that finds dark matter halos in cosmological simulations by scanning gridded density maps with a continuous-wavelet transform across many scales, then keeping only gravitationally bound candidates. It is, per the authors, the first MPI-parallelized wavelet-based halo finder with linear time scaling, and it broadly agrees with the standard friends-of-friends method on the SIMBA m50n512 simulation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Appendix B's quadratic peak model and the hand-removed 3/5 factor are the load-bearing calibration: if they are wrong, halo sizes, Eq. 5, and the FOF agreement shift; a synthetic recovery test can settle it.","rationale":"The reader and I identify the same load-bearing point: the cross-scale peak correction in Appendix B. It is load-bearing because the wavelet scale at which a 4D maximum survives determines halo size and boundary, the fitted Eq. 5 converts that scale to r_vir, and Section 4's n_ref dependence shows that the correction does not eliminate grid-resolution bias. The paper's own admission that the quadratic model is wrong and that the 3/5 factor is removed by hand turns a purportedly analytic correction into a tuned parameter, fit implicitly on the same simulation used for validation. I do not think this invalidates the method: the structural comparison with FOF after unbinding shows genuine correspondence, the algorithm is clearly described, the code is released, and the paper is honest about its limitations. But the quantitative claims—HMF excess, power-spectrum comparison, and the wavelet-scale-to-virial-radius relation—all inherit this calibration uncertainty, which justifies a conditional acceptance rather than a clean accept. Because the reader already reached CONDITIONAL and the concern raised here confirms rather than redirects that assessment, the verdict remains unchanged.","tokens_in":18345,"tokens_out":7479,"duration_ms":77430,"concrete_test":"Build a mock box containing ~500 NFW halos of known mass, center, and r200. Run CWTHF on the same particles with (i) rigid grid shifts of 0, 1/4, and 1/2 of a cell and (ii) n_ref = 300, 400, 550, 650 with otherwise default parameters. For matched halos, compare recovered r_vir or particle-overlap fraction to the input r200, and also rerun with the 3/5 factor restored in the Appendix B correction. If recovered radii drift by more than ~10% with grid shift or n_ref, or if restoring 3/5 changes matched radii by more than ~10%, the correction is numerically hand-tuned rather than a valid analytic bias removal, and the FOF comparison should be re-interpreted; if the recovered radii are stable and accurate, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The catalog-level agreement with FOF is produced after a calibration whose validity is not established. Section 2.2 and Appendix B correct peak positions and cross-scale peak values by fitting a 3D isotropic quadratic (Eqs. B2, B3) to grid values around each maximum. The paper then concedes that the actual CWT peak is 'sharper than that of a quadratic function' and compensates by 'removing the factor 3/5'—a one-parameter adjustment, not a derived result. This correction controls where a maximum survives in the 4D wavelet space, which sets the segmented halo boundary and feeds the fitted relation rvir = 0.091/kw - 0.012 (Eq. 5) that converts wavelet scales to physical sizes. Section 4 reports that increasing n_ref still 'consistently results in a reduction in halo size', so grid-resolution bias is not fully removed. If the quadratic model and the 3/5 removal are wrong, or are calibrated only for the m50n512 box, halo sizes, masses, and the medium/high-mass HMF excess all shift, and the agreement with FOF that supports the abstract's viability claim becomes partly a product of the correction rather than an independent measurement. This is a calibration risk, not a fatal flaw, but it is the most load-bearing assumption in the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CWTHF, a halo finder that identifies dark matter halos in cosmological simulations by computing the continuous wavelet transform (CWT) of a cloud-in-cell density grid and locating local maxima in the resulting 4D (3D plus scale) space. The algorithm segments the grid around these maxima, applies a density threshold and a self-boundness check, and assigns particles to halos. The method is MPI-parallelized and claimed to have O(N) time complexity. The authors test CWTHF on the SIMBA m50n512 dark-matter sub-simulation, compare the resulting catalog with a friends-of-friends (FOF) catalog built with yt, and report catalog counts, halo mass functions, halo power spectra, and a fitted relation between wavelet scale and virial radius. The central claim is that the agreement with FOF demonstrates the viability of CWT for halo finding, with the linear time complexity cited as the route to future gains.","tokens_in":18543,"tokens_out":5799,"duration_ms":53732,"significance":"If the central claim holds, CWTHF represents a novel and potentially scalable approach to halo identification, and the paper's release of the source code on Zenodo is a useful contribution to the community. The validation against an external FOF catalog on public SIMBA data, the parameter study in Section 4, and the actual performance measurements in Appendix A are all strengths. However, the paper's central comparison depends on a peak-correction calibration (Appendix B) that is explicitly conceded to be an analytic model with a manually adjusted factor, and this calibration affects which maxima survive cross-scale verification, the segmented halo boundaries, and the fitted wavelet-scale-to-virial-radius relation. The grid-resolution dependence noted in Section 4 further suggests that the halo sizes and statistics may not be converged at the default settings. These issues are fixable but must be addressed before the viability claim is fully established.","major_comments":[{"comment":"The cross-scale peak comparison, which decides which maxima are real halos and how large they are, rests on the assumption that the CWT near a peak is a 3D isotropic quadratic function. The paper itself concedes, in the last paragraph of Appendix B, that 'the precise form of the CWT peak is sharper than that of a quadratic function' and that 'removing the factor 3/5 will improve consistency across different resolutions.' This is a one-parameter ad hoc adjustment, not a derived result. Because this correction feeds directly into the cross-scale verification and the segmented halo boundary, and through Eq. (5) into the conversion from wavelet scale to virial radius, the catalog-level agreement with FOF reported in Section 3 could be partly a product of this calibration. I ask the authors to add a synthetic recovery test: insert halos with known density profiles, sizes, and masses into a particle distribution, run CWTHF, and show that the recovered sizes and masses, the cross-scale maxima, and the rvir-kw relation are unbiased and converge as n_ref and w_resolution increase. This test would also address the residual grid-resolution bias noted in Section 4.","section":"Section 2.2 and Appendix B, Eqs. (B2)-(B6)"},{"comment":"The paper states that 'increasing n_ref consistently results in a reduction in halo size' and that the HMF changes significantly for n_ref < 300. This means the default n_ref=400 may not be at numerical convergence, and the HMF/HPS comparisons in Section 3 are made at a resolution where halo sizes are still resolution-dependent. This is load-bearing because the abstract's viability claim relies on these comparisons. Please provide a convergence test in n_ref and w_resolution for the m50n512 catalog, showing that the HMF, the HPS, and the rvir-kw fit (Eq. 5) stabilize at the default settings within the quoted uncertainties. If convergence cannot be reached at practical grid sizes, the paper should state the systematic uncertainty on the halo-catalog statistics.","section":"Section 4 and Figure 12"},{"comment":"The comparison FOF catalog is post-processed by a self-boundness check that, as the paper notes, 'does not retain any particles' and discards entire FOF groups when a large fraction of particles are unbound. This crude unbinding is not the standard way to produce a bound FOF catalog, and it is different from the iterative unbinding used inside CWTHF. The resulting 182,362-halo FOF catalog is therefore not a standard reference, and the claim that the HMF discrepancy at the low-mass end is 'primarily due to the unbound halos identified by the FOF method' is not fully supported. Please either compare against a properly unbound FOF catalog (e.g., from AHF or ROCKSTAR on the same data) or clearly frame the comparison as being against FOF-plus-crude-unbinding, and adjust the conclusions accordingly.","section":"Section 3, Figures 9 and 10"}],"minor_comments":[{"comment":"The catalog name 'd4 r400 w20' is used without defining the meaning of the 'd4' component; footnote 3 gives the naming pattern only implicitly. Please define the notation explicitly.","section":"Section 3, first paragraph"},{"comment":"The fit of rvir versus 1/kw excludes 'halos from the three largest scales' with the justification that very large virial radii are not feasible in CDM. Please quantify this cutoff (e.g., by comparing with the maximum expected halo radius in the m50n512 box) and test the sensitivity of Eq. (6) to the exclusion.","section":"Section 3, Figure 11 and Eq. (5)"},{"comment":"The normalization of the Mexican hat wavelet is not explained; please state the normalization convention and cite the source (Wang & He 2021) more explicitly.","section":"Section 2.1, Eq. (4)"},{"comment":"The text states that 'increasing n_ref consistently results in a reduction in halo size, as questionable boundary regions are excluded.' Please specify which boundary regions are 'questionable' and how they are identified.","section":"Section 4, Figures 12 and 13"},{"comment":"The claim of being 'the first wavelet-based, MPI-parallelized halo finder' should be supported by a more thorough literature search or weakened to 'to our knowledge'.","section":"Section 1"},{"comment":"The performance statement 'A dataset that is eight times larger results in only approximately 0.6 times more computation time' is ambiguous; please clarify whether this is total time or time per particle and specify the exact numbers.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal as a computational methods paper. The main concern is the self-calibration of the peak correction in Appendix B, which is acknowledged by the authors to be approximate and manually adjusted. The authors should be encouraged to address this with a synthetic recovery test and a convergence study; the release of the code and the external FOF comparison are positive aspects. The current manuscript is not ready for acceptance without these additions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a solid methods paper that extends the authors' 2D wavelet finder to 3D, ships code on Zenodo, and validates against a yt-based FOF catalog on the public SIMBA m50n512 run. The central claim—that CWT can identify halos in 3D—is supported at a basic level: the CWT and FOF catalogs trace the same structures, and the halo mass function and power spectrum converge once the FOF catalog is unbound. That is real evidence, and the paper earns credit for it.\n\nThe genuinely new pieces are the grid-based fast CWT, the MPI parallelization, and the 3D implementation itself. The abstract's \"novel approach\" framing overstates the conceptual break—the wavelet formalism and Mexican hat kernel come from the authors' own prior papers—but the technical package is a legitimate extension, not a copy.\n\nThe soft spot is exactly where the stress-test note points: Appendix B. The peak-correction model assumes a 3D isotropic quadratic CWT peak, the paper then concedes the real peak is sharper, and the fix is to remove the 3/5 factor by hand. That one-parameter adjustment controls where maxima survive in 4D wavelet space, which sets halo boundaries and feeds Equation (5), the wavelet-scale-to-virial-radius relation. If that calibration is wrong or only works for m50n512, the HMF excess and the FOF agreement shift. The paper also reports that increasing n_ref consistently reduces halo size, so the grid-resolution bias is not fully removed. This is a calibration risk, not a fatal flaw, but it is load-bearing and it deserves a synthetic recovery test.\n\nOther soft spots are proportionate: a single simulation, a single reference finder, no stated FOF linking length, no error bars on the HMF/HPS comparisons, and the fit in Figure 11 excludes the three largest scales. The paper's own admissions are honest—Table 1 warns the defaults are box-specific, Appendix A says the code is not faster than FOF—and that honesty counts in its favor.\n\nThe citation pattern is self-heavy, but the cited results are the actual foundation of the method, so that is not a flaw. The validation is external, against an independent FOF implementation, so the circularity burden is mild.\n\nWho is this for? Simulation analysts who want a wavelet-based halo definition with a self-boundness check and a different take on halo boundaries. It deserves a serious referee. I would send it to review and push the authors to add a recovery test on a synthetic field, specify the FOF settings, and show how the 3/5 removal affects the final catalogs.","headline":"A useful 3D wavelet halo finder with shipped code and a real FOF comparison, but the peak-correction calibration in Appendix B is hand-tuned and should be stress-tested before the method is trusted at face value.","tokens_in":19206,"tokens_out":1296,"would_cite":true,"duration_ms":14894,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that dark matter halos in N-body simulations can be identified as cross-scale maxima of a continuous wavelet transform, producing catalogs comparable to friends-of-friends, with linear time complexity.","keywords":["dark matter halos","halo finder","continuous wavelet transform","friends-of-friends","N-body simulations","cosmological simulations","MPI parallelization","halo mass function"],"falsifier":"Re-run CWTHF on the same dataset with the peak-bias correction replaced by a high-resolution direct evaluation of the wavelet transform at the true peak (bypassing the quadratic approximation), and compare the resulting halo sizes and HMF to the FOF catalog; if halo boundaries or the rvir–kw relation shift noticeably when the 3/5 factor is reintroduced, the central claim that CWT catalogs match FOF is not robust.","tokens_in":18002,"feed_emoji":"🌌","tokens_out":6855,"duration_ms":58737,"temperature":0.7,"pith_summary":"The paper introduces CWTHF, a halo finder that identifies dark matter halos as peaks of a continuous wavelet transform computed on a cloud-in-cell density grid across multiple scales. The central claim is viability: CWTHF-produced halo catalogs agree with the conventional friends-of-friends (FOF) method at the catalog level, and the algorithm has linear time complexity O(N), promising gains for future exascale simulations. The paper shows that applying a self-boundness check to FOF groups removes most of the halo-mass-function discrepancy, implying that the remaining differences stem from physically unbound groups. The work also derives a calibration between wavelet scale and virial radius, rvir = 0.091/kw - 0.012, enabling physical interpretation of the wavelet scales. A sympathetic reader would care because the method offers a new, grid-based, parallelizable approach to halo finding with a well-defined boundary from the wavelet's positive region.","feed_headline":"Wavelet maxima reveal dark matter halos in linear time","feed_subtitle":"CWTHF builds self-bound halo catalogs whose mass function and power spectrum match friends-of-friends after unbinding.","key_machinery":"The central object is the continuous wavelet transform computed from the cloud-in-cell (CIC) grid using the isotropic Mexican-hat wavelet, defined as Ψ(w,x) = $w^{{3/2}}$(6 - $w^{2}$ $r^{2}$) $e^{{-w^2 r^2/4}}$, where w is the scale parameter (kw = w Lbox/1000 in dimensionless form). Because the CIC grid turns particles into pseudo-particles on a regular lattice, the transform reduces to a weighted sum of wavelets over grid points, giving linear time O(N) and making MPI domain decomposition along one axis straightforward. The argument is carried by the criterion that real halos are local maxima in the 4D space of 3D position plus scale — the cross-scale maximum — together with the positive region of the wavelet as the halo boundary. Two correction steps are load-bearing: fitting a 3D isotropic quadratic to the peak grid point to remove grid-position bias, and rescaling peak values across different grid resolutions to remove the averaging bias.","core_discovery":"CWTHF (Continuous Wavelet Transform Halo Finder) operates in 3D by binning particles onto a cloud-in-cell grid, computing the isotropic Mexican-hat wavelet transform at a sequence of scales, finding local maxima in each scale, and keeping only maxima that persist across adjacent scales. Around each cross-scale maximum the grid is segmented by the positive region of the wavelet, groups below a density threshold of 4ρ are discarded, and a self-boundness check removes unbound particles. Tested on a dark-matter-only sub-simulation with $512^{3}$ particles in a 50 h−1Mpc box, the CWT catalog contains 162,872 halos versus 253,112 FOF halos; when FOF groups are subjected to the same unbinding procedure the FOF count drops to 182,362 and the halo mass functions come into much closer agreement. The largest CWT halos match the largest FOF halos in position, mass, and particle content, and the halo power spectra agree once unbound FOF halos are removed. The paper also reports a fitted relation between the dimensionless wavelet scale and virial radius, rvir = 0.091/kw - 0.012 (equivalently rvir ≈ 1.8/w), giving the method a direct physical calibration.","pith_inferences":["The quadratic peak-shape assumption in Appendix B is the main technical risk; an obvious extension is to replace it with the exact wavelet peak profile or an empirically calibrated correction, which would remove the hand-tuned 3/5 factor and likely stabilize halo sizes further.","The rvir–kw calibration is derived from one simulation at one resolution; testing it on simulations with different box sizes, resolutions, and cosmologies would show whether the relation is universal or needs re-fitting.","Because the algorithm segments the grid by the positive wavelet region, the method implicitly defines halo boundaries by the scale of the wavelet; comparing CWT boundaries against other boundary definitions (e.g., splashback radii) could reveal whether the wavelet boundary carries physical meaning.","If linear-time complexity is realized in a compiled implementation, CWT-based halo finding could become competitive with FOF for trillions of particles, where FOF's linked-list overhead grows."],"forward_implications":["If the central claim holds, CWTHF provides a halo finder with linear time complexity whose memory use can be traded against the maximum scale kw high, letting users inspect large halos quickly with low memory.","The comparison implies that roughly 30% of FOF halos are unbound, and that an unbinding post-processing step reconciles FOF and wavelet catalogs; this supports the community practice of unbinding FOF outputs.","The CWT boundary combined with self-boundness check effectively turns CWTHF into a phase-space halo finder, addressing the 'linking bridge' problem of FOF.","The fitted scale–radius relation rvir ≈ 1.8/w gives a direct way to translate wavelet scales into physical halo sizes, useful for designing multiscale searches.","Because the algorithm's complexity is independent of grid resolution, larger simulations with fixed box length can be processed in roughly linear time."],"supporting_citations":[{"why":"Defines the friends-of-friends algorithm used as the reference halo finder in the comparison.","marker":"Davis et al. 1985"},{"why":"Previous 2D CWT halo-finding study whose methodology is extended to 3D in this paper.","marker":"Li et al. 2024"},{"why":"Supplies the self-boundness check procedure adopted by CWTHF.","marker":"Knollmann & Knebe 2009"},{"why":"Provides another phase-space halo finder whose unbinding approach motivates the self-boundness check.","marker":"Behroozi et al. 2013"},{"why":"Supplies the dark matter sub-simulation dataset used for testing CWTHF.","marker":"Davé et al. 2019"},{"why":"Used to generate the reference FOF catalog from the simulation.","marker":"Turk et al. 2011"},{"why":"Derives the Mexican-hat wavelet definition used in the CWT.","marker":"Wang & He 2021"}],"fun_headline_variants":["Wavelet maxima reveal dark matter halos in O(N) time","MPI-parallel wavelet finder matches FOF after unbinding","CWTHF: first wavelet-based MPI halo finder","Continuous wavelet transform locates halos in linear time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the continuous-wavelet field near a peak is well approximated by a 3D isotropic quadratic so that a grid-averaging bias can be corrected analytically — the paper itself concedes the real peak is sharper, and the 3/5 correction factor is removed by hand.","fun_headline_variants_meta":{"raw":{"variants":["Wavelet maxima reveal dark matter halos in O(N) time","MPI-parallel wavelet finder matches FOF after unbinding","CWTHF: first wavelet-based MPI halo finder","Continuous wavelet transform locates halos in linear time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000482,"raw_usage":{"total_tokens":2415,"prompt_tokens":1009,"completion_tokens":1406,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":625,"completion_tokens_details":{"reasoning_tokens":1348}},"tokens_in":625,"tokens_out":1406,"duration_ms":12342,"temperature":1.0,"reasoning_tokens":1348,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:03:33.884950+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run CWTHF on the same dataset with the peak-bias correction replaced by a high-resolution direct evaluation of the wavelet transform at the true peak (bypassing the quadratic approximation), and compare the resulting halo sizes and HMF to the FOF catalog; if halo boundaries or the rvir–kw relation shift noticeably when the 3/5 factor is reintroduced, the central claim that CWT catalogs match FOF is not robust.","supporting_citations":[],"review_version":1}