{"id":"f614313e-7735-4db7-9ea2-c79bf6fa853e","arxiv_id":"2501.16420","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"AMICO-WL is a weak-lensing cluster finder whose foreground-removal mode roughly doubles completeness at 70% purity on Euclid-like mocks, from 6.5% to 13.0%.","lead":"Galaxy clusters bend the light of galaxies behind them, so measuring those distortions can find clusters even where they emit nothing. A team has built a new analysis tool, AMICO-WL, to spot clusters this way in the huge upcoming Euclid and Rubin surveys, and it finds twice as many when foreground galaxies are first cut out.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Foreground-removal gain hinges on exact photo-zs; Euclid-like photo-z scatter across z_min=0.2 may erase the factor-of-two completeness advantage.","rationale":"The central claim is the quantitative completeness gain attributed to foreground removal. The paper is honest about the exact-redshift assumption, and the reader's weakest_assumption identifies it as load-bearing. This concern is more fundamental than the absence of error bars or the use of a single mock field because it is a systematic simplification that affects the mechanism of the proposed improvement: the z_min cut can only work as designed if source redshifts are known precisely. The concrete test directly targets this mechanism by adding photo-z scatter and checking whether the factor-of-two gain survives. The reader's verdict of CONDITIONAL is appropriate; no adjustment is needed.","tokens_in":24367,"tokens_out":8239,"duration_ms":80039,"concrete_test":"Re-run the FR02 and COMPLETE configurations on the same mock after assigning each source a photometric redshift drawn from p(z_ph|z_true) with sigma_z = 0.05(1+z) (or the Euclid requirement), applying the same z_min = 0.2 cut and filter parameters, and recompute completeness at the measured 70% purity using the blinking matching procedure. If the FR02-to-COMPLETE completeness ratio at 70% purity drops below about 1.5, the reported doubling does not survive realistic photo-z scatter.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (6.5% to 13.0% completeness at ~70% purity) is obtained on a mock catalogue in which every galaxy is assigned its exact redshift: Sect. 3 states 'we do not consider any photometric redshift uncertainty in our analysis.' The FR02 run uses these true redshifts to apply a hard cut at z_min = 0.2, cleanly separating foreground from background sources. In real Euclid data, photo-z errors with typical scatter sigma_z ~ 0.05(1+z) will scatter galaxies across this boundary, blurring the foreground/background separation that drives the reported doubling and changing the effective source density entering the filter normalization (Eq. 26) and the LSS noise contribution (Eq. 28). The purity calibration of Eq. (29) also assumes that positive and negative E-mode peaks have the same noise statistics; with photo-z misassignment the foreground contamination is no longer perfectly known, so this symmetry is broken. A secondary simplification is the neglect of source clustering (Sect. 3: 'we randomly distribute them in the map, neglecting source clustering'), which removes correlations in the shape-noise field that the filter and the SNR thresholding do not model. Both simplifications are acknowledged, but they directly affect the quantitative headline, not just its error bars.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AMICO-WL, a weak-lensing extension of the AMICO optimal matched filter code for galaxy cluster detection. The algorithm constructs E-mode and B-mode amplitude maps from galaxy ellipticities, derives SNR thresholds through a comparison of positive and negative E-mode peak statistics, and includes a cleaning procedure to handle blended detections. The method is tested on a 25 deg^2 Euclid-like mock built from the DUSTGRAIN-pathfinder simulation, using four configurations: COMPLETE and three foreground-removed runs (FR02, FR04, FR06) based on redshift cuts at z_min = 0.2, 0.4, 0.6. Detections are matched to haloes with M200 > 5e13 h^-1 Msun using a new 'blinking' procedure that removes individual redshift slices to infer the redshift of the lensing signal. The headline result is that, at ~70% purity, the FR02 run achieves 13.0% completeness versus 6.5% for the COMPLETE run, with a concurrent reduction in spurious detections.","tokens_in":24609,"tokens_out":8460,"duration_ms":74392,"significance":"If the headline result is robust, AMICO-WL provides a practical, well-tested weak-lensing cluster finder for Euclid, building on the established AMICO infrastructure and adding a principled SNR-threshold calibration and a novel matching technique for simulation validation. The paper is transparent about its mock assumptions and includes a careful spurious-detection taxonomy. Strengths include the use of the standard optimal filter formalism, the B-mode consistency check, the support for the matching procedure provided by the positional scatter distribution in Fig. 12, and the detailed analysis of the origin of false positives. The main risk is that the quantitative performance claims are derived under idealized assumptions—exact redshifts and no source clustering—that may not transfer to real Euclid data, and the absence of any uncertainty estimates makes it difficult to judge the significance of the reported factor-of-two gain.","major_comments":[{"comment":"The headline result (6.5% to 13.0% completeness at ~70% purity) rests on the FR02 configuration, which applies a hard redshift cut at z_min = 0.2 using the true redshifts of the mock galaxies. The manuscript states in Sect. 3 that 'we do not consider any photometric redshift uncertainty in our analysis.' Real Euclid photo-z errors, with typical scatter around sigma_z ~ 0.05(1+z), will scatter galaxies across this boundary, blurring the foreground/background separation, modifying the effective source density entering Eq. (26), and changing the LSS noise contribution in Eq. (28). This is a load-bearing simplification for the quantitative headline, and the paper should either include a photo-z error model or explicitly frame the factor-of-two gain as an idealized proof of concept rather than a Euclid-ready performance prediction.","section":"Sect. 3 and Sect. 4.2"},{"comment":"The mock galaxies are 'randomly distribute[d] in the map, neglecting source clustering.' The filter noise model (Eqs. 24-25) and the purity calibration based on positive/negative E-mode peak statistics (Eq. 29) assume uncorrelated shape noise and Gaussian LSS statistics. Source clustering introduces correlations in the noise field and in the peak counts, which can bias the SNR thresholds and the purity estimates. Since the paper claims the thresholding method sets a desired sample purity, a quantitative test of the sensitivity to source clustering (e.g., a clustered or log-normal source catalogue) is needed to support the validity of the purity calibration for survey-like data.","section":"Sect. 3 and Sect. 5.2"},{"comment":"The completeness and purity values are quoted without uncertainties, and all numbers come from a single 25 deg^2 mock field. In the '~70% PUR' rows of Table 2, the measured purity is not reported, only the SNR threshold, number of detections, and completeness. Because the Eq. (29) calibration is imperfect (the TOTAL row shows predicted 70% purity yielding measured purities of 61%, 71%, 64%, and 61% for COMPLETE, FR02, FR04, and FR06, respectively), the reader cannot verify that the headline 6.5% versus 13.0% comparison is actually at matched measured purity. The authors should report the measured purity in every row and provide error estimates, such as jackknife over subfields or field-to-field variance, for the completeness and purity values that support the central claim.","section":"Table 2 and Sect. 6.2"},{"comment":"The purity calibration of Eq. (29) assumes that negative E-mode peaks follow the same noise statistics as positive peaks except for the absence of cluster signals. The measured purities in the TOTAL row of Table 2 deviate from the predicted 70% by up to 9 percentage points for the COMPLETE, FR04, and FR06 runs, showing that the calibration is not accurate across configurations. This discrepancy indicates that the symmetry between positive and negative peaks is not fully realized in the simulations, and it directly affects the reliability of the SNR thresholds used to define the operating points. The calibration should be validated per configuration, or the paper should base its purity-matched comparison entirely on the measured purity rather than the predicted value.","section":"Sect. 5.2, Eq. (29)"}],"minor_comments":[{"comment":"The notation 'averageb' appears to be a typo for 'average b', and the sentence defining the bias term is confusing; please clarify that b is the bias of the estimator and that the constraint in Eq. (20) is imposed to make it zero.","section":"Sect. 2.2, Eq. (20)"},{"comment":"The parameter 'white noise N_epsilon' is not defined in the table caption; please specify whether it corresponds to the power spectrum P_epsilon of Eq. (24) or to its square root, and give its units.","section":"Table 1"},{"comment":"There are typographical issues in phrases such as 'lineofsight' appearing in the text; please correct these to 'line-of-sight' throughout.","section":"Sect. 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid methods contribution from an established group, and the blinking matching procedure is a genuinely useful validation tool. The main concern for the journal is whether the idealized mock assumptions (exact redshifts, no source clustering, single field) are sufficiently highlighted relative to the Euclid-oriented claims in the abstract; I would encourage the editor to require either a photo-z sensitivity test or a careful rewording of the headline claim. The manuscript is also a contender in the Euclid Weak Lensing Cluster Challenge, so the authors may be constrained in what they can disclose, but the referee report should still ask for the missing measured purities and error bars."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"AMICO-WL is a genuine and mostly careful piece of engineering: it takes the AMICO matched-filter infrastructure, adds a weak-lensing branch, a foreground redshift cut, and a peak-based purity calibration, and tests the result on a Euclid-like mock. The blinking-plane matching procedure is the most original part — it uses the simulation's redshift slices to give detections redshift information, and the positional scatter plot makes a convincing case that the matching is not dominated by random associations. The implementation looks sound: B-mode maps are consistent with noise, and the redshift distribution of matched halos peaks where lensing efficiency is highest.\n\nThe headline performance number — completeness doubling from 6.5% to 13% at 70% purity — should not be taken at face value. Three things stand in the way. First, the mock uses exact redshifts (stated in Sect. 3), so the foreground cut at z_min=0.2 is a perfectly clean split. Real Euclid photo-z scatter will blur that boundary, and the gain from foreground removal is exactly what the clean split buys. Second, the measurement is on a single 25 deg^2 realization with no error bars; the difference is based on a few dozen matched halos, so Poisson and cosmic variance are not negligible. Third, the purity calibration of Eq. (29) predicts 70% but the COMPLETE run delivers 61%; it works for FR02 (71%), but a self-calibration that misses by nine points in one of four runs is not yet a reliable product. The neglect of source clustering is a lesser concern, acknowledged in Sect. 3, but it feeds the same noise model that the purity calibration relies on.\n\nNone of this is fatal. The algorithm is a legitimate contender for Euclid weak-lensing cluster detection, and the blinking-plane matching is a useful validation tool in its own right. The paper reports all four runs honestly rather than hiding the FR04/FR06 results. What it needs is a photo-z-realistic test, error bars or multiple realizations, and a discussion of why the purity calibration fails for COMPLETE. I would send it to a serious referee, and I would expect the revision to make the quantitative claims appropriately conditional. This is the kind of paper a cluster working group should read now, before the Euclid cluster challenge papers come out.","headline":"A solid, honestly reported matched-filter weak-lensing cluster finder with a clever simulation-side matching tool; the headline completeness gain is real but conditional on exact redshifts and a single mock realization.","tokens_in":25207,"tokens_out":3024,"would_cite":false,"duration_ms":32010,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces AMICO-WL, an optimal matched-filter cluster detector for weak-lensing data, and shows that cutting foreground galaxies at redshift 0.2 doubles completeness at 70% purity on Euclid-like mock data.","keywords":["weak gravitational lensing","galaxy clusters","optimal matched filter","foreground removal","Euclid-like mock catalogues","purity and completeness","cluster detection","cosmic shear noise"],"falsifier":"Re-run the FR02 and COMPLETE configurations on a mock that assigns photometric-redshift scatter $\\sigma_z/(1+z)\\simeq 0.05$ to the galaxy redshifts before applying the $z_{\\min}=0.2$ cut; if the FR02 completeness advantage over COMPLETE at 70% purity does not remain significant, the headline gain does not transfer to real survey data.","tokens_in":24157,"feed_emoji":"🔭","tokens_out":8821,"duration_ms":68048,"temperature":0.7,"pith_summary":"Weak-lensing cluster detection traces all matter, not just baryons, so the paper extends the AMICO optimal matched-filter cluster finder to work on galaxy ellipticities instead of galaxy overdensities. The paper's central claim is that removing foreground galaxies with a redshift cut at $z_{\\min}=0.2$ makes the detector substantially more efficient: on a 25 square-degree Euclid-like mock field, at 70% purity completeness doubles from 6.5% to 13.0% while spurious detections decrease. This matters because upcoming wide surveys such as Euclid will measure the shapes of billions of galaxies, and a pure, complete weak-lensing cluster sample would complement optical, X-ray, and SZE selection with a mass-based selection independent of baryonic emission.","feed_headline":"Foreground cut doubles weak-lensing cluster detections","feed_subtitle":"At 70% purity, cutting foreground galaxies below z=0.2 lifts completeness from 6.5% to 13.0% in a Euclid-like mock.","key_machinery":"The load-bearing object is the linear optimal matched filter $\\Psi$, defined in Fourier space as $\\hat{\\Psi}(k)\\propto \\hat{\\tau}(k)/P_N(k)$, where $\\hat{\\tau}$ is the Fourier transform of the assumed Navarro-Frenk-White tangential-shear profile and $P_N=P_\\epsilon+P_\\gamma$ is the noise power spectrum combining shape noise and the cosmic-shear power spectrum from non-linear structure growth. The filter is applied to the tangential ellipticities of galaxies to build an amplitude map, and detections are peaks in the signal-to-noise map. Around that core the paper adds a foreground-removal redshift cut with re-optimised filter parameters, a signal-to-noise threshold calibrated by comparing positive and negative E-mode peaks so that the desired sample purity can be set in advance, and an iterative cleaning step that subtracts the filtered shear model of each detection before searching for the next peak. Performance is assessed with a lens-plane 'blinking' matching procedure that removes one $\\Delta z=0.1$ matter slice at a time, giving each detection a redshift anchor.","core_discovery":"On the paper's own terms, the discovery is that a simple pre-processing step—cutting every galaxy whose true redshift lies below $z_{\\min}$ out of the shear catalogue and re-optimising the filter for the truncated catalogue—transforms a weak-lensing cluster finder into one whose completeness at fixed purity roughly doubles, from 6.5% to 13.0% at 70% purity, while reducing spurious detections. The paper attributes the gain to removing the noise contribution of galaxies in the foreground of the target haloes, whose shapes dilute the tangential-shear signal and add variance to the maps. The paper also introduces a 'blinking' redshift-plane matching procedure that assigns redshift information to weak-lensing detections, shows that matched detections peak at lens redshifts $0.3<z<0.4$ where lensing efficiency is maximum, and classifies the remaining spurious detections into noise fluctuations, chains of low-mass haloes aligned along the line of sight, and structures missed by the reference halo mass cut.","pith_inferences":["A testable extension would rerun the same four configurations on a mock with photometric-redshift scatter $\\sigma_z/(1+z)\\simeq 0.05$; the completeness gap between FR02 and COMPLETE at fixed purity is likely to shrink once the sharp $z_{\\min}=0.2$ boundary is blurred.","Because the mock neglects source clustering, the shape-noise term $P_\\epsilon$ used in the filter and the negative-peak purity calibration may underestimate correlated-noise contamination on real data; adding clustering would quantify this.","The purity estimator compares positive and negative E-mode peaks and assumes they share the same noise and large-scale-structure statistics; on real data, shear systematics that break E/B symmetry would require external calibration of this estimator.","The foreground-removal idea could be made redshift-dependent, with a cut threshold matched to each lens plane, which the paper notes as a possible development; this could push the 70%-purity completeness beyond the reported 13%."],"forward_implications":["At a fixed sample purity of 70%, applying the $z_{\\min}=0.2$ foreground cut raises completeness from 6.5% to 13.0% compared with using the full galaxy catalogue.","The foreground cut also lowers the number of spurious detections: in the paper's taxonomy, non-vanished spurious detections fall from 17 to 6 and non-matched spurious detections from 65 to 42 in the FR02 run.","At the more conservative 80% purity level the completeness of the four configurations converges, because only the high signal-to-noise peaks from the most massive, best-lensed haloes survive.","The lens-plane blinking matching shows that detections concentrate at lens redshifts $0.3<z<0.4$, consistent with lensing geometry, and it supplies a redshift estimate that ordinary positional matching cannot give."],"supporting_citations":[{"why":"Supplies the optimal linear matched filter in Fourier space and the noise model that AMICO-WL adapts.","marker":"Maturi et al. (2005)"},{"why":"Refines the filter for weak-lensing halo detection with template halo parameters and large-scale-structure suppression.","marker":"Maturi et al. (2007)"},{"why":"Compared weak-lensing estimators and introduced the lens-plane removal idea that the blinking matching generalises.","marker":"Pace et al. (2007)"},{"why":"Provides the AMICO code infrastructure, map creation, and cleaning procedure that the weak-lensing branch reuses.","marker":"Bellagamba et al. (2018)"},{"why":"The Euclid Cluster Finder Challenge validated the parent AMICO algorithm's optical detection performance and motivates extending it to weak lensing.","marker":"Euclid Collaboration: Adam et al. (2019)"},{"why":"Supplies the light-cone ray-tracing algorithm that builds the convergence and shear planes used to simulate the mock observations.","marker":"Giocoli et al. (2015)"},{"why":"Provides the Euclid-like source redshift distribution with which the simulated shear catalogue is populated.","marker":"Boldrin et al. (2012)"},{"why":"Defines the NFW halo profile assumed as the shear template in both the filter kernel and the cleaning subtraction.","marker":"Navarro et al. (2004)"},{"why":"Provides the non-linear matter power spectrum used to compute the large-scale-structure noise term of the filter.","marker":"Takahashi et al. (2012)"}],"fun_headline_variants":["Foreground z-cut doubles weak-lensing cluster yield","Simple foreground cut doubles cluster completeness to 13%","AMICO-WL: foreground cut doubles detection completeness","Z-cutting foregrounds doubles cluster detections in weak lensing","Foreground cut doubles cluster completeness at 70% purity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every galaxy in the input catalogue has an exact, error-free redshift, so the $z_{\\min}$ cut cleanly separates foreground from background; the paper states that no photometric-redshift uncertainty is included.","fun_headline_variants_meta":{"raw":{"variants":["Foreground z-cut doubles weak-lensing cluster yield","Simple foreground cut doubles cluster completeness to 13%","AMICO-WL: foreground cut doubles detection completeness","Z-cutting foregrounds doubles cluster detections in weak lensing","Foreground cut doubles cluster completeness at 70% purity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001166,"raw_usage":{"total_tokens":4910,"prompt_tokens":1115,"completion_tokens":3795,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":731,"completion_tokens_details":{"reasoning_tokens":3716}},"tokens_in":731,"tokens_out":3795,"duration_ms":24975,"temperature":1.0,"reasoning_tokens":3716,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:22:44.792327+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the FR02 and COMPLETE configurations on a mock that assigns photometric-redshift scatter $\\sigma_z/(1+z)\\simeq 0.05$ to the galaxy redshifts before applying the $z_{\\min}=0.2$ cut; if the FR02 completeness advantage over COMPLETE at 70% purity does not remain significant, the headline gain does not transfer to real survey data.","supporting_citations":[{"cited_title":"2005, Astronomy & Astrophysics, 442, 851","cited_arxiv_id":null,"evidence_quote":"Supplies the optimal linear matched filter in Fourier space and the noise model that AMICO-WL adapts."},{"cited_title":"2007, Astronomy & Astrophysics, 462, 473","cited_arxiv_id":null,"evidence_quote":"Refines the filter for weak-lensing halo detection with template halo parameters and large-scale-structure suppression."},{"cited_title":"2007, Astronomy & Astrophysics, 471, 731","cited_arxiv_id":null,"evidence_quote":"Compared weak-lensing estimators and introduced the lens-plane removal idea that the blinking matching generalises."},{"cited_title":"B., Baldi, M., et al","cited_arxiv_id":null,"evidence_quote":"Supplies the light-cone ray-tracing algorithm that builds the convergence and shear planes used to simulate the mock observations."},{"cited_title":"F., Hayashi, E., Power, C., et al","cited_arxiv_id":null,"evidence_quote":"Defines the NFW halo profile assumed as the shear template in both the filter kernel and the cleaning subtraction."}],"review_version":1}