{"id":"8628f036-cf25-4041-8086-0dfdda320c90","arxiv_id":"2603.10636","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For Euclid-like weak-lensing cluster detection with multi-scale wavelets, one z_s,min=0.4 bin matches multi-bin tomography because spurious detections accumulate and cut purity.","lead":"A single source-redshift cut at z_min=0.4 detects weak-lensing galaxy clusters as well as combining multiple tomographic bins in Euclid-like mocks. Tomography's gains are limited mainly by spurious peaks stacking across bins, not by large-scale structure or photo-z errors alone.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Null tomography gain may be an artifact of fixed-threshold union across z_s,min bins rather than an intrinsic limit of source tomography.","rationale":"Only the abstract of 2603.10636 is available; the cached full text is the unrelated Q-StaR NoC paper (2603.10637), so this remains an abstract-level stress test, consistent with the reader’s LOW-confidence CONDITIONAL stance. The reader’s weakest assumption (mock/detector representativeness) is real but secondary: even if the wavelet multi-scale finder and Euclid-like n(z) are typical, the null gain can still be produced by the fixed-threshold union rule itself. That rule is the load-bearing step for both the single-bin equality and the ranking of limitations. The concrete test isolates the combination methodology without requiring new mocks. If coincidence or purity-matched re-thresholding does not improve the multi-bin tradeoff, the claim strengthens; if it does, the abstract’s conclusion needs qualification. Verdict stays CONDITIONAL pending full purity/completeness definitions, threshold choices, and combination details. No stronger rejection is warranted from the abstract alone.","tokens_in":17236,"tokens_out":635,"duration_ms":16487,"concrete_test":"From the purity–completeness curves (or equivalent tables) for the z_s,min=0.4 single bin versus the best multi-bin union at the same fixed S/N threshold, recompute purity after requiring spatial coincidence of peaks in at least two tomographic maps (or after raising the multi-bin threshold to match single-bin purity). If the coincidence (or re-thresholded) multi-bin curve lies above the single-bin curve over a non-trivial completeness range, the claimed equality and the dominance of spurious accumulation are combination-rule dependent.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim—that a single Euclid-like bin with z_s,min=0.4 matches any multi-bin combination, with spurious accumulation as the dominant limit—rests on the z_s,min-cut procedure: detections from separate lensing maps are combined at a fixed detection threshold. Under a simple union, false positives accumulate nearly independently while true peaks largely overlap, so purity falls by construction. The abstract attributes reduced gains partly to LSS and photo-z, but names spurious accumulation as dominant; that ranking is only meaningful if purity is evaluated after combination without re-thresholding, peak matching, or multi-bin coincidence requirements. If an alternative combination (e.g., requiring detection in ≥2 bins, or a joint multi-scale statistic) restores purity at comparable completeness, the equality of single- vs multi-bin performance is method-specific rather than a general result about tomographic weak-lensing peak finding. The progressive mock suite (NFW → N-body, true → photo-z) tests contamination but does not isolate whether the combination rule itself drives the null result.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The manuscript studies whether source-redshift tomography can increase the number of galaxy-cluster detections obtained solely from weak-lensing peaks. Using a wavelet multi-scale peak finder, the authors construct overlapping Euclid-like source bins by raising a minimum source redshift cut (z_s,min) and combine the resulting peak catalogues via a z_s,min-cut procedure. On a progressive suite of mocks (isolated NFW haloes, N-body embedded clusters, true and photometric redshifts), they report that a single bin with z_s,min = 0.4 matches the performance of any combination of up to four bins. Large-scale structure and photo-z errors reduce tomographic gains, but the dominant limitation is said to be the accumulation of spurious detections across bins, which lowers purity at fixed detection threshold.","tokens_in":17549,"tokens_out":714,"duration_ms":11604,"significance":"If the single-bin versus multi-bin equality is robust, the result is practically useful for Euclid-like surveys: it would justify simpler, non-tomographic peak pipelines without sacrificing catalogue size at fixed purity. The progressive mock design (NFW to N-body, true to photo-z) is a sound way to isolate contaminants. The claim is falsifiable and of direct interest to weak-lensing cluster cosmology. Significance is tempered by the fact that the result appears tied to one detector and one combination rule; a general statement about tomography would require broader validation.","major_comments":[{"comment":"The central claim that a single z_s,min=0.4 bin performs as well as multi-bin combinations rests on the z_s,min-cut combination rule (union of peaks at fixed detection threshold). Under a simple union, false positives accumulate nearly independently while true peaks largely overlap, so purity falls by construction. The abstract ranks spurious accumulation as the dominant limitation over LSS and photo-z; that ranking is only meaningful if purity is evaluated after combination without re-thresholding, peak matching, or multi-bin coincidence requirements. The manuscript should either (i) re-threshold or apply a coincidence cut after combination and re-compare single- vs multi-bin purity/completeness, or (ii) clearly restrict the claim to this specific combination rule rather than to tomography in general.","section":null},{"comment":"Only one peak finder (the recently introduced wavelet multi-scale method) is used. The null gain of tomography may be method-specific. At minimum, the paper should discuss whether the same single-bin optimum is expected for other common peak finders (e.g., aperture mass, Gaussian-smoothed SNR maps) or provide a limited cross-check on one alternative, so that the result is not over-generalised to all weak-lensing peak detection.","section":null},{"comment":"The abstract states that all combinations from one to four tomographic bins were considered and that z_s,min=0.4 is optimal, but does not specify how purity and completeness are defined after multi-bin combination, what detection threshold is held fixed, or how peaks are matched across bins and to true clusters. These definitions are load-bearing for the equality claim and must be stated explicitly (with tables or figures of purity vs completeness for single-bin vs best multi-bin) so the result can be reproduced and stress-tested.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is practical, not theoretical. On Euclid-like mocks they find that a single source bin with z_s,min=0.4 matches any combination of up to four overlapping tomographic bins for their wavelet multi-scale peak finder. They name accumulation of spurious peaks across bins as the main reason multi-bin does not win, ahead of LSS and photo-z. That is a useful, falsifiable methods result for people building WL cluster catalogs.\n\nWhat is new is the systematic enumeration: progressive mocks (isolated NFW → N-body embedded, true z → photo-z), all 1–4 bin combinations, and an explicit ranking of why tomography underperforms. They do not oversell it as a new probe; they diagnose a purity hit at fixed threshold. That honesty is a strength. Circularity is low—this is an empirical bake-off, not a derivation that forces the answer.\n\nThe soft spot that matters is the combination rule itself. The abstract’s z_s,min-cut is a fixed-threshold union of peaks from separate maps. Under a simple union, false positives add nearly independently while true peaks largely overlap, so purity falls by construction. If they never re-threshold after combination, never require multi-bin coincidence, and never try a joint multi-scale statistic, then “single bin is as good as multi-bin” is partly a statement about that union, not a general limit of source tomography. LSS and photo-z are tested; the combination rule is not isolated. That is the main caveat, not a fatal hole—just the right place for a referee to push.\n\nI only have the abstract (the cached full text is a different paper), so I cannot check purity/completeness curves, thresholds, or error bars. On the design as stated, the claim is clear and the mock ladder is sensible. This is for people who run or plan Euclid-like WL peak finders and care about binning strategy. It deserves a serious referee, not a desk reject. I would not put it in next week’s reading group unless we are deep in survey pipelines, and I would only cite it if I were choosing source cuts for a similar detector. Engage if you work that problem; otherwise file as solid methods guidance with one method-specific caveat.","headline":"Honest pipeline result for Euclid-like WL peaks: under their wavelet + z_s,min-cut setup, one bin at 0.4 matches multi-bin combos, but the null gain may be partly baked into how they union detections.","tokens_in":18161,"tokens_out":585,"would_cite":false,"duration_ms":11120,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A single source-redshift bin at z_min=0.4 finds as many weak-lensing clusters as multi-bin tomography on Euclid-like mocks.","keywords":["weak lensing","galaxy clusters","source redshift tomography","peak detection","Euclid","multi-scale wavelet","photometric redshifts"],"falsifier":"Repeat the single-bin versus multi-bin comparison on real Euclid weak-lensing maps or an independent simulation suite with a different peak finder, and test whether multi-bin combinations recover significantly more true clusters at fixed purity than the z_s,min=0.4 single bin.","tokens_in":18151,"feed_emoji":"🔭","tokens_out":795,"duration_ms":17644,"temperature":0.7,"pith_summary":"This paper asks whether splitting background galaxies into redshift bins and combining their weak-lensing peak maps can detect more galaxy clusters than one carefully chosen bin. Using a wavelet multi-scale peak finder on Euclid-like mocks, the authors build overlapping bins by raising the minimum source redshift and test every combination of one to four bins. They find that a single bin with minimum source redshift 0.4 matches the multi-bin combinations. Large-scale structure and photometric redshift errors shrink the expected gains, but the main limit is that false peaks accumulate when bins are merged, so purity falls at a fixed detection threshold. The result matters for Euclid-scale surveys: it suggests that simple single-bin maps already capture the detections that tomography was hoped to add, without the purity cost of combining bins.","feed_headline":"One redshift cut matches multi-bin cluster lensing","feed_subtitle":"On Euclid-like mocks, false peaks from combining bins erase tomography’s gains.","key_machinery":"The z_s,min-cut technique: overlapping source-redshift bins formed by progressively raising the minimum source redshift, each producing a lensing map that is searched with a wavelet multi-scale peak detector, then combining the resulting peak catalogues.","core_discovery":"On Euclid-like mocks, a single source-redshift bin with z_s,min=0.4 performs as well as any combination of up to four tomographic bins for weak-lensing cluster detection. Large-scale structure and photometric redshift errors reduce tomography’s potential gains, but the dominant limitation is the accumulation of spurious detections across bins, which lowers purity at a fixed detection threshold.","pith_inferences":["Peak finders that suppress noise differently from the wavelet method might still gain from tomography where this detector does not.","Real Euclid selection functions more complex than the mocks could reopen a multi-bin advantage if the purity model is incomplete.","Earlier weak-lensing forecasts that assumed ideal redshift bins without false-positive accumulation may have overstated tomography’s cluster-detection gains."],"forward_implications":["Single-bin weak-lensing cluster searches with a modest z_s,min cut can match multi-bin completeness and purity on Euclid-like data.","Analysis pipelines can avoid multi-bin peak combination without losing detections relative to the methods tested here.","Photometric redshift errors and large-scale structure already erase most of the theoretical tomographic gain before combination noise is counted.","If multi-bin methods are still used, spurious-peak accumulation must be controlled or purity will drop at fixed threshold."],"fun_headline_variants":["Single z_min=0.4 bin matches multi-bin weak lensing clusters","One source-redshift cut equals multi-bin cluster detections","Spurious peaks erase tomography gains in WL cluster finds","False detections across bins curb multi-redshift advantages","LSS and photo-z errors limit but do not dominate purity drop"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That these mocks and this particular wavelet detector plus bin-combination rule are representative enough that single-bin equality will hold for real Euclid data and other peak finders.","fun_headline_variants_meta":{"raw":{"variants":["Single z_min=0.4 bin matches multi-bin weak lensing clusters","One source-redshift cut equals multi-bin cluster detections","Spurious peaks erase tomography gains in WL cluster finds","False detections across bins curb multi-redshift advantages","LSS and photo-z errors limit but do not dominate purity drop"]},"model":"grok-4.5","effort":"low","cost_usd":0.003216,"raw_usage":{"total_tokens":1112,"prompt_tokens":816,"num_sources_used":0,"completion_tokens":89,"cost_in_usd_ticks":32160000,"prompt_tokens_details":{"text_tokens":816,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":207,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":816,"tokens_out":89,"duration_ms":2678,"temperature":1.0,"reasoning_tokens":207,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T23:24:56.802111+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Repeat the single-bin versus multi-bin comparison on real Euclid weak-lensing maps or an independent simulation suite with a different peak finder, and test whether multi-bin combinations recover significantly more true clusters at fixed purity than the z_s,min=0.4 single bin.","supporting_citations":[],"review_version":1}