{"id":"96b11e6e-b8dd-4b09-84a3-e0f7354396b3","arxiv_id":"2501.00301","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Spatial correlation with lower-redshift interlopers can estimate photometric contamination in Lyman-break galaxy samples: below 5.5% for GOODS-S z~6 and 62% for BoRG z~8.","lead":"This paper proposes a way to estimate how many high-redshift galaxy candidates are actually impostors, by measuring whether they cluster spatially with lower-redshift galaxies that could mimic them. It applies the idea to two Hubble surveys, finding negligible contamination in one sample and about 62% in another, matching earlier estimates.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 5.5% calibration assumes contaminants are a random subset of the faint z~1.3 population; if real Balmer-break interlopers cluster differently, both the GOODS-S limit and the BoRG 62% estimate can shift.","rationale":"The paper proposes a genuinely useful independent contamination check and provides a self-consistent mock calibration. The central quantitative claims, however, rest on the assumption that the contaminating population is a uniformly random subset of the faint interloper population, sharing its angular clustering. The reader's weakest_assumption identifies exactly this random-subset/clustering link, and I agree it is the right place to focus. It is load-bearing because both the 5.5% GOODS-S upper limit and the 62% BoRG central estimate are calibrated by injecting contaminants with uniform probability per faint low-z galaxy. Real colour selection is SED-dependent, so the true leak population may have a different ACF; if it is less clustered, the null GOODS-S measurement would permit contamination above 5.5%, and if it is more clustered, the published limit would be conservative but the quoted number would still lack a validated mapping. The BoRG result is additionally fragile because the expected faint z≈2 count is assumed to be a deterministic multiple of the observed bright count, so the simulated correlation coefficient at a given leakage fraction depends on an unmeasured field-to-field ratio. The concrete test I propose directly measures the leak-population ACF from real photometry, which would settle whether the assumption is acceptable. Since the reader already conditioned acceptance on addressable concerns and this is the most important such concern, the existing CONDITIONAL verdict remains appropriate.","tokens_in":19528,"tokens_out":19660,"duration_ms":211923,"concrete_test":"Apply the Bouwens et al. (2021) z≈5.6–6.5 dropout colour cuts to the Merlin et al. (2021) z≈1.2–1.5 GOODS-S catalog after convolving with the actual photometric-error distribution, identifying which faint z≈1.3 galaxies would leak into the z≈6 sample (the 8 cross-matched GOODS-S sources and the Vanzella et al. 2008 spectroscopically confirmed interloper can anchor the matching). Measure the angular ACF of this selected leak population and compare it with the ACF of the general faint (m≥24) z≈1.3 population used in §3.2. Let b ≡ ω_leak/ω_faint; if b is consistent with unity, the concern is retired for that field. If b<1 by more than the bootstrap error, recompute the §3.2 sensitivity threshold as 5.5%/b and check whether the recalibrated threshold rises above 5.5% at the 90% detection level; if it does, the GOODS-S upper limit is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Both headline numbers depend on the unstated equivalence between the true interloper population and a uniformly random draw from the faint population at the interloper redshift. In §3.2 the mock calibration selects contaminants by Monte Carlo, 'randomly select contaminants from interlopers' (§3), so every faint z≈1.30 mock galaxy has the same leakage probability f_leak, independent of SED, mass, redshift, or environment. That construction forces the injected contaminants to have exactly the angular two-point function of the general faint z≈1.30 population. The same function is then used in Eq. (14) to convert the observed cross-correlation into f_cont and in Fig. 4 to set the 5.5% sensitivity threshold. Real LBG colour selection does not work this way: Balmer/4000-Å break interlopers are selected through their SEDs, so the leak population is a colour- and mass-selected subset, not a random one. If its ACF is b times the parent ACF, the threshold and inversion both rescale by 1/b, and a less-clustered leak population pushes the null-detection limit above 5.5%. The paper validates only the average z≈1.30 ACF against the observations (§3.1), not the ACF of the objects that actually leak. The same uniform-leakage assumption enters the BoRG counts-in-cells simulation (§4.3), where the expected faint z≈2 count in each field is taken to be a deterministic multiple of the observed bright count and f_leak is applied uniformly; field-to-field scatter in the faint-to-bright ratio is not modelled. The 8/3 objects found in both the z≈1.3 and z≈6 catalogs (§2.1) are the closest real examples of leakers but are removed from the low-z catalog and never used to calibrate the leak ACF.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two estimators of contamination in photometrically selected Lyman-break galaxy samples based on spatial correlation with the parent population of lower-redshift interlopers. For a large contiguous field, the method uses the angular cross-correlation between z~6 galaxies and z~1.3 galaxies; applying this to CANDELS GOODS-S and XDF yields no significant signal, and mock observations from IllustrisTNG are used to claim that the contamination fraction is below 5.5% at 90% confidence. For a multi-field survey, the method uses counts-in-cell Pearson correlation between z~8 and z~2 galaxy counts; applying this to BoRG gives a leakage fraction f_leak=2.90±2.38% and a contamination fraction f_cont=62+13-39%, consistent with previous estimates. The paper is clearly written and the proposed framework is a conceptually attractive independent check of contamination, but several calibration and statistical issues need to be addressed before the headline limits can be accepted.","tokens_in":19926,"tokens_out":12117,"duration_ms":113163,"significance":"If the calibration assumptions are met, the paper offers a valuable independent check of contamination in LBG samples that does not rely on SED fitting or spectroscopic follow-up. The use of public IllustrisTNG simulations and external luminosity functions is a strength, and the method produces falsifiable predictions: for GOODS-S, contamination above ~5.5% should produce a detectable cross-correlation signal. The BoRG application, despite large uncertainties, provides a consistency test against earlier contamination estimates. However, the strength of the conclusions is currently limited by the assumption that interlopers are a uniformly random subset of the faint lower-redshift population, by an apparent overextension of the sensitivity limit to XDF, and by the statistical treatment of the BoRG correlation measurement.","major_comments":[{"comment":"The calibration assumes that contaminants are a uniformly random subset of the faint z~1.30 population. In the Monte Carlo, every faint z=1.30 mock galaxy has the same leakage probability, so the injected contaminants inherit exactly the angular clustering of the parent population; the same ACF is then used in Eq. (14) to convert the measured cross-correlation into f_cont and in Fig. 4 to set the 5.5% threshold. Real Balmer/4000-Å break interlopers are SED-selected and could be more or less clustered than the general faint population; if their ACF differs by a factor b, both the threshold and the inferred f_cont rescale by 1/b. The paper validates only the average z~1.3 ACF against the observations (Section 3.1), not the ACF of the objects that actually leak. Please quantify the sensitivity of the headline limits to this assumption, for example by using spectroscopically confirmed interlopers or an SED-selected mock interloper population.","section":"Section 3.1-3.2, abstract"},{"comment":"The 90%-confidence sensitivity threshold is derived from mock fields that match the GOODS-S geometry: the cutout boxes are 'approximately the size of the GOODS-S area.' The abstract and Section 3.3 extend the <5.5% conclusion to XDF, which has an area of 4.7 arcmin^2 versus 64.5 arcmin^2 for GOODS-S and a different sample size of z~6 galaxies. The same sensitivity threshold does not automatically apply to XDF. Either run the Monte Carlo with XDF-geometry mocks, including the smaller number of z~6 galaxies, or restrict the claim to GOODS-S.","section":"Section 3.1-3.2, abstract"},{"comment":"The BoRG contamination estimate rests on a Pearson correlation coefficient of 0.05±0.17, which Section 4.2 states is consistent with zero. The weighting scheme in Eq. (15) uses a Gaussian likelihood in r to compute the weighted mean f_leak=2.90±2.38%, yielding f_cont=62+13-39%. This is not a proper posterior: because the observed r is within 1σ of both zero and the simulated f_leak=0 value (0.01±0.17), the quoted 1σ interval that excludes zero contamination is not justified. The summary's phrase 'we detected evidence of number counts correlation' is also stronger than the data support. Please replace the weighting with a likelihood-based fit that accounts for the full noise distribution, or report the result as an upper limit.","section":"Section 4.4, Eq. (15), summary"},{"comment":"The simulation sets the expected faint z~2 count in each BoRG field to a deterministic multiple of the observed bright count, using the luminosity-function ratio from Marchesini et al. (2012), and then applies only Poisson noise. This ignores field-to-field cosmic variance and stochasticity in the faint-to-bright ratio, which are significant for 39 BoRG fields of roughly 4 arcmin^2 each. If the faint counts fluctuate independently of the bright counts, the f_leak-to-cP mapping and hence the contamination estimate change. Please include a model for this scatter or justify its neglect with a quantitative argument.","section":"Section 4.3"},{"comment":"The method assumes zero intrinsic cross-correlation between the z~6 and z~1.3 populations, but gravitational magnification of z~6 sources by foreground z~1.3 structures can produce a positive angular cross-correlation even with no contamination. The paper cites magnification bias as a concern for luminosity functions but does not quantify its contribution to Eq. (14). If lensing contributes at a level comparable to the 5.5% contamination threshold, the reported limit changes. A concrete test would be to cross-correlate the z~6 sample with a foreground sample at a redshift that cannot be an interloper (e.g., z~0.5), or to estimate the expected magnification signal from the z~1.3 dark-matter halos in the simulations.","section":"Section 2.2, Eq. (14)"}],"minor_comments":[{"comment":"In the third paragraph of Section 2.2, 'the cross-correlation function between galaxies at z~1.3 and z~2' should read 'z~6' rather than 'z~2'.","section":"Section 2.2"},{"comment":"Only ten bootstrap resamples are used to estimate the covariance matrix. Given the resulting noise and the ridge-regression stabilization, a brief sensitivity test on the ridge constant c=0.0001 would help the reader assess the robustness of the fit.","section":"Section 2.2"},{"comment":"The mock field size is stated as 0.12 x 0.2 degrees, which is approximately 86 arcmin^2, while the GOODS-S area is 64.5 arcmin^2; the match is only approximate and could be stated more precisely.","section":"Section 3.1"},{"comment":"The sentence 'C+19 discard two fields because they affected by star overdensity' contains a missing verb; it should read 'because they are affected by...'.","section":"Section 4.1"},{"comment":"For a methods paper, sharing the analysis code would aid reproducibility; the current statement only covers the data.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of MNRAS and the proposed method is timely given the large JWST samples expected. The main concern is the calibration assumption of random interlopers, which affects both headline numbers, together with the XDF overclaim and the statistical treatment of the BoRG result. I believe these issues are addressable within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is a genuine methodological contribution: it shows how spatial correlation with lower-redshift galaxies can serve as an independent contamination check for Lyman-break samples, in two survey geometries. The cross-correlation and counts-in-cells tools are established, but applying them to estimate contamination in z~6-8 LBG samples, with a simulation-calibrated sensitivity threshold, is new. The GOODS-S limit (<5.5% at 90% confidence) and the BoRG estimate (62+13-39%) are consistent with previous results, and the authors are appropriately cautious about the BoRG uncertainties.\n\nThe paper does good things: the IllustrisTNG mock catalogs include geometry and depth effects, the analysis uses external luminosity functions, and the authors flag the one spectroscopically confirmed interloper, giving a minimum contamination of 0.5%. That transparency is welcome.\n\nThe soft spots are real but not fatal. The main one is the assumption that contaminants are a random subset of the faint z~1.3 (or z~2) population. The Monte Carlo injection assigns equal leakage probability to every faint galaxy, so the injected contaminants have exactly the same angular clustering as the parent sample. Then the same clustering is used to convert the measured cross-correlation into f_cont and to set the 5.5% threshold. If real Balmer-break interlopers cluster differently—they are typically redder and more massive, so likely more clustered—the limit and the inversion shift. The paper should present the 5.5% as conditional on this assumption, and ideally test the sensitivity by injecting a biased subset. This is a caveat, not a fatal flaw, but the 90% confidence statement is stronger than the evidence.\n\nThe BoRG side has a similar issue: the expected faint z~2 counts are set to a deterministic multiple of the observed bright counts, so clustering between faint and bright z~2 galaxies is not modeled, and lensing magnification is ignored. The Pearson correlation has large error bars, so the 62% estimate is consistent with the null at about the 1-sigma level. The authors acknowledge the large uncertainty, but the systematic assumptions deserve more discussion.\n\nOverall, the paper deserves a serious referee. The method fills a practical need for cheap contamination estimates, especially for upcoming JWST pure-parallel surveys. The math is standard and the data handling is transparent. I would send it to review, with the expectation that the authors tighten the language about the sensitivity limit and explicitly test the random-subset assumption.","headline":"Useful, honest method paper for estimating LBG contamination via spatial correlation, with a caveat about the random-subset assumption.","tokens_in":20506,"tokens_out":5113,"would_cite":false,"duration_ms":49313,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Spatial clustering estimates contamination in Lyman-break galaxy samples","keywords":["Lyman-break galaxies","contamination fraction","cross-correlation function","counts-in-cells","high-redshift galaxies","photometric selection","GOODS-S","BoRG survey"],"falsifier":"A spectroscopic redshift survey of all the z~6 candidates in GOODS-S and XDF that finds an actual contamination fraction above 5.5% would contradict the paper's upper limit. Similarly, measuring the angular autocorrelation of the faint z~1.3 interlopers and finding that Balmer-break galaxies are significantly more clustered than the full faint population would invalidate the leakage-fraction mapping used in the Monte Carlo simulations.","tokens_in":19354,"feed_emoji":"🔭","tokens_out":2557,"duration_ms":26601,"temperature":0.7,"pith_summary":"This paper introduces a new way to estimate how many galaxies in photometrically selected high-redshift samples are actually lower-redshift interlopers. Instead of relying on spectroscopy or simulations of galaxy colors, the method measures whether the high-redshift candidates cluster spatially with the population of galaxies that could mimic them. If there is no clustering, contamination is low; if there is clustering, some of the candidates are likely interlopers. The method is applied to two survey strategies and validated against mock observations, giving an upper limit below 5.5% contamination for z~6 samples in GOODS-S and XDF, and a 62% contamination estimate for the z~8 BoRG sample that agrees with earlier work.","feed_headline":"New method measures galaxy-sample contamination via clustering","feed_subtitle":"Cross-correlation with lower-redshift galaxies limits z~6 contamination to <5.5% and finds ~62% in the BoRG z~8 sample.","key_machinery":"The cross-correlation function with a modified Landy-Szalay estimator, together with a counts-in-cells Pearson correlation for multi-field surveys. The method exploits the physical independence of galaxies at the two redshifts: any measured angular correlation must come from interlopers sharing the same parent population. Monte Carlo simulations with IllustrisTNG mocks convert a leakage fraction (probability that a faint interloper is misidentified as a high-redshift galaxy) into a contamination fraction and predict the signal expected at each contamination level.","core_discovery":"The paper claims that contamination in Lyman-break galaxy samples can be quantified from the angular clustering between the high-redshift candidates and galaxies at the redshift where a Balmer break mimics the Lyman break. For a single contiguous field, the cross-correlation function between z~6 galaxies and z~1.3 galaxies in GOODS-S and XDF is consistent with zero; Monte Carlo simulations based on IllustrisTNG mocks show that contamination above 5.5% would be detected 90% of the time, so the true contamination is below 5.5% at 90% confidence. For a survey of many independent pointings, the paper instead uses counts-in-cells: the number of z~8 candidates in each BoRG field is compared with the number of faint z~2 galaxies, and a Pearson correlation coefficient of 0.05 +/- 0.17 is measured. Matching this to simulations yields a contamination fraction of 62+13-39% for the BoRG z~8 sample, consistent with the previous 42% estimate of Bradley et al. (2012). These results demonstrate that the spatial-correlation approach works as an independent check of contamination.","pith_inferences":["The method could be extended to JWST-era samples at z>10, where contamination from intermediate-redshift galaxies is expected to rise sharply; a null cross-correlation signal would then set an upper limit that drives the practical survey depth needed.","The contamination fraction inferred from clustering assumes that interlopers trace the angular clustering of the general faint interloper population. If Balmer-break galaxies are more strongly clustered, the limit would be less constraining; checking this would require measuring the ACF of the interloper population itself.","Counts-in-cells uses a small number of fields (39) with Poisson noise; the 62% estimate for BoRG is consistent with the previous 42% only within the large error bars, and future surveys with more pointings will decide whether the technique can discriminate between competing contamination models.","The cross-correlation method could in principle be turned around: instead of contamination, it might measure the magnification bias between the two redshift populations, if strong lensing contributes to the apparent clustering signal."],"forward_implications":["If the method is correct, photometrically selected LBG samples in large contiguous surveys can be validated without expensive spectroscopy, simply by measuring the cross-correlation against the interloper population.","The depth dependence found here predicts that contamination fractions drop for deeper surveys because the high-redshift luminosity function is steeper than that of the interlopers; deeper future surveys should show lower contamination at fixed leakage.","Applying the counts-in-cells method to upcoming JWST pure-parallel surveys such as PANORAMIC should yield much tighter contamination constraints for z~8 and higher redshift samples.","The measured z~8 BoRG contamination of ~62% implies that current luminosity function estimates at z~8 from such shallow surveys may be substantially biased if not corrected."],"supporting_citations":[{"why":"Provided the clustering-based cross-correlation approach for refining photometric redshifts that this paper adapts to contamination estimation.","marker":"Ménard et al. 2013"},{"why":"Demonstrated cross-correlation techniques for estimating redshift distributions, providing the methodological foundation for using spatial correlation to infer contamination.","marker":"Rahman et al. 2016a,b"},{"why":"Supplies the formalism that connects the observed cross-correlation function to the true correlation functions and contamination fractions.","marker":"Awan & Gawiser 2020"},{"why":"Introduced the counts-in-cells method for measuring clustering in multi-field surveys, which the BoRG number-count analysis is based on.","marker":"Robertson 2010"},{"why":"Provides the previous contamination estimate (42%) for the BoRG z~8 sample, used as the baseline comparison for the new measurement.","marker":"Bradley et al. 2012"},{"why":"Supplies the z~6 galaxy catalog for the GOODS-S and XDF cross-correlation analysis.","marker":"Bouwens et al. 2021"},{"why":"Supplies the intermediate-redshift (z~1.3) galaxy catalog used as the interloper population in the GOODS-S and XDF analysis.","marker":"Merlin et al. 2021"},{"why":"Describes IllustrisTNG, the cosmological simulation from which mock catalogs are generated for the contamination calibration.","marker":"Springel et al. 2018"}],"fun_headline_variants":["Clustering reveals contamination in high-z galaxy samples","New method estimates galaxy sample contamination via clustering","Spatial correlation quantifies galaxy interloper contamination","Contamination in LBG samples measured from clustering","Galaxy sample purity assessed via cross-correlation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The depth of the analysis depends on the assumption that the interlopers are a random subset of the faint intermediate-redshift galaxy population, sharing that population's angular clustering; if actual Balmer-break interlopers are more or less clustered than the general faint population, the derived 5.5% limit and the leakage-to-contamination mapping would change.","fun_headline_variants_meta":{"raw":{"variants":["Clustering reveals contamination in high-z galaxy samples","New method estimates galaxy sample contamination via clustering","Spatial correlation quantifies galaxy interloper contamination","Contamination in LBG samples measured from clustering","Galaxy sample purity assessed via cross-correlation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1590,"prompt_tokens":1142,"completion_tokens":448,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":758,"completion_tokens_details":{"reasoning_tokens":377}},"tokens_in":758,"tokens_out":448,"duration_ms":5171,"temperature":1.0,"reasoning_tokens":377,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:54:18.154344+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A spectroscopic redshift survey of all the z~6 candidates in GOODS-S and XDF that finds an actual contamination fraction above 5.5% would contradict the paper's upper limit. Similarly, measuring the angular autocorrelation of the faint z~1.3 interlopers and finding that Balmer-break galaxies are significantly more clustered than the full faint population would invalidate the leakage-fraction mapping used in the Monte Carlo simulations.","supporting_citations":[],"review_version":1}