{"id":"9e28f066-4a9d-4a39-b718-2b27b7af9420","arxiv_id":"2602.20405","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A random-forest-guided set of color cuts in deep multi-band imaging selects z=1.1–1.6 emission-line galaxies at 1372 deg^-2 with 84% redshift-range success, more than doubling the net DESI ELG yield.","lead":"Using deep multi-band images, the authors find simple color cuts that pick out star-forming galaxies at redshift 1.1–1.6 about twice as efficiently as the current DESI target selection. If the gains are confirmed on independent sky areas, a future DESI-like survey could roughly double the high-redshift sample and shrink BAO distance errors by about a factor of two.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline yields are optimized and evaluated on the same spec-truth sample; without a held-out validation, the 84% purity and 1372 deg^-2 yield may be in-sample overestimates.","rationale":"I read the paper as a target-selection study whose main claim is that simple color cuts with deep grizy photometry can roughly double the net yield of secure z=1.1-1.6 ELGs relative to DESI, with f_reliable=89% and f_z=84%. For that claim to hold, the measured success rates and yield must not be inflated by the optimization procedure. The most insecure point is not the photometry or the color cuts themselves, but that the free parameters of the cuts are tuned to the very sample on which performance is reported. The loss function in Eq. 7 explicitly includes the yield target, so achieving Σ_yield=1372 deg^-2 is partly by construction. The paper does not provide cross-validation, error bars, or a truly independent test with the final cuts. The pilot sample is a valuable sanity check but is not a fully independent validation: it was used to guide the choice of color space and could not test the full range of cuts, and it lies in the same COSMOS field. The reader's weakest_assumption about representativeness is related, but I see the more immediate threat as in-sample overfitting; cosmic variance would add scatter, while overfitting would bias the central comparison. This does not invalidate the qualitative conclusion: two spectroscopic samples give similar (though not identical) optimized cuts, and the pilot sample yields 89%/77%/1375 deg^-2 with analogous cuts. But the exact quoted factors over DESI should be treated as upper estimates until a held-out validation is done. Therefore I do not change the reader's CONDITIONAL verdict; I would emphasize the need for a clean validation test.","tokens_in":19875,"tokens_out":5640,"duration_ms":55768,"concrete_test":"Apply the exact final cuts from §4.2 to the DESI-II pilot sample (Appendix A) without re-optimizing any parameter, and compute f_reliable, f_z, and Σ_yield using the same definitions; also perform a 5-fold cross-validation on the spec-truth sample, splitting by tile or spatial region. If the pilot-sample or held-out metrics fall materially below the reported 89%/84%/1372 (e.g., f_z < 75% or yield < 1100 deg^-2), the headline improvement is partly an artifact of in-sample optimization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison in Table 1 is computed on the same data used to set the selection. In Section 4.2, the four cut parameters (g_fiber<24.33, i-y-0.16>r-i, i-y>0.43, i-z>0.45) are optimized by minimizing L = -f_z*100 + w*(Σ_yield-1370)^2 (Eq. 7) directly against the spec-truth sample. The reported f_z=84% and Σ_yield=1372 deg^-2 are then the in-sample values of the objective, so the ≈2.4x and ≈2.1x improvements over DESI are not independent measurements. The DESI-II pilot sample (Appendix A) provides partial support but is not a clean held-out test: it was used to select the r-i/i-y/i-z color space, had restricted targeting that prevented exploration of positive offsets in the diagonal cut, and was also observed in the same COSMOS field. Thus the possibility remains that the cuts are overfit to the COSMOS field and to the specific spec-truth target selection (Eqs. 1a-c, 2a-c), and the true yield/purity on a new field could be lower.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes and tests color cuts for selecting z=1.1–1.6 emission-line galaxies (ELGs) using deep HSC grizy photometry and LSST-like data, with the goal of augmenting DESI-II ELG samples. A random forest trained on COSMOS2020 photometric redshifts guides the choice of r−i, i−y, and i−z colors; the cuts are then fine-tuned against a dedicated DESI 'spec-truth' spectroscopic sample of ~87,000 objects in COSMOS. The optimized selection is reported to achieve f_reliable=89%, f_z=1.1–1.6=84%, and Σ_yield=1372 deg^-2, versus 69%, 34%, and 660 deg^-2 for DESI ELGs (Table 1). The paper then uses a Fisher forecast to argue that combining this sample with DESI ELGs would triple the net ELG density and reduce BAO distance errors by roughly a factor of two (Section 2, Fig. 1).","tokens_in":20267,"tokens_out":6433,"duration_ms":60850,"significance":"The target-selection problem is timely and the empirical approach is appropriate. The paper’s strengths include the large dedicated spec-truth spectroscopic sample, explicit definitions of the success metrics, a test with degraded LSST Year-2 photometry (Appendix B), and an earlier pilot sample (Appendix A) that gives qualitatively similar results. Public data availability is another positive feature. However, the headline numbers are in-sample by construction: the cuts are optimized on the same spec-truth sample used to compute Table 1, and no uncertainties accompany the density estimates. The pilot sample provides partial independent support but is not a clean held-out test. With a proper held-out validation and uncertainty propagation, the method could make a useful contribution to DESI-II/LSST target selection.","major_comments":[{"comment":"The 89%/84%/1372 numbers are not independent measurements. The four free parameters (g_fiber, diagonal offset in i−y vs r−i, i−y cut, i−z cut) are obtained by minimizing L = −f_z×100 + w×(Σ_yield − 1370)^2 on the same spec-truth sample used to compute Table 1. The reported f_z=84% and Σ_yield=1372 deg^-2 are therefore values of the objective at the optimum, not out-of-sample performance. The comparison factors of 2.4× and 2.1× inherit this optimism. The pilot sample (Appendix A) gives 77% f_z and 1375 deg^-2, but it was used to select the r−i/i−y/i−z color space, its targeting prevented exploration of positive diagonal offsets, and it is in the same COSMOS field. I request a genuine held-out validation, e.g., split the spec-truth sample by tile/region, or apply cuts optimized on one half to the other half, and report both training and validation metrics.","section":"§4.2, Eq. (7), Table 1"},{"comment":"The surface-density estimates have no uncertainties and are derived from a single ~16 deg^2 HSC region, with masked-out areas ignored. The 1372 deg^-2 yield enters directly into the Fisher forecast; cosmic variance plus mask losses could shift the predicted BAO improvement. Please provide bootstrap/jackknife uncertainties over subfields or an analytic cosmic-variance term, and propagate them into the factor-of-two claim. Also state the effective area after masks rather than '~16 deg^2, ignoring masked-out regions'.","section":"§3.1, Eq. (5); §2 and Fig. 1"},{"comment":"Completeness of the spec-truth sample is asserted but not demonstrated. The statement that the broad cuts 'should include all of the ELGs at z=1.1–1.6' is load-bearing because the final cuts cannot select objects outside the spec-truth color-magnitude box. If a non-negligible population of z=1.1–1.6 ELGs lies outside those cuts (e.g., redder r−z or different g−r), the reported yield and purity are biased. Please quantify the completeness of the spec-truth selection against the COSMOS2020 photo-z sample, or test sensitivity by widening the spec-truth cuts and re-optimizing.","section":"§3.3, Eq. (1a)–(2c)"},{"comment":"The redshift range success rate f_z is computed using all objects with t>700 s, while f_reliable is restricted to 700<t<1400 s. The product f_reliable × f_z is then treated as the success rate for typical DESI exposures. If longer exposures preferentially select different galaxy populations (e.g., fainter or redder objects), the product may not represent a 1000 s exposure. Please check the sensitivity of f_z to the exposure-time cut, or report f_z for the 700<t<1400 s subsample as well.","section":"§4.2, Eqs. (2)–(4)"}],"minor_comments":[{"comment":"The target yield of 1370 deg^-2 is set by the same FishLSS forecast used later to claim a factor-of-two BAO improvement. Because Fisher errors scale roughly as n^-1/2 in the shot-noise regime, the improvement is essentially imposed by the choice of target. The abstract’s 'reducing uncertainties ... by a factor of ~2' should be framed as a conditional forecast, not an empirical validation.","section":"§2 vs §4.2"},{"comment":"No uncertainties are given for f_reliable, f_z, or Σ_yield. These are binomial fractions for f_reliable/f_z and a count density for Σ_yield; adding Poisson/binomial errors would make the comparison to DESI more meaningful.","section":"Table 1"},{"comment":"The random forest is reported to achieve an ROC AUC of 0.960, but no cross-validation or train/test split is described. Please clarify whether this AUC is in-sample or out-of-sample.","section":"§4.1"},{"comment":"The phrase 'only around 17.4% were also LOP targets' uses the abbreviation LOP without defining it. Please spell out the program name or rephrase.","section":"§4.3, p. 9"},{"comment":"The reliable-redshift cut uses 'log(flux [OII])' without stating the units of the [O II] flux or the base of the logarithm. Please specify.","section":"§3.3, Eq. (1)"},{"comment":"The initial color cuts are written as 'offset 1' and 'offset 2' set to zero. It would be clearer to state explicitly that these are starting points for the optimization in §4.2, not final cuts.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a practical and timely problem for DESI-II/LSST target selection. The central issue is that the headline metrics are optimized and evaluated on the same spec-truth sample, so the quantitative gains over DESI are likely optimistic. The pilot sample offers some reassurance but is not a clean held-out test. I do not see grounds for rejection, because the core method is sound and the problem can be addressed by re-analysis within the manuscript's scope: report training/validation splits, add uncertainties, and reframe the BAO claim as conditional. I would not require new observations; the existing data should suffice for a held-out test."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nHere's the short version: the headline numbers are in-sample. The four cuts (g_fiber<24.33, i−y−0.16>r−i, i−y>0.43, i−z>0.45) were tuned by minimizing Eq. (7) against the spec-truth sample, and the reported 84% redshift-range success and 1372 deg^-2 yield are the values of the objective function at the optimum. So the ~2.4x and ~2.1x improvements over DESI are not independent measurements.\n\nThat said, this is a competent and useful paper. What's new is the empirical validation: 68k DESI spectra in COSMOS with broad color coverage, used to test and refine a random-forest-guided color selection. The comparison to DESI ELGs is instructive, and the pilot sample in Appendix A provides partial confirmation: using different, more restrictive targeting cuts, it gets 77% and 1375 deg^-2. That tells me the effect is real, not pure overfitting. The robustness test to LSST Y2 depths (Appendix B) also produces nearly identical metrics.\n\nThe soft spots beyond the in-sample issue: no error bars on the yield estimates (a ~16 deg^2 field, so cosmic variance is non-trivial), and the DESI comparison is not fully controlled—the DESI numbers come from Raichoor et al. with different targeting and selection, though the authors do make an effort to match exposure times. Also, the whole exercise is an optimization; the BAO improvement is a Fisher forecast, not a measurement.\n\nThe paper is transparent—plot data on Zenodo, clear description of the optimization, honest about the photo-z failures. Citation pattern is appropriate. I think this should go to peer review, but the referee should press for a held-out validation. The cleanest fix would be to select a second field not used in any optimization step, or at least use the pilot sample as a genuinely independent check and show that the conclusions survive with error bars on the yields.\n\nIf you're planning DESI-II target selection, this is directly relevant; otherwise it's a solid example of how to use spectroscopic truth samples to validate photometric selections. Send it to a good referee; expect revision.","headline":"Useful empirical ELG selection, but the headline purity/yield numbers are in-sample—deserves peer review with a held-out validation demand.","tokens_in":21025,"tokens_out":3294,"would_cite":true,"duration_ms":32673,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Deep five-band photometry can select z ≈ 1.1–1.6 emission-line galaxies more than twice as efficiently as the current target selection, raising the net usable surface density from 660 to 1372 per square degree.","keywords":["emission-line galaxies","baryon acoustic oscillations","target selection","color cuts","deep multi-band imaging","photometric redshifts","random forest","large-scale structure"],"falsifier":"Apply the optimized color cuts to a large, independent spectroscopic sample spread over a different area of sky and measure the reliable-redshift fraction and net high-z surface density; a clear shortfall relative to about 89% and 1370 deg⁻² would contradict the central claim.","tokens_in":19826,"feed_emoji":"🔭","tokens_out":7006,"duration_ms":60451,"temperature":0.7,"pith_summary":"Deep multi-band imaging can do far better than the broadband three-band photometry used today at finding the star-forming galaxies that trace large-scale structure at z ≈ 1.1–1.6. The paper shows that a simple set of color cuts—using r−i, i−y, and i−z plus a g-band fiber magnitude limit—selects galaxies in that range with an 84% success rate and a net surface density yield of 1372 per square degree, compared to 34% and 660 for the current DESI ELG selection. If these numbers hold across the survey footprint, combining the new selection with the current one would roughly triple the effective high-redshift ELG density and reduce baryon acoustic oscillation distance-scale errors by about a factor of two. A reader should care because this is a low-cost way to sharpen dark-energy constraints without building new hardware.","feed_headline":"Deep imaging triples high-z galaxy yield for dark-energy surveys","feed_subtitle":"Simple color cuts from deep five-band imaging lift reliable z=1.1–1.6 galaxy yield from 660 to 1372 per square degree.","key_machinery":"The mechanism is the optimized color-cut set: g_fiber < 24.33, i−y−0.16 > r−i, i−y > 0.43, and i−z > 0.45. These cuts use the position of the 4000 Å break and the [O II] doublet, which move through the i and z bands at z ≈ 1.1–1.6, to reject low-redshift interlopers and galaxies whose [O II] would fall outside the spectrograph's wavelength window. The cuts were seeded by a random forest classifier—which identified r−i, i−y, and i−z as the most informative colors—and then fine-tuned against spectroscopic truth data. The same machinery remains effective when the photometry is degraded to early-survey depths, showing the selection is not fragile to modestly shallower imaging.","core_discovery":"The central claim is that simple color cuts in r−i, i−y, and i−z, applied to deep grizy photometry, select z = 1.1–1.6 emission-line galaxies with a redshift measurement success rate of 89%, a correct-redshift-range success rate of 84%, and a net surface density of 1372 deg⁻², versus 69%, 34%, and 660 deg⁻² for the current DESI ELG sample. The cuts were first designed by training a random forest on a deep, many-band photometric catalog, then calibrated and refined using a large spectroscopic truth sample of about 87,000 ELG-like objects. The result is a target class that is nearly pure in the desired redshift range and roughly twice as dense in useful targets per square degree, without requi","pith_inferences":["If the selection transfers beyond the calibration field, future surveys could concentrate spectroscopic fibers on the highest-redshift bins where dark-energy constraints are weakest, rather than spreading effort evenly across the full sample.","Because only about 17% of the new targets overlap the current ELG sample, the gain represents a largely new, fainter population; using it for clustering analyses will require measuring its bias, which the paper does not do.","The same color-cut logic could be extended with additional bands, such as u-band or narrow-band imaging, to push the selection beyond z ≈ 1.6 or to build photometric samples for CMB-lensing cross-correlations at z > 1—directions the paper notes but does not develop."],"forward_implications":["Combining the new selection with the current ELG sample would increase the net ELG number density at z = 1.1–1.6 by a factor of about 3, moving it out of the shot-noise-limited regime.","Fisher forecasts in the paper indicate this would reduce uncertainties on the BAO distance parameters α⊥ and α∥ by roughly a factor of 2 in the highest useful redshift bin, 1.3 < z < 1.5, where current constraints are weakest.","The selection remains close to optimal when photometry is degraded to two-year depths of a planned wide survey (net yield 1366 deg⁻² versus 1372 deg⁻²), meaning early survey data would already be sufficient.","Re-optimizing the cuts for a slightly broader redshift window (1.05 < z < 1.65) yields qualitatively similar performance, so the method can be tailored to a survey's preferred redshift range."],"fun_headline_variants":["Deep grizy cuts double high-z ELG density for DESI-like surveys","Deep imaging triples effective ELG sample, halves BAO error","Simple color cuts lift high-z ELG success to 84%","Rubin-like deep photometry selects 2x more reliable z~1.4 ELGs"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The result depends on the spectroscopic truth sample drawn from a single small (about 16 square degree) field being representative of the high-redshift ELG population and of the larger survey footprint; if that field's density or color distribution is atypical, the 1372 deg⁻² yield and the forecast BAO gain will not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Deep grizy cuts double high-z ELG density for DESI-like surveys","Deep imaging triples effective ELG sample, halves BAO error","Simple color cuts lift high-z ELG success to 84%","Rubin-like deep photometry selects 2x more reliable z~1.4 ELGs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001096,"raw_usage":{"total_tokens":4466,"prompt_tokens":854,"completion_tokens":3612,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":3529}},"tokens_in":598,"tokens_out":3612,"duration_ms":22647,"temperature":1.0,"reasoning_tokens":3529,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T21:21:11.518579+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the optimized color cuts to a large, independent spectroscopic sample spread over a different area of sky and measure the reliable-redshift fraction and net high-z surface density; a clear shortfall relative to about 89% and 1370 deg⁻² would contradict the central claim.","supporting_citations":[],"review_version":1}