{"id":"3cf6894f-3446-4aac-a4e5-4139fdb13ae7","arxiv_id":"2607.08562","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"XShooter redshifts for 58 lenses and 57 sources from DESI Legacy Survey candidates show no measurement bias, so the source-z distribution can calibrate large imaging surveys.","lead":"VLT/XShooter spectroscopy of 67 DESI Legacy Survey strong-lens candidates yields redshifts for 58 lenses and 57 sources. The measured source redshift distribution shows no obvious selection bias and can calibrate large imaging-survey lens analyses.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Representativeness claim rests on color–magnitude similarity, not a direct test of redshift completeness vs. true source-z.","rationale":"The Reader correctly isolates the softest link: the leap from “no obvious color/magnitude bias” (Fig. 4) to “the measured source-z distribution is representative.” That leap is load-bearing for the paper’s stated utility (calibration of large imaging surveys) yet is only qualitatively supported. The catalog itself (Table 2, quality flags, spectral figures) is solid and the observational efficiency numbers are useful; nothing is internally inconsistent. The concern is therefore not a reason to reject, but it does keep the verdict at CONDITIONAL until a quantitative completeness-vs-z test is performed. My concrete test is a direct, low-cost check that would settle whether the concern lands. I agree with the Reader’s identification of the weakest assumption and with the CONDITIONAL / high-confidence assessment.","tokens_in":26123,"tokens_out":615,"duration_ms":6447,"concrete_test":"Split the 67 observed systems into those with and without a secure source redshift (Qz≥3). For every system compute photometric redshift (or at least the expected observed-frame wavelengths of [O II], [O III], Hα) from the Legacy Survey photometry; then test whether the success fraction is statistically consistent with being constant across photometric-z bins (or across bins of line-wavelength proximity to bright sky lines). If the success rate drops by ≳20–30 % in any bin that is populated in the parent sample, the representativeness claim is weakened and a selection-function correction is required before the distribution can be used for calibration.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (Abstract; §3–4) is that the measured source-z distribution is “likely representative of the true one” because “we find no particular bias associated to the redshift measurement operation.” The only quantitative support is Fig. 4 (g−r / r−z colors of systems with successful Qz>2 redshifts vs. the observed sample) plus the statement that sources of any magnitude down to zAB=25.2 were recovered. Color–magnitude similarity does not constrain whether the spectroscopic success function S(z) is flat across the observed source range 0.28 < zs < 4.34. Emission-line detectability (the dominant Qz=4 channel for sources) depends on which lines fall into the NIR arm, sky-line density, and continuum S/N; these vary strongly with redshift. A color-matched subsample can still be incomplete at redshifts where [O II]/[O III]/Hα are buried or weak, so the histogram in Fig. 2 may not be an unbiased draw from the parent source-z distribution. The paper never constructs or applies a selection function, nor does it compare the successful vs. failed subsets in redshift-sensitive observables beyond broadband colors.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper reports VLT/XShooter spectroscopy of 67 grade-A strong-lens candidates from the DESI Legacy Imaging Surveys (Huang et al. 2021; Storfer et al. 2024) at Dec < −20°. Redshifts are measured for 58 lenses and 57 sources (45 systems with both), with quality flags Qz = 1–4 based on continuum and emission-line features. The sample yields a source redshift range 0.28 < zs < 4.34 (median zs = 1.54) and a lens distribution mostly at z < 1. The authors also note two sources with outflow signatures and seven with rotating disks. On the basis of magnitude and g−r / r−z color comparisons (Figs. 3–4) between systems with successful redshifts (Qz > 2) and the observed sample, they conclude that no particular bias is associated with the redshift measurement operation, so the measured source-z distribution is likely representative of the true parent distribution and can be used to calibrate analyses in large imaging surveys.","tokens_in":26403,"tokens_out":1039,"duration_ms":12265,"significance":"A homogeneous spectroscopic characterization of a southern DESI Legacy Survey lens sample is a useful community resource, especially given the forthcoming Euclid and LSST strong-lens catalogues. The 82 % / 85 % success rates, the tabulated redshifts with quality flags (Table 2), and the extensive per-system 1D/2D spectral figures constitute a concrete, reusable data product. The identification of outflows and rotating disks adds modest astrophysical interest. The central claim that the measured source-z histogram can be treated as representative for survey calibration is of clear practical value if it holds, but it currently rests on an incomplete empirical test.","major_comments":[{"comment":"The central claim (Abstract; end of §3; §4) that the measured source redshift distribution is “likely representative of the true one” is supported only by the visual similarity of g−r / r−z colors and magnitudes between systems with Qz > 2 and the observed sample (Figs. 3–4). Color–magnitude similarity does not constrain whether the spectroscopic success function S(z) is flat across 0.28 < zs < 4.34. Emission-line detectability (the dominant Qz = 4 channel) depends on which lines fall into the NIR arm, sky-line density, and continuum S/N—all of which vary strongly with redshift. The paper never constructs a selection function, nor does it compare successful versus failed subsets in any redshift-sensitive observable beyond broadband colors. Without such a test (or an explicit caveat that the histogram is an observed, not completeness-corrected, distribution), the claim that the sample “ca","section":null},{"comment":"§3 and Table 2 report 33 % of lens redshifts with Qz < 2 and several source redshifts with Qz = 1–2 (single-line or continuum-only). The text does not state whether the histograms in Fig. 2 or the “representative” claim include only Qz ≥ 3 systems or the full set. Because low-Qz redshifts are more susceptible to line misidentification and sky residuals (explicitly noted for several systems in the appendix figures), the paper should either restrict the science sample to a well-defined quality cut or quantify how the distribution changes when low-Qz objects are excluded.","section":null}],"minor_comments":[{"comment":"Storfer et al. is cited as 2022 in the Abstract and as 2024 in the body and references; the year should be made consistent.","section":null},{"comment":"Table 1 lists program IDs and OB counts but does not give the total on-source exposure or the fraction of grade-A nights; a one-line summary would help readers assess data quality.","section":null},{"comment":"Several appendix figure captions (e.g., DESI-024.5133-21.9299, DESI-033.8095-29.1570) discuss alternative line identifications or photometric-redshift tensions; these caveats should be cross-referenced in Table 2 or the main text so that users of the catalogue are aware of them.","section":null},{"comment":"Fig. 2 caption and axis labels would benefit from explicit statement of the quality cut (if any) applied to the plotted redshifts.","section":null},{"comment":"Minor typographical issues: “XSchooter” (p. 2), “HandK” for H and K lines, and occasional missing spaces around redshifts (e.g., z=0.295).","section":null}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a solid observational data paper whose main scientific claim is modestly overstated. Once the representativeness language is qualified and the quality-cut policy is clarified, it is suitable for A&A. The extensive spectral atlas is a genuine community asset and should not be undervalued."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a useful infrastructure paper: 58 lens and 57 source redshifts (45 systems with both) from XShooter on the southern Huang/Storfer grade-A DESI Legacy candidates, with clear Qz flags, success rates (82%/85%), and a long appendix of 1D/2D spectra that let you judge the line IDs yourself. That catalog is the real product. Parallel northern DESI, MUSE, and Keck efforts already exist; this fills the Dec < −20° gap and adds a few kinematic notes (two outflows, seven rotating disks).\n\nWhat they do well is the observational bookkeeping. Slit placement, reduction choices, quality grades, magnitude limits (lenses mostly need zAB < 20; sources go to 25.2), and the per-system figures are transparent. Citations to the parent imaging papers and the other spectroscopic campaigns are complete. No circularity: redshifts come from standard absorption/emission features against known wavelengths.\n\nThe soft spot is exactly the one the stress-test flags, and it is real but not fatal. The abstract and conclusion assert that “no particular bias” in the measurement step means the measured source-z histogram (median 1.54, range 0.28–4.34) is “likely representative” and can calibrate Stage-IV analyses. Support is only the visual color–magnitude comparison (Figs. 3–4) between successful Qz > 2 systems and the observed sample, plus the statement that sources of any magnitude were recovered. Color similarity does not test whether the spectroscopic success function is flat across redshift; emission-line detectability in the NIR arm varies with which lines land where and with sky-line density. They never build a selection function or compare successes vs. failures in redshift-sensitive observables. So the representativeness claim is qualitative and should be used with that caveat. Everything else in the paper stands without it.\n\nWho it is for: anyone building or calibrating strong-lens samples for Euclid/LSST cosmology or galaxy evolution who needs southern spectroscopic redshifts and quality flags. It is not a methods or theory paper. A serious editor should send it to referees; the catalog is solid enough to publish after the authors either quantify the selection function or soften the language around “representative.” I would cite the redshifts and flags; I would not lean on the unbiased-zs claim without further work.","headline":"Solid southern XShooter redshift catalog for DESI Legacy lenses; the “no bias / representative zs” claim is only qualitatively supported by color–magnitude plots and should be treated as provisional.","tokens_in":27020,"tokens_out":627,"would_cite":true,"duration_ms":6994,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"XShooter redshifts for 58 lenses and 57 sources show no measurement bias, so the source-z distribution can calibrate large imaging surveys.","keywords":["strong gravitational lensing","spectroscopic redshifts","DESI Legacy Surveys","XShooter","source redshift distribution","sample characterization","galaxy evolution","cosmology"],"falsifier":"A larger, deeper spectroscopic campaign of the same parent catalog that recovers redshifts for a substantial fraction of the currently failed or low-Qz sources and finds a systematically different redshift distribution would falsify the claim of representativeness.","tokens_in":27055,"feed_emoji":"🔭","tokens_out":642,"duration_ms":6026,"temperature":0.7,"pith_summary":"Huge imaging surveys are finding far more strong gravitational lenses than before, but those catalogs only become useful for cosmology and galaxy evolution once redshifts are known and selection effects are under control. This paper reports VLT/XShooter spectra of 67 grade-A lens systems found by neural-network searches in the southern DESI Legacy Imaging Surveys. Redshifts are secured for 58 lenses and 57 sources (45 systems with both). Lenses mostly lie below z ~ 1 and are measured from continuum absorption features; sources span 0.28 < z < 4.34 (median 1.54) and are measured mainly from emission lines. Color–magnitude distributions of successfully measured objects look like those of the full observed sample, so the authors conclude that the spectroscopic success function introduces no obvious bias. The measured source redshift distribution can therefore be treated as representative of the true parent population and used to calibrate forward models and training sets for Euclid, LSST and similar surveys. Along the way the spectra also flag two sources with outflow signatures and seven with rotating disks.","feed_headline":"No bias found in lens-source redshifts for survey calibration","feed_subtitle":"XShooter spectra of 67 DESI systems yield a representative source-z distribution for Euclid and LSST.","key_machinery":"Quality-flagged XShooter redshifts (Qz = 1–4) extracted from continuum absorption (lenses) and emission lines (sources), combined with a direct comparison of g−r / r−z colors and z-band magnitudes between successful measurements and the full observed set.","core_discovery":"Spectroscopic redshifts for 58 lenses and 57 sources drawn from neural-network-selected DESI Legacy Survey lenses show no color- or magnitude-dependent bias relative to the parent observed sample; the resulting source redshift distribution (median zs = 1.54) can therefore be used as a representative calibrator for large imaging-survey analyses.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["XShooter finds no bias in DESI lens-source redshifts","Unbiased zs from 58 lenses and 57 sources for survey calibration","DESI neural-net lenses yield representative source-z distribution","Median source z=1.54: XShooter sample ready for Euclid and LSST","No color or magnitude bias in measured DESI lens redshifts"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That matching color and magnitude distributions between systems with good redshifts and the full observed sample is enough to prove the spectroscopic success rate is uniform across the true source-redshift range.","fun_headline_variants_meta":{"raw":{"variants":["XShooter finds no bias in DESI lens-source redshifts","Unbiased zs from 58 lenses and 57 sources for survey calibration","DESI neural-net lenses yield representative source-z distribution","Median source z=1.54: XShooter sample ready for Euclid and LSST","No color or magnitude bias in measured DESI lens redshifts"]},"model":"grok-4.5","effort":"low","cost_usd":0.0076,"raw_usage":{"total_tokens":1778,"prompt_tokens":719,"num_sources_used":0,"completion_tokens":98,"cost_in_usd_ticks":76000000,"prompt_tokens_details":{"text_tokens":719,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":961,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":719,"tokens_out":98,"duration_ms":8565,"temperature":1.0,"reasoning_tokens":961,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T05:17:26.268369+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A larger, deeper spectroscopic campaign of the same parent catalog that recovers redshifts for a substantial fraction of the currently failed or low-Qz sources and finds a systematically different redshift distribution would falsify the claim of representativeness.","supporting_citations":[],"review_version":1}