{"id":"ad666ade-b249-4421-9005-0abe0cff7ea5","arxiv_id":"2411.12020","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"DESI DR1 large-scale structure catalogs and two-point clustering measurements are constructed with completeness and systematics corrections and validated against simulations to within 2% in inferred density.","lead":"DESI's first data release builds galaxy and quasar samples with corrections for how the telescope's fibers, imaging, and redshift measurements affect the observed density. The paper validates these samples by comparing their clustering measurements to realistic simulations, finding agreement within 2% in the inferred density field.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Priority veto mask assumes no QSO–LRG/ELG correlation; the admitted redshift overlap means masked regions track the same density field, and the Sec. 12 data-vs-altmtl comparison cannot detect a bias shared by both.","rationale":"The reader identified the same weakest assumption: the priority veto mask assumes a lack of correlation between QSOs and the lower-priority LRG/ELG samples. I agree this is the most load-bearing unresolved step. My extension is that the stated mitigation, 'simulations are important,' is not by itself sufficient: the Sec. 12 comparison uses altmtl mocks that share the same QSO population and the same priority-mask logic, so any bias induced by the mask-density correlation is present in both data and mocks and cancels in the chi^2 statistics. That makes the agreement in Table 9 a test of pipeline reproducibility rather than a test of whether the catalogs are unbiased for cosmological inference. The manuscript is unusually transparent: it states the assumption, admits it is not strictly true, and provides many robustness checks elsewhere. The covariance rescaling in Sec. 10.2 is documented and is secondary to the catalog-construction claim. The proposed test is feasible with existing mock infrastructure and would either retire the concern or demonstrate a specific, correctable bias. Because the concern is concrete and currently unverified, I would condition acceptance on running that test or an equivalent quantitative demonstration that the priority veto changes LRG/ELG clustering by less than 2% in inferred density.","tokens_in":54503,"tokens_out":4589,"duration_ms":53033,"concrete_test":"Run the 25 AbacusSummit altmtl mocks through the fiducial pipeline in two versions: (A) the real QSO catalog is used for the priority veto; (B) the QSO angular positions are replaced by random points drawn from the same QSO selection function before fiber assignment, preserving the vetoed area and its angular properties while destroying QSO–LRG/ELG clustering. Compare mean xi_l(s) and P_l(k) for LRG and ELG between A and B relative to the 25-mock scatter. If the difference in the inferred overdensity or in b_f (Eq. 12.1) exceeds 2%, or is >1 sigma of the mock scatter at the scales used in [8, 10], the priority-veto assumption fails and the Sec. 12 agreement cannot certify unbiased catalogs.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing step is the priority veto mask in Sec. 4.2. The text explicitly states its implicit assumption: 'the higher priority sample is not correlated with the sample being masked,' and then concedes 'This is not strictly true for QSO and LRG/ELG, as the samples overlap in redshift.' The veto removes 1666.7 deg^2 (20% of the footprint) from LRG/ELG wherever QSOs or strong lenses were assigned. Because QSOs at 0.8<z<2.1 trace the same large-scale density field as LRGs (0.4<z<1.1) and ELGs (0.8<z<1.6), the excluded regions are preferentially high-density regions of the lower-priority tracers. The random catalogs reproduce the mask geometry but not this density-dependent exclusion, so the Landy–Szalay estimator and window treatment in Sec. 10.1 do not remove the resulting bias. The key validation in Sec. 12 (Table 9) compares data to altmtl mocks that were processed with the same priority veto and the same correlated QSO field. A shared selection effect therefore cancels in the chi^2 comparison; the test demonstrates pipeline consistency but not that the priority veto leaves the LRG/ELG selection function unbiased. No separate quantification of the QSO-mask–density correlation is given. For ELGs in particular, with ~20% of the footprint vetoed and n(z) heavily overlapping QSOs below z=1.6, the effect could exceed the 2% overdensity-agreement claim even though Table 9 shows agreement after scale cuts and bias rescalings.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper describes the construction and validation of the DESI DR1 large-scale structure catalogs used for the 2024 cosmology papers: target selection and redshift cuts for BGS, LRG, ELG, and QSO; hardware, priority, and imaging veto masks; completeness and systematic weights; shuffled randoms; FKP weighting; correlation-function and power-spectrum estimators with window matrices and RIC/AIC corrections; covariances; and mock-based validation. The central quantitative claim is that the DR1 2-point clustering is generally in statistical agreement with simulations of DR1 to within 2% in the inferred real-space overdensity field, after the application of scale cuts and, in several cases, linear bias rescalings. The catalogs, measurements, and covariances are to be released publicly with DR1.","tokens_in":54873,"tokens_out":8910,"duration_ms":90971,"significance":"If the central claim is taken at face value, the paper delivers the calibrated public data products on which DESI 2024 BAO and full-shape analyses rest, and it does so with an unusual amount of transparency: catalog-level blinding, 128 altmtl fiber-assignment realizations, 25 AbacusSummit and 1000 EZmock realizations, detailed null tests for imaging systematics, and explicit recommendations for handling residual uncertainties. The comparison statistics in Table 9 are a useful benchmark. The strength of the claim is limited, however, by the fact that the mocks are calibrated to DESI EDR clustering, the RIC/AIC corrections are derived from the same mock suite, the Fourier-space covariance is rescaled to match the data, and the main data-mock comparison shares the same priority-veto selection. The 2% agreement is therefore best read as a pipeline-consistency and HOD-calibration check rather than an independent end-to-end validation.","major_comments":[{"comment":"Section 4.2 applies a priority veto that removes 1666.7 deg^2 (20% of the footprint) from the LRG and ELG samples wherever a QSO or strong-lens target was assigned a fiber. The text states the implicit assumption that the higher-priority sample is not correlated with the sample being masked and concedes that this is not strictly true because QSOs and LRG/ELG overlap in redshift. Since QSOs trace the same large-scale density field as LRGs (0.4<z<1.1) and ELGs (0.8<z<1.6), the vetoed regions are preferentially high-density regions of the lower-priority tracers. The random catalogs reproduce the mask geometry but not this density-dependent exclusion, so the Landy-Szalay estimator in Eq. (10.1) and the window treatment of Section 10.1.2 cannot remove the resulting bias. The principal validation, Section 12 and Table 9, compares data to altmtl mocks processed with the same priority veto and the same correlated QSO field, so a shared selection effect cancels in the chi^2 comparison; the test demonstrates pipeline consistency but not that the priority veto leaves the LRG/ELG selection function unbiased. The concern is compounded by footnote 25 of Section 11.1, which admits that the simulations are not full lightcones and that the redshift at which QSO and ELG targets overlap does not correspond to the redshift output of the simulations, so the simulated QSO-LRG/ELG angular correlation used in the veto validation may not match the data. I ask for a direct quantification, for example a comparison of altmtl mocks with the QSO-correlated veto against mocks with a randomized or shuffled QSO field, or a measurement of the cross-correlation between the veto mask and the LRG/ELG density, with the resulting bias propagated into the 2% statement.","section":"4.2, 11.1, 12"},{"comment":"Table 9 and the text of Section 12 show that the 'within 2%' agreement in the abstract and conclusions holds only after amplitude rescalings and scale cuts that vary by tracer. For example, the ELG 0.8<z<1.1 power-spectrum monopole has chi^2/dof = 220.9/80 before any adjustment; it drops to 151.6/79 with a 0.976 bias rescaling and reaches 87.8/79 only after also excluding k<0.02 h/Mpc, a range the text attributes to residual imaging systematics. The ELG 1.1<z<1.6 configuration-space monopole requires a 0.96 rescaling, a 4% bias difference. The QSO Fourier-space quadrupole has a scale-dependent shape mismatch that is not repaired by a linear bias factor and only becomes acceptable after cutting k<0.2 h/Mpc. The phrase 'generally, in statistical agreement to within 2%' should be replaced or qualified by a statement that the agreement is in shape after tracer-dependent bias rescalings and scale cuts, with the residual large-scale ELG excess and QSO shape mismatch explicitly flagged as systematic limitations.","section":"12, Table 9"},{"comment":"Because the mocks are calibrated to DESI EDR clustering (Section 11.1), the RIC/AIC corrections are constructed from the same AbacusSummit/EZmock realizations (Section 10.1.2), and the Fourier-space covariance is rescaled using the data via the configuration-space RascalC comparison (Section 10.2, Table 8), the chi^2 values in Table 9 do not provide an independent confirmation of the clustering amplitude. The paper should state explicitly in the abstract and conclusions that the agreement is a consistency check of the HOD-calibrated mock pipeline rather than a blind prediction, and it should report the sensitivity of the headline 2% number to the covariance rescaling factors (1.11-1.39). Without such a statement, a reader may over-interpret the validation as an external check of the selection function.","section":"11, 12, 10.2"}],"minor_comments":[{"comment":"The caption contains 'verticle grid lines' and should read 'vertical grid lines'.","section":"Figure 1"},{"comment":"The ELG2 P(k) 0-0.4 row with b_f=1 appears twice; one duplicate row should be removed.","section":"Table 9"},{"comment":"The text contains 'Fourer-space' and should read 'Fourier-space'.","section":"12.1"},{"comment":"The fiducial neutrino parameter is written 'P mnu = 0.06 eV'; this should be typeset as the sum of neutrino masses with a standard symbol to avoid ambiguity.","section":"1"},{"comment":"The blinded null-test acceptance criterion is described in one sentence; a compact equation or an explicit pointer to the rule in Ref. [14] would improve reproducibility.","section":"6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a DESI collaboration methods paper with many companion papers, and the priority-veto concern is acknowledged in the text. My recommendation for major revision is driven by the gap between the qualified Table 9 results and the abstract's 2% statement, and by the absence of a direct test of the QSO-correlated veto. The companion papers [10] and [19] may contain additional robustness tests that resolve some of these points; if so, a revised version should cite and summarize them explicitly rather than leaving the validation solely at the pipeline-consistency level."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is the DESI DR1 LSS catalog paper, and it is what a flagship survey methods paper should be: transparent, thorough, and honest about where the agreement with mocks breaks down. The genuinely new content is the DR1 samples, weights, randoms, veto masks, and the configuration- and Fourier-space 2-point measurements, plus the side-by-side comparison to altmtl mocks. The methodology is mostly inherited from SDSS/eBOSS and DESI EDR, as the paper acknowledges, but the validation is real and the paper does not oversell itself. Table 9 is the heart, and it shows agreement only after specific scale cuts and bias rescalings for several samples. That nuance matters, and the authors state it clearly.\n\nSoft spots, in order. First, the priority veto mask: Sec 4.2 explicitly admits the implicit assumption that QSO is uncorrelated with LRG/ELG is not strictly true because of redshift overlap. The mask removes 20% of the LRG/ELG footprint in regions where QSOs were assigned. The randoms reproduce the mask geometry but not any density-dependent exclusion, and the altmtl mocks are processed with the same veto and the same correlated QSO field, so the Sec 12 comparison is largely a pipeline-consistency test for this effect, not an external validation. The paper's defense - that the QSO redshift breadth dilutes the angular correlation - is plausible, but no quantitative estimate of the residual bias is given. I do not think this sinks the paper, since the cosmology papers carry the inference and the 2% claim is already broad, but it is a real gap and a referee should ask for a number.\n\nSecond, the Fourier-space covariance rescaling (Table 8) is a fudge: the mock covariance is multiplied by a factor derived from configuration-space mismatch. It is documented and probably acceptable, but the quoted chi-squared values have an empirical calibration baked in. Minor.\n\nThird, ELG and QSO get the roughest treatment: ELG requires k > 0.02 h/Mpc and shows residual low-k power; QSO has an unresolved quadrupole shape mismatch. The authors state these clearly and point to companion papers, which is the right move.\n\nAlso worth saying: the mocks are calibrated to DESI EDR clustering, so the 2% agreement is partly a consistency check of HOD fits rather than a fully independent prediction. The paper is upfront about this too.\n\nWho this is for: anyone using or reanalyzing DESI DR1 clustering, and anyone building the systematics pipeline for a future spectroscopic survey. It deserves serious peer review. My recommendation is accept with requested clarifications, especially a quantitative estimate of the priority-veto correlation.","headline":"A careful, transparent catalog paper whose 2% mock-agreement claim is real but narrower than it sounds; the priority-veto correlation is the one gap worth chasing.","tokens_in":719,"tokens_out":1917,"would_cite":true,"duration_ms":45792,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DESI DR1 galaxy and quasar catalogs reproduce simulated clustering to within 2% in bias, validating them for cosmological analysis.","keywords":["large-scale structure","galaxy clustering","quasar clustering","survey selection function","fiber assignment completeness","imaging systematics","baryon acoustic oscillations","redshift-space distortions"],"falsifier":"A direct check would be to compute the angular cross-spectrum between the QSO priority-veto region and the completeness-corrected LRG and ELG density fields in the released catalogs; a significant correlation at the scales used for cosmology would mean the effective selection function is biased. A second check is to repeat the data-mock comparison without the $k < 0.02\\,h\\,\\mathrm{Mpc}^{-1}$ cutoff for ELGs and without rescaling; if the reduced chi-square remains near 3 for the monopole, the claimed 2% consistency in Fourier space would not extend to the lowest-$k$ modes.","tokens_in":54288,"feed_emoji":"🔭","tokens_out":6051,"duration_ms":56153,"temperature":0.7,"pith_summary":"This paper establishes that the galaxy and quasar catalogs built from the survey's first data release, after corrections for survey completeness, imaging artifacts, and spectroscopic failures, produce two-point clustering measurements that agree with realistic simulations to within about 2% in the galaxy bias. The agreement holds in both configuration space and Fourier space across four tracer samples spanning redshifts 0.1 to 2.1. If this holds, the catalogs are fit for the companion cosmological analyses: baryon acoustic oscillation distances, redshift-space distortion growth rates, and primordial non-Gaussianity constraints. The paper also specifies the window functions and integral-constraint corrections needed to compare models to the released measurements.","feed_headline":"DESI DR1 clustering matches simulations to within 2 percent","feed_subtitle":"New galaxy and quasar catalogs pass systematics checks for BAO and full-shape cosmology.","key_machinery":"The load-bearing machinery is the large-scale structure catalog pipeline: matched sets of synthetic random catalogs, per-object weights combining fiber-assignment completeness ($w_{\\mathrm{comp}}$), imaging-systematic regression weights ($w_{\\mathrm{imsys}}$), and redshift-failure weights ($w_{\\mathrm{zfail}}$), plus FKP weights for optimal signal-to-noise. On top of these sit the two-point estimators, namely the Landy-Szalay correlation function and the FKP-based power spectrum, together with window matrices that encode the survey geometry, the small-angle ($\\theta < 0.05$ deg) cut that removes fiber-collision bias, and empirical corrections for radial and angular integral constraints calibrated on simulations. The validation comparison against 'altmtl' mocks, which run the same fiber-assignment realization loop as the data, is what carries the 2% agreement claim.","core_discovery":"The central discovery is a validation result: the raw two-point correlation function and power spectrum multipoles measured from the DESI DR1 large-scale structure catalogs are statistically consistent with the mean of 25 'altmtl' mock catalogs that reproduce the survey's target selection, fiber assignment, and redshift failures. For most tracers and scale ranges, the data-mock comparison yields reduced chi-square values near one; where discrepancies appear, they are absorbed by a linear bias factor of at most about 2% (e.g., 0.976 for ELGs, 0.983 for BGS). The paper argues these residuals are understood: ELG excess at $k < 0.02\\,h\\,\\mathrm{Mpc}^{-1}$ is attributed to residual imaging systematics and QSO scale-dependent mismatch to redshift uncertainty in the mocks. Hence, the effective selection function of the catalogs is accurate enough that observational systematics are subdominant to statistical errors in the 2024 cosmological analyses.","pith_inferences":["A testable extension is to cross-correlate the QSO priority-veto mask with the corrected LRG and ELG density fields; if a significant residual angular correlation survives, the effective selection function of those samples would need revision.","The paper's assumption that QSOs dilute angular correlations with LRGs and ELGs because of their broad redshift distribution could be checked directly on the released catalogs by computing the mask-density cross-spectrum.","If the 2% agreement persists when the same validation is applied to a larger, higher-completeness data release, it would strengthen the case that the systematic budget is under control; if not, the discrepancy would localize to the new single-pass regions.","The residual ELG excess at $k < 0.02\\,h\\,\\mathrm{Mpc}^{-1}$ suggests that imaging-systematic regression removes most but not all mode power; a future analysis could quantify how much of that excess is traced by the Galactic extinction difference map."],"forward_implications":["The public DR1 LSS catalogs and covariance matrices can be used for BAO, full-shape, and $f_{\\mathrm{NL}}$ analyses without additional per-systematic corrections.","Fiber-assignment incompleteness is adequately handled by the theta-cut plus window matrix for scales above $20\\,h^{-1}\\mathrm{Mpc}$; PIP-weighted catalogs extend this to smaller scales.","Residual imaging systematics mainly affect the lowest-$k$ Fourier modes of ELGs, so full-shape analyses must marginalize over an additional systematic component.","The 2% bias factors quantify the accuracy of the halo-occupation models used to build the mocks, giving a target for future simulation calibration.","The same pipeline, with weights and masks recomputed, applies to later data releases where single-pass regions shrink."],"supporting_citations":[{"why":"Supplies the LSS catalog construction method that defines the full, vetoed, and clustering catalogs.","marker":"[12]"},{"why":"Provides the alternate fiber-assignment realizations ('altmtl') used to build the validation mocks.","marker":"[13]"},{"why":"Describes the DR1 mock catalogs whose mean clustering is compared with the data.","marker":"[24]"},{"why":"Develops the theta-cut estimator and window-matrix rotation that remove fiber-assignment bias.","marker":"[23]"},{"why":"Quantifies ELG imaging systematics and supports the residual low-$k$ excess interpretation.","marker":"[18]"},{"why":"Documents spectroscopic success trends and the redshift-failure weight model used in the catalogs.","marker":"[20]"},{"why":"Defines the FKP weighting scheme used to optimize signal-to-noise in the 2-point measurements.","marker":"[59]"},{"why":"Provides the Landy-Szalay estimator used for the correlation function multipoles.","marker":"[86]"},{"why":"Companion paper that uses these catalogs for BAO and demonstrates post-reconstruction measurements are unbiased.","marker":"[8]"}],"fun_headline_variants":["DESI DR1 clustering matches mocks within 2%","DESI catalogs pass mock test for cosmology","DESI DR1 data-mock agreement buoys BAO analyses","DESI samples validated against simulations to 2%","DESI DR1 systematics check: clustering within 2% of mocks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the high-priority quasar sample does not leave a residual angular correlation with the lower-priority LRG and ELG samples after the priority veto mask is applied; the paper concedes this is not strictly true because the samples overlap in redshift, and it relies on simulations to capture the selection effect.","fun_headline_variants_meta":{"raw":{"variants":["DESI DR1 clustering matches mocks within 2%","DESI catalogs pass mock test for cosmology","DESI DR1 data-mock agreement buoys BAO analyses","DESI samples validated against simulations to 2%","DESI DR1 systematics check: clustering within 2% of mocks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1515,"prompt_tokens":1007,"completion_tokens":508,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":623,"completion_tokens_details":{"reasoning_tokens":422}},"tokens_in":623,"tokens_out":508,"duration_ms":4704,"temperature":1.0,"reasoning_tokens":422,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:58:53.767929+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct check would be to compute the angular cross-spectrum between the QSO priority-veto region and the completeness-corrected LRG and ELG density fields in the released catalogs; a significant correlation at the scales used for cosmology would mean the effective selection function is biased. A second check is to repeat the data-mock comparison without the $k < 0.02\\,h\\,\\mathrm{Mpc}^{-1}$ cutoff for ELGs and without rescaling; if the reduced chi-square remains near 3 for the monopole, the claimed 2% consistency in Fourier space would not extend to the lowest-$k$ modes.","supporting_citations":[],"review_version":1}