{"id":"77414de3-da13-4ec8-b57f-0f86b751f6f4","arxiv_id":"2412.12548","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An end-to-end simulation test of the DESI-Y1 3x2-pt pipeline recovers the fiducial cosmology within statistical errors and validates the analytical covariance against mock realisations.","lead":"The DESI-Lensing mock challenge builds realistic simulations of DESI-Y1 galaxies combined with KiDS, DES and HSC weak lensing data, then tests whether the joint 3x2-pt analysis pipeline recovers the true cosmological parameters. It validates the analytical error model and shows the pipeline recovers the input cosmology within statistical errors, a key dress rehearsal before the real DESI-Y1 lensing measurement is unblinded.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The new wp(R) covariance and its cross-covariances are never directly validated against the Buzzard ensemble (Sec. 3.2 admits this); the central 'recovery within statistical errors' claim therefore rests on an untested ingredient of the likelihood.","rationale":"The central claim is that the DESI-Y1 3x2-pt pipeline, featuring the newly derived wp cross-covariance, recovers fiducial cosmological parameters within the statistical error margin. For that claim to hold, the likelihood covariance must be accurate in every block used in the fit, especially the new wp block and its cross-terms with the angular statistics. The paper provides three types of support: (1) validation of the analytical covariance against the mock ensemble for xi+/- and gamma_t; (2) agreement of the covariance code with CosmoCov and the KiDS covariance code; and (3) global chi^2 and PTE statistics for the cosmological fits. None of these directly validates the wp block. The paper itself flags this in Sec. 3.2, stating that the numerical covariance for wp is unreliable due to the domain decomposition and that 'we could not perform an analogous test.' This is a load-bearing gap because a covariance error in the wp block would change the reported confidence contours and the PTE calibration, and therefore the meaning of 'within the statistical error margin.' It would not necessarily produce an anomalously large global chi^2, because the fit can absorb some covariance misspecification by adjusting nuisance parameters, and the PTE compares two fits that share the same possibly incorrect covariance. A concrete, feasible check is to estimate the numerical covariance of wp after removing the known domain-decomposition modulation and compare it to the analytic covariance; this directly tests the untested ingredient. The reader's conditional verdict is appropriate: the concern is specific, addressable, and currently unresolved, but it is not a demonstrated failure. Therefore no change to the reader's verdict is needed.","tokens_in":35897,"tokens_out":16199,"duration_ms":156205,"concrete_test":"Use the Buzzard mocks to compute the numerical covariance of wp(R) and its cross-covariance with gamma_t and xi+/- after removing the ADDGALS domain-decomposition effect: for each nside=4 domain, fit a constant multiplicative galaxy-density amplitude and divide it out of the wp and gamma_t measurements; then compare the residual covariance with the analytic Appendix D covariance on the fiducial scale cuts. If the diagonal ratios deviate by more than the ~10% 'error in the error' (N=160) or the off-diagonal structure disagrees, the parameter-recovery validation is not representative and the fiducial fits must be re-run with a corrected wp covariance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim is that the fiducial 3x2-pt analysis recovers the input cosmology within the statistical error margin. That claim is exercised through a likelihood whose covariance is the principal new ingredient: the projected-correlation wp block and its cross-covariance with xi+/- and gamma_t (Appendix D). Section 3.2 explicitly states that 'for the projected correlation function wp(R), the domain decomposition used in the construction of the Buzzard galaxy catalogues implies that our estimate of the numerical covariance is less reliable and we could not perform an analogous test.' The external-code comparisons in Appendix E cover only xi+/- and gamma_t configurations, not wp or its cross-terms. The remaining evidence for the wp covariance is the global chi^2 and PTE of the fits, but those are integrated statistics: they average over many parameters and data points, and a block-specific error in the wp covariance can shift the reported sigma-contours and PTE without producing an obviously bad global chi^2. Since the same (possibly wrong) covariance is used to define both the PTE numerator and denominator, a common scaling error does not cancel, and a non-scalar error in the wp block would directly miscalibrate the 'within statistical error margin' statement. The third-lens-bin exclusion and ~1sigma biases in Table 2 are consistent with, but do not isolate, an unvalidated wp covariance. This is a missing validation, not a demonstrated inconsistency; a direct mock-based check of the wp block is needed before the central claim is fully supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents an end-to-end validation of the DESI-Y1 3x2-pt cosmological analysis pipeline using the Buzzard N-body simulation suite. The authors construct mock galaxy and weak lensing catalogs matching DESI BGS/LRG lenses and KiDS-1000, DES-Y3, and HSC-Y1 sources, measure cosmic shear xi±(theta), galaxy-galaxy lensing gamma_t(theta), and projected clustering wp(R), and compute an analytic covariance including Gaussian, super-sample, and noise terms, with new cross-covariance terms involving wp(R) derived in Appendix D. The covariance is tested against the mock ensemble for xi± and gamma_t and against two external codes for xi± and gamma_t; cosmological parameter recovery is then tested with CosmoMC using fiducial scale cuts Rclus=7 h^-1 Mpc and Rggl=10 h^-1 Mpc, with the third lens bin excluded. The paper reports parameter biases in the Omega_m-S8 plane below about 1 sigma and concludes that the fiducial 3x2-pt configuration recovers the underlying cosmology within the statistical error appropriate for DESI-Y1.","tokens_in":36210,"tokens_out":7731,"duration_ms":67863,"significance":"If correct, the paper would provide a valuable end-to-end validation of a key analysis pipeline and introduces a new analytic ingredient—the cross-covariance between the projected correlation function and angular shear/GGL statistics—that will be used in the DESI-Y1 analysis. The strengths include the realistic mock construction, the detailed analytic covariance formalism, the public data release for the figures, and the explicit comparison to external covariance codes at the 1% level for the angular statistics. However, the central validation claim is weakened by the admitted lack of a direct mock-based test of the wp covariance block, which is a principal new ingredient and an input to the parameter-recovery likelihood. The paper is therefore a solid methodological contribution whose headline claim requires an additional validation step before it can be taken as fully demonstrated.","major_comments":[{"comment":"The analytic covariance for wp(R) and its cross-covariances is not directly validated against the Buzzard ensemble; Sec. 3.2 states this explicitly, and the external-code comparisons in Appendix E cover only the xi± and gamma_t configurations, not the wp block or its cross-terms. The global chi^2 and PTE statistics used in Secs. 5.5-5.6 are integrated quantities that cannot localize a block-specific error in the wp covariance, and because the same covariance enters both the data and fiducial-model fits in Eq. (21), a block-specific miscalibration would directly affect the reported 'within statistical error margin' statement. Since the headline claim rests on this likelihood ingredient, a direct validation of the wp covariance (e.g., against an alternative simulation without the domain-decomposition issue, or against an independent analytic code including the wp block) is needed.","section":"Sec. 3.2, Apps. D-E"},{"comment":"The third lens redshift bin (0.3<z<0.4) is excluded from all fits because of the Buzzard light-cone transition, as stated in Sec. 5, but the abstract and Sec. 5.6 present the recovery claim without this qualification. The validation is therefore performed on four lens bins rather than the five-bin DESI-Y1 configuration, so the headline claim is stronger than the test actually performed. The paper should state this exclusion explicitly wherever the 'within statistical error' claim is made and, ideally, quantify the impact of the excluded bin on the reported biases.","section":"Sec. 5, Sec. 5.6"},{"comment":"The analytic covariance is computed using fiducial linear bias factors b=(1.35,1.51,1.65,2.21,2.44) that are fitted to the mock projected correlation functions, while the same wp data are later analyzed in the likelihood with galaxy bias parameters that are free. This partial dependence of the covariance on the analyzed data introduces a potential circularity that is not quantified; the paper should test the sensitivity of the parameter-recovery conclusions to the bias values assumed in the covariance, for example by repeating the fiducial fit with bias parameters shifted by their fitted uncertainties.","section":"Sec. 2.1 and Sec. 3.1"}],"minor_comments":[{"comment":"The maximum wp scale cut is set to half the projected separation of the average nside=8 pixel size; this is an ad hoc choice and a sensitivity test varying this cut (e.g., to one-third or two-thirds of the pixel scale) would strengthen the scale-cut validation.","section":"Sec. 5.1"},{"comment":"The PTE definition invokes a rescaling of the covariance by the number of mock realizations M, but the exact form of the rescaled covariance C^M is not written out; specifying this rescaling explicitly would improve reproducibility.","section":"Eq. (21)"},{"comment":"The individual lens redshift bins are not labeled directly in the panels of Fig. 17; adding the redshift ranges to each panel would improve readability.","section":"Fig. 17"}],"recommendation":"major_revision","confidential_remarks":"This is a pipeline-validation paper with a strong abstract that overstates the covariance validation: the wp covariance block is not directly tested against the mock ensemble, and the third lens bin is excluded post hoc. The missing wp validation is the main risk to the headline claim and should be addressed before acceptance, either by adding a direct validation or by explicitly qualifying the claim in the abstract and conclusions. The analytic formalism and end-to-end setup are valuable and the paper is likely to be influential for DESI-Y1, but the central claim needs the additional test."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Chris,\n\nThe one thing to know: this is a carefully done end-to-end validation of the DESI-Y1 3x2-pt analysis pipeline, and the genuinely new bit—the analytic covariance for wp(R) and its cross-terms with the angular statistics—is also the one bit that is not directly checked against the mocks. The paper says so itself in Sec. 3.2. That is a missing validation, not a demonstrated error, and it is the main reason to be conditional rather than fully convinced.\n\nWhat the paper does well: the covariance formalism is extended in a clear way (Appendices A–D), and the code comparisons in Appendix E are a real plus—agreement with CosmoCov and the KiDS covariance code to better than 1%. The mock validation for the shear diagonal (<1%) and gamma_t (<5%) is solid. The parameter recovery tests use the PTE framework from DeRose et al. 2022, which handles projection effects decently, and the mocks themselves are realistic: HOD lenses, photo-z errors, source weights, magnification, multiplicative shear bias. The paper is also unusually upfront about its own limitations—the excluded third lens bin, the domain-decomposition problem, the HSC ~1sigma biases.\n\nWhere it is soft: the wp covariance and its cross-covariances are the principal new ingredient, and the only evidence for them is the global chi^2/PTE of the fits plus the code comparisons, which don't actually cover the wp block. That is exactly the stress-test concern, and it lands. A block-specific error in the wp covariance could shift the reported contour sizes without ruining the global PTE, since the same covariance enters both numerator and denominator of the PTE. Also mildly circular: the bias factors used in the covariance are fitted from the mock wp, and the fiducial scale cuts are chosen with knowledge of the mock fits. That is not fatal—the external code comparisons and fiducial-model PTE provide independent grounding—but it means the 'recovery within statistical errors' claim is not as cleanly established as the abstract suggests.\n\nWho it's for: anyone working on 3x2-pt analysis with DESI or similar combined lensing+clustering datasets. The covariance extension is a real reference point, and the validation framework is a good template. If I were refereeing, I'd ask for a direct mock-based check of the wp block (or a clear argument why the domain decomposition makes it impossible and why the integrated tests are sufficient), then I'd be happy to accept.\n\nRecommendation: send it to peer review. It's a serious, honest paper with one specific addressable gap.","headline":"Solid mock challenge for the DESI-Y1 3x2-pt pipeline; the new wp covariance is the key ingredient and the one part not directly mock-validated.","tokens_in":37058,"tokens_out":2374,"would_cite":true,"duration_ms":20141,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that a full mock-based analysis pipeline for DESI Year 1 lensing and clustering data recovers the input cosmology within the statistical error margins.","keywords":["3x2-pt correlation functions","weak gravitational lensing","projected correlation function","analytical covariance","super-sample covariance","DESI Year 1","cosmological parameter recovery","mock challenge"],"falsifier":"A direct check would be to compare the analytical $w_p(R)$ covariance to the scatter of $w_p$ measurements across a suite of mock catalogues built without the domain-decomposition artefact or with that effect corrected; if the analytical errors deviate from the mock scatter by more than the few-percent level seen for $\\xi_\\pm$ and $\\gamma_t$, the parameter-recovery claim would be undermined. A cheaper test is to check whether $w_p$-only fits across many realisations produce a $\\chi^2$ distribution consistent with the assumed number of degrees of freedom.","tokens_in":35607,"feed_emoji":"🔭","tokens_out":4681,"duration_ms":40471,"temperature":0.7,"pith_summary":"This paper is an end-to-end dress rehearsal for the combined analysis of DESI Year 1 galaxy clustering with weak lensing data from KiDS, DES, and HSC. The authors build realistic mock catalogues from N-body simulations, compute the full analytical covariance of the 3x2-pt correlation functions, and fit the mock data with a Bayesian pipeline. Their central claim is that the fiducial cosmological parameters of the simulation are recovered within the statistical error margin of the experiment, meaning the pipeline is unbiased at the precision DESI-Y1 will deliver. The new ingredient is the analytical covariance attached to the projected correlation function $w_p(R)$ and its cross-correlations with the shear statistics, which has not been validated against the mock ensemble because the simulations cannot provide a reliable numerical covariance for it.","feed_headline":"DESI mock challenge recovers cosmology within error bars","feed_subtitle":"Simulated DESI, KiDS, DES and HSC data confirm the 3x2-pt pipeline is ready for Year 1 analysis.","key_machinery":"The central object is the analytical covariance matrix of the combined 3x2-pt data vector. It is built from Gaussian, noise, and super-sample contributions following the formulations of Krause & Eifler and Joachimi et al., with a new extension: the covariances between the projected correlation function $w_p(R)$ and the angular shear statistics $\\xi_\\pm(\\theta)$ and $\\gamma_t(\\theta)$, derived in the Limber approximation in Appendix D. This machinery sets the error bars in the parameter fits; the validation of the shear and galaxy-galaxy lensing diagonal errors against the mock ensemble, agreeing to a few percent, is what gives the parameter-recovery claim its weight.","core_discovery":"On the paper's own terms, the discovery is that the 3x2-pt pipeline for DESI-Y1, combining cosmic shear $\\xi_\\pm$, galaxy-galaxy lensing $\\gamma_t$, and projected clustering $w_p$, is ready for real data: fits to simulated data vectors return the input values of $\\Omega_m$ and $S_8$ within roughly half a $\\sigma$ to one $\\sigma$, with probability-to-exceed statistics comparable to earlier DES-Y3 validation. The parameter biases in the fitted data are consistent with the noise expected from eight mock realizations plus prior volume effects. The paper therefore concludes that the fiducial fitting configuration produces an acceptable recovery of the underlying cosmology at DESI-Y1 precision.","pith_inferences":["The $w_p$ covariance, being the one part of the error model not directly checked against the mock ensemble, carries residual risk: if the true noise or super-sample terms for projected clustering are misestimated, the reported parameter errors on the clustering component would be off by that amount.","Because the mocks omit intrinsic alignments and other astrophysical effects, this validation establishes a floor on pipeline performance rather than a guarantee that real-data fits will be unbiased at the same level.","The same analytic covariance machinery is the natural foundation for the paper's stated next step, a joint analysis with full 3D clustering including redshift-space distortions, where the added growth-rate information would enter through new cross-covariance terms."],"forward_implications":["The DESI-Y1 3x2-pt analysis can proceed using the analytical covariance, including the new $w_p$ cross-terms, without needing a large mock ensemble to calibrate the error matrix.","Cosmic shear alone recovers the fiducial parameters with the smallest bias, while the galaxy-galaxy lensing plus clustering combination shows somewhat larger, though still acceptable, biases, meaning the choice of small-scale cuts matters.","Including the projected correlation function of the spectroscopic lenses adds clustering signal-to-noise without degrading the cosmological parameter recovery.","The validated framework directly supports the upcoming joint DESI-Y1, KiDS, DES and HSC cosmology analysis."],"supporting_citations":[{"why":"Supplies the Gaussian covariance formalism that the paper extends to the projected correlation function.","marker":"Krause & Eifler 2017"},{"why":"Provides the KiDS covariance framework, the super-sample covariance treatment, and the code against which the new covariance is cross-checked.","marker":"Joachimi et al. 2021"},{"why":"Describes the Buzzard N-body simulation suite that provides the mock catalogues and the fiducial cosmology.","marker":"DeRose et al. 2019"},{"why":"Details the construction of the DESI lens and weak lensing source mocks, including weights, photo-z errors and shear calibration.","marker":"Lange et al. 2024"},{"why":"Provides the Bayesian inference platform used for the cosmological parameter fits.","marker":"Lewis 2013"},{"why":"Defines the probability-to-exceed statistic used to quantify parameter recovery and supplies the DES-Y3 validation benchmark.","marker":"DeRose et al. 2022"},{"why":"Hosts the CosmoCov code used in the appendix comparison that validates the analytical covariance implementation.","marker":"Fang et al. 2020a"}],"fun_headline_variants":["DESI 3x2pt mock challenge recovers cosmology within error","Simulated DESI data confirm 3x2pt pipeline readiness","Mock challenge validates DESI-Y1 lensing analysis","DESI lensing pipeline passes end-to-end mock test","3x2pt analysis ready for DESI-Y1 after mock validation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analytical error model for the projected clustering measurement and its correlation with the shear measurements is assumed to be correct, but it could not be checked directly against the simulated data because the simulations' domain decomposition spoils the numerical covariance estimate for $w_p$.","fun_headline_variants_meta":{"raw":{"variants":["DESI 3x2pt mock challenge recovers cosmology within error","Simulated DESI data confirm 3x2pt pipeline readiness","Mock challenge validates DESI-Y1 lensing analysis","DESI lensing pipeline passes end-to-end mock test","3x2pt analysis ready for DESI-Y1 after mock validation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000171,"raw_usage":{"total_tokens":1277,"prompt_tokens":954,"completion_tokens":323,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":235}},"tokens_in":570,"tokens_out":323,"duration_ms":3535,"temperature":1.0,"reasoning_tokens":235,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:58:48.650875+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct check would be to compare the analytical $w_p(R)$ covariance to the scatter of $w_p$ measurements across a suite of mock catalogues built without the domain-decomposition artefact or with that effect corrected; if the analytical errors deviate from the mock scatter by more than the few-percent level seen for $\\xi_\\pm$ and $\\gamma_t$, the parameter-recovery claim would be undermined. A cheaper test is to check whether $w_p$-only fits across many realisations produce a $\\chi^2$ distribution consistent with the assumed number of degrees of freedom.","supporting_citations":[],"review_version":1}