{"id":"911243e7-9fa7-4244-8f36-229ae510aaa7","arxiv_id":"2608.03876","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Augur is a validated DESC Fisher-forecasting pipeline whose Y1/Y10 3x2pt constraints match nested sampling and external forecast codes when derivative step sizes and redshift sampling are chosen in a stable regime.","lead":"This paper presents Augur, a new public software tool that computes Fisher forecasts for LSST dark energy analyses using the DESC's own cosmology, covariance, and likelihood libraries. The authors test it against existing forecast codes and full Bayesian posterior sampling, showing that in a carefully chosen stable configuration the forecasts agree.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Recommended 500-knot n(z) sampling conflicts with Fig. 6 caption requiring ~1000 knots; if unconverged, headline validation rests on an unstable configuration.","rationale":"The reader's verdict accepts with moderate confidence, focusing on Fisher-formalism assumptions. I find a more concrete, checkable weak point: the recommended 500-knot n(z) sampling is in tension with the paper's own Figure 6 caption, which says stable solutions require nearly 1000 knots. Since the central claim is explicitly about a stable configuration, an unstable n(z) sampling would undermine the headline validation. This is internal to the paper and can be settled by a single recomputation. I do not see a need for rejection; conditional acceptance is appropriate until the knot-count issue is resolved. The paper is otherwise careful and transparent, with explicit stability criteria and comparisons to multiple pipelines.","tokens_in":29603,"tokens_out":11579,"duration_ms":118086,"concrete_test":"Run the fiducial Y1 and Y10 Augur Fisher calculations with all settings fixed (Δs=0.05, same covariance, same power spectrum) but with n(z) represented by 500 and by 1000 or more knots (e.g., 2000). Compute the max relative difference in Fisher matrix elements and the DEFOM/LSSFOM between the two. If any Fisher element changes by more than 10% or the FOM changes by more than the paper's 20% threshold, the adopted 500-knot configuration is not converged; the paper should either adopt a converged knot count or demonstrate that the headline comparisons are insensitive to this choice.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central validation (Figs. 4/5/8) is conditional on operating in a numerically stable configuration, which the text defines as normalized step size Δs=0.05 and n(z) sampled with 500 knots (Secs. 5.1, 5.2). However, the caption of Fig. 6 states that 'stable solutions are generally found when increasing the number of knots in the n(z) distribution to nearly 1000 evaluations, regardless of the analysis year considered.' Section 5.2 claims convergence at 'several hundred' knots and adopts 500, but the Fig. 6 caption suggests 500 may still be in the non-asymptotic regime. If the DEFOM or Fisher-matrix elements at 500 knots differ from the converged (>1000-knot) values by more than the paper's own stability thresholds (10% on Fisher elements, 20% on FOMs), then the 'stable configuration' used for the headline comparisons is not actually stable, and the conclusion that Augur reproduces nested-sampling posteriors and SRD forecasts would need re-qualification. This is an internal inconsistency in the recommended hyperparameter choice, not a critique of the Fisher formalism itself.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Augur, a Fisher forecasting package for the LSST DESC software ecosystem, built on CCL, TJPCov, and firecrown. The authors perform a validation exercise for 3x2pt harmonic-space analyses with LSST Y1/Y10-like setups, testing the stability of Fisher matrices against derivative step size, n(z) knot sampling, matter power spectrum prescription, and bandpower binning. They compare Augur's Fisher contours, correlation matrices, and figures of merit against a CosmoSIS+firecrown Fisher implementation, the DESC SRD pipeline, and direct PolyChord nested sampling, reporting good agreement within the thresholds defined in Section 4. They also test the Fisher bias formalism with one-dimensional parameter shifts. The central claim is that, when operated in a numerically stable configuration, Augur reliably reproduces both external Fisher pipelines and full posterior sampling.","tokens_in":29920,"tokens_out":5113,"duration_ms":48595,"significance":"If the validation holds, Augur provides a useful, publicly available forecasting component for DESC cosmological analyses. The paper's strongest contributions are its external check against nested sampling (Figures 4 and 5), the systematic mapping of numerical hyperparameter sensitivity (Figures 1-3 and 6-7), and the explicit documentation of stability criteria for Fisher-matrix users. The work is methodologically valuable for the DESC software stack and for future Fisher analyses of LSST data. The code is publicly available, which aids reproducibility. The comparison with CosmoSIS is limited by the shared firecrown likelihood, but the paper is appropriately cautious in describing it as a validation of the Fisher implementation rather than of the physical model.","major_comments":[{"comment":"There is an internal contradiction in the recommended n(z) knot configuration. Section 5.2 states that 'convergence is achieved when the n(z) distribution is sampled with several hundred knots' and adopts 500 knots, while the caption of Figure 6 states that 'stable solutions are generally found when increasing the number of knots in the n(z) distribution to nearly 1000 evaluations, regardless of the analysis year considered.' Since the headline validations in Figures 4, 5, and 8 use 500 knots, this is load-bearing. Please report the DEFOM and Fisher-matrix differences between 500 and 1000+ knots against the Section 4 thresholds; if 500 is not in the converged regime, rerun the comparisons at a converged knot number and update the recommended configuration.","section":"Section 5.2 / Figure 6"},{"comment":"The claimed Y1 DEFOM agreement with the SRD pipeline is not supported by the numbers in Figure 8. The caption lists DEFOM = 36.95 for the DESC SRD Pipeline and 41.36 for Augur (and CosmoSIS), a relative difference of about 12%, which exceeds the 10% agreement threshold stated in Section 4 and the 'within 10%' claim in Section 5.4 and the Conclusion. Please correct either the numbers or the text, and ensure the reported Y1 agreement is consistent with the stated threshold.","section":"Section 5.4 / Figure 8"},{"comment":"The stability criterion is written as (F^{Δs}_{αβ} - F^{Δs'}_{αβ})/F^{Δs'}_{αβ} ≤ 0.1 without an absolute value. As written, any negative relative difference satisfies the inequality, so it does not enforce 'does not change by more than 10%' in the sense intended. If the implementation uses an absolute value, this should be stated explicitly. This matters because the stable configuration chosen in Section 5.1 is selected using this criterion.","section":"Section 4, stability criterion"}],"minor_comments":[{"comment":"The legend text in the figure says 'CosmoSIS+ firecrown Y1 Nested Sampling', but the caption and text attribute the nested sampling to PolyChord. Since the nested sampling run uses the firecrown likelihood but is not a CosmoSIS Fisher run, the legend label should be changed to 'PolyChord nested sampling' for clarity.","section":"Figure 4"},{"comment":"The caption for Figure 6b contains a stray 'the' at the end: 'for the number of knots evaluated in the n(z) distribution. the'.","section":"Figure 6"},{"comment":"The paper correctly notes that Eq. (4) assumes a parameter-independent covariance. This is an important limitation, particularly for future analyses using model covariances; it would be helpful to state explicitly in the conclusions that this assumption was not varied in the validation.","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the 500-knot vs. 'nearly 1000-knot' convergence statement is real and is reflected in Major Comment 1. The Y1 DEFOM discrepancy (12% vs. the claimed 10%) is also a concrete numerical inconsistency. Both are fixable within the paper's scope. The nested-sampling validation is a strong external check and the paper's central approach is sound; with these corrections, the paper would be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nBottom line: this is a careful validation paper for a new DESC Fisher-forecasting tool, and it deserves a serious referee. The core claim—that Augur reproduces the degeneracy structure of nested-sampling posteriors and the DEFOMs of the DESC SRD pipeline—is supported by real tests: correlation-matrix differences below 0.05, DEFOM agreement at 10% (Y1) and 2% (Y10). The main contribution is not new formalism; it's a maintained, integrated tool plus practical stability diagnostics. The normalized step size (5% of prior width) and the n(z) knot convergence test are genuinely useful for anyone running Fisher forecasts in this ecosystem.\n\nThe soft spot is the n(z) knot issue. Section 5.2 claims convergence at 'several hundred' knots and adopts 500 as the fiducial configuration. But the caption of Fig. 6 says stable solutions are 'generally found when increasing the number of knots ... to nearly 1000 evaluations.' That's an internal inconsistency. If 500 knots is not in the converged regime, then the validation in Figs. 4/5/8 may depend on an under-resolved n(z). The authors should either run the headline comparisons at 1000 knots or show that the Fisher-matrix elements at 500 knots are within their 10% stability threshold of the 1000-knot values. This is load-bearing, not a nitpick.\n\nOther minor concerns: the CosmoSIS comparison shares the same firecrown likelihood, so it's more a check of the derivative implementation than a fully independent pipeline test; the SRD comparison is more independent. The nested-sampling comparisons would be stronger with error bars on the correlation-matrix differences, since PolyChord itself has sampling noise. And pinning a commit hash and shipping configs would help reproducibility.\n\nOverall, the central argument holds up: in a stable numerical configuration, Augur gives Fisher results consistent with direct sampling and with legacy pipelines. The Fisher-formalism caveats (local quadratic approximation, fixed covariance) are acknowledged and standard for the field. This is a solid, useful paper for the DESC community and for anyone comparing forecasting tools.\n\nRecommendation: send to peer review. With a clear revision addressing the n(z) knot inconsistency, I'd accept.","headline":"Worth a serious referee: the validation is real, but the 500-knot claim conflicts with the 1000-knot caption in Fig. 6.","tokens_in":30451,"tokens_out":2785,"would_cite":true,"duration_ms":24292,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Augur Fisher forecasts reproduce full posterior contours for LSST 3x2pt analyses, the paper shows, once its numerical settings are stabilized.","keywords":["Fisher forecast","3x2pt analysis","LSST","angular power spectra","numerical derivatives","nested sampling validation","dark energy figure of merit","cosmological parameter constraints"],"falsifier":"Re-run the Augur-versus-PolyChord comparison at a shifted fiducial cosmology, say w0 = -1.3 with the other parameters rescaled consistently, and check whether the correlation-matrix difference stays below 0.1 and the DEFOM within 20%. If those criteria fail, the claimed stable configuration is specific to the fiducial used here rather than a general property of the pipeline.","tokens_in":29530,"feed_emoji":"🔭","tokens_out":6584,"duration_ms":60720,"temperature":0.7,"pith_summary":"Augur is a public Fisher-forecasting library for the LSST Dark Energy Science Collaboration that computes parameter constraints from the curvature of the log-likelihood, using numerical derivatives of 3x2pt angular power spectra (galaxy clustering, galaxy-galaxy lensing, and cosmic shear). The paper's central claim is that, when operated in a stable numerical configuration — 5-point-stencil derivatives with step sizes around 5% of the prior extent and redshift distributions sampled with about 500 knots — Augur's Fisher contours reproduce the shape, orientation, and area of posterior contours obtained by direct nested sampling, with correlation-matrix differences below 0.05. The same configuration reproduces the Dark Energy Figure of Merit of the DESC Science Requirements Document pipeline to 10% (Year 1) and 2% (Year 10), and agrees near-exactly with an independent Fisher implementation sharing the same likelihood. This matters because survey designers can then test modeling choices cheaply, without running full MCMC or nested sampling, while staying inside the software stack that will be used for the actual data analysis. The paper also documents where the approximation breaks down: sub-percent step sizes produce overly optimistic figures of merit and incorrect degeneracy directions.","feed_headline":"Augur Fisher forecasts match full posterior contours","feed_subtitle":"With stable step sizes, the DESC forecasting tool agrees with nested sampling and SRD figures of merit.","key_machinery":"The engine is the Fisher information matrix evaluated from numerical derivatives of the harmonic-space angular power spectrum data vector, F_alpha_beta = (df/dtheta_alpha)^T C^{-1} (df/dtheta_beta), evaluated at the fiducial cosmology under a Gaussian likelihood with parameter-independent covariance. The practical workhorse is the 5-point stencil derivative; the step size is normalized to 5% of each parameter's uniform prior extent, and the redshift distribution is supplied as a spline over roughly 500 knots. These two choices control the trade-off between truncation error and numerical noise that dominates Fisher computations for cosmological two-point functions, and the paper's stability t","core_discovery":"The paper establishes a reliable operating point for Augur: normalized derivative step size of 5% of the uniform prior range for every parameter, evaluated with the 5-point stencil or numdifftools, and n(z) represented by roughly 500 spline knots. At this point the Fisher matrix approximates the local likelihood well enough that its full correlation matrix sits within 0.05 of the nested-sampling correlation matrix for both Y1 and Y10 setups, its w0-wa degeneracy direction matches sampling, and its Dark Energy Figure of Merit agrees with the DESC SRD pipeline to 10% (Y1) and 2% (Y10). The paper also characterizes unstable regimes: at smaller step sizes, oscillatory features in non-linear matt","pith_inferences":["The validation is fiducial-specific: nothing in the paper demonstrates that the 5%-step / 500-knot recipe transfers to other fiducial cosmologies or to likelihoods with parameter-dependent covariance, so users of Augur should re-run the stability tests for each new model.","A natural next test would be repeating the nested-sampling comparison at a shifted fiducial, such as w0 = -1.3, to map how far the linearity region extends; the paper's own procedures provide the tool for that test.","The breakdown of Fisher bias beyond roughly 1 sigma suggests a hybrid workflow: use Augur for fast screening, then invoke nested sampling or MCMC only in parameter regions where the bias checks fail the 10% criterion.","The sensitivity of Halofit's oscillatory derivatives to step size points toward automatic differentiation or analytic derivative models as a way to push stable Fisher forecasts to smaller scales than finite differencing allows."],"forward_implications":["Users can replace expensive MCMC or nested-sampling scans with Augur Fisher forecasts for LSST Y1/Y10 3x2pt setups, provided they adopt the documented stable step sizes and n(z) sampling.","Forecasts, likelihoods, and covariances now share one ecosystem, so Fisher studies can be compared directly with the actual analysis pipeline rather than a separate forecasting code.","Fisher bias estimates can be trusted for small model mis-specifications, up to about 1 sigma in w0 and 0.5 sigma in wa, providing a fast first-pass diagnostic for systematics.","Forecasts produced with sub-percent derivative step sizes or sparse n(z) sampling should be treated as unreliable; the paper's thresholds give a concrete way to check stability before interpreting results.","Bin-averaging bandpowers stabilizes derivative behavior and changes the DEFOM through scale-cut effects, so binning choices must be reported and tested alongside any Fisher forecast."],"supporting_citations":[{"why":"Supplies the Fisher-forecast approach that Augur implements for cosmological parameter constraints.","marker":"Tegmark 1997"},{"why":"Provides the 5-point stencil derivative method whose step-size behavior is the central stability target of the validation.","marker":"Bhandari et al. 2021"},{"why":"The Core Cosmology Library supplies the matter power spectrum and angular power spectrum predictions that form Augur's data vector.","marker":"Chisari et al. 2019a"},{"why":"PolyChord nested sampling provides the direct posterior with which Augur's Fisher contours and correlation matrices are compared.","marker":"Handley et al. 2015"},{"why":"CosmoSIS provides the independent Fisher implementation used for cross-validation of Augur's derivative methodology.","marker":"Zuntz et al. 2015"},{"why":"The DESC Science Requirements Document defines the Y1/Y10 3x2pt setup, scale cuts, binning, and benchmark figures of merit.","marker":"The LSST Dark Energy Science Collaboration et al. 2018"},{"why":"CosmoLike powers the DESC SRD pipeline whose forecasts and covariance are the reference for the Augur comparison.","marker":"Krause and Eifler 2017"},{"why":"Halofit is the non-linear matter power spectrum prescription whose oscillatory derivative behavior drives step-size sensitivity.","marker":"Takahashi et al. 2012"},{"why":"Provides the linear power spectrum used in the fiducial EH+Halofit configuration.","marker":"Eisenstein and Hu 1999"},{"why":"HMCode2020 is the alternative non-linear prescription used in the stability comparisons.","marker":"Mead et al. 2021"}],"fun_headline_variants":["Augur Fisher forecasts match nested sampling within 5%","Fisher forecasting tool Augur passes DESC validation tests","Stable Augur settings give Fisher contours close to sampling","Augur: Fisher forecasts agree with full posterior to 0.05","New DESC tool Augur calibrates Fisher forecasts for LSST"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The whole validation rests on the log-posterior being locally quadratic at the fiducial cosmology and on the data covariance being independent of the cosmological parameters; if either fails, the Fisher contours will misstate the posterior even in the numerically stable regime.","fun_headline_variants_meta":{"raw":{"variants":["Augur Fisher forecasts match nested sampling within 5%","Fisher forecasting tool Augur passes DESC validation tests","Stable Augur settings give Fisher contours close to sampling","Augur: Fisher forecasts agree with full posterior to 0.05","New DESC tool Augur calibrates Fisher forecasts for LSST"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000149,"raw_usage":{"total_tokens":1005,"prompt_tokens":696,"completion_tokens":309,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":440,"completion_tokens_details":{"reasoning_tokens":226}},"tokens_in":440,"tokens_out":309,"duration_ms":3414,"temperature":1.0,"reasoning_tokens":226,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T10:33:48.239408+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the Augur-versus-PolyChord comparison at a shifted fiducial cosmology, say w0 = -1.3 with the other parameters rescaled consistently, and check whether the correlation-matrix difference stays below 0.1 and the DEFOM within 20%. If those criteria fail, the claimed stable configuration is specific to the fiducial used here rather than a general property of the pipeline.","supporting_citations":[],"review_version":1}