{"id":"44d2c41a-c3fe-46c2-8ba0-f72220e2e23d","arxiv_id":"2508.21182","paper_version":4,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For slitless-spectroscopy surveys, a 5% catastrophic redshift failure rate biases the growth rate and primordial amplitude by 6-16% (~2.2σ) unless the failure rate is measured and included in the clustering model.","lead":"Galaxy surveys measure distances by splitting light into spectra, and small mistakes in those measurements can distort the inferred clumpiness of the universe. Using simulated galaxy catalogs, this paper finds that for space-based slitless surveys such as Euclid, typical errors could shift key cosmological parameters by 6 to 16 percent, and shows that knowing the rate of catastrophic redshift failures is the key to correcting them.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The (1-fc)^2 correction and 2.2σ bias numbers assume catastrophic failures are a random, unclustered subset; realistic environment-selected interlopers would make the correction scale-dependent and undermine the quantitative mitigation claims.","rationale":"The paper is internally consistent: the mock contamination follows the model described in Sec. 2.3, and the (1-f_c)^2 correction is validated for that model in Eq. (4.7). The most load-bearing step is the translation of that model to a practical recommendation for Euclid. The reader's weakest assumption (hypothetical slitless model) is on target, but it can be sharpened. The specific failure mode that would break the central quantitative claims is not simply an unknown f_c value; it is the assumption that catastrophics are randomly selected and hence uncorrelated with the target field. Real slitless interlopers are line-confusion systems with their own clustering and bias, so the power-spectrum contamination is not a pure (1-f_c)^2 amplitude suppression. Since the paper's validation mocks use random assignment, they cannot reveal this. A targeted mock test with environment- or mass-dependent f_c would settle whether the proposed correction and fixed-f_c mitigation survive. Because the paper's qualitative direction (measure f_c, include it in the model) is likely robust, the verdict remains CONDITIONAL; no change from the reader's assessment.","tokens_in":19948,"tokens_out":18140,"duration_ms":199075,"concrete_test":"Construct contaminated mocks with the same f_c=5% but make the probability of a catastrophic failure depend on a property correlated with large-scale structure—for example, assign f_c preferentially to halos in the lowest mass quartile or to halos in high-density environments—while keeping the total f_c and the displacement distribution unchanged. Measure R(k)=P_contam/P_clean over k∈[0.02,0.2] and refit the baseline EFT model. If R(k) deviates from (1-f_c)^2 by more than ~10% or is scale-dependent, the correction in Eq. (4.6) is not a valid general mitigation; the fixed-f_c recovery in §5.2.2 would then not restore unbiased constraints for realistic Euclid interlopers.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claims (Sec. 5.2.2: 2.20σ bias in ln(10^10 A_s), recovery to <0.1σ with EFT+free f_c, 0.85σ with fixed f_c) all rest on the assumption that catastrophic failures are a random, unclustered 5% subset of the sample. Section 3 assigns Δv_error to a random 5% of halos without any dependence on halo mass, environment, or redshift, and the effective-field-theory correction Eq. (4.6) multiplies the power spectrum by (1-f_c)^2. That factor follows only if the catastrophic population is completely uncorrelated with the true field (P_cross≈0 and the interloper term f_c^2 P_c is absorbed as shot noise). For a slitless survey such as Euclid, catastrophic failures are dominated by line confusion (e.g., [OII] at z~0.8 misidentified as Hα at z~1.2) and by sky residuals; these interlopers are not a random subset—they have their own bias, redshift distribution, and possibly an environment-dependent selection function. In that case the observed power spectrum contains an f_c^2 P_interloper term and a 2f_c(1-f_c) P_cross term that are not described by Eq. (4.6), and the correction can become scale-dependent. The paper's Eq. (4.7) validates (1-f_c)^2 only for its own random-selection mocks, so it cannot detect this failure. Consequently, the headline 6-16% bias and the efficacy of the fixed-f_c mitigation are not robust to the most plausible departures from the assumed catastrophic model.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses 500 Quijote halo mocks at z=1 to study how spectroscopic redshift errors (Gaussian/Lorentzian smearing and catastrophic failures) propagate into full-shape power-spectrum cosmological constraints. Two independent fitting pipelines (ShapeFit and Full-Modeling with EFT) are applied. The authors find that redshift uncertainty is largely absorbed by EFT counterterms, keeping parameter biases below 5%, while a hypothetical slitless-like error model (LRG-like Gaussian smearing with sigma_v=85.7 km/s plus a 5% catastrophic failure rate) biases the fractional growth rate df and ln(10^10 As) by 6-16% (~2.2 sigma). They propose a (1-f_c)^2 correction to the power spectrum; freeing f_c removes the bias but degrades the As constraint by 60%, while fixing f_c to its input value restores constraining power with a ~0.85 sigma residual bias. The paper also finds no bias in w0 and wa and up to 80% degradation in the neutrino mass constraint under QSO-like smearing.","tokens_in":20431,"tokens_out":7646,"duration_ms":84684,"significance":"The study is well constructed: it uses a large mock suite, controlled contamination, and two established fitting pipelines, and it clearly demonstrates that unmodeled catastrophic redshift errors can shift amplitude-related parameters at a level comparable to the statistical precision of a Euclid-like survey. The result that redshift uncertainty is absorbed by EFT counterterms is consistent with earlier work and usefully confirmed in a controlled setting. If the central claims hold, the paper provides a practical warning for slitless surveys and a simple, testable correction scheme. The extension to w0waCDM and massive neutrinos is also valuable. However, the quantitative mitigation claims are conditioned on a random, unclustered catastrophic-failure model and on exact knowledge of f_c; these conditions are not robust to the most plausible departures expected for real slitless spectroscopy, so the headline numbers should be interpreted as illustrative rather than as a Euclid forecast.","major_comments":[{"comment":"The (1-f_c)^2 correction assumes that catastrophic failures are a randomly selected, unclustered subset of the sample. The mocks assign Δv_error to a random 5% of halos with no dependence on mass or environment, and the validation in Eq. (4.7) therefore only tests this specific model. Realistic slitless interlopers (line confusion, sky residuals) have their own bias and redshift distribution, producing 2f_c(1-f_c)P_cross and f_c^2 P_interloper terms that Eq. (4.6) omits; these are generically scale-dependent. The claim in Sec. 5.2.2 that fixing f_c restores unbiased constraints (0.85σ) is thus not robust to the most plausible departure from the adopted model. I recommend a test with clustered interlopers (e.g., catastrophics drawn from a biased subsample or a different effective redshift) or an explicit caveat limiting the correction to random catastrophics.","section":"Sec. 3 and Sec. 4.3.2, Eq. (4.6)-(4.7)"},{"comment":"The 'fixed f_c' mitigation presumes f_c is known exactly, yet the paper's own recommendation is that f_c must be accurately estimated. The free-f_c fit exhibits a 60% degradation in the ln(10^10 As) error and a strong f_c-As degeneracy, implying that f_c is weakly constrained by the data. A modest misestimate of f_c could therefore produce a non-negligible residual bias. The paper should quantify the sensitivity, e.g., by fixing f_c to its input value offset by ±1% or by the expected calibration uncertainty, and reporting the resulting bias in ln(10^10 As) and df. Without such a test, the mitigation advice is incomplete.","section":"Sec. 5.2.2, Table 2"},{"comment":"The slitless-like error model is an ad hoc combination of a specific Gaussian sigma_v=85.7 km/s, f_c=5%, and the ELG-like log-normal catastrophic displacement distribution. The headline numbers (6-16% biases, 2.2σ, 60% degradation, 0.85σ recovery) are all conditional on this model, yet the abstract presents them without the 'hypothetical' qualifier used in Sec. 2.3. The quantitative impact and the optimal mitigation will change if the true Euclid catastrophic rate, velocity scale, or clustering of interlopers differs. Please either scan a range of f_c and sigma_v (or catastrophic displacement scales) or state more prominently that the quoted numbers are illustrative rather than forecasts.","section":"Sec. 2.3, Table 1"}],"minor_comments":[{"comment":"The abstract says the fixed-f_c model leaves a 'modest bias of 1.0σ', while Sec. 5.2.2 and Table 2 report ~0.85σ. Please make these consistent.","section":"Abstract vs. Sec. 5.2.2"},{"comment":"Catastrophic shifts can be as large as 10^6 km/s, which at z=1 corresponds to a comoving displacement much larger than the simulation box. The treatment of halos shifted outside the box (periodic wrapping? exclusion?) is not described and could affect the effective 'randomization' of catastrophics. Please clarify.","section":"Sec. 3"},{"comment":"There are several typos: 'quarupole' in the Figure 3 caption, 'slitles-like' in Sec. 5.3.2, and 'contract' instead of 'contrast' in Sec. 5.2.2. A proofreading pass is needed.","section":"Figure 3 and throughout"},{"comment":"The statement that redshift uncertainty keeps parameter biases below 5% is based on the scatter of best-fit values in Figures 4-5, but the statistical uncertainty on the ensemble-mean bias is not reported. Reporting the mean and standard error of the 200 fits would strengthen the claim.","section":"Sec. 5.2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid, well-executed mock study, but its central quantitative claims and the proposed mitigation are conditional on a random-catastrophic model and exact knowledge of f_c. The stress-test concern about clustered interlopers is valid and not addressed. I do not see a fatal flaw in the methodology, and the qualitative conclusion that slitless surveys must model f_c is likely correct. I recommend major revision to make the robustness limits explicit and to add the requested tests, rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a solid systematics paper that deserves referee time. The genuinely new piece is the combined treatment of redshift uncertainty and catastrophic failures in the EFT full-shape framework, with a quantitative comparison of freeing fc versus fixing it in the correction model. The result that redshift uncertainty is absorbed by counterterms is consistent with earlier work (Simon et al.), but it is good to see it validated on 500 Quijote realizations with two independent pipelines. The (1-fc)^2 amplitude correction is not new, but checking its accuracy in controlled mocks and demonstrating the precision cost of freeing fc is a useful contribution. The extension to w0wa and neutrino mass is a reasonable bonus, and the neutrino degradation result is worth knowing.\n\nNow the soft spots, in rough proportion. First, the slitless-like error model is hypothetical: LRG-like Gaussian smearing plus 5% log-normal catastrophics from DESI ELG repeat observations. The headline 6–16% biases, the 2.2σ shift, and the 0.85σ recovery with fixed fc all inherit that assumption. If Euclid's catastrophics are dominated by clustered line-confusion interlopers rather than a random 5% subset, the (1-fc)^2 correction can become scale-dependent, and the validation in Eq. (4.7) would not catch it. The stress-test note makes this point correctly. That said, the paper is careful to call the correction an approximation and to frame the conclusions as \"at minimum, measure fc and include it.\" So the central argument holds up; only the quantitative headline numbers are model-dependent.\n\nSecond, no code or data artifacts are released. That makes independent reproduction harder, though the mock construction is described in enough detail that someone could rebuild it. Third, the abstract's opening sentence says errors \"can bias... at sub-percent level\" before the paper reports 6–16% biases for slitless-like errors. That is an internal inconsistency in framing and should be fixed.\n\nWho gets value from this: anyone planning full-shape cosmological analyses for Euclid or other slitless surveys, and people building systematics models for DESI-style analyses. It is not a paradigm shift, but it is a careful, actionable study. I would send it to peer review and expect the authors to address the fc-model dependence and release code.\n\nRecommendation: engage with it, cite it if you work in this area, and use it as a reference for why fc estimation matters. The paper is honest about its limitations in the text, even if the abstract oversells somewhat.","headline":"A careful, useful mock-based study of redshift errors for full-shape EFT analyses; the central mitigation advice is sound, but the headline 6–16% numbers are tied to a hypothetical slitless error model.","tokens_in":20889,"tokens_out":1801,"would_cite":true,"duration_ms":21849,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that catastrophic redshift failures, not redshift uncertainty, are the main danger for full-shape cosmological fits of slitless surveys, biasing growth and amplitude estimates by 6–16% (~2.2σ).","keywords":["spectroscopic redshift errors","catastrophic redshift failures","full-shape galaxy clustering","galaxy power spectrum","effective field theory of large-scale structure","slitless spectroscopy","cosmological parameter bias","neutrino mass constraints"],"falsifier":"Take the actual redshift-validation repeat observations from Euclid's first data release, measure the catastrophic failure rate fc and the full displacement distribution, then rebuild the contaminated mocks; if fc comes out well below 5% or the displacement distribution is not symmetric, the reported 6–16% biases and 2.2σ shifts would not reproduce. Alternatively, run the same fits with a bispectrum or an alternative nuisance-parameter model; if a >1σ bias in ln(10^10 A_s) persists after applying (1-fc)^2 with fc fixed, the paper's mitigation recipe would be falsified.","tokens_in":1820,"feed_emoji":"🔭","tokens_out":2291,"duration_ms":77312,"temperature":0.7,"pith_summary":"What is the paper trying to establish? Spectroscopic surveys like Euclid measure galaxy redshifts with slitless spectroscopy, which combines small random redshift errors with a small but nonzero fraction of catastrophic line misidentifications. The paper shows, using mock galaxy catalogs, that the random errors are mostly harmless: their damping of the power spectrum is absorbed by the effective-field-theory counterterms already in the full-shape model, leaving parameter biases below 5%. The catastrophic failures are the real problem: by shifting a few percent of galaxies to wrong redshifts, they suppress the power spectrum amplitude by roughly (1-fc)^2 and bias the fractional growth rate df and log primordial amplitude ln(10^10 A_s) by 6–16%, at the ~2.2σ level for a slitless-like fc=5%. The paper's practical conclusion is that space-based slitless surveys must measure their catastrophic rate fc and include it in the clustering model—either by multiplying the predicted spectrum by (1-fc)^2 with fc fixed from calibration, or by fitting fc with a tight prior—or their growth and amplitude constraints will be systematically wrong.","feed_headline":"Slitless redshift failures bias growth and amplitude by up to 16%","feed_subtitle":"Euclid-style surveys need to measure catastrophic redshift rates or their cosmological fits will be off by 2-sigma.","key_machinery":"The mechanism is a two-part decomposition of redshift errors. Redshift uncertainty acts as a Gaussian line-of-sight velocity smearing that the effective-field-theory counterterms (alpha2, alpha4) absorb, so it leaves little imprint on cosmological parameters. Catastrophic failures, by contrast, remove a fraction fc of galaxies to very wrong redshifts, which suppresses the observed galaxy power spectrum by an approximately constant factor (1-fc)^2. This factor is the load-bearing object of the paper: it introduces a degeneracy between the catastrophic rate fc and the primordial amplitude ln(10^10 A_s), which is what drives the reported 2.2σ biases when fc is left unmodeled.","core_discovery":"The paper claims that the impact of spectroscopic redshift errors on full-shape galaxy clustering separates cleanly into two regimes. Redshift uncertainty is a line-of-sight velocity smearing whose scale-dependent damping is degenerate with the EFT counterterms (alpha2, alpha4), so standard fits recover unbiased cosmological parameters (biases <5%). Catastrophic failures cannot be absorbed: they suppress the measured power spectrum by an approximately constant factor (1-fc)^2, and when fc is 5% (slitless-like) this suppression masquerades as a lower primordial amplitude and lower growth rate—moving df and ln(10^10 A_s) by 6–16%, about 2.2σ. The paper validates a multiplicative correction (1-","pith_inferences":["Editorial extension: the paper's headline numbers are conditioned on an assumed 5% catastrophic rate and a symmetric long-tailed displacement distribution; if Euclid's real catastrophic rate differs, the reported amplitude suppression and 2.2σ shifts would scale accordingly, making the measurement of fc the decisive practical step.","Editorial extension: the strong fc–ln(10^10 A_s) degeneracy suggests that a narrow Gaussian prior on fc from repeat observations, rather than a flat prior or a fixed value, could recover most of the constraining power while still marginalizing over calibration uncertainty; the paper tests only the two extremes.","Editorial extension: higher-order statistics such as the bispectrum are a natural testable extension, since they may break the fc–amplitude degeneracy that causes the 60% degradation, restoring unbiased constraints without needing to fix fc.","Editorial extension: applying the same mock-contamination pipeline to future repeat-observation catalogs from a slitless survey would turn the hypothetical model into an empirical per-tracer error model, which is the direct path to validating the mitigation recipe."],"forward_implications":["For DESI-like galaxy populations with fc around 1%, redshift catastrophics are negligible for full-shape fits, so standard EFT analyses remain valid.","For Euclid-like slitless surveys, full-shape analysis must estimate fc and either apply the (1-fc)^2 correction or condition on fc to avoid 2.2σ biases in growth rate and primordial amplitude.","Redshift uncertainty alone is not a major threat to baseline cosmological parameters because EFT counterterms absorb its damping, though it does inflate neutrino-mass uncertainties.","Dark energy parameters w0 and wa are not biased by redshift errors, but the same errors can weaken summed-neutrino-mass constraints by up to 80% in the worst case considered.","BAO distance measurements are robust against these redshift errors, with shifts below roughly 0.3σ, so standard ruler cosmology is not the main concern."],"supporting_citations":[{"why":"Characterizes the limitations of Euclid's slitless spectroscopic measurements, motivating the combined uncertainty-plus-catastrophics scenario.","marker":"[17]"},{"why":"Introduces the Gaussian damping term used to model redshift uncertainty in the power spectrum.","marker":"[20]"},{"why":"Supplies the ELG catastrophic-failure model (log-normal displacement, fc) from repeat observations used to build contaminated mocks.","marker":"[21]"},{"why":"Analyzes redshift interlopers for slitless surveys, establishing the contamination regime the paper addresses.","marker":"[23]"},{"why":"Supplies the Quijote N-body simulation suite used to construct the clean and contaminated mock catalogs.","marker":"[42]"},{"why":"Defines the fiducial Planck 2018 cosmology used in the simulations and fits.","marker":"[43]"},{"why":"Provides the EFT-based power-spectrum code used for the full-shape likelihood evaluations.","marker":"[55]"},{"why":"Introduces the ShapeFit compression scheme used to report the fractional-growth-rate biases.","marker":"[57]"},{"why":"Sets the baseline full-shape analysis setup and priors adopted for the fits.","marker":"[62]"},{"why":"Shows the leading redshift-uncertainty corrections are degenerate with EFT counterterms, supporting the paper's absorption argument.","marker":"[63]"}],"fun_headline_variants":["Catastrophic redshift errors bias growth and amplitude by 16%","Redshift catastrophes: 5% errors, 2.2σ bias in cosmology","Slitless survey errors shift growth and amplitude by up to 16%","Euclid-style redshift failures cause 2.2σ bias in key parameters"],"cache_read_input_tokens":22528,"weakest_assumption_plain":"The central numbers assume a hypothetical slitless scenario in which 5% of galaxies get catastrophically wrong redshifts, with the wrong-redshift scatter taken from a long-tailed distribution fit to one galaxy type; if the real mission's catastrophic rate or scatter is different, the predicted biases and the best mitigation strategy would change.","fun_headline_variants_meta":{"raw":{"variants":["Catastrophic redshift errors bias growth and amplitude by 16%","Redshift catastrophes: 5% errors, 2.2σ bias in cosmology","Slitless survey errors shift growth and amplitude by up to 16%","Euclid-style redshift failures cause 2.2σ bias in key parameters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000648,"raw_usage":{"total_tokens":2883,"prompt_tokens":887,"completion_tokens":1996,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":631,"completion_tokens_details":{"reasoning_tokens":1920}},"tokens_in":631,"tokens_out":1996,"duration_ms":15873,"temperature":1.0,"reasoning_tokens":1920,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:31:06.486702+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the actual redshift-validation repeat observations from Euclid's first data release, measure the catastrophic failure rate fc and the full displacement distribution, then rebuild the contaminated mocks; if fc comes out well below 5% or the displacement distribution is not symmetric, the reported 6–16% biases and 2.2σ shifts would not reproduce. Alternatively, run the same fits with a bispectrum or an alternative nuisance-parameter model; if a >1σ bias in ln(10^10 A_s) persists after applying (1-fc)^2 with fc fixed, the paper's mitigation recipe would be falsified.","supporting_citations":[{"cited_title":"Collaboration, V.L","cited_arxiv_id":null,"evidence_quote":"Characterizes the limitations of Euclid's slitless spectroscopic measurements, motivating the combined uncertainty-plus-catastrophics scenario."},{"cited_title":"Hou, A.G","cited_arxiv_id":null,"evidence_quote":"Introduces the Gaussian damping term used to model redshift uncertainty in the power spectrum."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ELG catastrophic-failure model (log-normal displacement, fc) from repeat observations used to build contaminated mocks."},{"cited_title":"Villaescusa-Navarro, C","cited_arxiv_id":null,"evidence_quote":"Supplies the Quijote N-body simulation suite used to construct the clean and contaminated mock catalogs."},{"cited_title":"Collaboration, N","cited_arxiv_id":null,"evidence_quote":"Defines the fiducial Planck 2018 cosmology used in the simulations and fits."},{"cited_title":"Noriega, A","cited_arxiv_id":null,"evidence_quote":"Provides the EFT-based power-spectrum code used for the full-shape likelihood evaluations."},{"cited_title":"Brieden, H","cited_arxiv_id":null,"evidence_quote":"Introduces the ShapeFit compression scheme used to report the fractional-growth-rate biases."},{"cited_title":"Collaboration, A.G","cited_arxiv_id":null,"evidence_quote":"Sets the baseline full-shape analysis setup and priors adopted for the fits."}],"review_version":1}