{"id":"ae091a20-91d7-4686-988a-c2275b117526","arxiv_id":"2507.00157","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LSST supernova peculiar velocities could measure the growth rate fσ8 to 10% precision at redshift 0.02-0.14 even with photometric typing, based on realistic simulations.","lead":"This paper simulates 10 years of supernova observations for the Rubin Observatory's LSST survey and predicts that type Ia supernovae alone can measure the cosmic growth rate fσ8 to 10% precision at low redshift. The forecast tells cosmologists how well an independent, low-redshift probe can test gravity theories and complement galaxy clustering measurements.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 10% forecast rests on a Photo-typed sample whose 0.02% contamination is produced by training and testing SNN on the same simulation; Sect. 5 shows >2% contamination breaks the likelihood, so the realism of the classifier is the load-bearing assumption.","rationale":"The reader's weakest assumption is that the simulated sample and photometric classification accurately represent the real LSST survey. My stress-test identifies the same load-bearing point: the Photo-typed sample's 0.02% contamination is the linchpin of the 'most realistic' scenario, because the paper's own Sect. 5 shows that contamination above ~2% breaks the maximum-likelihood method. The in-sample training/evaluation of SNN makes this 0.02% value optimistic in a way that is not tested by the paper's internal consistency checks. The Appendix C tests (random PVs recover zero, true PVs recover input) validate the likelihood machinery but do not validate the realism of the classification. The Parsnip comparison similarly uses the same simulation for training and testing. This concern does not overturn the paper: it is a forecast, the authors explicitly acknowledge the dependence on simulation assumptions, and they propose calibration on first-year LSST data. The reader already assigned CONDITIONAL, and my analysis supports that verdict without changing it. Secondary concerns, such as the post-hoc choice of 0.02 < z < 0.14 based on simulation residuals and the persistent ~7% low offset in the 0.06-0.10 bin, reinforce conditionality but are less central than the contamination threshold: even an unbiased window and independent realizations would not save the forecast if real contamination exceeds ~2%. The recommended verdict is therefore UNCHANGED, i.e., CONDITIONAL acceptance pending demonstration that the photometric classification remains below the contamination tolerance on data or on a genuinely independent simulation.","tokens_in":28646,"tokens_out":7127,"duration_ms":89554,"concrete_test":"Retrain SNN on an independent mock built with different non-Ia SED libraries (e.g., PLAsTiCC templates not used in this paper), different SN rate priors (e.g., doubling the 91bg-like and Iax fractions), and a different SN Ia intrinsic scatter model; apply the trained classifier to the paper's 8 LSST test realizations, then apply the Sect. 2.6 SALT quality cuts. Measure the post-cut contamination and re-run the Sect. 3 maximum-likelihood fit. If contamination exceeds ~2% or the recovered f_sigma8 shifts by more than the statistical error, the 10% 'most realistic' forecast is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline precision (10% in 0.02 < z < 0.14, Sect. 4.4) is obtained from the Photo-typed sample, which reaches a post-cut contamination of only 0.021 ± 0.007% (Sect. 2.7.3). This near-zero contamination is the key reason the most realistic scenario avoids the bias mechanism demonstrated in Sect. 5, where contamination above ~2% makes the maximum-likelihood fit biased and sometimes drives f_sigma8 to zero. The problem is that the SNN classifier is trained on a simulation built with the same SED templates, SN rates, and selection cuts as the test sample (Sect. 2.5), so the performance is measured in-sample. The Parsnip comparison in Appendix B does not resolve this because Parsnip is also trained and evaluated on the same simulations. The paper itself flags the 0.02% contamination as lower than previously seen and not investigated (Sect. 2.7.3), and states that all results 'strongly depend on the assumptions made for the simulation, especially on the rates for the different SN types' (Sect. 2.3.4). If real LSST data have higher non-Ia rates, broader SED diversity, or different classifier calibration, the post-cut contamination could exceed the ~2% threshold and the central claim would fail. This is a correctness risk attached to the realism of the simulated sample, not an internal inconsistency: the pipeline is self-consistent, but the 'most realistic' label is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a forecast for measuring the cosmic growth-rate parameter fσ8 using peculiar velocities of Type Ia supernovae from LSST. The authors build simulated LSST supernova light curves on top of the Uchuu UniverseMachine mocks, using a realistic observing strategy (OpSim), DIA detection efficiencies, host-galaxy spectroscopic redshift efficiencies from DESI and 4MOST, and photometric classification with SuperNNova. They consider three scenarios: a pure 'Full' sample, a 'Spec-z' sample with realistic host spectroscopy, and a 'Photo-typed' sample including machine-learning classification and contamination from non-Ia supernovae. Using a maximum-likelihood velocity-field estimator, they recover fσ8 in several redshift bins and report a 10% precision in 0.02<z<0.14 for the Photo-typed sample, 9% for Spec-z, and 8% for Full, with tomographic errors around 12–15%. They also test the effect of contamination and find that above ~2% the estimator becomes biased and can return fσ8≈0.","tokens_in":28944,"tokens_out":4659,"duration_ms":53723,"significance":"If the forecast is correct, LSST SNe Ia would provide an independent low-redshift growth-rate measurement that complements RSD surveys such as DESI and Euclid, and would be competitive with ZTF PV constraints at z<0.06. The paper's controlled recovery tests (Appendix C) are a genuine strength: random PV inputs return null results and true PV inputs recover the injected fσ8, demonstrating that the pipeline is internally consistent. The use of the Uchuu mocks and a realistic survey simulator is also a step beyond Fisher-matrix forecasts. The main weakness is that the headline Photo-typed scenario relies on a post-cut contamination of 0.021% that is produced by training and testing the classifier on the same simulation; the paper's own stress test (Sect. 5) shows the likelihood breaks above ~2% contamination. The realism of the classifier is therefore the load-bearing assumption for the 'most realistic' label.","major_comments":[{"comment":"The headline 10% precision in the Photo-typed scenario rests on a post-cut contamination of 0.021 ± 0.007% (Sect. 2.7.3). The paper's own contamination test (Sect. 5, Fig. 9) shows that the maximum-likelihood fit becomes biased above ~2% contamination and sometimes drives fσ8 to zero. The SNN classifier was trained and tested on the same simulation (Sect. 2.5), and the Parsnip cross-check (Appendix B) is also evaluated in-sample. The paper explicitly flags that the 0.02% contamination is 'lower than has been seen before' and that a 'proper investigation of this is beyond the scope of this work' (Sect. 2.7.3), while Sect. 2.3.4 states that all results 'strongly depend on the assumptions made for the simulation, especially on the rates for the different SN types'. Since a modest degradation of classifier performance to a contamination of ~0.5–2% would materially change the forecast, the central claim is not yet robust to out-of-sample classification. I recommend adding an out-of-sample validation (e.g., train on one Uchuu realization and test on another, or calibrate against DES/ZTF photometric classification performance) and, following the method of Sect. 5, presenting the expected fσ8 precision as a function of contamination.","section":"Sect. 2.7.3 and Sect. 5"},{"comment":"The redshift window 0.02<z<0.14 is chosen after inspecting the PV residuals shown in Fig. 6, and the same window is then applied to all samples. The low-redshift cut avoids saturation and the local peculiar-velocity field; the high-redshift cut avoids the growing PV bias. However, because the window is selected from the same simulated data that are used to quote the forecast, the paper should either justify the window from independent considerations (such as the saturation limit, the Hubble-flow criterion, and a pre-specified bias threshold from simulations) or quantify the sensitivity of the headline precision to the exact window boundaries. Additionally, the persistent recovery ratio of about 0.93 in the 0.06<z<0.10 bin across all three samples, with mean reduced chi-squared values around 1.2–1.3, is suggestive of a small uncorrected systematic (possibly a residual Malmquist-bias effect at z≈0.08) rather than pure sample variance; eight realizations are insufficient to distinguish these explanations.","section":"Sect. 4.1 and Table 1"},{"comment":"The contamination stress test constructs contaminated samples by applying SALT quality cuts to non-Ia SNe, rather than by using the photometric classifier (SNN). This is a reasonable first-order test, but it does not reproduce the redshift- and type-dependent misclassification patterns of a machine-learning classifier, which may preferentially retain faint 91bg-like or core-collapse SNe at particular redshifts. The ~2% threshold should therefore be understood as a property of the likelihood under a specific contamination composition, not as a universal guarantee that any contaminant fraction below 2% is safe. The statement in Sect. 4.4 that a 'low percentage of contamination in the HD does not bias fσ8 from PVs' would be stronger if supported by a contamination test that uses the actual SNN-selected contaminants at varying thresholds.","section":"Sect. 5"}],"minor_comments":[{"comment":"The phrase 'Assuming that the velocity field is irrational' should read 'irrotational'; the intended meaning is that the velocity field is curl-free so that it can be written in terms of a divergence scalar.","section":"Sect. 3.2"},{"comment":"The sentence 'The low contamination can be a results of the fact that we are analyzing only low-redshift SNe' contains a grammatical error ('a results'), and the claimed explanation is not tested. Please rephrase and, if possible, provide a brief diagnostic (e.g., contamination as a function of redshift or SALT-fit properties).","section":"Sect. 2.7.3"},{"comment":"The text says 'the averages are always compatible with the fiducial value, except for the redshift range 0.06<z<0.1' and then repeats similar phrasing in Sects. 4.3 and 4.4. This repetition is unnecessary; a single statement that the same 0.06–0.1 bin shows a sub-2σ offset in all scenarios would be clearer.","section":"Sect. 4.2 and Table 1"},{"comment":"The notation 'σ2 8,f id' and '(f σ8)/(f σ8)fid' is typeset awkwardly; please use a consistent subscript convention (e.g., 'fid') in both the equation and the text. Also, 'kmax = 1 hMpc−1' should include a space before 'hMpc−1'.","section":"Eq. (20) and surrounding text"},{"comment":"The approximation of the Roman K213 magnitude as WISE W1 for the 4MOST color cuts is stated and then dismissed with a test that changes the color cuts. Please report the outcome of that test quantitatively (e.g., the change in Rhost or in the final Spec-z sample size) so that the reader can assess the impact.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid forecasting exercise with a self-consistent pipeline and useful recovery tests. The main issue is that the 'most realistic' scenario is calibrated entirely in-sample: the near-zero contamination is produced by training and testing the classifier on the same simulation, and the paper's own Sect. 5 shows the measurement is fragile above ~2% contamination. The authors are appropriately cautious in some places (Sect. 2.3.4) but the abstract and conclusions state the 10% precision as a likely outcome. I would encourage the editor to request an out-of-sample classifier validation or a contamination-dependent forecast before publication; this is a major revision, not a rejection, because the internal consistency of the pipeline is not in doubt."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a competent forecast, not a breakthrough. The genuinely new numbers are LSST-specific: 10% on fsigma8 in 0.02<z<0.14 with photometric typing, 14-18% in tomographic bins, and a ~2% contamination threshold above which the maximum-likelihood fit breaks. The pipeline extends the Carreres et al. ZTF framework to LSST with realistic OpSim cadence, DIA detection, DESI/4MOST host-redshift efficiencies, and SNN photometric classification. That is real work and it is clearly described.\n\nWhat it does well: the controlled tests in Appendix C are the right sanity checks. Random PVs return null; true PVs recover the input. The contamination test in Section 5 is useful and honestly executed; it quantifies where the method breaks instead of hand-waving. The paper is also candid about its own assumptions, which matters for a simulation-based forecast.\n\nSoft spots: the 10% headline depends on the Photo-typed sample reaching 0.02% contamination after quality cuts. That number comes from an SNN trained and tested on simulations built from the same SED templates, rates, and selection. The Parsnip comparison in Appendix B does not fix this, since Parsnip is evaluated on the same simulations. The authors themselves flag the contamination as lower than previously seen and leave it uninvestigated; their own Section 5 shows that above ~2% contamination the likelihood is biased. If real LSST data have higher non-Ia rates, broader SED diversity, or worse classifier calibration, the headline precision is not guaranteed. That is a correctness risk in the realism of the input, not an internal inconsistency.\n\nThe redshift window 0.02<z<0.14 is chosen after inspecting PV residuals; that is a mild post-selection on the analysis range, though the cuts are physically motivated by saturation, Hubble flow, and Malmquist bias. The persistent ~7% low offset in 0.06<z<0.10 across all samples is attributed to sample variance with only 8 realizations; that is plausible but not demonstrated. Host mismatch is neglected too; DES found ~1.7% in their sample.\n\nWho it is for: people planning low-redshift growth-rate probes, LSST DESC members, and anyone combining PV with RSD. It deserves a serious referee. I would send it out, asking the authors to sharpen the classifier-realism discussion and to state plainly that the 10% is conditional on the simulation assumptions.","headline":"A solid and honest LSST forecast paper whose headline 10% growth-rate precision is plausible but rests on a photometric classifier trained and tested on the same simulation, which no current test validates against real data.","tokens_in":29569,"tokens_out":2031,"would_cite":true,"duration_ms":25048,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.80.-k"],"model":"deepseek-v4-flash","headline":"This paper forecasts that LSST supernova peculiar velocities alone can measure the cosmic growth rate to 10 percent precision at low redshift, even when typing is photometric.","keywords":["peculiar velocities","growth rate fσ8","Type Ia supernovae","LSST forecast","maximum likelihood","photometric classification","structure growth","tests of gravity"],"falsifier":"Apply the same classifier, the same quality cuts, and the same maximum-likelihood fit to the first years of real LSST data and measure the contamination fraction from a spectroscopically typed subsample. If the real contamination exceeds roughly 2%, or if the recovered $f\\sigma_8$ in $0.02 < z < 0.14$ deviates from the value expected under the fiducial cosmology by more than the forecast uncertainty, the central claim fails.","tokens_in":28354,"feed_emoji":"🔭","tokens_out":15742,"duration_ms":146540,"temperature":0.7,"pith_summary":"The paper's central claim is that the ten-year LSST supernova sample, used on its own, can measure the cosmic growth rate $f\\sigma_8$ — the product of the linear growth rate of structure and the amplitude of matter fluctuations — to 10% precision in the low-redshift range $0.02 < z < 0.14$, using only the peculiar velocities of Type Ia supernovae. This is the forecast for the most realistic scenario, in which supernovae are typed photometrically by a neural network and the sample retains some contamination from non-Ia events. The authors build an end-to-end simulation — galaxy mocks with a correlated large-scale velocity field, the LSST observing strategy, detection and selection effects, light-curve fitting, and photometric classification — and recover the input $f\\sigma_8$ without bias in every scenario they consider. The result matters because peculiar-velocity measurements give their strongest growth signal at low redshift, exactly where galaxy-clustering estimates of $f\\sigma_8$ are limited by cosmic variance; if the forecast holds, LSST becomes an independent low-redshift probe of gravity and dark energy.","feed_headline":"10 percent cosmic-growth measurement forecast from LSST supernovae","feed_subtitle":"A realistic forecast says supernova motions alone can test gravity at low redshift, complementing galaxy clustering.","key_machinery":"The load-bearing object is the maximum-likelihood estimator for the peculiar-velocity field: a multivariate Gaussian likelihood whose covariance matrix adds per-supernova measurement errors to an analytical velocity covariance $C^{vv}$, built from the velocity-divergence power spectrum $P_{\\theta\\theta}$, a nonlinear model fitted to N-body simulations, and a small-scale damping term that models redshift-space distortions. Individual velocities are extracted from Hubble-diagram residuals through a first-order expansion that converts standardized distance-modulus residuals into line-of-sight peculiar velocities. Feeding this likelihood is a simulation stack — galaxy mocks from the Uchuu simulation supplying a correlated large-scale velocity field, the LSST Operations Simulator providing the ten-year observing calendar, difference-imaging detection efficiencies, and the SuperNNova classifier for photometric typing — which is what makes the forecast conditional on realistic survey behavior.","core_discovery":"On the paper's own terms, the discovery is that realistic survey selections do not erase the peculiar-velocity signal carried by LSST supernovae. After adding saturation limits, a detection-efficiency model for difference imaging, host-galaxy spectroscopic redshift completeness from the DESI and 4MOST surveys, SALT3 light-curve fitting with quality cuts, and neural-network photometric classification, the maximum-likelihood analysis recovers the fiducial $f\\sigma_8$ over $0.02 < z < 0.14$ without bias: 8% precision for an idealized fully spectroscopic sample, 9% when host redshifts come from DESI and 4MOST, and 10% in the most realistic photo-typed scenario, with per-bin errors near 14–18% in tomographic bins. The paper further maps where the method breaks: contamination below roughly 2% leaves the estimate unbiased, while above that level the Gaussian likelihood is no longer valid, the fitted standardization parameters drift, and some realizations return $f\\sigma_8 \\simeq 0$.","pith_inferences":["If the forecast holds, the same velocity field could be cross-correlated with galaxy density fields from DESI and 4MOST — a step the paper lists as future work — probably tightening low-redshift growth constraints beyond what velocities alone provide.","The 0.02% post-cut contamination is the fragile link: the paper flags it as lower than previously seen and uninvestigated, and its own tests show the method breaks at about 2%. A direct check is to measure the contamination of a spectroscopically confirmed subsample in early LSST data before trusting the 10% number.","Because the mocks come from a single $z=0$ snapshot, the forecast implicitly assumes no growth evolution inside $0.02 < z < 0.14$; a tomographic result inconsistent with the fiducial shape would be ambiguous between new growth physics and a limitation of the single-snapshot simulation.","The forecast's dependence on assumed supernova rates means first-year LSST rate measurements will likely move the predicted sample size; if the real rates of 91bg-like and Iax events are higher than assumed, contamination after classification rises and the 10% precision would degrade toward the paper's own break point."],"forward_implications":["In the most realistic (Photo-typed) scenario, LSST supernova peculiar velocities measure $f\\sigma_8$ at 10% precision over $0.02 < z < 0.14$, giving a low-redshift growth probe that complements redshift-space distortion constraints from DESI and Euclid.","Because the Spec-z scenario reaches 9% precision, host-galaxy spectroscopy from DESI and 4MOST improves the measurement only marginally, implying the probe does not require spectroscopic follow-up of every supernova.","The contamination study sets a concrete requirement for the analysis: keep non-Ia contamination below about 2%, beyond which the maximum-likelihood estimate becomes biased.","Tomographic bins with roughly 14–18% errors allow the growth index $\\gamma$ to be constrained, offering a direct test of general relativity against alternative gravity models.","The survey-duration forecast shows precision improving from about 20% after the first year to roughly 12% after five years in the $0.02 < z < 0.14$ bin, making the probe competitive within the survey's nominal lifetime."],"supporting_citations":[{"why":"Establishes the maximum-likelihood formalism for growth-rate measurement from peculiar velocities and the derivation of the velocity covariance matrix the paper uses.","marker":"Johnson et al. 2014"},{"why":"Predecessor simulation-plus-likelihood pipeline for ZTF that this work extends to LSST; source of the peculiar-velocity estimator and covariance validation.","marker":"Carreres et al. 2023"},{"why":"Supplies the Uchuu N-body galaxy mocks whose correlated large-scale velocity field seeds the simulated supernova redshifts.","marker":"Ishiyama et al. 2021"},{"why":"Provides the empirical nonlinear velocity-divergence power spectrum model used to construct the velocity covariance.","marker":"Bel et al. 2019"},{"why":"Supplies the small-scale damping function that models redshift-space distortions inside the velocity covariance.","marker":"Koda et al. 2014"},{"why":"The SALT3 spectral energy distribution model used both to generate and to fit the Type Ia supernova light curves.","marker":"Kenworthy et al. 2021"},{"why":"The SuperNNova classifier that performs photometric typing in the realistic scenario, including its training and hyperparameters.","marker":"Möller & de Boissière 2019"},{"why":"Provides the PLAsTiCC SED models and assumed rates for the 91bg-like and Iax contaminant populations.","marker":"Kessler et al. 2019"},{"why":"Gives the volumetric SN Ia rate used to assign supernovae to host galaxies and set the sample's redshift distribution.","marker":"Frohmaier et al. 2019"}],"fun_headline_variants":["LSST supernovae forecast 10% cosmic-growth precision","10% growth-rate forecast from LSST supernova motions","Supernova motions to pin cosmic growth at 10% error","LSST SNe to deliver 10% cosmic growth measurement","Peculiar velocities of LSST SNe test growth at 10%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecast stands or falls on whether the simulated LSST survey faithfully represents the real one — especially the assumed rates of different supernova types and the photometric classifier's real-world performance — because after quality cuts the simulated Photo-typed sample has only 0.02% contamination, which the paper itself notes is lower than previously seen and leaves uninvestigated, while its own contamination study shows the maximum-likelihood measurement becomes biased above roughly 2%.","fun_headline_variants_meta":{"raw":{"variants":["LSST supernovae forecast 10% cosmic-growth precision","10% growth-rate forecast from LSST supernova motions","Supernova motions to pin cosmic growth at 10% error","LSST SNe to deliver 10% cosmic growth measurement","Peculiar velocities of LSST SNe test growth at 10%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000896,"raw_usage":{"total_tokens":3900,"prompt_tokens":1024,"completion_tokens":2876,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":2787}},"tokens_in":640,"tokens_out":2876,"duration_ms":24915,"temperature":1.0,"reasoning_tokens":2787,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:23:13.241335+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the same classifier, the same quality cuts, and the same maximum-likelihood fit to the first years of real LSST data and measure the contamination fraction from a spectroscopically typed subsample. If the real contamination exceeds roughly 2%, or if the recovered $f\\sigma_8$ in $0.02 < z < 0.14$ deviates from the value expected under the fiducial cosmology by more than the forecast uncertainty, the central claim fails.","supporting_citations":[{"cited_title":"2014, Monthly Notices of the Royal Astronomical Society, 444, 3926","cited_arxiv_id":null,"evidence_quote":"Establishes the maximum-likelihood formalism for growth-rate measurement from peculiar velocities and the derivation of the velocity covariance matrix the paper uses."},{"cited_title":"A., et al","cited_arxiv_id":null,"evidence_quote":"Supplies the Uchuu N-body galaxy mocks whose correlated large-scale velocity field seeds the simulated supernova redshifts."},{"cited_title":"2014, Monthly Notices of the Royal Astronomical Society, 445, 4267","cited_arxiv_id":null,"evidence_quote":"Supplies the small-scale damping function that models redshift-space distortions inside the velocity covariance."},{"cited_title":"D., Jones, D","cited_arxiv_id":null,"evidence_quote":"The SALT3 spectral energy distribution model used both to generate and to fit the Type Ia supernova light curves."},{"cited_title":"E., et al","cited_arxiv_id":null,"evidence_quote":"Gives the volumetric SN Ia rate used to assign supernovae to host galaxies and set the sample's redshift distribution."}],"review_version":1}