REVIEW 5 major objections 7 minor 4 cited by
Fully photometric supernova cosmology—machine classification, SN+host photo-z, and bias-corrected distances—reaches a w0-wa figure of merit near 150 on simulated LSST deep-field data, with small residual biases.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A fully photometric LSST-DDF supernova cosmology pipeline (classification + host photo-z + BBC) reaches FoM ≈150 but shows small significant biases in w0 and wa.
T0 review reviewed 2026-08-03 challenge →
load-bearing objection Transparent, well-executed simulation study of a fully photometric LSST-DDF SN cosmology pipeline; the headline FoM≈150 is real for these idealized simulations, but the bias validation is weakened by a CMB prior computed from the true cosmology and by true-posterior photo-z PDFs. the 5 major comments →
A Fully Photometric Approach to Type Ia Supernova Cosmology in the LSST Era: Host Galaxy Redshifts and Supernova Classification
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
On the paper's own terms, the central discovery is a demonstration: a fully photometric SN Ia cosmology analysis of simulated LSST DDF data, with no spectroscopy at high redshift, produces a bias-corrected Hubble diagram whose w0-wa dark-energy figure of merit is about 150—roughly 2.8 times the DES-SN5YR value of 54. The same analysis, repeated on 25 independent realizations, reveals small but real biases in the recovered parameters (w-bias ~0.008–0.026, w0-bias up to ~0.06, wa-bias ~0.12–0.27 depending on binning and systematics), and a linear fit to Hubble residuals shows a slope with 6.5σ significance when all samples are combined. The authors traced the bias to the non-Ia contamination:
What carries the argument
The analysis chain is the central object: SALT3 light-curve fits with redshift floated as a fifth parameter and the host photo-z PDF used as a prior (the joint SN+host photometric redshift); SCONE, a convolutional neural-network photometric classifier that outputs a probability the event is a Type Ia; and BEAMS with Bias Corrections (BBC), which weights each event by classification probability and Hubble-residual likelihood while applying simulation-based distance bias corrections in a 4D space of redshift, stretch, color, and host mass. The key new term is χ²_syst, a fit prior that keeps events stable under systematic perturbations, reducing the common-event sample loss from ~12% to ~4%. A
Load-bearing premise
The whole quoted precision rests on the simulated host photo-z probability distributions being faithful stand-ins for what real LSST photo-z codes will produce—the paper uses ideal 'true posterior' PDFs and its own full host catalog has a 46% outlier fraction, so if real photo-z PDFs are less well calibrated, the FoM and bias results do not transfer.
What would settle it
Repeat the 25-realization analysis with host photo-z PDFs produced by a realistic photo-z code trained on an independent spectroscopic sample (preserving the same outlier fraction and scatter) rather than the true posterior; if the recovered w-bias moves outside ~0.01–0.03 or the w0-wa FoM drops well below 150, the central claim fails. A cheaper check: raise the SCONE probability threshold to PIa > 0.01 and test whether the 6.5σ Hubble-residual slope and wa-bias vanish, which the paper's own perfect-classifier test suggests they should.
If this is right
- If the pipeline's performance transfers to real data, a few years of LSST Deep Drilling Field observations can deliver w0-wa constraints with FoM ~150 using only photometric information, roughly tripling DES-SN5YR's constraining power.
- A spectroscopically confirmed low-redshift sample (z < 0.08, ~700 selected events here) plus a fully photometric high-z sample is sufficient; no high-z spectroscopy or classification is required.
- The residual biases, though small, are statistically significant across 25 realizations, so the method is not ready for real LSST data without further development.
- Calibration uncertainties, not photo-z uncertainties, dominate the degradation of the figure of merit; the photo-z–color anticorrelation partially self-corrects distance errors.
- A rebinned Hubble diagram achieves nearly the same precision as an unbinned one with roughly four times less compute, which matters for the much larger WFD samples to come.
Where Pith is reading between the lines
- The paper's own clue—bias disappears when true non-Ia are removed—suggests the residual bias lives in the contamination model (the BEAMS contaminant likelihood and its coupling to classification probabilities), not in the photo-z or standardization; a natural test is to perturb the non-Ia Hubble residual model and see how the w-bias responds.
- Because the simulated photo-z PDFs are ideal true posteriors with a 46% outlier fraction on the full host catalog, real photo-z estimators could easily be worse; an extension is to rerun the same pipeline with photo-z PDFs sampled from an independently trained estimator on the same host library.
- The joint SN+host photo-z improves the outlier fraction roughly fivefold over host-only photo-z (from 0.46 to 0.10 for SNe Ia), implying that light-curve information is doing most of the redshift work; pushing further on the SN term (e.g., full flux-level likelihoods) may reduce remaining biases.
- A direct forecast implication the paper leaves implicit: the ~150 FoM assumes no classifier systematic; since classifier systematics contributed roughly a third of DES-SN5YR's error budget, a realistic LSST analysis would need to add that term, likely lowering the achievable FoM.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an end-to-end simulated cosmology analysis for LSST Deep Drilling Field SNe Ia, using photometric classification (SCONE), SN+host photometric redshifts, and BEAMS with Bias Corrections (BBC). The pipeline is applied to 25 independent ELAsTiCC-based simulations with realistic light curves, non-Ia contamination, host mis-associations, and low-z spectroscopically confirmed samples. The authors report a bias-corrected Hubble diagram, statistical and systematic covariance matrices, and wCDM and w0waCDM constraints, with a headline figure of merit of about 150 for the stat+syst w0waCDM analysis, compared with DES-SN5YR's 54. They also report small but statistically significant residual biases in the recovered dark-energy parameters and describe several code improvements and diagnostic tests.
Significance. If the quoted performance transfers to real data, this would be an important step toward fully photometric supernova cosmology in the LSST era, demonstrating that photometric classification and host photo-zs can be combined with BBC to produce competitive dark-energy constraints. The paper is transparent about its idealizations and missing systematics, and it ships a concrete simulated-data pipeline with 25 realizations, making the internal consistency of the bias tests checkable. The main value is as a realistic mock challenge and methodology demonstration, but the headline FoM and bias estimates are conditional on several idealized inputs that are not yet validated against real LSST data.
major comments (5)
- [§IIID4, Tables VII–VIII] The CMB prior is circular for the bias tests: the R-shift prior is explicitly 'computed from the same cosmological parameters that were used to generate the SNe Ia'. Because the reported w-bias and w0/wa biases are obtained from fits that include this prior, the posterior is pulled toward the true cosmology by construction. The quoted 3–8σ biases therefore do not measure the SN+photo-z pipeline bias. Please re-run the bias tests with a prior centered on a different cosmology (e.g., shifted w0/wa) or report the SN-only bias without the CMB prior. This is required to support the claim that the residual biases are 'small'.
- [§II.B, Fig. 4, Table IV] The photo-z PDFs are 'true posteriors' derived from the same underlying redshift–photometry relation as the host catalog, rather than outputs of realistic photo-z estimators. The improvement in fout from 0.46 (full catalog) to 0.10 (SN+host Ia) and the resulting FoM rely on these idealized PDFs. The only photo-z systematic considered is a coherent 0.01 shift, which does not represent calibration errors, overconfidence, or outlier tails. The manuscript itself lists 'photo-z pdf variations that are more complex than a fixed shift' as missing. Please add stress tests with miscalibrated or broadened PDFs, or a realistic photo-z estimator, and discuss how the FoM and bias estimates would change.
- [§IIID3, Table VI] The common-event cut, which removes ~4% of the sample when requiring the same events to pass all systematic variations, is not modeled in the BBC biasCor simulation. The manuscript acknowledges this 'could be a potential source of unmodelled bias.' Because the central result includes residual biases, this unmodelled selection effect needs to be quantified or explicitly shown to be subdominant; otherwise it is a plausible contributor to the observed biases.
- [§III.B, §V] The 'Missing systematics' list includes several effects that dominate real SN analyses: intrinsic-scatter astrophysical model, photometric classification training sets, complex photo-z PDF variations, rate model redshift dependence, and cosmology used in bias corrections. The stat+syst FoM of ~150 and the bias estimates therefore exclude major parts of the systematic budget. Please either include estimates of these effects or clearly label the FoM as a conditional upper limit under the current systematics model.
- [§IV.D] The bias tests show that the bias disappears when all true non-Ia events are excluded, and that a cut of PIa>0.01 rejects 75% of non-Ia. However, no cosmology results are shown after applying this cut. To support the conclusion that contamination drives the residual bias, please present the w and w0/wa biases for the sample with the PIa>0.01 cut, or at least with a validated classifier-purity threshold.
minor comments (7)
- [Abstract] 'DESVYR' should be 'DES-SN5YR'.
- [Eq. (6), §IIIA] The notation in χ²_syst is under-specified: please define the index k, the meaning of p_k^ref, and σ_k^ref (fitted uncertainty from the reference fit).
- [§II.B] '11 quantiles corresponding to an integrated cumulative density function’s probabilities of 0, 10%...100%' is slightly ambiguous; consider calling this an 11-point CDF representation rather than quantiles.
- [Fig. 13 caption] Typo: 'BOA' should be 'BAO'.
- [§IV.B] 'FoM = degrades from 196 to 156' is a typo; should read 'degrades from 196 to 156'.
- [References] 'Kessler et al. 2025, in developement' has a typo ('development'). Also, ensure all arXiv identifiers and journal names are formatted consistently.
- [Fig. 4] The panels show data below z_true=0.4 but the σIQR and fout metrics are quoted for 0.4<z<1.4; please note this range directly in the panel labels or caption to avoid confusion.
Circularity Check
Bias validation is partly closed-loop: the CMB prior and BBC distance corrections are built from the same cosmology and fitted redshifts used in the simulation, and the photo-z PDFs are truth-matched posteriors.
specific steps
-
self definitional
[Sec. III.D.4, 'Cosmology Fitting and Figure of Merit' (after Eq. 13)]
"We approximate a Cosmic Microwave Background (CMB) prior using the R-shift parameter (e.g., see Eq. 69 in Komatsu et al. (2009)) computed from the same cosmological parameters that were used to generate the SNe Ia. The R-uncertainty is σR = 0.006, tuned to approximate the constraining power of Planck Collaboration et al. (2020)."
The CMB prior is centered on the exact (ΩM, w) values used to simulate the SNe and to build the BBC bias-correction simulation. The wCDM and w0waCDM fits in Tables VII-VIII therefore include an external constraint whose central value is the simulation input; any bias in the photometric pipeline is partially pulled back to the input by the prior. The reported 3-8σ bias significances are thus not an independent test of pipeline recovery of the truth; they are significances computed with the truth injected into the prior used for fitting.
-
self definitional
[Sec. III.D.2, 'BEAMS with Bias Corrections (BBC)', paragraph on subtle issues]
"The second issue concerns the µbias computation, where µtrue is computed at SALT3-fitted zphot rather than the true redshift."
In Eq. 10, the corrected distance is µ + Δµbias, with Δµbias = µ − µtrue. If µtrue is evaluated at the fitted zphot, the distance-modulus part of the correction becomes µ + (µ − µtrue(zphot)) = µtrue(zphot) (up to fitted nuisance terms and binning). The 'bias-corrected' distance is thus constructed to equal the fiducial cosmology at the fitted, possibly wrong, redshift. Any photo-z error is cancelled by construction, so the resulting Hubble diagram and the bias tests cannot reveal redshift-dependent distance biases from photo-z errors; the validation loop absorbs the very quantity it aims to test.
-
other
[Sec. II.B, 'Host Galaxies', photo-z PDF paragraph]
"The photo-z PDFs are true posteriors, derived jointly from the underlying redshift-photometry relation of the transient-specific host catalogs, rather than being the output of any specific photo-z estimators that come bundled with their own assumptions of priors/training sets that may be unrealistically well-matched to the test set relative to what we will have with real, non-simulated data."
These PDFs are built from the same underlying catalog and redshift-photometry relation that generated the simulated host redshifts. When used as the host prior in Eq. 5 (χ2_host = −2 log P_host(zphot)), the photo-z constraint in the 'fully photometric' analysis is, by construction, a perfect Bayesian posterior of the simulation. The large improvement in Fig. 4 (host fout 0.46 → SN+host fout 0.10 for Ia) is therefore an expected consequence of using truth-matched priors, not a validated prediction of what a real photo-z estimator will deliver. The paper is transparent about this idealization, but the quoted FoM ≈ 150 and the bias tests are conditional on perfectly calibrated photo-z PDFs.
full rationale
The paper is an LSST simulation forecast, so some closed-loop elements are inherent to simulation-based validation; using an input cosmology to generate data and then asking whether the pipeline recovers it is normal. However, the validation loop here is tighter than a simple end-to-end test. (1) The R-shift CMB prior is computed from the same cosmological parameters used to generate the SNe, so the fitted dark-energy parameters are pulled toward the simulation input; the significance of the residual w-bias is computed with that truth-injecting prior. (2) In BBC, µtrue is evaluated at the SALT3-fitted zphot, which by the equation µ_corrected = µ + Δµbias = µtrue(zphot) cancels photo-z-induced distance errors by construction; this is a specific, quotable reduction rather than a vague concern. (3) The host photo-z PDFs are 'true posteriors' derived from the same underlying catalog and redshift-photometry relation as the simulated data, so the fout=0.10 achieved by the joint SN+host fit is not a test against realistic photo-z estimators. The paper itself lists some of these limitations (e.g., 'photo-z pdf variations that are more complex than a fixed shift' and 'cosmology parameters used in simulated bias corrections' as missing systematics), which supports the reading that the quoted FoM and bias significances are partially self-referential. I do not find a load-bearing self-citation chain: BBC (KS17), SCONE, and the SN+host photo-z formalism are established tools with independent support, and the paper does not invoke a self-authored uniqueness theorem. The central claim is therefore not entirely circular—there are real simulated selection effects, classification contamination, and a genuine BBC implementation—but the validation of the fully photometric pipeline is partially closed-loop, meriting a score of 6 rather than a lower score.
Axiom & Free-Parameter Ledger
free parameters (4)
- σ_R CMB prior width =
0.006
- zphot systematic shift =
0.01
- 1σ prior scale in χ²_syst (Eq. 6) =
1σ prior
- Low-z sample size =
4200 generated
axioms (6)
- domain assumption SALT3 SED model + NIR extension adequately models Type Ia supernova spectral energy distributions
- ad hoc to paper Simulated photo-z PDFs are true posteriors, not outputs of realistic photo-z estimators
- ad hoc to paper CMB prior using R-shift is computed from the same cosmology as the simulated SNe
- domain assumption Volumetric SN Ia and CC rates from Dilday 2008, Hounsell 2018, Strolger 2015, etc.
- domain assumption Host galaxy libraries (ELAsTiCC HOSTLIBs) and WGTMAPs preserve transient-host correlations
- domain assumption Flat wCDM/CPL cosmology with Ωm=0.315, w=-1 used to generate simulations
Cite this review
Pith. "Pith review of A Fully Photometric Approach to Type Ia Supernova Cosmology in the LSST Era: Host Galaxy Redshifts and Supernova Classification." pith.science (2026). https://pith.science/paper/VI3XQQB7
@misc{pith2026251206319,
author = {Pith},
title = {Pith review of: A Fully Photometric Approach to Type Ia Supernova Cosmology in the LSST Era: Host Galaxy Redshifts and Supernova Classification},
year = {2026},
howpublished = {\url{https://pith.science/paper/VI3XQQB7}},
note = {Machine review of arXiv:2512.06319}
}
read the original abstract
The upcoming Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) is expected to discover nearly a million Type Ia supernovae (SNeIa), offering an unprecedented opportunity to constrain dark energy. The vast majority of these events will lack spectroscopic classification and redshifts, necessitating a fully photometric approach to maximize cosmology constraining power. We present detailed simulations based on the Extended LSST Astronomical Time Series Classification Challenge (ELAsTiCC), and a cosmological analysis using photometrically classified SNeIa with host galaxy photometric redshifts. This dataset features realistic multi-band light curves, non-SNIa contamination, host mis-associations, and transient-host correlations across the high-redshift Deep Drilling Fields (DDF) (~ 50 deg^2). We also include a spectroscopically confirmed low-redshift sample based on the Wide Fast Deep (WFD) fields. We employ a joint SN+host photometric redshift fit, a neural network based photometric classifier (SCONE), and BEAMS with Bias Corrections (BBC) methodology to construct a bias-corrected Hubble diagram. We produce statistical + systematic covariance matrices, and perform cosmology fitting with a prior using Cosmic Microwave Background constraints. We fit and present results for the wCDM dark energy model, and the more general Chevallier-Polarski-Linder (CPL) w0wa model. With a simulated sample of ~6000 events, we achieve a Figure of Merit (FoM) value of about 150, which is significantly larger than the DESVYR FoM of 54. Averaging analysis results over 25 independent samples, we find small but significant biases indicating a need for further analysis testing and development.
Figures
Forward citations
Cited by 4 Pith papers
-
Machine Learning Closure Audits for LSST Photometric Supernova Cosmology
Out-of-fold LightGBM recovers up to R²=0.982 of bias-corrected Hubble residual variance in LSST SN Ia mocks and R²=0.725 on DES 5YR, with consistent SHAP rankings, while redshift bins explain <1%.
-
Calibration-Induced Systematics in SALT3 Training and Their Impact on Dark Energy Constraints from Stage IV Supernova Surveys
Calibration uncertainties during supernova light-curve fitting cause roughly 50% degradation in dark energy figure of merit for Stage IV surveys, dominating over 13% degradation from model training errors and showing ...
-
Variational Graph Neural Networks for Uncertainty Quantification in Inverse Problems
A decoder-only variational GNN recovers elastic moduli and 3D hyperelastic loads from displacement fields while reporting cognitive and statistical uncertainty at lower cost than full Bayesian nets.
-
Impact of Calibration Systematics on Dark Energy Constraints from LSST Type Ia Supernovae
Linear passband tilts shift best-fit w0 and wa by ~0.025 sigma and enlarge the w0-wa contour area by ~5% per 1%/100nm increase; quadratic tilts yield less conclusive results.
Reference graph
Works this paper leans on
-
[1]
Photometric Classification: SCONE 10 D
Host Galaxy Association 10 C. Photometric Classification: SCONE 10 D. BEAMS with Bias Corrections (BBC) 10
-
[2]
σIQR = RMS of the inter quartile distribution of ∆z(1+z) divided by 1.349,
-
[3]
Covariance Matrix 12 arXiv:2512.06319v1 [astro-ph.CO] 6 Dec 2025 2
Pith/arXiv arXiv 2025
-
[4]
BEAMS with Bias Correc- tion s
Cosmology Fitting and Figure of Merit 12 IV. Cosmology Results 13 A. wCDM Results 13 B. w0waCDM Results 13 C. Comparison with DES-SN5YR and improvements with DESI 14 D. Bias Tests 16 V. Conclusions 16 Acknowledgments 16 References 18 I. INTRODUCTION The discovery of dark energy and the accelerating ex- pansionoftheUniverse, firstidentifiedthroughtheobser-...
1999
-
[5]
∆z(1+z) = mean bias on∆z(1+z),
-
[6]
Total generated
fout = the outlier fraction of events satisfying |∆z(1+z)|> 0.1. Our metric values are ∆z(1+z) = 0.001, σIQR = 0.127 and fout = 0.46 (see Fig. 4(a)). Weight Maps Unlike the PLaSTiCC HOSTLIB used in Mitra et al. (2023), the use ofELAsTiCC HOSTLIBs enables modelling ofthehostgalaxycorrelations. Hereweweighteachtran- sient host as a function of log of stella...
2023
-
[7]
at least three bands with maximum signal to noise ratio SNR> 4 2.|x1|< 3.0 3.|c|< 0.3
-
[8]
stretch uncertaintyσx1 < 1.0
-
[9]
time of peak brightness uncertaintyσt0 < 2.0 days
-
[10]
No explicit cuts are applied onSCONE classification
Pfit > 0.0510. No explicit cuts are applied onSCONE classification
-
[11]
Fitted bands satisfy 2800 <⟨λcen⟩/(1 +zphot) < 17000 the valid range of the rest frame wavelengths fortheSALT3model; samecutisappliedfor zphot± σ(zphot)
-
[12]
The statistics after cuts are presented in the last col- umn of Table III, and the transient type fractions (for high-z) are shown in Fig
valid bias correction for all systematic variations (see Section IIID). The statistics after cuts are presented in the last col- umn of Table III, and the transient type fractions (for high-z) are shown in Fig. 2. The high-z true-Ia fraction is 15% for physical rates (Fig. 2a), increases to∼ 50% after two detection triggers (Fig. 2b), and increases again ...
2023
-
[13]
astrophysical model of intrinsic scatter, which dom- inates the systematic error budget in DES-SN5YR
-
[14]
photometric classification (e.g, different classifiers and training sets), which contributes∼ 1/3 to the total DES-SN5YR systematic (see Table 7 in Vin- cenzi et al. (2024))
2024
-
[15]
photo-z pdf variations that are more complex than a fixed shift
-
[16]
simulated redshift dependence of volumetric rate
-
[17]
cosmology parameters used in simulated bias cor- rections. The first 7 cuts are applied only to the nominal anal- ysis without systematics; for each systematic variation we do not apply these 7 cuts but instead process the list of events passing cuts from the nominal analysis. For example, if an event has fitted SALT3 color parameter c = 0.299, and migrat...
-
[18]
2016, Sullivan et al
Host Galaxy Association Host galaxy matching is done via the Directional Light Radius (DLR) method (Gupta et al. 2016, Sullivan et al. 2010). A normalized, dimensionless parameterdDLR is defined as the ratio of the angular separation∆θ between the supernova (SN) and DLR defined as the galaxy’s cen- troid to the galaxy’s effective radius in the direction o...
2016
-
[20]
(2007), Newling et al
BEAMS The BEAMS framework was introduced by Kunz et al. (2007), Newling et al. (2012), Hlozek et al. (2012), Knights et al. (2013). In the BEAMS framework, the 11 Supernova Classification with a Convolutional Neural Network 11 0 2 4 6 Directional Light Radius 100 101 102 103 (a) No Host (2%) Match (98%) Double Match all (12%) Double Match correct (11%) 18...
2007
-
[21]
(2011) method of measuring distances independently of cosmological parameters
BBC BBC is a fitting framework that accounts for selec- tion effects using a simulation, incorporates BEAMS to account for non-Ia contamination, and implements the Marriner et al. (2011) method of measuring distances independently of cosmological parameters. BBC reads the SALT3 fitted parameters (high-z and low-z) from the data and biasCor simulation (sec...
2011
-
[22]
This covariance matrix (C) is C = Cstat + Csyst (11) where Cstat is the diagonal covariance representing mea- surementerrors
Covariance Matrix The uncertainties in SN distance measurements are en- capsulated in a covariance matrix, which quantifies both statistical and systematic uncertainties. This covariance matrix (C) is C = Cstat + Csyst (11) where Cstat is the diagonal covariance representing mea- surementerrors. Csys isthesystematiccovariancematrix (Conley et al. 2011), C...
2011
-
[23]
Cosmology Fitting and Figure of Merit In cosmology fitting, we minimize the followingχ2: χ2 = ∆µT C−1∆µ, (13) where C is the covariance matrix (Eq. 11), and∆µ is the vector of differences between observed and theoreti- cal distance moduli. We use the same fast minimization code13 that was used in Mitra et al. (2023). We approxi- mate a Cosmic Microwave Ba...
Pith/arXiv arXiv 2023
-
[2021]
and SuperNNova Möller & de Boissière (2020b) We use SCONE as our photometric classifier to de- termine the probability (PIa) that each event is a SN Ia. SCONErequiresphotometricdataonly, withouttheneed for accurate redshift estimation, and it has relatively low computational and dataset size requirements for achiev- ing high accuracy. For training the net...
This paper was first reviewed by deepseek-v4-flash on August 3, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.